Storage system and method for locating an anomaly of a storage system
Patent Information
- Application Number
- CN202611339873.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-31
- Publication Date
- 2026-09-29
AI Technical Summary
[0006]本发明提供了存储系统及存储系统的异常定位方法,以至少解决相关技术中无法高效、完整地还原存储系统异常时的传输时序和交互过程,存在故障定位过程效率低下,难以精准溯源的问题
[0009]通过本发明,由于在扩展器的上行端口和下行端口均设置记录单元以实时生成时序数据,并由第一处理器将其存储于内部集成的存储介质中,从而在硬件底层实现了全路径数据流的完整记录与存储;同时,主机端第二处理器通过软件架构内的探针实时监控日志数据,并在识别到异常事件时主动生成读取指令,按需从扩展器中提取对应的时序数据,进而自动解析生成时序全景图,通过时序全景图将异常位置直观映射至具体的物理端口或链路环节,从而实现对存储系统异常位置的精准定位、显著提升故障排查准确率。由此,解决了无法高效、完整地还原存储系统异常时的传输时序和交互过程,导致故障定位过程效率低下,难以精准溯源的问题。
Smart Images

Figure CN122838162A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to storage systems and methods for locating anomalies in storage systems. Background Technology
[0002] In storage systems, the SAS (Serial Attached SCSI) interface is widely used due to its high reliability. To meet the demands of massive storage, SAS systems typically employ a multi-level expander cascade topology, connecting a large number of hard drives to the host. However, this complex cascaded architecture makes the I / O transmission path lengthy, and any failure in any link or device can lead to I / O failures, posing a significant challenge to fault location.
[0003] Fault location methods in related technologies mainly rely on software logs for troubleshooting, but they have the following obvious limitations: Coarse-grained information: Software logs typically record summary fault descriptions or statistical information (such as "CRC error count exceeded" or "OPEN_REJECT occurred"), with time precision mostly at the millisecond or microsecond level. However, the underlying interactions of the SAS protocol occur at the nanosecond level, and logs cannot record the abnormal details of protocol timing in real time and with precision.
[0004] Limited location capabilities: Logs typically only record anomaly conclusions for the local device or its own layer. For example, the initiator's logs may only record that it received an error, but cannot determine on which physical link the error occurred (e.g., from the initiator to the first-level expander, or between two expanders). These highly aggregated, conclusive logs alone are insufficient to reconstruct the complete transmission scenario at the time of the failure.
[0005] Furthermore, while external protocol analyzers can be used for fault location in related technologies, and can capture low-level protocol interactions with nanosecond-level precision, they must be connected in series on a physical link and can only monitor one link at a time. In multi-level cascaded scenarios, it is impossible to capture the end-to-end full path timing simultaneously, making it difficult to determine whether the anomaly originates from this link segment or is propagated from upstream. Summary of the Invention
[0006] This invention provides a storage system and a method for locating anomalies in a storage system, thereby addressing the problems in related technologies, such as the inability to efficiently and completely reconstruct the transmission timing and interaction process when a storage system malfunctions, resulting in low efficiency in fault location and difficulty in accurately tracing the source of the fault.
[0007] This invention provides a storage system, comprising: a host end and at least one expansion end; the expansion end is provided with at least one expander and at least one storage device; the expander includes a first processor, an uplink port, a downlink port, and at least one device port, the device port being used to connect to the storage device, both the uplink and downlink ports being provided with recording units to generate time-series data based on the data stream of the corresponding ports, the first processor integrating a storage medium to store the time-series data recorded by the recording units; the host end includes at least one second processor, the software architecture of the second processor being provided with at least one probe, the probe being used to monitor the log data of the software architecture; when the second processor identifies an abnormal event based on the log data, it generates a read instruction, reads time-series data from the storage medium in the expander according to the read instruction, generates a time-series panorama based on the at least one time-series data read, and locates the abnormal location of the storage system based on the time-series panorama.
[0008] The present invention also provides an anomaly localization method for a storage system. The method is applied to the second processor of the storage system described in the above embodiment. The method includes: monitoring the log data of the software architecture through probes set within the software architecture; generating a read instruction when an abnormal event is identified based on the log data; reading timing data from the storage medium within the expander based on the read instruction; generating a timing panorama based on at least one piece of time-series data read; and locating the abnormal location of the storage system based on the timing panorama.
[0009] This invention achieves complete recording and storage of the entire data flow at the hardware level by setting recording units on both the uplink and downlink ports of the extender to generate time-series data in real time, and storing this data in the internally integrated storage medium by the first processor. Simultaneously, the second processor on the host side monitors log data in real time through probes within the software architecture, and actively generates read commands when abnormal events are detected. It extracts the corresponding time-series data from the extender as needed, and automatically parses and generates a time-series panorama. This panorama visually maps the abnormal location to a specific physical port or link, enabling precise location of storage system anomalies and significantly improving the accuracy of fault diagnosis. Therefore, it solves the problem of inefficient and incomplete reconstruction of transmission timing and interaction processes during storage system anomalies, leading to inefficient fault location and difficulty in accurate source tracing. Attached Figure Description
[0010] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1A block diagram of a storage system provided in an embodiment of the present invention; Figure 2 A topology diagram of a storage system provided in one embodiment of the present invention; Figure 3 An interaction timing diagram provided for one embodiment of the present invention; Figure 4 This is a flowchart of an anomaly localization method for a storage system provided in an embodiment of the present invention. Detailed Implementation
[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.
[0013] It should be noted that, in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0014] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0015] Before introducing this invention, a brief overview of the basic concepts and technical points involved in this invention is provided below: In the SAS architecture, three core protocols are defined: SSP (Serial SCSI Protocol), SMP (Serial Management Protocol), and STP (Serial ATA Tunneling Protocol).
[0016] 1. The SSP protocol is used to execute SCSI commands and manage tasks, responsible for transmitting SCSI commands and reading / writing data between the Initiator and the Target (i.e., the SAS hard drive). Specifically: a) The Initiator sends an SSP to the hard drive: This is used to perform I / O (input / output) operations, including issuing SCSI read / write commands, transferring data, and receiving status responses from the hard drive.
[0017] b) The Initiator sends an SSP to the Expander: The Expander itself can be used as an SSP target port. The Initiator can perform certain configuration or status queries on the Expander through the SSP. However, in practical applications, the Expander is mainly used as a forwarding node for SSP frames rather than the final target.
[0018] 2. The SMP protocol is specifically designed for managing SAS topologies and Expander devices. Its characteristics include: a) Only the Initiator can initiate an SMP connection.
[0019] b) The Initiator sends SMP to the Expander: This is used to perform management functions for the Expander and PHY (Physical Layer Interface), including topology discovery, PHY status reporting, routing table configuration, zoning management, firmware upgrades, and other management operations.
[0020] c) SAS hard drives typically do not contain an SMP Target; only Expanders contain an SMP Target to respond to SMP requests.
[0021] In this embodiment of the invention, because the interaction methods of SSP and SMP instructions are different, their parsing methods need to be described separately.
[0022] 1) SMP command frame format and unique identifier The SMP transport layer defines only two frame types: SMP Request and SMP Response. The communication process is as follows: the Initiator sends an SMP Request frame to the Expander, and the Expander returns an SMP Response frame after processing it. The entire process is completed within a single, complete connection.
[0023] Each SMP request corresponds to a specific SMP Function, identified by the Function Code field (e.g., DISCOVER = 0x01, REPORT GENERAL = 0x02, REPORT PHY = 0x03, etc.). The key field for achieving a unique match between SMP requests and responses is the Function Code.
[0024] 2) Unique identifier for the frame format of SSP instructions The headers of SSP COMMAND and TASK frames contain an important field: IPTT (InitiatorPort Transfer Tag). IPTT is a 16-bit tag value assigned by the initiator to uniquely identify a SCSI command. IPTT plays a crucial context identification role in the SSP protocol. 1) Command Identifier: The Initiator sends a COMMAND frame with IPTT=N as the start of a command processing interaction. When the Target processes the command, it will carry the same IPTT=N in the response frame.
[0025] 2) Response matching: The initiator determines which command has been completed by matching the IPTT in the response frame and associating the response with the previously issued command.
[0026] 3) Lifecycle Management: Based on the SSP protocol, the SAS Initiator uses the receipt of a response frame with IPTT value N as a marker that the command processing is complete. Afterward, IPTT N can be recycled and reused to process new commands.
[0027] 4) Error recovery: In the event of abnormal situations such as ACK / NAK loss, IPTT is used to determine whether a duplicate response frame has been received, so as to avoid incorrectly matching the retransmitted response to the new command, thereby preventing hard disk failure.
[0028] Before transmitting data, the SSP protocol requires establishing a connection. For two SAS devices to communicate, they must first open a connection using an OPEN Address Frame, and then select the appropriate PHY (Physical Layer Interface) to establish the connection. The SSP protocol's transmit and receive format involves two layers of elements: Primitives and Frames. The relationship between primitives and frames is as follows: Primitives are the "signals," responsible for establishing, maintaining, and dismantling the link; frames are the "cargo," carrying the actual SCSI commands and data. They work together—after a connection is established using primitives, frames can be transmitted over a stable link.
[0029] In the SAS architecture, the Expander plays a role similar to a switch / router in a network, situated in the middle layer of the system, connecting upwards to the Initiator and downwards to the Target (hard drive). The Expander's specific role in the SSP command forwarding process is as follows: 1) ECM Smart Router The Expander device includes an ECM (Expander Connection Manager), which is responsible for resolving the target address in the OPEN addressframe, querying the routing table, and performing arbitration decisions. When a "frame or primitive" is received on the input of a certain PHY, the internal routing table is consulted to determine on which PHY to output and forward it.
[0030] 2) Frame processing mechanism For SSP frames that pass through, Expander does not parse the frame content at all above the port layer, but only performs transparent conversion to primitive conversion processing mechanism 3). The Expander does not simply "pass through" the primitives; rather, it involves multiple transformation relationships: One-to-one direct forwarding: Upon receiving the ALIGN primitive, ALIGN is output directly (for link maintenance). Upon receiving the SYNC primitive, SYNC is output directly (while remaining in idle state). One-to-many forwarding: Upon receiving the BROADCAST primitive, forward it to all other ports except the source port, thus enabling the broadcast event to propagate throughout the domain. Primitive transformation (converting output when input error occurs): If an OPEN Address Frame is received but the target PHY is busy, return the OPEN_REJECT(RETRY) primitive. If an OPEN Address Frame is received but no match is found in the routing table, return the OPEN_REJECT(NODESTINATION) primitive. If an OPEN Address Frame is received but the target PHY does not exist, return the OPEN_REJECT(DESTINATION NOT EXIST) primitive. If a link error (such as Disparity Error) occurs at the input end, it is converted to an ERROR primitive to notify the peer end. Link rate mismatch → Expander inserts ALIGN primitives to negotiate rate.
[0031] Many-to-one aggregation: Multiple input port connection requests compete for the same output port → ECM arbitrates, the winner gets the connection, and other requests return OPEN_REJECT.
[0032] Furthermore, for ease of subsequent understanding, this invention requires a brief description of the basic concepts involved. Primitive primitives are special control signals defined by the SAS link layer, used for link management, flow control, and connection control. They are DWORD (double-word, 4-byte) level codes that are continuously transmitted on the link.
[0033] SSP frames are the true information carriers of SCSI commands, data, and status. The SSP transport layer defines five frame types: COMMAND frames: transmit SCSI command requests; RESPONSE frames: transmit response status after command execution; DATA frames: transmit read / write data; XFER_RDY frames: the target notifies the initiator that it is ready to receive data; TASK frames: task management commands (such as ABORT TASK).
[0034] A controller enclosure (or main enclosure) is a chassis unit within a SAS storage system that contains the core controller (CPU, memory, HBA card). The controller enclosure is responsible for running the storage operating system, handling I / O requests, managing the SAS domain topology, and providing external service interfaces (such as FC, iSCSI, and SAS front-end ports). The CPU inside the controller enclosure acts as the initiator within the SAS domain.
[0035] An expansion enclosure (also known as a JBOD, Just a Bunch of Disks) is a chassis unit in a SAS storage system that contains only an expander chip and hard drive bays, but not the core controller. Expansion enclosures are cascaded to the main enclosure or a higher-level expansion enclosure via SAS cables. The expander chip inside the enclosure is responsible for amplifying and distributing the upstream SAS signals to each hard drive bay within the enclosure. The main function of expansion enclosures is to increase the number and capacity of hard drives in a storage system; they can be cascaded in multiple levels to support hundreds or thousands of hard drives.
[0036] Uplink port: This refers to an aggregation of a group of PHYs on the SAS Expander used to connect to higher-level devices (such as the HBA card of the Initiator, or the downlink port of the previous-level Expander). In the SAS protocol, multiple PHYs can be aggregated into a wide port to provide higher aggregate bandwidth and link-level redundancy. The uplink wide port is the entry point for data flowing from the Initiator to the Expander, and also the exit point for data returning from the Expander to the Initiator. In the physical topology, the device connected to an Uplink wide port of an Expander is the parent node of that Expander in the SAS domain topology tree. (This article describes a wide port consisting of 4 PHYs.)
[0037] Downlink port: This refers to an aggregation of PHYs on the SAS Expander used to connect to the next-level Expander. In contrast to the uplink wide port, the downlink wide port is the exit point for data to continue forwarding from the current Expander downstream, and also the entry point for data returning from downstream. In a multi-level cascaded topology, the device connected to the downlink wide port is a child node (the next-level Expander) of the current Expander in the SAS domain topology tree. The downlink wide port is also composed of multiple PHYs, providing high-bandwidth cascaded links. (This article describes a wide port consisting of four PHYs.)
[0038] A device port refers to a single PHY (or a single PHY not aggregated with other PHYs into a wide port) on the SAS Expander used for direct connection to a terminal hard drive device. Unlike wide ports, narrow ports consist of only a single PHY and do not involve multi-PHY aggregation. Each narrow hard drive port corresponds to the connection of one physical hard drive, and the PHY on the port is directly connected to the SAS physical port of the hard drive via a SAS cable or backplane wiring. (This article describes one narrow port consisting of one PHY).
[0039] The IPTT (Initiator Port Transfer Tag) is a key field in the SSP (Serial SCSI Protocol) frame header of the SAS protocol. The IPTT is assigned by the Initiator to each SCSI command, serving as a unique identifier for that command throughout its entire lifecycle. For a complete SSP command interaction, all related frames—including COMMAND frames, XFER_RDY frames (write operations), DATA frames, and RESPONSE frames—carry the same IPTT value. The IPTT is globally unique within the SAS domain (within the Initiator port scope), and therefore, in this invention, it is used as the core association key for all related frames and primitives that associate the same SCSI command across devices and ports.
[0040] The following description, with reference to the accompanying drawings, describes a storage system and its anomaly localization method according to embodiments of the present invention. Addressing the problems mentioned in the background section, the present invention provides a storage system that generates real-time time-series data by setting recording units on both the uplink and downlink ports of the extender. This data is then stored internally by a first processor in an integrated storage medium, thus achieving complete recording and storage of the entire data flow at the hardware level. Simultaneously, a second processor on the host side monitors log data in real-time using probes within its software architecture. Upon detecting an anomaly, it actively generates a read command, extracting the corresponding time-series data from the extender as needed. This data is then automatically parsed to generate a time-series panorama, which visually maps the anomaly location to a specific physical port or link, thereby achieving precise localization of the storage system's anomaly location and significantly improving the accuracy of fault diagnosis. This solves the problem of inefficient and incomplete reconstruction of the transmission timing and interaction processes during storage system anomalies, leading to inefficient fault localization and difficulty in accurately tracing the source.
[0041] Specifically, Figure 1 This is a block diagram of the storage system according to an embodiment of the present invention.
[0042] like Figure 1 As shown, the storage system includes a host terminal 200 and at least one expansion terminal 100.
[0043] The expansion terminal 100 is provided with at least one expander and at least one storage device. The expander includes a first processor 101, an uplink port 102, a downlink port 103 and at least one device port 104. The device port 104 is used to connect to the storage device. Both the uplink port 102 and the downlink port 103 are provided with recording units to generate time-series data according to the data stream of the corresponding port. The first processor 101 integrates a storage medium to store the time-series data recorded by the recording units. The host 200 includes at least one second processor 201. The software architecture of the second processor 201 is equipped with at least one probe, which is used to monitor the log data of the software architecture. When the second processor 201 identifies an abnormal event based on the log data, it generates a read instruction, reads timing data from the storage medium in the expander based on the read instruction, generates a timing panorama based on the at least one time-series data read, and locates the abnormal location of the storage system based on the timing panorama.
[0044] It should be noted that the host terminal in this embodiment of the invention is the main cabinet described in the above embodiments, and the expansion terminal is the expansion cabinet described in the above embodiments.
[0045] It is understood that, at the hardware link level, the extension end of this embodiment of the invention is equipped with at least one extender and at least one storage device. The extender, as the core hub of data routing, includes a first processor 101, an uplink port 102, a downlink port 103, and at least one device port 104, wherein the device port 104 is used to connect to the storage device. In this embodiment of the invention, to break the "black box" state of the underlying devices, this embodiment is equipped with a recording unit on both the uplink port 102 and the downlink port 103 of the extender, which is used to monitor the data flow of the corresponding port in real time and generate nanosecond-level timing data; at the same time, the first processor 101 integrates a storage medium for caching and storing the timing data generated by the above-mentioned recording unit, thereby completely preserving the complete process of IO interaction at the hardware level.
[0046] At the host software level, the host includes at least one second processor 201, whose software architecture incorporates at least one probe. This probe non-intrusively monitors the software architecture's log data in real time, acting as a listener for abnormal events. When the second processor 201 identifies an abnormal event (such as an I / O timeout or error) based on the log data monitored by the probe, it immediately generates a read command. The host then extracts the corresponding timing data from the storage medium within the expander based on this read command. Finally, the second processor 201 performs cross-level correlation between the host's software logs and the timing data at the expander's underlying layer, generating an end-to-end full-path timing panorama. Based on this timing panorama, it accurately locates the abnormal position and root cause of the storage system failure.
[0047] Therefore, this embodiment of the invention achieves real-time capture and caching of underlying hardware-level data streams by setting recording units on the uplink port 102 and downlink port 103 of the expander, and having the first processor 101 temporarily store the time-series data in the internal storage medium. When the probe in the host (second processor 201) detects an abnormal event at the software architecture level, it can immediately generate a read command to accurately retrieve the full amount of time-series data before and after the anomaly from the expander. This lowers the threshold for troubleshooting complex storage systems, allows clear observation of the specific physical port and interaction links where the anomaly occurred, visualizes the fault location, significantly improves the location accuracy, and ensures the integrity and timeliness of the fault scene data.
[0048] Specifically, to meet the demand for a large number of hard drives for massive data storage, the SAS protocol supports flexible topology expansion capabilities. A typical SAS storage topology employs a cascading architecture, with the host acting as the initiator via an HBA (Host Bus Adapter), extending the link through multiple levels of expanders, and ultimately connecting to a large number of target devices, i.e., hard drives. Specifically, this cascading architecture includes the following layers: First level: Initiator layer: The host HBA card acts as the initiator of the SAS domain, responsible for initiating SCSI commands, managing the topology discovery and device enumeration of the SAS domain, and is the starting point for IO requests.
[0049] The second layer: Expander layer: The Expander is a key expansion device in the SAS topology. Each Expander provides multiple physical ports (PHYs) and supports wide port aggregation to improve bandwidth. Through the cascading of Expanders, a single Initiator can be expanded to connect dozens or even hundreds of hard disk devices. The Expander internally includes Expander Function and Table Routing mechanism, which are responsible for frame forwarding and routing decisions.
[0050] The third layer: Target layer: The terminal hard disk device acts as the SAS Target, receiving and executing SCSI commands issued by the Initiator to complete data read and write operations. This multi-level cascading architecture enables the storage system to connect a large number of hard disk devices with limited Initiator port resources, building a high-density storage array and effectively meeting the storage capacity expansion needs of enterprise applications.
[0051] As one possible way to achieve this, such as Figure 2 As shown, the host terminal in this embodiment of the invention is Figure 2 The main cabinet shown has an expansion end as follows: Figure 2 The expansion cabinet (JBOD) shown has an expander as follows: Figure 2 The Expander shown.
[0052] This storage system adopts a master-slave cascaded architecture in its physical topology, mainly including one main enclosure and at least one expansion enclosure (e.g., JBOD-A, JBOD-B, and JBOD-C). In this embodiment of the invention, the master end is connected to three expansion enclosures, named JBOD-A, JBOD-B, and JBOD-C respectively; each JBOD contains two expanders for connecting two ports of a disk; each expander contains: one upstream wide port, one downstream wide port, and N hard disk ports.
[0053] To facilitate subsequent understanding, this embodiment of the invention provides a brief overview of the inherent characteristics of data flow under complex topologies within a storage system, and the underlying logic for achieving full-path time-series reassembly based on these characteristics, as follows: 1. Hardware-level deployment of the timing recording unit inside the extender In SAS storage systems, the expander serves as the core switching node, containing multiple physical layer interfaces (PHYs). Taking an expander with 36 PHYs as an example, this expander integrates 36 independent timing recording units. Each timing recording unit is bound to its corresponding PHY, used to capture and record the characteristics of the raw data stream passing through that physical link in real time at the hardware level.
[0054] 2. Dynamic routing transmission of SSP commands in multi-level cascaded topologies When Serial SCSI (SSP) commands are transmitted over a wide port, an internal arbitration mechanism is used to dynamically select the physical link. Taking the host sending an I / O command to the target hard drive as an example, the data transmission process is as follows: (1) Sending path (Initiator → Hard disk): Between the initiator and each level of cascaded extender, one of the multiple physical links on the wide port is dynamically selected to establish a connection and transmit data; when the data arrives at the extender that is directly connected to the target hard disk, a connection is directly established through the narrow port to transmit the data to the target hard disk.
[0055] like Figure 2 As shown, taking the main cabinet sending IO commands to DiskC as an example: the main cabinet's initiator (initiator1) and initiator (initiator2) send IO commands to the hard drive DiskC located on expansion cabinet C (JBOD-C) through expander A1 (ExpanderA1), expander B1 (ExpanderB1), and expander C1 (ExpanderA1), respectively. The specific path is as follows: a) Between Initiator1 and Expander A1: Select one of the four links in the wide port to establish a link and transmit data.
[0056] b) Between Expander A1 and Expander B1: Select one of the four links in the wide port to establish a link and transmit data.
[0057] c) Between Expander B1 and Expander C1: Select one of the four links in the wide port to establish a link and transmit data.
[0058] d) Between Expander C1 and DiskC: There is only one narrow port for direct connection to transfer data to the hard drive.
[0059] (2) Return path (hard disk → Initiator): After the data returned by the target hard disk enters the extender through the narrow port, it is dynamically selected through the wide port between each level of cascaded extender and sent back to the initiator hop by hop.
[0060] Based on the above embodiment, the hard drive returns data to Initiator1, and the specific path is as follows: a) Between DiskC and Expander C1: There is only one narrow port, which directly establishes a connection and transfers data to the hard drive.
[0061] b) Between Expander C1 and Expander B1: Select one of the four links in the wide port to establish a link and transmit data.
[0062] c) Between Expander B1 and Expander A1: Select one of the four links in the wide port to establish a link and transmit data.
[0063] d) Between Expander A1 and Initiator1: Select one of the four links in the wide port to establish a link and transmit data.
[0064] In addition, in this embodiment of the invention, the host will also maliciously send SSP (data read / write) commands to Expander. The link is selected through wide port internal arbitration, and the transmission process can be referred to the above embodiment.
[0065] 3. SMP command path locking transmission in multi-level cascaded topologies.
[0066] The transmission mechanism of SMP commands (configuration management commands) differs from that of SSP. SMP commands employ a strict request-response model; once the path is selected during connection establishment, it remains locked throughout the entire interaction, ensuring a return along the same path. For example, consider a host sending an SMP command to a target extender: (1) Sending path (Initiator → Target Expander): Between the initiator and each level of cascade expander, one of the multiple physical links of the wide port is selected to establish a connection; the selected physical path is exclusively locked in this SMP transaction.
[0067] (2) Return path (target Expander → Initiator): The response data returned by the target expander returns along the original path of the physical path that was exclusively locked when the request was sent, until it reaches the initiator.
[0068] like Figure 2 As shown, taking the main cabinet sending an SMP command to ExpanderC as an example: the main cabinet's initiator1 sends the SMP command to ExpanderC1 through ExpanderA1 and ExpanderB1 respectively. The specific transmission establishment process is as follows: 1) Send SMP Request data: (Initiator1 → Expander C1) a) Between Initiator1 and Expander A1: Select one of the four links in the wide port to establish a connection and transmit data; this path is exclusively used by this SMP transaction; b) Between Expander C1 and Expander B1: Select one of the four links in the wide port to establish a connection and transmit data; this path is exclusively used by this SMP transaction; c) Between Expander B1 and Expander C1: Select one of the four links in the wide port to establish a link and transmit data.
[0069] 2) Return SMP Response data (Expander C1 → Initiator1): a) Between Expander C1 and Expander B1: The exclusive path is returned when an SMP Request is sent to transmit data.
[0070] b) Between Expander B1 and Expander A1: The exclusive path is returned when an SMP Request is sent to transmit data.
[0071] c) Between Expander A1 and Initiator1: The exclusive path returned when sending an SMP Request is used to transmit data.
[0072] 4. Distributed time-series recording and host-side global reassembly mechanism Based on the inherent characteristics of the aforementioned transmission paths, the selection of wide-port links by the extender during SSP data forwarding exhibits dynamic uncertainty, increasing the complexity of single-point link status analysis. Furthermore, considering that the extender itself is a small processor, it is unsuitable for performing extremely complex protocol parsing and reassembly tasks. Therefore, this embodiment of the invention adopts an architecture of distributed recording and centralized reassembly: Each extender operates independently, responsible only for extracting primitives and frames from the raw data stream and recording their detailed raw characteristics. To fully represent the end-to-end interaction process from the host-side initiator to the target hard drive, the host-side second processor 201 runs a unified application that globally combines and processes the local time-series data captured from each extender, ultimately integrating the complete time-series interaction diagram of all extenders, thereby achieving accurate reconstruction of the underlying data flow of the complex storage system.
[0073] In one embodiment of the present invention, the first processor 101 includes at least one communication interface, and at least one extension terminal forms a cascaded structure. The cascaded structure includes multiple levels of extension terminals. The uplink port 102 of the extender in the first level extension terminal is connected to at least one communication interface of the second processor 201, and the downlink port 103 of the extender in the next level extension terminal is connected to the uplink port 102 of the extender in the next level extension terminal. Both the uplink port 102 and the downlink port 103 include at least one physical layer interface. The recording unit generates timing data based on the data stream from the physical layer interface. The data stream from the physical layer interface in the uplink port 102 is the incoming data, and the data stream from the physical layer interface in the downlink port 103 is the outgoing data.
[0074] Specifically, such as Figure 2 As shown, in this embodiment of the invention, the expansion terminals (i.e., expansion cabinets JBOD-A, JBOD-B, and JBOD-C) are networked using a multi-level cascaded structure. The expanders (Expander A1 / A2) within the first-level expansion terminal (JBOD-A) are configured with an uplink port 102, which is directly connected to the second processor 201 within the main cabinet via a physical link. Figure 2The communication interface of Initiator 1 / Initiator 2 in the system establishes a communication handshake with the host. Within and between the expansion terminals, signal transmission follows a strict hierarchical logic: the downlink port 103 of the expander in the next higher level (e.g., JBOD-A) is connected via cable to the uplink port 102 of the expander in the next lower level (e.g., JBOD-B). Thus, this embodiment of the invention provides a communication interface through the first processor and cascades multiple expansion terminals into a multi-level structure. This invention can overcome the port number limitation of a single-level expander and easily construct a large-scale, multi-level storage system architecture. Simultaneously, since each expansion terminal has time-series recording and storage capabilities, when a system anomaly occurs, the host can extract time-series data step-by-step along the cascaded links, thereby achieving end-to-end full path tracing of ultra-long physical links spanning multiple expander nodes, ensuring accurate location of faulty nodes even in complex topologies.
[0075] Furthermore, each port includes at least one PHY. For example, to achieve high-bandwidth transmission, a wide port is typically composed of four physical layer interfaces bundled together. To clarify the direction of data monitoring, this embodiment of the invention defines the data flow as follows: Inbound data: refers to the data stream that enters the current extender through the physical layer interface in the uplink port 102 (i.e., instructions / data from the previous level device or host). Outgoing data: refers to the data stream that leaves the current extender through the physical layer interface in the downlink port 103 (i.e., instructions / data sent to the next level device or hard drive).
[0076] In one embodiment of the present invention, the recording unit is further configured to record the data stream of the corresponding port, obtain the interface identifier and transmission direction of the physical layer interface corresponding to the data stream, extract the data content and timestamp of the data stream, identify the data type of the data stream, and generate time-series data based on at least one of the interface identifier, transmission direction, data content, timestamp and data type, wherein the transmission direction includes the inflow direction and the outflow direction.
[0077] Specifically, both uplink port 102 and downlink port 103 are equipped with recording units for real-time capture and analysis of data streams. The specific workflow is as follows: The recording unit monitors the electrical signal activity of each physical layer interface in real time. When a data stream is detected, the recording unit obtains the interface identifier corresponding to the data stream (i.e., identifies which PHY port the data comes from) and the transmission direction (determines whether it is flowing in or out). Through deep analysis of the data stream, key data content and precise timestamps are extracted. Simultaneously, by parsing the data frame header, the data type (e.g., whether it is a SCSI command frame, status frame, or raw data frame) is identified. By associating and encapsulating the above dimensions (such as interface identifier, transmission direction, data content, timestamp, and data type), standardized time-series data is generated. This time-series data can completely reconstruct the operating state of the storage system at any given time, providing accurate data support for subsequent fault diagnosis and performance analysis. Therefore, the recording unit in this embodiment of the invention can monitor the PHY's transmitted and received data streams in a bypass manner, identify them as primitives and frames, and form each record; each record contains at least the following entry information: 64-bit nanosecond-level timestamp, PHY_ID, transmission direction (transmit / receive), data type (frame or primitive), and data content.
[0078] In this embodiment of the invention, the first processor 101 integrates a storage medium to store the timing data recorded by the recording unit. The storage medium can be a static random access memory (SRAM) and serves as a ring buffer (TraceBuffer) to store the native timing record data of all PHYs.
[0079] The first processor 101 in this embodiment of the invention is further configured to, if no read instruction is detected, overwrite the historical time sequence data in the storage medium with the current time sequence data when the storage medium is detected to be full; if a read instruction is detected, pause the storage of the current time sequence data until at least one historical time sequence data is read from the storage medium.
[0080] Specifically, the application embodiment adds a timing collection module to the firmware of SAS Expander. On the one hand, it can read the timing data stored in SRAM, and on the other hand, it provides timing collection functions, supporting the main program to collect timing data through different transmission channels. During operation, the recording unit of each PHY detects the input and output data streams and generates data records in SRAM. If the SRAM is full and has not been collected by the application, it will be overwritten in a loop and then stored again. When the application triggers the collection action, it supports out-of-band or in-band interfaces to collect the logs, and recording will continue after the collection is completed.
[0081] In this embodiment of the invention, the software architecture of the second processor 201 includes at least one probe for monitoring the log data of the software architecture. The software architecture of the second processor 201 in this embodiment includes an application layer and an operating system layer. The application layer adds probes at the start and end nodes of read and write operations to record the start and end times of the read and write operations. The operating system layer attaches probes to the target sub-layer to record the operation data of the target sub-layer.
[0082] Specifically, 1) Application layer: Add probes to the start and end nodes that initiate read and write operations in the application to record the start and end times of each instruction in detail.
[0083] 2) Operating System Layer (Linux System Layer): Related logs can be added to the target sub-layer. The target sub-layer includes the device layer, SCSI middleware layer, and HBA driver layer, etc., specifically including: a) At the block device layer, attach a probe to the static trace point `block:block_rq_issue` to record the initiation time when the block layer dispatches the IO request to the device driver, the target device master / slave device number (dev_t), the operation type (read / write / flush, etc.), the starting sector number, and the number of sectors; attach a probe to the static trace point `block:block_rq_complete` to record the completion timestamp and return status code when the IO request is completed. By pairing the initiation and completion events of the request, the end-to-end latency of the block layer IO can be calculated.
[0084] b) SCSI Middleware Layer: A probe is mounted on the static trace point `scsi:scsi_dispatch_cmd_start` to record the timestamp of SCSI command dispatch to the HBA driver, CDB opcodes (such as READ(10), WRITE(10), INQUIRY, etc.), LUN, command length, and IPTT. A probe is mounted on the static trace point `scsi:scsi_dispatch_cmd_done` to record the timestamp of command completion, SCSI host status code (host_byte, such as DID_OK for success, DID_ERROR for error, DID_TIMEOUT for timeout, etc.), driver status code (driver_byte), and the actual length of data transmitted. A probe is mounted on the static trace point `scsi:scsi_dispatch_cmd_timeout` to capture the timeout occurrence and the command waiting time when a timeout occurs.
[0085] c) Driver side: Mount probes at the entry points of command submission functions (such as mpt3sas_scsih_queue_command, pm80xx_queue_command, etc.) of the HBA driver to record the time when commands are enqueued into the HBA firmware queue; mount probes at the entry points of interrupt handling functions (such as the ISR functions of each driver) to record the time when HBA hardware interrupts occur; mount probes at the entry points of driver error handling functions to record the trigger time of the error recovery process.
[0086] Furthermore, locations in existing stored procedures that have already been identified as anomalies and added to the logs can also be added as probes. By adding existing probe locations and combining them with existing application logging, a complementary approach can be achieved to maximize the discovery of anomalies.
[0087] It should be noted that the storage system in this embodiment of the invention includes both a general-purpose open-source Linux storage system and a dedicated closed-source Linux system. eBPF probes can be added to both systems, but the location of the probes varies depending on the application. Those skilled in the art can configure the system according to the specific circumstances, without making any specific limitations.
[0088] In one embodiment of the present invention, before reading timing data from the storage medium within the expander according to the read instruction, the second processor 201 sends an instruction set instruction to at least one expander to capture the link information of the expander, extract the link status of the storage device connected to the expander in the link information, generate the current topology of the storage system according to the cascade structure of at least one expander and the link status of the storage device connected to the expander, and sends the read instruction to the expander in the current topology.
[0089] Understandably, before reading timing data from the storage medium within the expander according to the read instruction, this embodiment of the invention can also establish a timing data collection channel between the main storage program within the second processor 201 and each expander. Specifically, the second processor 201 has a dedicated module for interacting with each level of expanders to collect timing data. This embodiment of the invention periodically sends existing SMP instructions to the expanders to capture the link information (e.g., attach SAS address) of each expander, which is used to identify the topology of the SAS domain and subsequently generate the cascading order of each expander level. In this embodiment of the invention, periodic probing aims to detect the link status in real time. When link changes are caused by hard drive or cable plugging / unplugging, the changes in link status caused by hard drive or cable plugging / unplugging can be detected in real time. This embodiment of the invention utilizes this topology query mechanism to record the acquired information in a preset format, providing basic support for subsequent data parsing and extraction, topology identification, and timing graph generation.
[0090] In one embodiment of this invention, during storage system operation, an eBPF (Extended Berkeley Packet Filter) probe is used as an anomaly detection sensor to perform real-time statistics and monitoring of I / O transmissions. The detection points of the eBPF probe include, but are not limited to: (1) IO instruction timeout: Determine whether an abnormal timeout has occurred by calculating the start and end timestamps of each layer of IO request; (2) IO exception return code: Monitor the exception response returned after the IO path issues the command; (3) HBA driver anomaly: Detect underlying transmission faults in the HBA driver layer; (4) System log anomalies: Capture anomaly information already recorded in the existing storage system log.
[0091] Furthermore, when the eBPF probe detects any of the above-mentioned abnormal events, it immediately triggers a linkage mechanism as an event pusher, specifically including: (1) Record the anomaly context: Record the detailed information of the detected anomaly into the log, which will serve as the starting point for subsequent panoramic time series analysis; (2) Trigger full collection: Notify the log collection module of the storage main program to collect all logs related to this IO. The collection priority is as follows: exception information detected by eBPF, time-series data (TraceBuffer) of each level of Expander, and existing logs recorded by the storage main program (such as HBA card logs).
[0092] In the log collection process, the main program first probes the topology to identify all Expander objects, then sends TraceBuffer read commands (including Offset and Length) to each object to read time-series data in packet-based transmission. The main program then concatenates these packet data to form complete time-series data. Because TraceBuffers are easily overwritten under high I / O loads, the main program and eBPF probes must work closely together to ensure the timeliness and completeness of time-series data collection.
[0093] The above embodiment illustrates log collection proactively triggered by eBPF after detecting an IO anomaly. In addition, other collection methods are also supported: 1) CLI command to actively trigger: Provides CLI commands for collecting timing logs under the main program, which are used to actively trigger collection.
[0094] 2) Periodic collection: When the IO pressure is not high, periodically collect a time-series log.
[0095] 3) Proactive Complementary Collection of Anomalies: When the storage master program detects an anomaly, it proactively initiates the collection of time-series logs to complement the data in eBPF.
[0096] Therefore, in this embodiment of the invention, when the probe detects an abnormal IO event, it immediately uses the log collection module of the main storage program, with the main storage program responsible for collecting and summarizing the timing information of the entire path. Since the amount of recorded timing data is also large when the IO data volume is large, the TraceBuffer may be overwritten in a very short time. This embodiment of the invention, through the collaboration of the main storage program and the eBPF probe, can promptly collect the most complete amount of timing data.
[0097] In one embodiment of the present invention, the first processor 101 parses the data offset and data length in the read instruction, extracts timing data from the storage medium according to the data offset and data length, and transmits the extracted timing data to the second processor 201.
[0098] Understandably, after acquiring the timing data of each level of extender, this embodiment of the invention can first perform preliminary parsing of the collected fault event data packets to extract abnormal I / O information recorded by the eBPF probe, including the initiator port transmission tag (IPTT), error code, start and end timestamps, and timeout duration. This abnormal I / O information will serve as the entry point for subsequent panoramic timing diagram analysis, used to analyze the interaction timing before and after the abnormal I / O.
[0099] Simultaneously, the host side parses and obtains the topology information, reconstructing the device connection relationships under the SAS domain topology: (1) Identify the initiator: Obtain the SAS address of each Initiator (such as Initiator1 and Initiator2) from the stored main program, and use it as the initiator identifier of the SSP / SMP command; (2) Identify the cascading relationship of the expanders: Determine the cascading level of each expander according to the topology information (e.g., the expander directly connected to the initiator is the first level (abbreviated as Expander A), the expander connected to the initiator through the first level expander is the second level (abbreviated as Expander B), the expander connected to the initiator through the second level expander is the third level (abbreviated as Expander C), and so on. There may also be Expander D, Expander E, etc. This invention does not make specific limitations on this, and obtain the SAS address of each expander; All PHY interfaces of each Expander are classified according to the port dimension: on the one hand, the correspondence between PHY_ID and PORT_ID is listed, that is, to determine the PHY connected to the uplink wide port, the PHY connected to the downlink wide port, and the PHY connected to the downlink narrow port; on the other hand, the correspondence between PHY_ID and hard disk slot SLOT_ID is listed. (3) Based on the topology information queried by the main program, identify the PHY_ID and SAS address of the hard disk connected to each SLOT_ID under each Expander.
[0100] In one embodiment of the present invention, the second processor 201 is further configured to parse the interface identifier, timestamp, and data type in the timing data, generate a port timing table based on the interface identifier and timestamp, determine the target interaction instruction corresponding to the timing data based on the data type, generate a first timing diagram based on the port timing table and the timing data corresponding to the data read / write instruction if the target interaction instruction is a data read / write instruction, generate a second timing diagram based on the port timing table and the timing data corresponding to the configuration management instruction if the target interaction instruction is a configuration management instruction, generate a second timing diagram based on the port timing table and the timing data corresponding to the configuration management instruction, and generate a timing panorama based on the first timing diagram and the second timing diagram.
[0101] It is understandable that, due to the different interaction methods of SSP (data read / write command) and SMP (configuration management command), the timing splicing stage of this embodiment of the invention adopts a classification processing strategy. If the target interaction command is a data read / write command (i.e., an SSP command), since it has a dynamic link selection mechanism within the wide port, the second processor 201 will perform full path reassembly based on the aforementioned generated port timing table and the timing data corresponding to the data read / write command to generate a first timing diagram reflecting the data read / write interaction process; (2) If the target interaction instruction is a configuration management instruction (i.e. SMP instruction), since it has path locking characteristics in the entire request-response interaction process, the second processor 201 will perform time reassembly of the original path return based on the port timing table and the timing data corresponding to the configuration management instruction, and generate a second timing diagram reflecting the configuration management interaction process.
[0102] Furthermore, the second processor 201 associates and merges the generated first timing diagram with the second timing diagram to generate a complete timing panorama that covers the entire path interaction state of the entire storage system, providing a global perspective for visualization support for subsequent root cause analysis of faults.
[0103] In one embodiment of the present invention, the second processor 201 is further configured to generate a total timing table based on a timestamp, split the total timing table into at least one sub-timing table based on an interface identifier, classify at least one sub-timing table to a port dimension based on the correspondence between the interface identifier and the port, so as to obtain a first timing table for uplink port 102, a second timing table for downlink port 103 and a third timing table for device port 104, and generate a port timing table based on the first timing table, the second timing table and the third timing table.
[0104] It is understood that, in the embodiments of the present invention, the time-series record data segment of each Expander can be parsed one by one, and each Expander can be parsed as follows: 1. Parse the timing log, arrange it in chronological order according to timestamps, and reconstruct the "general timing table" that may be distributed on any PHY_ID, any frame or primitive, and currently shows no pattern; 2. First, split the timing sequence according to the PHY_ID dimension. For example, taking an expander with 36 PHYs as an example, split the total sequence table into 36 "sub-timing tables", with each table corresponding to a PHY_ID; 3. Based on the mapping relationship between each PHY_ID and the physical port, the acquired timing data is categorized by port dimension. The categorization process specifically includes: determining the physical port to which each PHY_ID belongs based on the pre-configured port mapping table; merging the timing data corresponding to all PHY_IDs belonging to the same physical port; and generating uplink wide port timing tables, downlink wide port timing tables, and timing tables for each hard disk port based on the merged data, and performing subsequent interactive communication and timing analysis on a port-by-port basis.
[0105] In one embodiment of the present invention, the second processor 201 is further configured to identify the transmission direction and data content in the timing data corresponding to the data read / write instruction, determine the data content corresponding to the uplink port 102, downlink port 103 and device port 104 according to the transmission direction, extract the first timing table, the second timing table and the third timing table in the port timing table, traverse the first timing table of the uplink port 102 of the first-level extender, extract at least one first frame in the data read / write instruction interaction in the data content of the uplink port 102, determine the target device of the data read / write instruction according to the first frame, query the command frame of the first frame according to the type of the target device, the first timing table, the second timing table and the third timing table, generate a frame timing diagram according to the command frame of at least one first frame, query the timing data related to the command frame according to the frame timing diagram and the port timing table, and generate a first timing diagram according to the timing data related to the command frame.
[0106] Understandably, SSP command (data read / write command) interactions involve the interaction of various frames (such as COMMAND frames, possibly XFER_RDY frames, DATA frames, and RESPONSE frames, etc.). Furthermore, a complete SSP command interaction involves frames that share a common identifier: the same IPTT (Initiator Port Transfer Tag) assigned by the Initiator. The IPTT is a key field in the SSP (Serial SCSI Protocol) frame header of the SAS protocol. The IPTT is assigned by the Initiator to each SCSI command as a unique identifier throughout its lifecycle. The fact that all frames share the same IPTT in a single interaction leverages the global uniqueness of the IPTT within the SAS domain to achieve full-path frame-level association across multiple Expander levels.
[0107] In this embodiment of the invention, after obtaining the timing data of each level of extender, the SSP command of the Serial SCSI protocol can be restored from the entire link. The specific implementation steps are as follows: Step 1. Instruction classification and anomaly identification based on target address The second processor 201 first iterates through the timing table of the uplink wide port of the first-level extender, extracts all COMMAND frames that initiate SSP command interactions, and parses the initiator port transmission tag (IPTT) parameter, arranging them in timestamp order. Then, it parses the destination SAS address (destination_sas_addr) field in the COMMAND frame to determine the target of the SSP command, specifically in three scenarios: (1) Target is an intermediate extender: If the target SAS address is equal to the SAS address of an extender node, then the node is determined to be the target extender. If the target extender has a preceding extender (i.e., an intermediate extender), then the transmission path of the SSP instruction is: sequentially through the uplink wide port input and downlink wide port output of the intermediate extender, and finally reaching the target extender; (2) The target is a terminal storage device: If the target SAS address is equal to the SAS address of the hard disk node connected to a certain extender, then the hard disk is determined as the target, and the extender directly connected to the hard disk is determined as the target extender. If there is a front-end intermediate extender, the transmission path of the SSP command is as follows: it passes through the uplink wide port input and downlink wide port output of the intermediate extender in sequence, then through the uplink wide port of the target extender, and finally through the downlink narrow port output to the target hard disk; (3) Target address anomaly identification: If the target SAS address cannot be matched, the system determines that the SAS link is abnormally interrupted (which may be caused by SAS signal quality degradation, hard disk hot-swapping or SAS address configuration error), and locates it as an independent abnormal event as a potential root cause of the fault.
[0108] Step 2. Cross-node frame-level timing reassembly based on identifier (IPTT) After determining the target extender or target hard disk of the SSP instruction, the second processor 201 uses the extracted IPTT parameters as the search identifier to traverse the uplink / downlink wide port timing tables of all intermediate extenders involved in the SSP instruction, as well as the uplink wide port timing table and downlink hard disk narrow port timing table of the target extender, to find all frame records containing the same IPTT parameters.
[0109] Subsequently, the frame records are timestamped and sorted according to the cascading order of the extenders, as well as the "request transmission order from Initiator to Target" and the "response transmission order from Target to Initiator," thereby reconstructing a complete full-path frame-level timing sequence spanning multiple extenders. Each frame record in this sequence is precisely labeled with its respective extender, port type (uplink / downlink / hard disk port), and timestamp, enabling accurate measurement of transmission delay for each hop.
[0110] Step 3. Frame-based low-level primitive-level interaction reconstruction After generating the full-path frame-level timing sequence, the second processor 201 further extracts the specific physical layer interface (PHY) timing diagram from the port timing sequence traversing the link. Centered on each data frame in the frame timing diagram, it retrieves the underlying primitive sequences related to the frame transmission both forward and backward. In actual execution, this embodiment of the invention can extract handshake primitives before frame transmission (such as the OPEN address frame, AIP primitives, OPEN_ACCEPT / OPEN_REJECT primitives during the connection establishment phase, and the RRDY primitive sequence during the credit exchange phase), as well as acknowledgment and close primitives after frame transmission (such as ACK / NAK primitives, DONE primitives, and CLOSE primitives). These primitive sequences are concatenated with the data frames to finally generate a complete "primitive-frame interaction" full-path SSP timing panorama containing all SSP commands, thereby achieving pixel-level reconstruction of the underlying physical interaction process.
[0111] Specifically, in one embodiment of the present invention, the second processor 201 can perform global parsing and sorting of the acquired underlying timing data. By traversing the timing table of the first-level Expander uplink wide port, all COMMAND frames (the first frame in an SSP command interaction) are extracted, and the IPTT parameters are identified and arranged sequentially according to the timestamp order. Secondly, the present invention can parse the destination_sas_addr field in the Command frame to find the Target of the SSP command, which is described in three scenarios as follows: The target is the Expander itself: If destination_sas_addr equals the SAS Address of an Expander node, then the target of the command is that Expander, which is the "target Expander" of the SSP instruction; if there is a preceding Expander upstream of the target Expander (called the "intermediate Expander"), then the SSP instruction first passes through the "upstream wide port input of the intermediate Expander, then through the downstream wide port output", and then to the target Expander; The target is the hard drive connected to a certain level of Expander: If destination_sas_addr is equal to the SAS Address of a Target (hard drive) node in a certain Expander node, then the target of the command is that hard drive, and the Expander directly connected to that hard drive is called the target Expander; if there is a preceding Expander (called the "intermediate Expander") upstream of the target Expander, then the SSP command first passes through the "upstream wide port input of the intermediate Expander, then through the downstream wide port output", then through the "upstream wide port of the target Expander, then through the downstream narrow port output connected to the target hard drive", and finally to the target hard drive; Target not found: The SAS link may have been interrupted due to the following reasons: 1. Poor SAS signal quality; 2. Hard drive removal / removal; 3. Incorrect SAS address configuration.
[0112] Furthermore, embodiments of the present invention can iterate through each Expander involved in each SSP instruction to find frames with the same IPTT; the following steps are performed: 1) After determining the target expander or target for each Cmd, search for the same IPTT frame records in the "uplink wide port timing" and "downlink wide port timing" of "all involved intermediate expanders", and search for all involved frames one by one in the "uplink wide port timing" and / or "downlink hard disk narrow port timing" of the target expander; 2) Arrange these frames according to the "cascading order of the expanders" and the "transmission order from the initiator to the target, and then the transmission order from the target to the hard disk response", and sort out the order of the timestamps to form a complete frame-level timing sequence that spans multiple expanders and the entire path; each frame record in the sequence is marked with its own expander, its own port type (uplink / downlink / hard disk port) and precise timestamp. (The complete frame-level transmission path of an SSP command is precisely reconstructed as follows: For example, the COMMAND frame is sent from the Initiator → received at the uplink port 102 of Expander1 → forwarded at the downlink port 103 of Expander1 after routing → received at the uplink port 102 of Expander2 → sent to the target hard drive at the hard drive port of Expander2 after routing → the target hard drive returns a RESPONSE frame → enters through the hard drive port of Expander2 → forwarded through the uplink port 102 of Expander2 → enters through the downlink port 103 of Expander1 → forwarded through the uplink port 102 of Expander1 → finally returns to the Initiator. The timestamp and delay of each hop can be precisely measured.)
[0113] 3) From the timing of all ports that pass through the link, first find the timing diagram of the PHY. Taking each frame on the timing diagram as the center, search forward and backward for primitive sequences related to the transmission of that frame. Add primitives before frame transmission (such as: OPEN address frame in the connection establishment phase, possible AIP primitives (Arbitration In Progress), OPEN_ACCEPT or OPEN_REJECT primitives, RRDY primitive sequence in the credit exchange phase, etc.) and primitives after frame transmission (such as: ACK or NAK primitives in the acknowledgment phase, primitives after frame transmission is completed, DONE primitives, and CLOSE primitives in the connection closing phase) to the timing diagram to form a complete "primitive-frame interaction" full-path SSP timing panorama containing all SSP commands.
[0114] In one embodiment of the present invention, the second processor 201 is further configured to identify the transmission direction and data content in the timing data corresponding to the configuration management instruction, determine the data content corresponding to the uplink port 102 and the downlink port 103 according to the transmission direction, extract the first timing table in the port timing table, traverse the first timing table of the uplink port 102 of the first-level extender, extract at least one reply frame in the configuration management instruction interaction in the data content of the uplink port 102, determine the extender of the target of the configuration management instruction according to the reply frame, traverse other extenders involved in the configuration management instruction, query the data content to find that the target extender has the same function code as other extenders, and generate a second timing diagram according to at least one reply frame and the same function code.
[0115] The SMP instructions in this embodiment of the invention adopt a request-response model and do not involve the connection establishment process of the OPEN address frame. They use the SMP function code and the target SAS address as the association clues. The specific steps are as follows: Step 1. SMP instruction source extraction and target path determination The second processor 201 first traverses the timing table of the uplink wide port of the first-level expander, extracts all SMP Request frames, and arranges them strictly according to timestamp order. Then, it parses the destination SAS address (destination_sas_addr) field in the SMP Request frame to determine the target node. Since SMP commands are all sent from the initiator to the expander, this target node is the "target Expander". If there is a preceding expander (i.e., an "intermediate expander") upstream of the target expander, the system determines that the transmission path of the SMP command is: first through the uplink wide port input of the intermediate expander, then through its downlink wide port output, and finally reaching the target expander.
[0116] Step 2. Cross-node full-path time-series reassembly based on Function codes After determining the transmission path of the SMP instruction, the second processor 201 uses the SMP Function code as a search identifier to traverse the port timing tables of all expanders (including intermediate and target expanders) involved in the SMP instruction, searching for frame records with the same Function code. By matching the Function codes, the system can accurately locate the SMP Request frame and its corresponding SMP Response frame in each SMP interaction. Finally, the system concatenates and reassembles the extracted Request and Response frames according to their timestamps to generate a complete SMP timing diagram that reflects the entire path interaction process of the SMP instruction. Due to the path-locking characteristic of SMP interaction, this timing diagram can clearly and accurately present the original return process of the management instruction in a multi-level cascaded topology.
[0117] Specifically, based on the above embodiments, this embodiment of the invention can traverse the timing table of the uplink wide port of the first-level Expander, extract all SMP Requests, and arrange them in chronological order; parse the destination_sas_addr field in the Command frame to find the Target of the SSP command; since SMP commands are all sent from the Initiator to the Expander, the Target is the "target Expander"; if there is a preceding Expander (called an "intermediate Expander") uplink to the target Expander, the SMP command first passes through the "uplink wide port input of the intermediate Expander, then through the downlink wide port output", and then to the target Expander; traverse "each Expander involved in each SMP command" one by one, including the intermediate Expander and the target Expander, to find the same Function code; find the SMP Request and SMP Response in each SMP interaction; and form the full-path SMP timing diagram according to the timestamp order.
[0118] In one embodiment of the present invention, the second processor 201 is further configured to extract the timestamps of the first timing diagram and the second timing diagram, sort the first timing diagram and the second timing diagram according to the timestamps, and generate a timing panorama based on the sorting result. The timing panorama is a two-dimensional coordinate diagram of device identifier and time. The device identifier includes at least one of port identifier and interface identifier.
[0119] It is understood that the second processor 201 in this embodiment of the invention first merges the generated full-path SSP timing diagram with the full-path SMP timing diagram, and then globally sorts the merged timing data strictly according to the order of timestamps. When generating the visualization chart, a two-dimensional coordinate system is constructed: the device and port are used as the horizontal axis (X-axis), and the devices and ports include, in order: the initiator, the upstream port 102 and downstream port 103 of each level of expander (such as Expander1), the hard disk ports 1 to N, and the target hard disk; the time axis is used as the vertical axis (Y-axis). Through this two-dimensional coordinate system, the merged timing data is mapped into a two-dimensional timing panorama that completely covers the entire path interaction state of the storage system. As a unique, objective, and unambiguous source of facts, the timing panorama contains the root cause of the problem, providing a basis for analysis of IO anomaly problems and greatly improving the accuracy of root cause location. In one embodiment of the present invention, the second processor is further configured to: associate and match the abnormal time-series data captured by the probe with the time-series panorama; if an abnormal event is matched, identify all associated frames and underlying primitives belonging to the same abnormal command, and perform abnormal display processing on the associated frames and underlying primitives in the time-series panorama.
[0120] After generating the time-series panorama, the abnormal time-series data captured by the second processor 201 probe is correlated and matched with the panorama. For matched abnormal events, the system automatically identifies all associated frames and underlying primitives belonging to the same abnormal command and performs special processing on their display style in the time-series panorama: the display color of the relevant frames and primitives is switched to a conspicuous abnormal highlight color (such as red or orange), and their graphic borders are thickened. Through the above visualization highlighting mechanism, operations and maintenance personnel can quickly and accurately locate the specific physical port and interaction link where the abnormality occurred in the complex panorama time-series panorama, thereby significantly improving the efficiency of fault diagnosis.
[0121] For ease of understanding, the embodiments of the present invention can be described in detail with reference to specific examples, such as... Figure 3 The diagram illustrates the SAS link timing from Initiator 1 through Expander A and Expander B, ultimately reaching one of the Targets connected to Expander B. For ease of understanding, each Expander is divided into three parts, representing: 1. Connect to the upstream wide port of the parent Expander; 2. Connect to the downstream wide port of the next-level Expander; 3. Connect each hard drive via a narrow port.
[0122] The uplink and downlink wide ports consist of four independent PHY circuits. For simplicity, they are treated as a single object on the timing diagram, and their corresponding PHY_IDs should be indicated on the jump arrows.
[0123] Hard drive ports: This may include 12 or 25 platters (i.e., 25 narrow ports). Since it's impossible to divide the jump process into 25 objects in a timing diagram, the corresponding slot_ID and PHY_ID need to be indicated on the jump arrows to ensure accurate location of a specific hard drive and its corresponding physical interface.
[0124] Because the interaction between the downlink port of the previous Expander and the uplink port of the next Expander in the SSP command chain is uncertain, the timing data captured by a protocol analyzer for the input and output of the same link may not be the same. This timing panorama allows us to trace the detailed flow of an SSP command: from the Initiator, through multiple Expander stages to the Target, and the detailed timing data returned by the Target.
[0125] In another embodiment, the SAS protocol's transmission rate has reached as high as 12 Gbps or even 24 Gbps. In such a high-speed transmission environment, even capturing only the underlying data within a 1-millisecond time window can contain tens of thousands of primitives and frames. Faced with such massive and complex time-series data, relying solely on manual fault location would present extremely high analytical difficulties and a tedious workload. The SAS protocol system itself is highly complex, resulting in a high technical barrier for engineers when analyzing problems, and involves a large amount of repetitive protocol parsing work. Given the significant advantages of artificial intelligence (AI) in processing standard protocol parsing and recognizing patterns in massive amounts of data, this invention proposes introducing an AI agent into the storage system. By using the AI agent for deep analysis and automatically identifying normal interaction sequences in the time-series diagram, the aim is to significantly reduce the fault location threshold for engineers and significantly improve the efficiency of fault diagnosis in complex storage systems. The specific steps are as follows: 1. Construct a multi-dimensional SAS protocol knowledge base The AI agent first performs structured parsing of the SAS protocol family documents (including SAM, SPC, SPL, SES, etc.), identifying the document's chapter structure and building a hierarchical index. In the knowledge extraction phase, the agent focuses on extracting the following key information: (1) Basic element definition: including the name, encoding value, transmission direction, triggering condition and exception handling rules of primitives; the definition of frame header fields of SSP frame and SMP frame (field name, offset, length, etc.); and the precise meaning of various status codes and reason codes (such as OPEN_REJECT reason code, SCSI status code, SMP return code).
[0126] (2) Protocol interaction mechanism: including the complete interaction sequence and state transition conditions of connection establishment, data transmission, and connection closure; various timeout thresholds (such as OPEN_TIMEOUT and CLOSE_TIMEOUT) and corresponding processing actions specified in the protocol; and the handling process for exceptions.
[0127] During the knowledge reorganization phase, the AI agent, following the analytical thinking patterns of engineers, reorganizes the above information into a multi-dimensional indexed knowledge base. Specifically, this includes: indexing by primitive type, aggregation by interaction flow of connection lifecycle, classification by device role (initiator, extender, target), layering by protocol layer (physical layer, link layer, port layer, transport layer, application layer), and organization by SCSI command opcodes.
[0128] 2. Intelligent association and interaction between sequence diagrams and protocol knowledge bases The AI agent can deeply analyze primitives and frame data in time sequence diagrams and perform real-time matching with the knowledge base. For primitives, the agent automatically extracts relevant introductions, reference chapters, and analysis suggestions; for frame data, the agent automatically unpacks and parses the data, extracting internal opcodes and configuration formats, and converting the underlying hexadecimal values into a format that engineers can intuitively recognize.
[0129] Meanwhile, the system provides an interactive window for engineers to interact with the AI knowledge base. When engineers encounter questions while analyzing time sequence diagrams, they can ask the AI agent at any time. The agent will instantly display all relevant knowledge points and reference chapters surrounding the concept, providing a convenient dictionary-like search experience.
[0130] 3. Provides intelligent time series diagram interaction and analysis assistance. When engineers are reviewing the time series overview, the system provides convenient analysis functions similar to a code editor: (1) Explanation of hovering concept: When the mouse hovers over an event marker (such as a primitive or frame) in a sequence diagram, the AI agent automatically retrieves and displays the complete protocol information of the event from the knowledge base.
[0131] (2) Hover Highlighting and Summary: When the mouse hovers over a primitive or frame event, the AI agent automatically identifies the SSP or SMP instruction to which it belongs, highlights the complete interaction process, and automatically summarizes the content, interaction status and transmission / reception time of the instruction.
[0132] (3) Intelligent filtering and hiding: For complex timing diagrams, the system provides multi-dimensional filtering functions. It supports instruction hiding (hiding instructions that are not of interest) and hard disk filtering (keeping only the serial link timing of specific hard disks); at the same time, the AI agent automatically identifies and retrieves abnormal events, supports masking normal timing, and enables engineers to focus on abnormal links.
[0133] 4. Intelligent analysis and classification of time-series interaction processes AI agents can automatically analyze and classify time series graphs by continuously learning from time series data. (1) Model training and learning: In normal scenarios, the system automatically feeds time series data to the AI for training via CLI commands; at the same time, engineers can manually select normal or abnormal time series to feed to the agent for learning. The AI can accurately identify time series that contain intermediate abnormal retries (such as open_retry) but are eventually recovered as normal time series.
[0134] (2) Automatic Analysis and Classification Display: The AI automatically classifies the current time series into three categories: time series that have been trained and confirmed to be normal, time series with clear abnormal events, and suspicious time series that cannot be identified due to complex interactions and lack of training. The system provides different colors or classification labels for these three types of time series. By automatically eliminating most normal time series through AI, engineers only need to focus on analyzing time series marked as "unrecognizable" and "abnormal", thereby significantly reducing the scope of analysis and greatly improving work efficiency.
[0135] According to the storage system proposed in this invention, recording units are set on both the uplink and downlink ports of the extender to generate time-series data in real time. The first processor stores this data in an internally integrated storage medium, thus achieving complete recording and storage of the entire data flow at the hardware level. Simultaneously, the second processor on the host side monitors log data in real time through probes within the software architecture. Upon detecting abnormal events, it actively generates read commands, extracts the corresponding time-series data from the extender as needed, and automatically parses and generates a time-series panorama. This panorama visually maps the abnormal location to a specific physical port or link, enabling precise location of storage system anomalies and significantly improving the accuracy of fault diagnosis. This solves the problem of inefficient and incomplete reconstruction of transmission timing and interaction processes during storage system anomalies, leading to inefficient fault location and difficulty in accurate source tracing.
[0136] Next, with reference to the accompanying drawings, an anomaly localization method for a storage system according to an embodiment of the present invention is described, which is applied to the second processor of the storage system described above.
[0137] Figure 4 This is a flowchart illustrating an anomaly location method for a storage system provided in an embodiment of the present invention.
[0138] like Figure 4 As shown, the anomaly localization method of this storage system includes the following steps: In step S101, the log data of the software architecture is monitored by probes set within the software architecture.
[0139] In step S102, when an abnormal event is identified based on the log data, a read instruction is generated, and timing data is read from the storage medium within the expander according to the read instruction.
[0140] In step S103, a time series panorama is generated based on at least one time series data read, and the abnormal location of the storage system is located based on the time series panorama.
[0141] It should be noted that the description of the features in the preceding embodiments can store the relevant descriptions of the corresponding embodiments of the system, which will not be repeated here.
[0142] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0143] Any of the components, modules, units, parts, methods, and operations described herein can be implemented using software, firmware, hardware (e.g., fixed logic circuitry), manual processing, or any combination thereof. Alternatively or additionally, any functionality described herein can be performed at least in part by one or more hardware logic components, such as, but not limited to, a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), an application-specific standard product (ASSP), a system-on-a-chip (SoC), a complex programmable logic device (CPLD), a microprocessor (MCU), etc. The terms "system," "computing device," or "apparatus" as used herein encompass various means, devices, and machines for processing data, including, for example, one or more programmable processors, computers, SoCs, or combinations thereof. The apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of these. The aforementioned computer program (also known as a program, software, software application, app, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, object, or other unit suitable for a computing environment.
[0144] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0145] The storage system provided by this invention has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and core ideas of this invention. It should be noted that those skilled in the art can make various improvements and modifications to this invention without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this invention.
Claims
1. A storage system, characterized in that, include: The host and at least one extension; The expansion terminal is provided with at least one expander and at least one storage device; The extender includes a first processor, an uplink port, a downlink port, and at least one device port. The device port is used to connect to the storage device. Both the uplink port and the downlink port are provided with a recording unit to generate time-series data based on the data stream of the corresponding port. The first processor integrates a storage medium to store the time-series data recorded by the recording unit. The host includes at least one second processor, and the software architecture of the second processor is equipped with at least one probe, which is used to monitor the log data of the software architecture. When the second processor identifies an abnormal event based on the log data, it generates a read instruction, reads the timing data from the storage medium within the expander based on the read instruction, generates a timing panorama based on at least one of the read timing data, and locates the abnormal location of the storage system based on the timing panorama.
2. The storage system according to claim 1, characterized in that, The first processor includes at least one communication interface, and the at least one extension terminal forms a cascaded structure. The cascaded structure includes multiple levels of the extension terminal. The uplink port of the extender in the first level of the extension terminal is connected to at least one communication interface of the second processor, and the downlink port of the extender in the previous level of the extension terminal is connected to the uplink port of the extender in the next level of the extension terminal.
3. The storage system according to claim 1 or 2, characterized in that, Both the uplink port and the downlink port include at least one physical layer interface. The recording unit generates timing data based on the data stream of the physical layer interface. The data stream of the physical layer interface in the uplink port is the incoming data, and the data stream of the physical layer interface in the downlink port is the outgoing data.
4. The storage system according to claim 1, characterized in that, The recording unit is also used to record the data stream of the corresponding port, obtain the interface identifier and transmission direction of the physical layer interface corresponding to the data stream, extract the data content and timestamp of the data stream, identify the data type of the data stream, and generate the time-series data according to at least one of the interface identifier, the transmission direction, the data content, the timestamp and the data type, wherein the transmission direction includes the inflow direction and the outflow direction.
5. The storage system according to claim 1 or 4, characterized in that, The first processor is further configured to, if the read instruction is not recognized, overwrite the historical time-series data in the storage medium with the current time-series data when the storage medium is detected to be full; if the read instruction is recognized, pause the storage of the current time-series data until at least one historical time-series data has been read from the storage medium.
6. The storage system according to claim 1, characterized in that, The software architecture of the second processor includes an application layer and an operating system layer. The application layer adds the probes at the start and end nodes of read and write operations to record the start and end times of the read and write operations. The operating system layer attaches the probes to the target sub-layer to record the operation data of the target sub-layer.
7. The storage system according to claim 6, characterized in that, The target sublayer includes at least one device layer, an intermediate layer, and a driver layer, wherein the device layer and the intermediate layer determine the mounting position of the probe according to a pre-set target tracking point type, and the driver layer determines the mounting position of the probe according to a pre-set function type.
8. The storage system according to claim 1, characterized in that, Before reading the timing data from the storage medium within the expander according to the read instruction, the second processor sends an instruction set instruction to at least one of the expanders to capture the link information of the expanders, extract the link status of the storage devices connected to the expanders from the link information, generate the current topology of the storage system based on the cascade structure of at least one expander and the link status of the storage devices connected to the expanders, and sends the read instruction to the expanders in the current topology.
9. The storage system according to claim 1, characterized in that, The first processor parses the data offset and data length in the read instruction, extracts the timing data from the storage medium according to the data offset and data length, and transmits the extracted timing data to the second processor.
10. The storage system according to claim 1, characterized in that, The second processor is further configured to parse the interface identifier, timestamp, and data type in the timing data. It generates a port timing table based on the interface identifier and timestamp, determines the corresponding target interaction instruction based on the data type, and if the target interaction instruction is a data read / write instruction, generates a first timing diagram based on the port timing table and the timing data corresponding to the data read / write instruction. If the target interaction instruction is a configuration management instruction, generates a second timing diagram based on the port timing table and the timing data corresponding to the configuration management instruction, and generates the timing panorama based on the first and second timing diagrams.
11. The storage system according to claim 10, characterized in that, The second processor is further configured to generate a total timing table based on the timestamp, split the total timing table into at least one sub-timing table based on the interface identifier, classify at least one sub-timing table to the port dimension based on the correspondence between the interface identifier and the port, so as to obtain a first timing table for the uplink port, a second timing table for the downlink port and a third timing table for the device port, and generate the port timing table based on the first timing table, the second timing table and the third timing table.
12. The storage system according to claim 10, characterized in that, The second processor is further configured to identify the transmission direction and data content in the timing data corresponding to the data read / write instruction, determine the data content corresponding to the uplink port, the downlink port, and the device port based on the transmission direction, extract the first timing table, the second timing table, and the third timing table in the port timing table, traverse the first timing table of the uplink port of the first-level extender, extract at least one first frame in the data read / write instruction interaction in the data content of the uplink port, determine the target device of the data read / write instruction based on the first frame, query the command frame of the first frame based on the type of the target device, the first timing table, the second timing table, and the third timing table, generate a frame timing diagram based on the command frame of at least one first frame, query the timing data related to the command frame based on the frame timing diagram and the port timing table, and generate a first timing diagram based on the timing data related to the command frame.
13. The storage system according to claim 10, characterized in that, The second processor is further configured to identify the transmission direction and data content in the timing data corresponding to the configuration management instruction, determine the data content corresponding to the uplink port and the downlink port according to the transmission direction, extract the first timing table in the port timing table, traverse the first timing table of the uplink port of the first-level extender, extract at least one reply frame in the configuration management instruction interaction in the data content of the uplink port, determine the extender of the target of the configuration management instruction according to the reply frame, traverse other extenders involved in the configuration management instruction, query the same function code of the target extender and the other extenders in the data content, and generate a second timing diagram according to at least one reply frame and the same function code.
14. The storage system according to claim 10, characterized in that, The second processor is further configured to extract the timestamps of the first timing diagram and the second timing diagram, sort the first timing diagram and the second timing diagram according to the timestamps, and generate the timing panorama according to the sorting result. The timing panorama is a two-dimensional coordinate graph of device identifier and time. The device identifier includes at least one of port identifier and interface identifier.
15. A method for anomaly localization in a storage system, characterized in that, The method is applied to a second processor of a storage system as described in any one of claims 1-14, wherein the method comprises: Monitor the log data of the software architecture using probes set within the software architecture; When an abnormal event is identified based on the log data, a read instruction is generated, and the time-series data is read from the storage medium within the expander according to the read instruction; A time-series panorama is generated based on at least one of the time-series data read, and the abnormal location of the storage system is located based on the time-series panorama.