Transmission Link Fault Detection Method and Program Product for Multi-Level Cascade Scenarios

By obtaining the real-time link information and instruction sending and receiving of the multi-stage cascade extender and executing the fault detection program, the problem of insufficient transmission link failure detection in multi-stage cascade scenarios is solved, and accurate fault location and efficiency improvement is achieved.

CN119902939BActive Publication Date: 2025-05-30INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510381447.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-05-30
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

The prior art lacks detection of transmission link faults in multi-stage cascade scenarios, which is prone to false alarms as receiving terminal failures, and it is difficult to accurately detect the cause and location of the fault.

Method used

By obtaining the real-time link information of the multi-stage cascade extender and the command sending and receiving status of the target port in the target device, the corresponding fault detection program is executed to determine the fault detection result of the transmission link.

Benefits of technology

Accurate fault detection of multi-stage cascade extenders and transmission links is realized, which avoids false alarms on hard disks and improves fault positioning efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119902939B_ABST
    Figure CN119902939B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and a program product for detecting transmission link faults in a multi-stage cascaded scenario, which relates to the technical field of fault detection and includes: when there is a transmission link fault detection mechanism in a target device, obtaining the real-time link information of multi-stage cascaded expanders in the target device and the instruction sending and receiving conditions of at least one target port; when it is detected that the real-time link information is abnormal, executing a corresponding first link fault detection program to determine a first fault detection result; when it is detected that the instruction sending and receiving conditions fail, detecting the failure type of the instruction sending and receiving conditions and executing a second link fault detection program corresponding to the failure type to determine a second fault detection result, solving the technical problem that it is impossible to detect whether it is a transmission link fault or a hard disk fault, achieving the technical effects of accurately locating the faulty device or faulty link or hard disk fault on the instruction sending and receiving transmission path and giving an alarm, avoiding false alarms for hard disks, and improving the fault location efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of fault detection technology, and in particular to a transmission link fault detection method and program product in a multi-stage cascade scenario. Background Art

[0002] In related technologies, with the rapid development of electronic devices, in order to meet the needs of electronic device capacity and performance, multiple key electronic devices will be connected in a multi-level cascade manner, which will form a very complex topology, causing the proportion of transmission link failures to gradually increase. When electronic devices fail to send or receive instructions or time out when processing multi-threaded tasks, the cause of the failure is no longer just the failure of the execution terminal. For example, when the IO from the host to the hard disk fails or times out, in addition to the hard disk failure, it may also be a transmission link failure.

[0003] However, in the related art, fault detection of electronic devices with multi-level cascade scenarios still focuses on the receiving terminal, ignoring the transmission link failure in the multi-level cascade scenario, and is easily misreported as a receiving terminal (such as a hard disk) failure. There is a lack of accurate detection of transmission link failures and specific fault causes and fault locations, which makes it difficult to adapt to the fault detection needs of electronic devices in multi-level cascade scenarios and needs to be urgently resolved. Summary of the invention

[0004] The present invention provides a transmission link fault detection method and program product for a multi-stage cascade scenario, so as to at least solve the problems in the related art that fault detection of electronic devices with multi-stage cascade scenarios still focuses on receiving terminals, ignores transmission link faults in multi-stage cascade scenarios, is easily misreported as receiving terminal faults, lacks accurate detection of specific fault causes or fault locations in transmission link faults, and is difficult to adapt to fault detection needs of electronic devices in multi-stage cascade scenarios.

[0005] The present invention provides a transmission link fault detection method for a multi-stage cascade scenario, comprising: in the case where a transmission link fault detection mechanism exists in a target device, obtaining real-time link information of a multi-stage cascade expander in the target device and the instruction receiving and sending status of at least one target port; when an abnormality is detected in the real-time link information, executing a corresponding first link fault detection program to determine a first fault detection result of the multi-stage cascade port and / or the transmission link of the multi-stage cascade expander; when a failure is detected in the instruction receiving and sending status, detecting the failure type of the instruction receiving and sending status, and executing a second link fault detection program corresponding to the failure type to determine a second fault detection result of the target device.

[0006] The present invention also provides a computer program product, including: an acquisition module, configured to acquire real-time link information of a multi-stage cascaded expander and instruction transceiver conditions of at least one target port in the target device when a transmission link fault detection mechanism exists in the target device; a first detection module, configured to execute a corresponding first link fault detection program when detecting that the real-time link information is abnormal, so as to determine a first fault detection result of multi-stage cascaded ports and / or a transmission link of the multi-stage cascaded expander; a second detection module, configured to detect a failure type of the instruction transceiver conditions when detecting that the instruction transceiver conditions fail, and execute a second link fault detection program corresponding to the failure type, so as to determine a second fault detection result of the target device.

[0007] The present invention also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of the transmission link fault detection method for any one of the above multi-stage cascaded scenarios when executing the computer program.

[0008] The present invention also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of the transmission link fault detection method for any one of the above multi-stage cascaded scenarios are implemented.

[0009] Through the present invention, when a transmission link fault detection mechanism exists in a target device, real-time link information of a multi-stage cascaded expander and instruction transceiver conditions of at least one target port in the target device can be acquired, and then a corresponding first link fault detection program is executed when the real-time link information is abnormal, so as to accurately detect whether a multi-stage cascaded port and / or a transmission link of the multi-stage cascaded expander fails; when detecting that the instruction transceiver conditions fail, a corresponding second link fault detection program is executed according to the failure type of the instruction transceiver conditions, so as to accurately determine a second fault detection result of the target device. Therefore, the problem in the related art that the fault detection of an electronic device with a multi-stage cascaded scenario still mostly focuses on a receiving terminal, ignores the transmission link fault in the multi-stage cascaded scenario, is easily misreported as a receiving terminal fault, lacks accurate detection of specific fault reasons or fault locations in the transmission link fault, and is difficult to meet the fault detection requirements of the electronic device in the multi-stage cascaded scenario can be solved, and the technical effect of comprehensively determining a fault object and a fault type when the real-time link information of the multi-stage cascaded expander is abnormal and the transceiver instruction fails, accurately locating a fault transmission node or a fault transmission link where the real-time link information is abnormal, etc., and can accurately locate a fault device, a fault link or a hard disk fault on the transmission path of the transceiver instruction and give an alarm when the transceiver instruction fails, avoiding false alarms for the hard disk, and improving the fault location efficiency is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] To more clearly illustrate the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0011] Figure 1 Flowchart of a transmission link fault detection method for a multi - level cascaded scenario provided by an embodiment of the present invention;

[0012] Figure 2 Schematic diagram of the link from CPU to hard disk in an embodiment of the present invention;

[0013] Figure 3 Schematic diagram of two access paths of a SAS hard disk in an embodiment of the present invention;

[0014] Figure 4 Schematic diagram of the framework of SAS protocol layering in an embodiment of the present invention;

[0015] Figure 5 Schematic diagram of link anomaly query in an embodiment of the present invention;

[0016] Figure 6 Flowchart of transmission link fault detection for a dual - control host in a multi - level cascaded scenario in an embodiment of the present invention;

[0017] Figure 7 Schematic diagram of fault alarm in an embodiment of the present invention;

[0018] Figure 8 Block diagram of a computer program product provided according to an embodiment of the present invention;

[0019] Reference numerals:

[0020] Among them, 10 - computer program product; 100 - acquisition module, 200 - first detection module, 300 - second detection module. Detailed implementation manners

[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0022] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0023] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0024] An embodiment of the present invention provides a method for detecting transmission link faults in a multi-stage cascaded scenario. The method will be described in detail in combination with the execution process of the method for detecting transmission link faults in a multi-stage cascaded scenario.

[0025] Specifically, Figure 1 FIG. is a flowchart of a method for detecting transmission link faults in a multi-stage cascaded scenario according to an embodiment of the present invention.

[0026] As Figure 1 shown, the method for detecting transmission link faults in a multi-stage cascaded scenario includes the following steps:

[0027] In step S101, when there is a transmission link fault detection mechanism in the target device, obtain the real-time link information of the multi-stage cascaded expander in the target device and the instruction transceiver status of at least one target port.

[0028] It can be understood that the target device here refers to an electronic device with multi-stage cascaded components that needs to detect transmission link faults, such as a terminal device (such as a computer), a system device, etc.; the transmission link fault detection mechanism here can be understood as a detection mechanism established in advance in the target device by the present invention for detecting whether there is a transmission link abnormality in the case of various device abnormalities.

[0029] And, the link here can be understood as the path between the multi-stage cascaded components in the target device, as well as between the multi-stage cascaded components and various interfaces; at least one target port here can be understood as at least one target port in the interfaces established by the instruction transceiver terminal.

[0030] For example, in a computer with SAS (Serial Attached SCSI, a serial Small Computer System Interface, which is a commonly used interface in computer storage currently) hard drives installed, a storage system is often designed to process the massive data transmitted between the server and the hard drives reliably and quickly. The interfaces of this storage system include, but are not limited to, an external pcie interface for the CPU to process service data, a common SAS interface for the hard drives, and so on. During the operation of the computer, the pcie interface needs to be converted to the SAS interface and connected to each hard drive through a SAS expander. To provide high reliability, each SAS hard drive supports dual ports. At this time, there will be two paths from the CPU to the hard drive, which can be called links. As Figure 2 shown, Figure 2 This is a schematic diagram of the link from the CPU to the hard drive (disk) in an embodiment of the present invention.

[0031] In some embodiments, to meet the requirements of larger capacity and higher performance, and for the target device to meet various processing requirements, multiple multi-stage cascaded components may be installed.

[0032] For example, a computer with SAS hard drives installed will expand and connect more hard drives by cascading more levels of Expanders. Figure 3 This is a schematic diagram of two access paths of the SAS hard drive in an embodiment of the present invention. As Figure 3 shown, taking a 3-level Expander as an example, the access path of the host to the hard drive can be, but is not limited to, expressed as:

[0033] (1) As an intermediate node between the host and the hard drive, the Expander connects different devices through SAS connectors and cables;

[0034] (2) The Expander and the Disk are numbered respectively. When the host wants to access Disk-1 (Disk-1), it needs to pass through Expander1. When it wants to access Disk-2 (Disk-2), it needs to pass through Expander1 and Expander2;

[0035] (3) If 10 levels of Expanders are cascaded, when the host wants to access the hard drive expanded by Expander10, the link it needs to pass through includes a total of 10 Expanders and the cables of the devices from Expander1 to Expander10.

[0036] In the process of instruction interaction in a multi-level scenario, after passing through multiple Expander nodes, a failure of any intermediate transmission node (such as Expander1, Disk-1, etc.) may cause an IO (Input / Output) transmission failure. For example, abnormal transmission signals, a single physical link failure between the Expander and the hard disk resulting in a single disk IO failure, a link failure between Expanders resulting in IO failures of multiple subsequent hard disks, etc.

[0037] As a possible implementation, in order to detect whether the target device has a network terminal failure, a hard disk failure, a failure of a switch, a router, an Expander, a Disk-X involved in the middle, or a transmission link failure between these components when these abnormalities occur in the target device, the embodiments of the present invention may first obtain the real-time link information of the multi-level cascaded expanders in the target device and the instruction sending and receiving conditions of at least one target port.

[0038] Among them, the target device needs to set the transmission failure detection mechanism in the embodiments of the present invention. The multi-level cascaded expanders can be understood as multi-level switches, routers in a network system, or Expanders, Disk-X, etc. in a computer here. At least one target port can be understood as the interface of the transmission terminal here. For example, the two interfaces of a SAS hard disk.

[0039] The embodiments of the present invention can obtain this information as data support, which is convenient for detecting the actual failures of the target device by using the transmission link failure detection mechanism, and helps to accurately locate the information and failure problems of the multi-level cascaded expanders and the link.

[0040] Optionally, in an embodiment of the present invention, before obtaining the real-time link information of the multi-level cascaded expanders in the target device and the instruction sending and receiving conditions of at least one target port, it further includes: collecting software error codes and firmware error codes in the historical transmission link failures of the target device; establishing a transmission link failure detection mechanism in combination with the software error codes, firmware error codes, preset expansion instructions, corresponding parsing strategies, and the management module of the target device.

[0041] Based on the related descriptions of other embodiments, it can be understood that the present invention will obtain the real-time link information of the multi-level cascaded expanders in the target device and the instruction sending and receiving conditions of at least one target port when there is a transmission link failure detection mechanism in the target device. Among them, the transmission link failure detection mechanism can be pre-established by the present invention in the target device, that is, the target device needs to have this transmission link failure detection mechanism to implement the transmission link failure detection method of the present invention.

[0042] In some embodiments, the present invention can consider software error codes and hardware error codes in various historical transmission link failures and utilize this information as the information support for the transmission link failure detection mechanism. Among them, the software error code and the hardware error code can be respectively understood here as abnormal information that can locate device failures as transmission link failures in the software implementation part and the hardware (firmware) implementation part of the target device.

[0043] For example, in a host SAS (SAS3.0, SAS4.0) card, the SAS protocol standard divides the levels into Physical Layer, Link Layer, Port Layer, Transport Layer, and Application Layer.

[0044] In the software implementation part of each layer under the state machine (a mathematical model used to describe the state transitions of an object under different conditions and the corresponding behaviors), the abnormal information that can locate a transmission link failure includes but is not limited to:

[0045] (1) Abnormal logs related to transmission recorded in the driver log, such as errors of the SAS card itself and errors in receiving IOs;

[0046] (2) Abnormalities recognized by the transport layer, such as: invalid XFER_RDY (transmission ready signal or status), invalid ResponseFrame (an invalid ResponseFrame means that the received response frame does not conform to the expected format, content, or protocol regulations and cannot be correctly parsed or recognized), invalid DataFrame, frame loss due to errors received from the hard disk or Expander (an invalid DataFrame means that there is a problem with this data frame and it cannot be correctly processed or used);

[0047] (4) Abnormalities recognized by the port layer, such as: loss of connection between Initiator and Target;

[0048] (4) Abnormalities recognized by the link layer, such as: determined frame error, connection establishment error, ACK (Acknowledgement Signal) signal timeout;

[0049] (5) Abnormalities recognized by the physical layer: such as, invalid Dword (4-byte data unit), polarity error;

[0050] These error types are partly the SAS standard protocol and partly dependent on vendor implementation. Embodiments of the present invention can add these error codes to the transmission link fault detection mechanism, which can be used as the trigger conditions for triggering the transmission link fault detection of each Expander or hard disk.

[0051] Further, on the Expander, the firmware implementation part can locate the exception information that is a transmission link error, including but not limited to:

[0052] (1) Error records encountered during the operation of the state machines involved in each layer of IO transmission in the firmware;

[0053] (2) Exception records related to the Expander processor, such as: abnormal firmware operation, abnormal restart;

[0054] (3) The status register values related to the link fed back during the processing of the Expander hardware mechanism, such as: whether the PHY (physical, physical layer interface, the channel for physically transmitting data) is Ready (completed preparation), whether it is disabled (whether it is set to the prohibited working state), whether it is connected, on-off count, link on-off CRC (Cyclic Redundancy Check) error, polarity value, routing table loss;

[0055] (4) Information on the interaction of the Expander responding to the chassis management module, such as: chip temperature exceeding the threshold, recording abnormal restart, recording the disk kicking, port lifting, and restart instructions issued by the chassis management module;

[0056] The SAS standard protocol defines CRC errors, rates, etc., and most rely on the mechanism of the processor signal. Since some target devices, such as computers equipped with SAS3.0 Expanders, do not support accurate positioning of signal fault types, embodiments of the present invention can add this information to the transmission link fault detection mechanism to provide as detailed information as possible for the chassis management module. Combining this information, the chassis management can clearly locate which device is faulty and provide feasible troubleshooting steps.

[0057] In addition, there are various layers of mechanisms between the host main program of the computer and the firmware of the SAS Controller (Serial Attached SCSI Controller) to complete the interaction with the Expander and the hard disks: for example, sending SSP commands (Serial SCSI Protocol commands, a communication protocol instruction set based on serial SCSI technology) to the hard disks to achieve the sending and receiving of IO commands for all hard disks; sending SES commands (SCSI Enclosure Services commands, a set of commands used to communicate with the chassis management module in the SCSI protocol) to the Expander to perform chassis management functions such as detecting the status of hard disks, the temperature and voltage of devices, and reporting alarms; sending SMP commands (SCSI Management Protocol commands, a management protocol in the SCSI protocol) to the Expander to identify topology information, PHY rate, CRC count of the PHY, etc.

[0058] Therefore, the embodiments of the present invention can also add a preset extended command to the SMP commands for the interaction between the host and the Expander. Here, the preset extended command can be understood as a certain command for detecting the above error information added on the basis of the original SMP query commands of the computer, so as to achieve the host's query of the most detailed link fault information from the Expander on the basis of the existing SMP command transmission channel and command code.

[0059] Moreover, the embodiments of the present invention also take into account that the host main program in the computer is equipped with a chassis management module and an IO sending and receiving module. Therefore, the embodiments of the present invention can also determine the transmission link fault detection mechanism in the embodiments of the present invention in combination with the parsing strategy corresponding to a certain extended command and the management module of the target device.

[0060] It should be noted that the management commands used in the examples in the embodiments of the present invention to manage the Expander are SAS 3.0, which belongs to the category of extended chassis management; for models such as storage and servers that use BMC chips to manage the chassis status, BMC can also be used to detect the information of each device. That is, change "adding SMP commands to query link information" to "adding commands for BMC to interact with the Expander to query link information".

[0061] The function of the IO sending and receiving module is as follows: when each SAS disk has two independent SAS ports, that is, there are two IO paths, the host can select one of the IO paths to send IO commands to the hard disk; when the sending and receiving of this IO path fails, it can select the other IO path to continue sending and receiving commands; and send SMP commands when the link is on or off to quickly identify the on or off state of the Expander link.

[0062] Moreover, the function of the chassis management module (the chassis management modules of the dual-controller hosts are integrated into a whole through the data synchronization module) is as follows: query device status information, such as the link on / off information of the Exp, temperature-related information, and the on / off information of the PHY link (including the physical PHY ports between Expanders and between the Expander and the hard disk); and issue control commands as needed, such as: disk ejection, wide port ejection, and triggering Expander reset; in addition, comprehensively determine whether to report a certain alarm according to the status of the dual-controller; for example: when the temperature of a certain temperature sensor on the device exceeds the monitored threshold, report an over-temperature alarm; when an IO failure of a certain hard disk is received by the IO transceiver module, report the failure of the hard disk, etc.

[0063] Based on the above, in the embodiment of the present invention, a channel for parsing the SMP instruction code and reporting it to the chassis management module is newly added at the host end; a function of receiving the SMP instruction, parsing, and identifying link error information is newly added to the chassis management module; an alarm channel for intermediate devices (SAS cards, a certain level of Expanders) on the IO path; combined with the functional operations of the chassis management, identify the failure of which device and trigger an alarm, etc.

[0064] In the embodiment of the present invention, abnormal information that can locate a transmission link failure (error) in multiple situations can be added to the transmission link failure detection mechanism, and as much detailed information as possible is provided for the chassis management module, so that the chassis management can combine this information to clearly locate the failure of which device and provide feasible troubleshooting steps. In order to more reliably and efficiently implement the failure detection process, the present invention also adds certain SMP extension instructions, corresponding parsing strategies, and a management module for the target device, which can implement more detection functions and management operations.

[0065] Optionally, in an embodiment of the present invention, before obtaining the real-time link information of the multi-level cascaded expanders in the target device and the instruction transceiver situation of at least one target port, it further includes: traversing the routing table pre-established by the multi-level cascaded expanders and determining the topology structure corresponding to the routing table; obtaining the real-time link information according to the topology structure.

[0066] In some embodiments, the number of multi-level cascaded expanders in the target device is extremely large, and there are also many components connected by the multi-level cascaded expanders. In order to facilitate obtaining the real-time link information of the multi-level cascaded expanders in the target device and implementing the subsequent fault location process, the present invention can traverse the routing table pre-established by the multi-level cascaded expanders and establish the topology structure corresponding to the routing table before obtaining the real-time link information of the multi-level cascaded expanders in the target device and the instruction transceiver situation of at least one target port, and respectively determine the fault determination rules for the narrow link connecting the hard disk and the wide port connecting the cable, so as to obtain the real-time link information according to the topology structure and perform the subsequent detection process.

[0067] For example, after the target device such as a computer is powered on, the embodiments of the present invention can initialize it. First, after the Expander starts, a routing table is established, and then the host starts. It traverses the routing tables of the SMPs of each level of Expander, sorts out the complete topology structure, and reports it to the chassis management module. At the same time, the host side registers certain SMP extension instructions and waits for the trigger opportunity.

[0068] It should be noted that since a SAS hard disk has two links, at least two controller hosts are included in a complete SAS topology, including two IO paths. The transceiver module of the host and the chassis management module both need to integrate these two paths.

[0069] The embodiments of the present invention can establish a topology structure corresponding to the routing table pre-established by the multi-level cascaded expander. By utilizing the characteristic that the topology structure includes the complex link relationships between the multi-level cascaded expanders and other multiple components in the target device, the real-time link information of the multi-level cascaded expanders in the target device can be obtained, which can ensure that the obtained information can cover all the real-time link information, and further ensure the accuracy of the positioning result.

[0070] Step S102, when it is detected that the real-time link information is abnormal, execute the corresponding first link fault detection program to determine the first fault detection result of the multi-level cascaded ports and / or transmission links of the multi-level cascaded expander.

[0071] In some embodiments, after obtaining the real-time link information of the multi-level cascaded expander of the target device, the present invention can execute the corresponding first link fault detection program when it is detected that the real-time link information is abnormal.

[0072] Here, executing the corresponding first link fault detection program can be understood as that when the abnormal situations existing in the detected real-time link information are different, the embodiments of the present invention will execute different first link fault detection programs.

[0073] For example, in a computer equipped with SAS hard disks, multiple hard disks are linked through a multi-level cascaded expander. During the process of reading hard disk information, if a fault occurs, it may report that a certain expander is abnormal, it may report that a device linked to a certain expander is abnormal, it may report that the link between a certain expander and another expander is abnormal, or it may report that the link between the expander and the hard disk is abnormal, etc. At this time, some detection programs in the first link fault detection programs corresponding to these abnormal information can be executed.

[0074] However, the link fault detection programs corresponding to these exception messages may not be applicable to various other faults that may occur during the transmission of instructions. Therefore, the embodiments of the present invention can execute other link fault detection programs in the corresponding first link fault detection program in the event of other faults.

[0075] When the embodiments of the present invention detect that the real-time link information of the multi-stage cascade expander of the target device is abnormal, they can execute the corresponding first link fault detection program, avoiding directly judging that it is a transmission link fault of a certain fixed type in case of an error, but performing different first link fault detection programs according to the actual situation, which can effectively improve the accuracy of the link fault detection result.

[0076] Optionally, in an embodiment of the present invention, before executing the corresponding first link fault detection program, it further includes: detecting the status change report information of the multi-stage cascade expander and the sending time of the status change report information; judging whether the real-time link information is abnormal according to the status change report information and the sending time; if the status change report information meets the preset abnormal conditions and / or the sending time is inconsistent with the preset sending time, it is determined that the real-time link information is abnormal, otherwise, it is determined that the real-time link information is not abnormal.

[0077] Based on the relevant descriptions of other embodiments, it can be understood that when the present invention faces the situation that the real-time link information of the multi-stage cascade expander is abnormal, it can execute the corresponding first link fault detection program, that is, it is necessary to determine that the real-time link information of the multi-stage cascade expander is abnormal before performing the fault detection.

[0078] In some embodiments, when determining whether the real-time link information of the multi-stage cascade expander is abnormal, the embodiments of the present invention can, but are not limited to, determine it according to the status change report information fed back by the multi-stage cascade expander and the sending time of the status change report information.

[0079] Among them, when the status change report information meets the preset abnormal conditions, or the sending time of the status change report information is inconsistent with the expected sending time, or the status change report information meets the preset abnormal conditions and the sending time is also inconsistent, the embodiments of the present invention can determine that the real-time link information is abnormal and can execute the corresponding first link fault detection program.

[0080] Among them, the preset abnormal conditions can be understood here as the status change report information fed back by the multi-stage cascade expander being abnormal information, that is, the status change report information fed back when the multi-stage cascade expander or the devices and links connected thereto have abnormal changes (faults). The preset sending time can be understood here as the time when the multi-stage cascade expander feeds back the status change report information in the periodic example measurement. For example, the status change report information is fed back exactly at zero minutes every hour, etc.

[0081] It should be noted that there may be many faults involved in the multi - level cascaded expander, and the configurations of each device are different. Therefore, both the preset abnormal conditions and the preset sending time can be set or adjusted by those skilled in the art according to the actual situation. In the embodiments of the present invention, only exemplary descriptions are given without specific limitations.

[0082] For example, in a storage system, when some abnormal changes occur in certain states or configurations related to the multi - level cascaded expander Expander, such as a link negotiation failure, the Expander notifies this situation to the host and other related devices by broadcasting, and it can be determined that there is an abnormality in the real - time link information.

[0083] Similarly, it should be noted that there may be many faults involved in the multi - level cascaded expander. Therefore, those skilled in the art can also add or delete the criteria for determining the abnormality of the real - time link information according to actual needs.

[0084] The embodiments of the present invention can determine whether the real - time link information has an abnormality according to the status change report information of the multi - level cascaded expander and the sending time of the status change report information. Thus, combined with the increased software error codes and hardware error codes, more extensive and accurate fault detection of the target device can be achieved.

[0085] Optionally, in an embodiment of the present invention, when it is detected that there is an abnormality in the real - time link information, a corresponding first link fault detection program is executed, including: when the status change report information is an abnormal status change report information and / or the sending time is inconsistent with the preset sending time, execute the first link fault query program; when there are other abnormal situations in the real - time link information except that the status change report information is an abnormal status change report information and / or the sending time is inconsistent with the preset sending time, execute the first link error query process.

[0086] Based on the related descriptions of other embodiments, it can be understood that when the present invention detects that there is an abnormality in the real - time link information, it does not directly execute a certain fixed transmission link fault detection, but will execute the corresponding first link fault detection program according to the abnormal situation.

[0087] In the actual execution process, the present invention will execute the first link fault query program when the real - time link information has an abnormal status change report information and / or the sending time is inconsistent with a certain sending time. And when there are other abnormal situations in the real - time link information except these abnormal situations,

[0088] For example, when the host receives a BroadcastChange broadcast event from the Expander (a message passing mechanism that sends a broadcast message to notify other components or objects interested in the event when a specific state change or event occurs in a component or object in the system), such as the plugging or unplugging of a hard disk, the plugging or unplugging of a SAS cable, the powering on or off of the Expander or the hard disk, or the failure and reconnection of link negotiation due to a SAS signal failure, the present invention can execute the corresponding first link failure query process. For example, traverse each Expander to identify which link has changed in terms of connection or disconnection.

[0089] Alternatively, when the host has not received a BroadcastChange broadcast event from the Expander for a long time and performs periodic measurement, that is, the sending time of the status change report information is inconsistent with a certain sending time (delayed), the embodiments of the present invention can also execute the corresponding first link failure query process.

[0090] When there is an abnormality in the real-time link information and it still cannot be restored after the exception handling strategies at each layer (such as: instruction retry, reset operation on the expander), the embodiments of the present invention can execute the corresponding first link error query process. For example, check whether there is a fault record in the SAS card driver. If there is a fault, report the SAS card fault to the chassis management module, etc.

[0091] The embodiments of the present invention can execute the corresponding first link failure query process or the first link error query process, etc. under different abnormal conditions of the real-time link information, ensuring that the query process targets the current event, avoiding confusing the cause of the fault in cases with similar problems, and effectively improving the accuracy of the detection result.

[0092] Step S103, when it is detected that there is a failure in the instruction transceiver situation, detect the failure type of the instruction transceiver situation, and execute the second link failure detection program corresponding to the failure type to determine the second fault detection result of the target device.

[0093] Those skilled in the art of this technology can understand that in the case of a target device with multiple cascaded expanders, it can process tasks in multiple threads. During the task processing process, the main control end needs to send certain execution instructions (control instructions) to the main execution end, and correspondingly, the main execution end will also give a response after receiving and executing the instructions. In this process, the transceiver of the instructions is extremely important. When the transceiver of the instructions fails in a multi-cascaded scenario, the detection of the transmission link failure is extremely important.

[0094] As a possible implementation, considering that the status change report information of the multi-level cascaded expander may spontaneously report an exception in the event of an abnormality, then, in the embodiments of the present invention, when it is detected that there is a failure in the instruction transceiver situation, the failure type of the instruction transceiver situation can be detected, that is, it is detected whether the instruction transceiver failure is caused by an existing exception.

[0095] Then, the embodiments of the present invention can execute the corresponding second link fault detection program according to the failure type of the transceiver instruction, and no longer simply determine it as an error of the receiving terminal, but perform certain fault detection on both the transmission link and the receiving terminal, and then determine the second fault detection result of the target device.

[0096] For example, in a computer equipped with SAS hard disks, according to the SAS protocol layering (physical layer, link layer, port layer, transport layer, and application layer), it is passed up and down between adjacent layers, and each layer has an independent state machine and error handling mechanism.

[0097] Figure 4 It is a schematic diagram of the framework of the SAS protocol layering for an embodiment of the present invention. As Figure 4 shown, from the perspective of the upper-layer application, only the dotted part is passed, but in fact, each layer of the intermediate device is passed, which is a very long path. As long as there is a failure at one node during transmission, the transmission fails; a normal link is a key prerequisite for normal IO transmission. In IO services, whether a single or multiple IO instruction transceivers are abnormal, whether it is an IO exception for one disk or multiple disks, the link failure should be preferentially determined.

[0098] If, when the IO instruction from the host to the hard disk is incorrect, it is not comprehensively determined in connection with the link state at that moment, and it is not excluded whether a link failure has occurred at that moment, or a long timeout caused by a link transmission conflict, it may be misjudged as a hard disk failure.

[0099] Or, when the chassis management module uses the Send / Receive Diagnostic Page to query the device status, it ignores the delay caused by passing through the SCSI command layer and the impact of periodic column measurement on the IO performance, and cannot obtain the change of the status in real time. Then, some scenarios (such as: the hard disk link changes rapidly multiple times within the two query time intervals) cannot be recognized, which will lead to incorrect detection results.

[0100] Or, the SMP query of the link layer of the host for some fault types is less. When the driver end detects some transmission faults (transmission timeout, receiving abnormal stop data), it does not immediately identify which Expander is faulty and cannot be associated with the chassis management module, which will also lead to incorrect detection results.

[0101] Most importantly, when the chassis management module receives an I / O exception, it may simply tend to identify it as a hard disk failure and report a hard disk alarm. In an increasingly complex topology, with the link failure phenomenon becoming more and more prominent, it is easy to cause false alarms for hard disks. If there is a wide port signal failure between Expanders, there is a risk that most disks of the subsequent Expander will have I / O failures, and the alarm will be incorrect. False alarms may trigger subsequent error repair strategies, greatly increasing the maintenance cost.

[0102] Based on this, in an embodiment of the present invention, when it is detected that there is a failure in the instruction transceiver situation, such as when an exception occurs for a certain I / O at the host end, instead of directly reporting a hard disk failure alarm by the chassis management module, a time window is opened to perform a corresponding second link failure detection program. For example, using the SMP instruction with fast response at the host transport layer, quickly obtain the most detailed link status information that can be supported and provided by the SAS Expander chips of all nodes on the I / O path, and accurately determine whether it is an alarm of a certain device (SAS card, a certain Expander, a certain section of SAS cable) on the I / O path or a hard disk failure.

[0103] Figure 5 It is a schematic diagram of link exception query for an embodiment of the present invention, as Figure 5 shown: Taking a three-level Expander as an example, the host sends an SMP instruction, first queries the information of Expander1, then queries Expander2, and then queries Expander3; the chassis management module comprehensively determines whether it is a device problem (Expander1 failure or Expander2 failure or Expander3 failure) or a connection problem (connection problem between the host and Expander1 or connection problem between Expander1 and Expander2 or connection problem between Expander2 and Expander3 or connection between Expander3 and the hard disk, etc.) according to the SMP information; after these checks show no abnormalities, then suspect a hard disk failure and perform hard disk detection and alarm.

[0104] Thus, accurate alarm of devices is realized, false alarms for hard disks are avoided, and the fault location efficiency is improved.

[0105] Optionally, in an embodiment of the present invention, when it is detected that there is a failure in the instruction transceiver situation, the failure type of the instruction transceiver situation is detected, including: detecting whether there is an abnormality in the real-time link information before the transceiver situation fails at at least one target port; if there is an abnormality in the real-time link information before the transceiver situation fails at at least one target port, then determine that the failure type is a passive type, otherwise, determine that the failure type is an active type.

[0106] During the actual execution process, the status change report information of the multi-level cascaded expander may spontaneously report an exception in the event of an abnormality. Then, when detecting the failure type of the instruction transceiver failure, the present invention can first detect whether there is an abnormality in the real-time link information before the instruction transceiver failure.

[0107] If there is an abnormality in the real-time link information before the instruction transceiver failure, it can be deduced that the reason for the instruction transceiver failure may be due to a failure in the multi-level cascaded expander or its connected transmission link. At this time, the instruction transceiver failure is caused passively, and the failure type of the instruction transceiver failure can be determined as the passive type.

[0108] If there is no abnormality in the real-time link information before the instruction transceiver failure, it can be deduced that the reason for the instruction transceiver failure may not be due to a failure in the multi-level cascaded expander or its connected transmission link. Then, the failure type of the instruction transceiver failure can be determined as the active type.

[0109] The embodiments of the present invention can detect the failure type of the instruction transceiver situation, so as to execute a corresponding targeted second link fault detection program according to the failure type, which can effectively improve the fault detection efficiency and reduce the fault detection time and cost.

[0110] Optionally, in an embodiment of the present invention, when it is detected that there is a failure in the instruction transceiver situation, the failure type of the instruction transceiver situation is detected, and the second link fault detection program corresponding to the failure type is executed, including: in the case of the passive type, execute the second link fault query process to determine the second fault detection result of the target device; when the type is the active type, execute the second link error query process to determine the second fault detection result of the target device.

[0111] In some embodiments, when the failure type of the instruction transceiver failure is the passive type, the present invention can execute the second link fault query process and determine the second fault detection result of the target device at this time according to the result of the second link fault query process.

[0112] When the failure type of the instruction transceiver failure is the active type, the present invention can execute the second link error query process and determine the second fault detection result of the target device at this time according to the result of the second link error query.

[0113] Among them, the difference between the second link fault query process and the second link error query process can be determined in connection with the more likely reason for the instruction transceiver failure: the second link fault query process mainly but not limited to query the link of the multi-level cascaded expander, and the second error query process mainly but not limited to query the electronic devices involved in the transceiver path.

[0114] Embodiments of the present invention can execute corresponding second link fault query processes and second link error query processes under different transceiver failure types, determine the key points of fault detection based on the reasons for the failure types, reduce the detection time and cost, and ensure the accuracy of the detection results.

[0115] Optionally, in an embodiment of the present invention, when the type is passive, execute the second link fault query process to determine the second fault detection result of the target device, including: when the type is passive, traverse multiple levels of cascaded ports and their transmission links to determine the on / off changes of the transmission links; based on the pre-established routing table and the on / off changes, identify the on / off connection ends of the transmission links; determine the second fault detection result of the target device according to the on / off connection ends.

[0116] Based on the relevant descriptions of other embodiments, it can be understood that when the type of instruction transceiver failure is passive, the present invention will execute the second link fault query process to determine the second fault detection result of the target device.

[0117] In some embodiments, when the present invention executes the second link fault query process to determine the second fault detection result, it can, but is not limited to, first traverse multiple levels of cascaded ports and their transmission links to determine the on / off changes of the transmission links; then, based on the routing table pre-established by the multi-level cascade expander and the on / off changes, identify the on / off connection ends of the transmission links; finally, the second fault detection result of the target device can be determined according to the on / off connection ends.

[0118] For example, when the computer receives a link on / off BroadcastChange event sent by the Expander and has a passive response transceiver instruction failure, the second link fault query process at this time can, but is not limited to, be expressed as follows:

[0119] (1) Traverse each Expander to identify which link has on / off changes;

[0120] (2) Combine the routing table to identify the belonging of the PHY, and comprehensively determine: if it is a narrow port connected to the hard disk, it can be immediately reported to the chassis management module; if the on / off is the wide port cascaded by the Expander, query the on / off conditions of all links corresponding to this port multiple times within the maximum connection event range, and then report it to the chassis management module (the wide port consists of multiple links. In the normal scenario, due to actions such as plugging and unplugging, kicking the port, and resetting, the expected is that all links change on / off together. Due to the deviation of the negotiation time of different links and the fast SMP query instruction, a certain time window needs to be given to confirm multiple times until all links change stably (only at this time can it reflect whether the wide port on / off is normal).

[0121] (3) After the query is completed, report it to the chassis management module of the local host.

[0122] In an embodiment of the present invention, the on-off link ends can be obtained by detecting the on-off changes of each expander in the multi-stage cascaded expander and the links connected thereto, and then the specific faulty interfaces can be determined in combination with the routing table. Thus, when the type of failure in receiving and sending instructions is a passive type, comprehensive fault detection can be performed on the multi-stage cascaded expander and its links, and it can quickly and accurately detect where there is a problem in the entire path of instruction reception and transmission, obtaining accurate fault detection results and fault location results.

[0123] Optionally, in an embodiment of the present invention, when the type is an active type, a second link error query process is executed to determine the second fault detection result of the target device, including: when the type of failure is an active type, a query instruction is sent to the multi-stage cascaded expander; based on the reception and transmission information of the query instruction, the second fault detection result of the target device is generated.

[0124] Based on the related descriptions of other embodiments, it can be understood that when the type of failure in receiving and sending instructions is an active type, the present invention will execute a second link error query process to determine the second fault detection result of the target device.

[0125] In some embodiments, when the present invention executes the second link error query process to determine the second fault detection result, it can, but is not limited to, generate the second fault detection result of the target device according to the reception and transmission information of the query instruction by sending a query instruction to the multi-stage cascaded expander.

[0126] For example, the present invention first adds certain SMP extension instructions in the host SAS reception and transmission module of the computer: instructions for obtaining the link information of each level of Expander. Then, when a failure occurs in receiving and sending instructions when the multi-stage cascaded expander does not report abnormal status change report information, a second link fault query process is actively initiated. The process can, but is not limited to, be expressed as follows:

[0127] (1) In the order from near to far from the host, the intermediate devices (SAS card or Expander) involved in the IO path from the host to the hard disk {host application layer, SAS card, Expander1, Expander2, Expander3, hard disk} are checked one by one: First, check whether there is a fault record in the SAS card driver. If there is a fault, report the SAS card fault to the chassis management module; thus, check whether the application layer program executes incorrectly, avoiding the situation where the fault cause is due to software defects and cannot cover some extreme scenarios;

[0128] (2) Send an SMP instruction to Expander1 to query the standard defined instruction code and the newly added extension instruction code of each PHY. The specific process can, but is not limited to, be expressed as follows:

[0129] If no Response from the Expander is received, initially suspect that the Expander may have lost power, or is a normal restart after receiving a chassis management restart command, or is a firmware abnormal restart. Query multiple times within the time period "exceeding the restart time length required by the Expander" to determine whether the Expander has restarted (only determine that the Expander has restarted, and whether it is a normal restart is identified by the chassis management module);

[0130] ① If an abnormal Response from the Expander is received, it is determined that the Expander is abnormal;

[0131] ② If a normal Response from the Expander is received, parse the link status information and determine whether a fault has occurred; if a fault has occurred, combine the routing table information to determine the connection problem;

[0132] ③ If it is detected that the Expander has restarted or continuously fails to respond, report the Expander information (Expander SAS address) and the IO object information (hard disk information) to the chassis management module together;

[0133] ④ If a connection anomaly is detected, report the objects connected by the link (the interconnected objects include: the SAS card and the Expander connected by the wide port, the previous-level Expander and the next-level Expander connected by the wide port, the Expander and the hard disk connected by the narrow port) and the IO object information (hard disk information) to the chassis management module together;

[0134] (3) Perform the above operation on Expander2; and so on; if there is no abnormality in each level of Expander, suspect a hard disk failure; further determine by sending a series of detection commands to the hard disk; and report the IO object information (hard disk information) to the chassis management module together;

[0135] After the SMP query is completed, report to the chassis management module of the local host.

[0136] Thus, after the host SAS transceiver module adds certain SMP commands (commands to obtain the link information of each level of Expander) in the embodiment of the present invention, when the IO command sent to the hard disk is abnormal, it can quickly respond in the form of SMP commands, quickly obtain all the link-related fault information that can be extracted on the SAS Expander chip, and report it to the chassis management module.

[0137] In an embodiment of the present invention, the second fault detection result of the target device can be generated according to the transceiver information result of the query instruction sent to the multi-level cascade expander. Thus, when the failure type of the transceiver instruction is an active type, it can be sequentially detected whether it is an application layer error, whether it is a multi-level cascade expander problem, whether it is a link problem, and whether it is a link object problem. Thus, the topology is associated with "abnormal positioning of transceiver instructions", and a detailed link exception query is achieved, and the second fault detection result of the target device with extremely high accuracy is obtained.

[0138] Optionally, in an embodiment of the present invention, when the type is an active type, the second link error query process is executed to determine the second fault detection result of the target device, and it further includes: obtaining the instruction transceiver information of other target ports except at least one target port; generating the second fault detection result of the target device according to the instruction transceiver information of other target ports except at least one target port.

[0139] As a possible implementation manner, considering that there may be more than one target port in the target device, and in some cases, there may be a link fault only in a certain path for transceiver instructions. To more conveniently and accurately detect the fault cause of the target device, the embodiment of the present invention can also generate the second fault detection result of the target device in combination with the instruction transceiver information of other target ports except at least one target port in the target device.

[0140] For example, Figure 6 is a flowchart of the transmission link fault detection of the dual-controller host in a multi-level cascade scenario according to an embodiment of the present invention, as Figure 6 shown:

[0141] Step S601: The dual-controller host is powered on and initialized, that is, after the Expander starts, a routing table is established (SAS protocol standard process), and then the host starts, traverses the routing tables of the SMPs of each level of the Expander, sorts out the complete topology structure, and reports it to the chassis management module; at the same time, a certain SMP extension instruction is registered at the host end, waiting for the trigger opportunity;

[0142] Step S602: All the Expanders of the dual-controller detect the link-related information of themselves in real time. Both controllers of the host are always ready to execute the transceiver of the IO instruction for the hard disk, and both are detecting the trigger opportunity of "initiating the SMP detection process";

[0143] Step S603, the two controllers are respectively abbreviated as controller 1 and controller 2, corresponding to one port and another port respectively. According to the current service (task) scheduling policy, select the IO path of one of the controller hosts to execute the transceiver of the IO instruction for the hard disk;

[0144] Step S604, here the embodiment of the present invention first selects controller 1 to execute the transceiver of the IO instruction for the hard disk;

[0145] Step S605: Determine whether there is an abnormality in the I / O instruction transmission and reception executed by Controller 1. If there is no abnormality, continue to use Controller 1 to execute the I / O instruction transmission and reception for the hard disk;

[0146] Step S606: If an abnormality occurs in the I / O instruction transmission and reception executed by Controller 1, it will trigger the SMP query process for the Controller 1 path and report the query information to the chassis management module;

[0147] Step S607: After the chassis management module receives the link information of Controller 1, switch to the I / O path corresponding to Controller 2 to resend the failed I / O and wait for whether the link information of Controller 2 is received;

[0148] Step S608: Controller 2 executes the I / O instruction transmission and reception for the hard disk;

[0149] Step S609: Determine whether there is an abnormality in the I / O instruction transmission and reception executed by Controller 2. If there is no abnormality, continue to use Controller 2 to execute the I / O instruction transmission and reception for the hard disk. At this time, the embodiment of the present invention will no longer initiate the SMP query process. In the actual scenario, a common situation is a link failure on a certain I / O path;

[0150] Step S610: If an abnormality also occurs in the I / O instruction transmission and reception executed by Controller 2, it will trigger the SMP query process for the Controller 2 path. At this time, the embodiment of the present invention will initiate the SMP query process and report the query result to the chassis management module of this controller;

[0151] Step S611: After the chassis management module receives the SMP information of Controller 2, wait for whether the link information of Controller 1 is received;

[0152] Step S612: Within the timeout period, the chassis management comprehensively determines the cause of the I / O abnormality based on all the obtained link information (only Controller 2's, or Controller 1's and Controller 2's):

[0153] If the link information of Controller 1 is received, receive and parse it, start a certain timeout period (longer than the I / O path switching and the transceiver query event for the other controller path), and wait for whether the other controller receives the link information;

[0154] After the timed retry, based on all the received link information (defined according to the actual fault, which may only be the link information of Controller 1, or the link information of both controllers), combined with the status information queried by the chassis management module itself and the control information that has been initiated, comprehensively determine whether a real fault has occurred and how to give an alarm;

[0155] Among them, the conditions included in the comprehensive determination include, but are not limited to, the following conditions:

[0156] ①Identify the connection attribute of the PHY in combination with the routing table, whether it is a narrow port connected to a hard disk or a wide port connected by Expander cascading;

[0157] ②Along with the on / off information of the SMP, whether link error information is received at the same time;

[0158] ③Whether chassis management control commands have been issued, such as ejecting the disk, raising the port, issuing control commands to the link, control commands issued by the SMP, etc.;

[0159] ④Whether there are state changes in chassis management, such as changes in the board in-position state, cable in-position changes, power state changes, etc.;

[0160] The process of comprehensive judgment can be but is not limited to expressed as follows:

[0161] ①If it is a narrow port connected to a hard disk, the chassis management module combines whether it has issued an eject disk command itself to determine whether it is an abnormal on / off;

[0162] ②If it is a wide port connected by Expander cascading, in addition to the above information, the chassis management module also needs to consider whether all the PHYs within the wide port are on / off together to comprehensively determine whether it is an abnormal on / off; because the wide port consists of multiple links, in a normal control scenario, all the links will be on / off at the same time; due to bit errors or abnormal signal quality, some links negotiate to disconnect after sending IO, resulting in IO errors for an indefinite number of random hard disks after this wide port; it is necessary to identify this scenario and not report that multiple hard disks fail simultaneously, as such false alarms will lead to incorrect handling operations and have a great impact;

[0163] In step S613, if an Expander access exception is received, the chassis management module combines the management information queried by itself, such as: determining whether a plug / unplug action has been performed through the board in-position judgment, and also the power-off operation, and issuing a restart command to the Expander, to comprehensively determine the result (whether it is a normal restart or an abnormal restart);

[0164] If error information of the hard disk link is received, it includes but is not limited to the following three situations: According to the routing table information, query the error location, such as the SAS card, the Expander device on this IO path, the wide port cable link connection between two Expander devices, the link connection from the Expander to the hard disk, the hard disk of this target IO, etc.; If there is a fault in a certain node between the host and the hard disk, report a node fault alarm; If there is no alarm in the node, then determine whether the hard disk has a fault according to the hard disk fault determination strategy;

[0165] Step S614: If there is a transmission link failure, report the fault alarm of the device (such as SAS card, Expander, cable), and provide reference steps for fault repair; among them, for abnormal Expander access, corresponding alarms can be made according to the comprehensive judgment result; in the error information of the hard disk link, if the error location is queried according to the routing table information, corresponding alarms can also be directly made, and if there is a fault at a certain node between the host and the hard disk, report the node fault alarm.

[0166] Step S615: If there is no alarm at the node (no transmission link alarm), then determine whether the hard disk has a fault according to the hard disk fault determination strategy, report the fault alarm of the target hard disk with failed IO transceiver, and provide reference steps for fault repair.

[0167] And, Figure 7 is a schematic diagram of fault alarm for an embodiment of the present invention. As Figure 7 shown, after receiving the SMP abnormal information, the embodiment of the present invention can also combine the decision fault points (such as Expander device, cable, hard disk), comprehensively determine the fault object and decide how to alarm, and provide possible fault troubleshooting and repair steps for users:

[0168] (1) Send an instruction to the Expander to query the information of the device where the Expander is located, related to the link, such as:

[0169] Hard disk slot information (present, connection status, hard disk address);

[0170] Cable presence signal;

[0171] Whether the power supply to the hard disk and Expander is normal;

[0172] Expander operating status, reset reason;

[0173] (2) According to the service requirements, send an instruction to the Expander to send Expander control operations:

[0174] Kick the hard disk from the slot;

[0175] Operations on Expander Phy;

[0176] Hard disk plug and unplug, SAS cable plug and unplug;

[0177] (3) Receive the information from the host side:

[0178] IO exception events on the host side;

[0179] Wait for the link on / off information and error code information queried by SMP.

[0180] In an embodiment of the present invention, when receiving the link information of a target port, it can wait to receive the link information of other target ports, so as to comprehensively determine the fault object and fault type according to the combined link information of multiple target ports and in combination with the original fault determination rules; and it can alarm for link-related faults, accurately report the fault object and fault type, and give possible repair methods, which has extremely strong practicability.

[0181] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0182] An embodiment of the present invention also provides a computer program product.

[0183] Figure 8 It is a block diagram of a computer program product provided according to an embodiment of the present invention.

[0184] As Figure 8 shown, the computer program product 10 includes: an acquisition module, a first detection module, and a second detection module.

[0185] Among them, the acquisition module is used to obtain the real-time link information of the multi-stage cascaded expander in the target device and the instruction sending and receiving situation of at least one target port when there is a transmission link fault detection mechanism in the target device.

[0186] The first detection module is used to execute the corresponding first link fault detection program when detecting that the real-time link information is abnormal, so as to determine the first fault detection result of the multi-stage cascade port and / or the transmission link of the multi-stage cascaded expander.

[0187] The second detection module is used to detect the failure type of the instruction sending and receiving situation when detecting that the instruction sending and receiving situation fails, and execute the second link fault detection program corresponding to the failure type to determine the second fault detection result of the target device.

[0188] Optionally, in an embodiment of the present invention, it further includes: a collection module and a establishment module.

[0189] Among them, the collection module is used to collect the software error code and firmware error code in the historical transmission link fault of the target device before obtaining the real-time link information of the multi-stage cascaded expander in the target device and the instruction sending and receiving situation of at least one target port;

[0190] The establishment module is used to establish a transmission link fault detection mechanism in combination with the software error code, firmware error code, preset extension instruction, corresponding parsing strategy, and the management module of the target device.

[0191] Optionally, in an embodiment of the present invention, it further includes: a third detection module and a judgment module.

[0192] The third detection module is used to detect the status change report information of the multi-stage cascaded expander and the sending time of the status change report information before executing the corresponding first link fault detection program.

[0193] The judgment module is used to judge whether there is an abnormality in the real-time link information according to the status change report information and the sending time.

[0194] The determination module is used to determine that there is an abnormality in the real-time link information when the status change report information meets the preset abnormal conditions and / or the sending time is inconsistent with the preset sending time; otherwise, it is determined that there is no abnormality in the real-time link information.

[0195] Optionally, in an embodiment of the present invention, the first detection module includes: a first execution unit and a second execution unit.

[0196] The first execution unit is used to execute the first link fault query program when the status change report information is abnormal status change report information and / or the sending time is inconsistent with the preset sending time.

[0197] The second execution unit is used to execute the first link error query process when there are other abnormal conditions in the real-time link information except that the status change report information is abnormal status change report information and / or the sending time is inconsistent with the preset sending time.

[0198] For the description of the features in the corresponding embodiment of the computer program product 10, reference can be made to the relevant description of the corresponding embodiment of the transmission link fault detection method in the multi-stage cascaded scenario, which will not be elaborated here one by one.

[0199] An embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above embodiments of the transmission link fault detection method in the multi-stage cascaded scenario.

[0200] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the transmission link fault detection method in the multi-stage cascaded scenario when running.

[0201] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), external hard drives, magnetic disks, or optical discs that can store computer programs.

[0202] The embodiments of the present invention also provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the transmission link fault detection method for a multi-stage cascade scenario.

[0203] The embodiments of the present invention also provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the transmission link fault detection method for a multi-stage cascade scenario.

[0204] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0205] The above has introduced in detail a transmission link fault detection method and program product for a multi-stage cascade scenario provided by the present invention. Specific examples are used herein to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A transmission link failure detection method for a multi-stage cascade scenario, characterized in that: The following steps are involved: In the case where the target device has a transmission link failure detection mechanism, obtaining real-time link information of a multi-stage cascade expander in the target device and command receiving and sending status of at least one target port; When an abnormality is detected in the real-time link information, executing a corresponding first link fault detection program to determine a first fault detection result of the multi-stage cascade port and / or the transmission link of the multi-stage cascade expander; When a failure in the instruction receiving and sending is detected, detecting a failure type of the instruction receiving and sending, and executing a second link failure detection program corresponding to the failure type to determine a second failure detection result of the target device; Wherein, when the failure of the instruction receiving and sending is detected, detecting the failure type of the instruction receiving and sending includes: detecting whether the real-time link information is abnormal before the failure of the receiving and sending of the at least one target port; if the real-time link information is abnormal before the failure of the receiving and sending of the at least one target port, determining the failure type as a passive type, otherwise, determining the failure type as an active type; Wherein, when the failure of the instruction receiving and sending is detected, the failure type of the instruction receiving and sending is detected, and the second link fault detection procedure corresponding to the failure type is executed, including: when the failure type is the passive type, the second link fault query process is executed to determine the second fault detection result of the target device; when the failure type is the active type, the second link error query process is executed to determine the second fault detection result of the target device; Wherein, when the failure type is the passive type, executing the second link fault query process to determine the second fault detection result of the target device includes: when the failure type is the passive type, traversing the multi-stage cascade port and its transmission link to determine the on-off change of the transmission link; based on the routing table pre-established by the multi-stage cascade expander and the on-off change, identifying the on-off link end of the transmission link; determining the second fault detection result of the target device according to the on-off link end; Among them, when the failure type is the active type, a second link error query process is executed to determine the second fault detection result of the target device, including: when the failure type is the active type, a query instruction is sent to the multi-stage cascade expander; based on the sending and receiving information of the query instruction, the second fault detection result of the target device is generated.

2. The method according to claim 1, characterized in that Before acquiring the real-time link information of the multi-stage cascade expander in the target device and the instruction sending and receiving status of the at least one target port, the method further includes: Collecting software error codes and firmware error codes in historical transmission link failures of the target device; The transmission link fault detection mechanism is established in combination with the software error code, the firmware error code, the preset extension instruction and the corresponding parsing strategy and the management module of the target device.

3. The method according to claim 1, characterized in that Before executing the corresponding first link failure detection program, the method further includes: Detecting status change report information of the multi-stage cascade expander and the sending time of the status change report information; Determining whether the real-time link information is abnormal according to the state change report information and the sending time; If the state change report information satisfies the preset abnormal condition and / or the sending time is inconsistent with the preset sending time, it is determined that the real-time link information is abnormal; otherwise, it is determined that the real-time link information is not abnormal.

4. The method according to claim 3, characterized in that: When the real-time link information is detected to be abnormal, executing the corresponding first link fault detection program includes: When the state change report information is abnormal state change report information and / or the sending time is inconsistent with the preset sending time, executing a first link fault query procedure; When the real-time link information has other abnormal situations except that the state change report information is abnormal state change report information and / or the sending time is inconsistent with the preset sending time, the first link error query process is executed.

5. The method according to claim 1, characterized in that: Before acquiring the real-time link information of the multi-stage cascade expander in the target device and the instruction sending and receiving status of the at least one target port, the method further includes: Traversing the routing table pre-established by the multi-stage cascade expander, and determining the topological structure corresponding to the routing table; The real-time link information is acquired according to the topology structure.

6. A computer program product, characterized in that include: An acquisition module, used to acquire real-time link information of a multi-stage cascade expander in the target device and the instruction receiving and sending status of at least one target port when a transmission link failure detection mechanism exists in the target device; A first detection module, configured to execute a corresponding first link fault detection program when an abnormality is detected in the real-time link information, so as to determine a first fault detection result of the multi-stage cascade port and / or the transmission link of the multi-stage cascade expander; A second detection module is used to detect the failure type of the instruction receiving and sending situation when it is detected that there is a failure in the instruction receiving and sending situation, and execute a second link failure detection program corresponding to the failure type to determine a second failure detection result of the target device; The second detection module includes: a detection unit, configured to detect whether the real-time link information is abnormal before the transceiver status of the at least one target port fails; a judgment unit, configured to judge that the failure type is a passive type if the real-time link information is abnormal before the transceiver status of the at least one target port fails, otherwise, judge that the failure type is an active type; The second detection module includes: a first determination unit, configured to execute a second link fault query process to determine a second fault detection result of the target device when the failure type is the passive type; a second determination unit, configured to execute a second link error query process to determine a second fault detection result of the target device when the failure type is the active type; The first determination unit includes: a traversal subunit, which is used to traverse the multi-stage cascade ports and their transmission links to determine the on-off changes of the transmission links when the failure type is the passive type; an identification subunit, which is used to identify the on-off link end of the transmission link based on the routing table pre-established by the multi-stage cascade expander and the on-off changes; a determination subunit, which is used to determine the second fault detection result of the target device according to the on-off link end; Among them, the second determination unit includes: a sending subunit, used to send a query instruction to the multi-stage cascade expander when the failure type is the active type; a generating subunit, used to generate a second fault detection result of the target device based on the sending and receiving information of the query instruction.

7. The computer program product according to claim 6, characterized in that Also includes: A collection module, used for collecting software error codes and firmware error codes in historical transmission link failures of the target device before acquiring the real-time link information of the multi-stage cascade expander in the target device and the instruction sending and receiving status of the at least one target port; An establishment module is used to establish the transmission link fault detection mechanism in combination with the software error code, the firmware error code, the preset extension instruction and the corresponding parsing strategy and the management module of the target device.

8. The computer program product according to claim 6, characterized in that Also includes: A third detection module, used for detecting the state change report information of the multi-stage cascade expander and the sending time of the state change report information before executing the corresponding first link fault detection program; A judging module, used for judging whether the real-time link information is abnormal according to the state change report information and the sending time; The determination module is used to determine that the real-time link information is abnormal if the state change report information meets the preset abnormal condition and / or the sending time is inconsistent with the preset sending time; otherwise, determine that the real-time link information is not abnormal.

9. The computer program product according to claim 8, characterized in that The first detection module includes: A first execution unit, configured to execute a first link fault query program when the state change report information is abnormal state change report information and / or the sending time is inconsistent with a preset sending time; The second execution unit is used to execute the first link error query process when the real-time link information has other abnormal situations except that the state change report information is abnormal state change report information and / or the sending time is inconsistent with the preset sending time.

10. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the transmission link failure detection method for a multi-stage cascade scenario as described in any one of claims 1 to 5.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the transmission link failure detection method for a multi-stage cascade scenario as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Monitoring method, device and apparatus for extender

    CN111124818A

  • Link fault diagnosis method and device, electronic equipment and readable storage medium

    CN114448915A