Fault detection method and data management platform for storage systems

By monitoring path identification information in real time, the problem of business downtime caused by single-controller paths in multi-controller storage systems was solved, enabling timely fault detection and alarm, and ensuring system stability.

CN121210218BActive Publication Date: 2026-03-03INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511771513.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-03-03
Estimated Expiration
2045-11-28

AI Technical Summary

Technical Problem

In multi-controller storage systems, when only one controller remains connected to the host, business downtime is easily caused by a fault, and existing technologies struggle to detect and resolve this issue in a timely manner.

Method used

By monitoring the path identification information between the host and the storage system, it can determine in real time whether the storage system is in a single-controller path state, identify any redundant nodes or path failures, and promptly send alarm information to avoid business downtime.

Benefits of technology

It enables timely detection of storage system faults, reduces the probability of business downtime, and ensures the stability and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121210218B_ABST
    Figure CN121210218B_ABST
Patent Text Reader

Abstract

The application discloses a fault detection method and a data management platform of a storage system. There is a data transmission path between a host and multiple controller nodes of the storage system. Each controller node corresponds to at least one path. The method determines whether the identification information of the current available paths is the same through real-time monitoring. If the identification information is the same, the storage system is in a single control path state, that is, there is no redundant path between the controller nodes. When the controller node fails, is upgraded or is powered off, the host cannot read and write the storage system, leading to business downtime. Therefore, it is determined that the storage system has a non-redundant node fault. In addition, if there is only one current available path, it is determined that the storage system has a non-redundant path fault. Compared with periodical testing, the fault can be found in time, the probability of business downtime is greatly reduced, the stability of the business is ensured, and the problem that periodical testing is difficult to timely perceive the non-redundant fault of the storage system and easily leads to business downtime is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of storage technology, and in particular to fault detection of storage systems. Background Technology

[0002] In a multi-controller storage control system, multiple controllers are redundant to ensure the reliability of services; on the other hand, multiple controllers work together to carry out services and achieve a high throughput.

[0003] If only one controller remains connected to the host in the storage system, and that controller fails due to software or hardware issues (such as a reboot or power failure), the storage system will be unable to provide services, resulting in a short-term or long-term outage. Similarly, if only one controller remains connected to the host, and this is not apparent to the user, upgrading that node or replacing components after a power outage will also cause service outages. Summary of the Invention

[0004] This application provides a fault detection method and data management platform for a storage system, which at least solves the problem in related technologies where regular testing makes it difficult to detect in a timely manner that only one controller exists in the storage system and its connection path to the host, which can easily lead to business downtime.

[0005] This application provides a fault detection method for a storage system. Data transmission paths exist between the host and multiple controller nodes of the storage system, with each controller node corresponding to at least one path. The method includes: the host obtaining the identification information of currently available paths; paths with the same controller node have the same identification information; and currently available paths are paths that the host can read from and write to the storage system. If there is only one type of identification information for currently available paths, the host determines that the storage system is in a single-controller-path state. If the storage system is in a single-controller-path state and there are multiple currently available paths, the host determines that the storage system has a fault with a node without redundancy. If the storage system is in a single-controller-path state and there is only one currently available path, the host determines that the storage system has a fault with a path without redundancy.

[0006] This application also provides a data management platform, including: a storage system including multiple controller nodes; a host having data transmission paths between it and the multiple controller nodes of the storage system, each controller node corresponding to at least one path; a memory for storing computer programs; and a processor for implementing the steps of a fault detection method for any of the storage systems when executing the computer programs.

[0007] Through this application, the fault detection method of the above-mentioned storage system monitors in real time whether the identification information of the currently available paths is the same. If they are all the same, the storage system is in a single-controller path state, that is, there is only one controller in the storage system with a connection path to the host. There are no redundant paths between the controller nodes. When the controller node fails, is upgraded, or is powered off, the host cannot read or write to the storage system, resulting in business downtime. Therefore, it is determined that the storage system has a non-redundant node fault. In addition, if there is only one currently available path, a path failure will cause business downtime. Therefore, it is determined that the storage system has a non-redundant path fault. Compared with periodic testing, faults can be detected in time, which greatly reduces the probability of business downtime, ensures business stability, and solves the problem that periodic testing is difficult to detect in time the non-redundant fault of the storage system, which can easily lead to business downtime. Attached Figure Description

[0008] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 This is a hardware structure block diagram of a fault detection method for a storage system according to an embodiment of this application;

[0010] Figure 2 This is a flowchart of a fault detection method for a storage system according to an embodiment of this application;

[0011] Figure 3 This is a flowchart illustrating how to determine the state of a single-control path according to an embodiment of this application;

[0012] Figure 4 This is a schematic diagram of the state of a dual-controller storage system according to an embodiment of this application;

[0013] Figure 5 This is a flowchart illustrating a path deployment method according to an embodiment of this application;

[0014] Figure 6 This is a flowchart of a single-control path status verification according to an embodiment of this application;

[0015] Figure 7 This is a structural block diagram of a fault detection device for a storage system according to an embodiment of this application. Detailed Implementation

[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0017] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," "third," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0018] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0019] The specific application environment architecture or specific hardware architecture on which the execution of the fault detection method of the storage system depends is described here.

[0020] The methods and embodiments provided in this application can be executed on a server device or a similar computing device. Taking running on a server device as an example, Figure 1 This is a hardware structure block diagram of a fault detection method for a storage system according to an embodiment of this application. Figure 1 As shown, the server device may include one or more ( Figure 1 Only one is shown in the image. A processor 102 (which may include, but is not limited to, a central processing unit (CPU), microprocessor (MCU), or programmable logic device (FPGA), etc.) and a memory 104 for storing data are also shown. The server device may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the server equipment described above. For example, the server equipment may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0021] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the fault detection method of the storage system in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thus implementing the above-described method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to server devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0022] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the server device. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0023] The embodiments of this application provide a fault detection method for a storage system. The method is described in detail below in conjunction with the execution flow of the fault detection method for a storage system.

[0024] The following explains the technical terms used in this application:

[0025] NVMe (Non-Volatile Memory Express): A non-volatile memory host controller interface specification that optimizes SSD access performance;

[0026] SCSI (Small Computer System Interface): A small computer system interface used for communication between a host and storage devices;

[0027] RTPG (Report Target Port Groups): A SCSI protocol command used by the host to obtain information about storage system port groups, primarily for determining the path status of ports;

[0028] WWPN (World Wide Port Name): A globally unique port identifier, represented in hexadecimal and separated by colons, for example: :50:06:04:81:D6:F3:45:42, which serves as a unique port identifier for communication.

[0029] This embodiment provides a fault detection method for a storage system, wherein there are data transmission paths between the host and multiple controller nodes of the storage system, and each controller node corresponds to at least one path. Figure 2 This is a flowchart of a fault detection method for a storage system according to an embodiment of this application, such as... Figure 2 As shown, the method includes the following steps:

[0030] Step S202: The host obtains the identification information of the currently available paths. The identification information of paths on the same controller node is the same. The currently available paths are the paths that the host can read and write to the storage system.

[0031] Specifically, when constructing the link path between the host's LUN and the controller node, in order to distinguish the paths corresponding to different controllers, the path identification information of the same controller node is the same, and the path identification information of different controller nodes is different. The host obtains the identification information of the currently available path to determine whether the controller node corresponding to the currently available path is the same.

[0032] Step S204: If there is only one type of identification information for the currently available path, the host determines that the storage system is in a single-control path state.

[0033] Specifically, if the identification information of each currently available path is the same, it indicates that the controller nodes corresponding to the currently available paths are the same, and there are no redundant paths between the controller nodes, then the storage system is determined to be in a single-controller-path state.

[0034] Step S206: When the storage system is in a single-controller path state and there are multiple currently available paths, the host determines that the storage system has a fault with no redundant nodes.

[0035] Specifically, the storage system is in a single-controller path state with no redundant nodes. If the controller node fails, is upgraded, or is powered off, the host cannot read or write the volume corresponding to the LUN in the storage system, causing the service to crash. Therefore, it is determined that the storage system has a fault of no redundant nodes.

[0036] In step S208, when the storage system is in a single-controller path state and there is only one currently available path, the host determines that the storage system has a non-redundant path fault.

[0037] Specifically, the storage system is in a single-controller path state and there is only one currently available path. There is no redundant path, which makes it easy for path failures to cause business downtime. It is determined that the storage system has a non-redundant path failure.

[0038] By following the steps above, and by monitoring in real time whether the identification information of the currently available paths is the same, if they are all the same, the storage system is in a single-controller path state. That is, only one controller in the storage system has a connection path with the host, and there are no redundant paths between the controller nodes. This means that when the controller node fails, is upgraded, or is powered off, the host cannot read or write to the storage system, resulting in business downtime. Therefore, it is determined that the storage system has a non-redundant node fault. In addition, if there is only one currently available path, a path failure will cause business downtime. Therefore, it is determined that the storage system has a non-redundant path fault. Compared with periodic testing, faults can be detected in time, which greatly reduces the probability of business downtime, ensures business stability, and solves the problem that periodic testing is unable to detect the non-redundant fault in the storage system in time, which can easily lead to business downtime.

[0039] To accommodate different scenarios, as an optional implementation, step S202 above includes:

[0040] Step S2022: When the port group identifiers of the paths of different controller nodes are different, the host obtains the port group identifier of the currently available path and obtains the identification information of the currently available path.

[0041] In step S2024, if the port group identifiers of different controller nodes are the same, the host obtains the target global port name or target device name of the currently available path to obtain the identification information of the currently available path.

[0042] In the above implementation, the host includes multiple LUNs, each LUN corresponding to a volume on a controller node. Each LUN has multiple accessible paths, and SCSI protocol commands can be used to determine whether different paths belong to the same controller node. Generally, paths with the same port group ID exhibit consistent status changes, and the storage system assigns different port group IDs to different controllers. Host multipathing determines whether a path is a single-controller path by querying the port group ID of each valid path. However, some storage system vendors may set the same port group ID for different controller nodes. In this case, the target global port name (WWPN) or target device name of the storage system can be used for further determination. For example, on a Linux host system, the `lsscsi` command can be used to query the target WWPN or target device name corresponding to a specific path. The IP-SAN protocol corresponds to the target device name, meaning that paths on the same controller node have the same target device name; the FC-SAN protocol corresponds to the target global port name (target wwpn), meaning paths on the same controller node have the same target global port name (target wwpn).

[0043] Taking querying the port group identifier (port group id) as an example, such as Figure 3 As shown, the process first queries the first currently available path, retrieves and records its port group ID, and then obtains the next currently available path. If no other currently available path exists, the path is in a single-controller path state. If other valid available paths exist, the port group IDs of these paths are retrieved and compared with the port group ID of the first path. If they are different, different controller node paths exist, and the traversal query stops; otherwise, the traversal continues. This process continues until all currently available paths have been traversed. If different port group IDs exist, it indicates that multiple controllers exist, meaning the path is not in a single-controller path state; otherwise, the path is in a single-controller path state.

[0044] To determine whether a fault path exists, as an optional implementation, after the host determines that the storage system is in a single-controller path state, the above method further includes:

[0045] Step S302: The host sends a single-control path status alarm message to the storage system;

[0046] Step S304: The host obtains the status of the controller nodes of the storage system. If the number of controller nodes in the online state is greater than 1, it is determined that there is a fault path.

[0047] In the above implementation, the host sends a single-controller path status alarm message to the storage system to query the status of the controller nodes of the storage system. If the number of controller nodes in the online state is greater than 1, it indicates that there is no currently available path between the online controller node and the host, and therefore there must be a faulty path.

[0048] Taking a dual-controller storage system as an example, such as Figure 4 As shown, the host uses the multipath software multipath, which displays three basic states: The first is the normal state, where all nodes are online and the path is available; the second scenario is the single-controller path state, where both controller node 1 (Node1) and controller node 2 (Node2) are online, but the link of controller node 2 (Node2) fails, making it unavailable. In this scenario, the user needs to be notified to check the cause of the path failure and repair the faulty path; the third scenario is that only one controller node is online, meaning that controller node 2 (Node2) fails, which in turn makes the path unavailable. This scenario mainly occurs when nodes restart or when the user replaces hardware and actively powers down the device.

[0049] To determine whether a path is available, as an optional implementation, the method further includes the following steps before the host obtains the identification information of currently available paths:

[0050] Step S402: The host obtains the status information of each path;

[0051] Step S404: If the path status information is in an available state, the host determines that the path is currently available.

[0052] In step S406, if the path status information is unavailable or offline, the host determines that the path is not currently available.

[0053] In the above implementation, the available state refers to a path that is in the active state and can carry services, i.e., the optimal or non-optimal state. The unavailable state is the unavailable state, and the offline state is the offline state. If the path's status information is available, then it is a currently available path; otherwise, it is not a currently available path.

[0054] To achieve path redundancy, as an optional implementation, before the host obtains the identification information of the currently available path, the above method further includes:

[0055] Step S502: After the volume mapping between the host and the storage system is completed, the host checks whether there is a path between the host and each controller node;

[0056] Step S504: If paths exist between the host and each controller node, the host deployment is confirmed to be successful.

[0057] In step S506, if there is no path between the host and the target controller node, the host sends a prompt message. The prompt message is to add a path connection between the host and the target controller node, where the target controller node can be any controller node.

[0058] In the above implementation, when mapping a host to a volume, it checks whether the host has paths on each controller node. If there are connections on all nodes, the paths are completely redundant, and the deployment is successful. If there are controller nodes without connections, the host sends a prompt message: Add path connection information between the host and the target controller node.

[0059] To deploy redundant paths, as an optional implementation, step S506 above includes:

[0060] Step S5062: The host obtains the status of the target controller node;

[0061] Step S5064: When the target controller node is online, the host sends a first prompt message, which is information to add a path connection between the host and the target controller node;

[0062] In step S5066, if the target controller node is offline, the host sends a second prompt message. The second prompt message is information on adding a path connection between the host and the target controller node after the target controller node recovers.

[0063] In the above implementation, the host obtains the status of each controller node. If all controller nodes are online, a first prompt message is sent to the user: it is recommended to add a path connection to a certain node. If the user chooses to add, the system checks again after the addition; if the user does not add, the deployment can be set to fail or succeed depending on the importance of the business. If the node is offline, a second prompt message is sent to the user: it is recommended to add the host connection after the node recovers.

[0064] To determine the deployment result, as an optional implementation, after the host sends a notification message, the method further includes:

[0065] Step S602: When adding path connections between the host and the target controller node, the host re-detects whether there are paths between the host and each controller node, and if there are paths between the host and each controller node, the host determines that the deployment is successful.

[0066] In step S604, without adding a path connection between the host and the target controller node, the host deployment is determined to have failed.

[0067] In the above implementation, after the host sends a prompt message to add a path connection between the host and the target controller node, if the user chooses to add it, the system will check again after adding it. If there is a path between the host and each controller node, the host will determine that the deployment is successful. If the user does not add any more, the host will usually determine that the deployment has failed. Of course, the system can also set the deployment to fail or succeed based on the importance of the business.

[0068] Deployment Process: When deploying services, the connection status of paths can be checked in advance. If a host is indeed not connected at a certain node, the user is prompted during deployment. After the user confirms, redundant paths are added, such as... Figure 5 As shown, firstly, during host-volume mapping, it checks if the host has paths on all nodes. If all nodes have connections, the paths are completely redundant, and the execution succeeds. If any node has no connection, it further checks if the node is online. If it is offline, it prompts the user to add a host connection after the node recovers. If all nodes are online, it prompts the user to add a path connection to a certain node. If the user chooses to add a connection, it checks again after adding it; if the user does not add another connection, the deployment can be set to fail or succeed based on the importance of the business.

[0069] To avoid misjudgment, as an optional implementation, step S304 above includes:

[0070] Step S3042: The host obtains the detection results of whether the path between the storage system and each controller node is available;

[0071] Step S3044: If the detection result shows that there are currently available paths between the host and each controller node, the host determines that the storage system is not in a single-controller path state and updates the identification information of the corresponding path.

[0072] Step S3046: If the detection result shows that the number of controller nodes that are currently available and online between the host and a single controller node is greater than 1, the host determines that there is a faulty path.

[0073] In the above embodiments, such as Figure 6As shown, after receiving a single-controller path status alarm from the multipath software, the storage system first checks its current status. If it is in a single-controller-node-on-line state (meaning other controller nodes are temporarily unavailable due to power failure or restart), no alarm needs to be reported to notify the user. If the check finds that all controller nodes are online and multiple controller node paths are available, then the host multipath software's judgment is incorrect or there is a delay in the host multipath software's synchronization information; updating the corresponding path's identifier information is sufficient. If the number of online controller nodes is greater than one and there is indeed a case where only a single controller node has a available path, then an alarm needs to be reported to the user to remind them to repair the faulty path in a timely manner.

[0074] To facilitate the transmission of alarm information, as an optional implementation, step S302 above includes:

[0075] In step S3022, the host sends an in-band management command to the storage system. The in-band management command includes single-controller path status alarm information.

[0076] In the above implementation, the host sends an in-band management command to the storage system, using the SCSI Mode Select (10) command format to define the SCSI in-band management command. According to the SPC4 protocol document, the opcode of the SCSI Mode Select (10) command is 55h. According to the protocol document, mode page 00h is a vendor-defined field; mode page 20h to 3eh and all subpage codes below it are vendor-defined fields. The in-band management mode page is defined as 0x20, and the subpage code is 0x00. After the storage successfully executes the command, it returns a success message to the host. First, the SCSI in-band management command CDB uses a byte-10 length CDB structure, as shown in Table 1, where PARAMETER LIST LENGTH = 128+4=132B. The mode page header is 4B, and the Mode parameters (multipath information) are 128B. If the value is not specified, the PARAM_LIST_LEGNTH_ERR error code is returned. Secondly, because the subpage field is used, according to the SPC4 protocol specification, the sub_page mode page format is used as the data out buffer format, as shown in Table 2. PARAMETER DATALENGTH indicates the length of the multipath information, which is fixed at 128 bytes. Otherwise, the PARAM_LIST_LEGNTH_ERR error code is returned. The Mode Parameters are vendor-defined data, as shown in Table 3. Among them, multipath information: multipath information is set independently by the self-developed multipath software; path alarm information: indicates that a volume is in a single-controller path state. This part can be flexibly set; if there are many volumes, it can be omitted here. Reserved: filled with 0s. If it is not 0, the command terminates and returns a parametervalue invalid error to the host.

[0077] Table 1

[0078]

[0079] Table 2

[0080]

[0081] Table 3

[0082]

[0083] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0084] Embodiments of this application also provide a fault detection device for a storage system, applied to a host computer. Figure 7 This is a structural block diagram of a fault detection device for a storage system according to an embodiment of this application, such as... Figure 7 As shown, the device includes:

[0085] The first acquisition unit 702 is used to acquire the identification information of the currently available path. The identification information of the paths of the same controller node is the same. The currently available path is the path that the host can read and write to the storage system.

[0086] Specifically, when constructing the link path between the host's LUN and the controller node, in order to distinguish the paths corresponding to different controllers, the path identification information of the same controller node is the same, and the path identification information of different controller nodes is different. The host obtains the identification information of the currently available path to determine whether the controller node corresponding to the currently available path is the same.

[0087] The first determining unit 704 is used to determine that the storage system is in a single-control path state when there is only one type of identification information for the currently available path.

[0088] Specifically, if the identification information of each currently available path is the same, it indicates that the controller nodes corresponding to the currently available paths are the same, and there are no redundant paths between the controller nodes, then the storage system is determined to be in a single-controller-path state.

[0089] The second determining unit 706 is used to determine whether there is a non-redundant fault in the storage system when the storage system is in a single-control path state and there are multiple currently available paths.

[0090] Specifically, the storage system is in a single-controller path state with no redundant nodes. If the controller node fails, is upgraded, or is powered off, the host cannot read or write the volume corresponding to the LUN in the storage system, causing the service to crash. Therefore, it is determined that the storage system has a lack of redundancy.

[0091] The third determining unit 708 is used to determine, when the storage system is in a single-control path state and there is only one currently available path, that the host has a fault of no redundant path in the storage system.

[0092] Specifically, the storage system is in a single-controller path state and there is only one currently available path. There is no redundant path, which makes it easy for path failures to cause business downtime. It is determined that the storage system has a non-redundant path failure.

[0093] By using the above devices, the identification information of currently available paths is monitored in real time. If they are all the same, the storage system is in a single-controller path state, meaning that only one controller in the storage system has a connection path with the host. There are no redundant paths between controller nodes. This means that when the controller node fails, is upgraded, or is powered off, the host cannot read or write to the storage system, causing business downtime. Therefore, it is determined that the storage system has a non-redundant node fault. In addition, if there is only one currently available path, a path failure will cause business downtime. Therefore, it is determined that the storage system has a non-redundant path fault. Compared with periodic testing, faults can be detected in time, greatly reducing the probability of business downtime, ensuring business stability, and solving the problem that periodic testing is unable to detect non-redundant faults in the storage system in time, which can easily lead to business downtime.

[0094] To accommodate different scenarios, as an optional implementation, the first acquisition unit described above includes:

[0095] The first acquisition module is used to acquire the port group identifier of the currently available path and obtain the identification information of the currently available path when the port group identifiers of the paths of different controller nodes are different.

[0096] The second acquisition module is used to obtain the target global port name or target device name of the currently available path when there are different controller nodes with the same port group identifier, and obtain the identification information of the currently available path.

[0097] In the above implementation, the host includes multiple LUNs, each LUN corresponding to a volume on a controller node. Each LUN has multiple accessible paths, and SCSI protocol commands can be used to determine whether different paths belong to the same controller node. Generally, paths with the same port group ID exhibit consistent status changes, and the storage system assigns different port group IDs to different controllers. Host multipathing determines whether a path is a single-controller path by querying the port group ID of each valid path. However, some storage system vendors may set the same port group ID for different controller nodes. In this case, the target global port name (WWPN) or target device name of the storage system can be used for further determination. For example, on a Linux host system, the `lsscsi` command can be used to query the target WWPN or target device name corresponding to a specific path. The IP-SAN protocol corresponds to the target device name, meaning that paths on the same controller node have the same target device name; the FC-SAN protocol corresponds to the target global port name (target wwpn), meaning paths on the same controller node have the same target global port name (target wwpn).

[0098] Taking querying the port group identifier (port group id) as an example, such as Figure 3 As shown, the process first queries the first currently available path, retrieves and records its port group ID, and then obtains the next currently available path. If no other currently available path exists, the path is in a single-controller path state. If other valid available paths exist, the port group IDs of these paths are retrieved and compared with the port group ID of the first path. If they are different, different controller node paths exist, and the traversal query stops; otherwise, the traversal continues. This process continues until all currently available paths have been traversed. If different port group IDs exist, it indicates that multiple controllers exist, meaning the path is not in a single-controller path state; otherwise, the path is in a single-controller path state.

[0099] To determine whether a fault path exists, as an optional implementation, the above-mentioned apparatus further includes:

[0100] The first sending unit is used to send single-control path status alarm information to the storage system after the host determines that the storage system is in a single-control path state;

[0101] The third determining unit is used to obtain the status of the controller nodes of the storage system, and if the number of controller nodes in the online state is greater than 1, it determines that there is a fault path.

[0102] In the above implementation, the host sends a single-controller path status alarm message to the storage system to query the status of the controller nodes of the storage system. If the number of controller nodes in the online state is greater than 1, it indicates that there is no currently available path between the online controller node and the host, and therefore there must be a faulty path.

[0103] Taking a dual-controller storage system as an example, such as Figure 4 As shown, the host uses the multipath software multipath, which displays three basic states: The first is the normal state, where all nodes are online and the path is available; the second scenario is the single-controller path state, where both controller node 1 (Node1) and controller node 2 (Node2) are online, but the link of controller node 2 (Node2) fails, making it unavailable. In this scenario, the user needs to be notified to check the cause of the path failure and repair the faulty path; the third scenario is that only one controller node is online, meaning that controller node 2 (Node2) fails, which in turn makes the path unavailable. This scenario mainly occurs when nodes restart or when the user replaces hardware and actively powers down the device.

[0104] To determine whether a path is available, as an optional implementation, the above-described apparatus further includes:

[0105] The second acquisition unit is used to acquire the status information of each path before the host acquires the identification information of the currently available paths;

[0106] The fourth determining unit is used to determine whether a path is currently available when the path's status information indicates that the path is available.

[0107] The fifth determining unit is used to determine that a path is not currently available when the path's status information is unavailable or offline.

[0108] In the above implementation, the available state refers to a path that is in the active state and can carry services, i.e., the optimal or non-optimal state. The unavailable state is the unavailable state, and the offline state is the offline state. If the path's status information is available, then it is a currently available path; otherwise, it is not a currently available path.

[0109] To achieve path redundancy, as an optional implementation, the above-mentioned apparatus further includes:

[0110] The first detection unit is used to detect whether there is a path between the host and each controller node before the host obtains the identification information of the currently available path and after the volume of the host and the storage system has been mapped.

[0111] The sixth determining unit is used to determine successful deployment if paths exist between the host and each controller node.

[0112] The second sending unit is used to send a prompt message when there is no path between the host and the target controller node. The prompt message is information to add a path connection between the host and the target controller node, where the target controller node can be any controller node.

[0113] In the above implementation, when mapping a host to a volume, it checks whether the host has paths on each controller node. If there are connections on all nodes, the paths are completely redundant, and the deployment is successful. If there are controller nodes without connections, the host sends a prompt message: Add path connection information between the host and the target controller node.

[0114] To deploy redundant paths, as an optional implementation, the second sending unit described above includes:

[0115] The third acquisition module is used to acquire the status of the target controller node;

[0116] The first sending module is used to send a first prompt message when the target controller node is online. The first prompt message is information to add a path connection between the host and the target controller node.

[0117] The second sending module is used to send a second prompt message when the target controller node is offline. The second prompt message is information on adding a path connection between the host and the target controller node after the target controller node recovers.

[0118] In the above implementation, the host obtains the status of each controller node. If all controller nodes are online, a first prompt message is sent to the user: it is recommended to add a path connection to a certain node. If the user chooses to add, the system checks again after the addition; if the user does not add, the deployment can be set to fail or succeed depending on the importance of the business. If the node is offline, a second prompt message is sent to the user: it is recommended to add the host connection after the node recovers.

[0119] To determine the deployment result, as an optional implementation, the above-mentioned apparatus further includes:

[0120] The second detection unit is used to re-detect whether there is a path between the host and each controller node after the host sends the prompt information and the path connection between the host and the target controller node is increased. If there is a path between the host and each controller node, the host determines that the deployment is successful.

[0121] The seventh determination unit is used to determine deployment failure without increasing the path connection between the host and the target controller node.

[0122] In the above implementation, after the host sends a prompt message to add a path connection between the host and the target controller node, if the user chooses to add it, the system will check again after adding it. If there is a path between the host and each controller node, the host will determine that the deployment is successful. If the user does not add any more, the host will usually determine that the deployment has failed. Of course, the system can also set the deployment to fail or succeed based on the importance of the business.

[0123] Deployment Process: When deploying services, the connection status of paths can be checked in advance. If a host is indeed not connected at a certain node, the user is prompted during deployment. After the user confirms, redundant paths are added, such as... Figure 5 As shown, firstly, during host-volume mapping, it checks if the host has paths on all nodes. If all nodes have connections, the paths are completely redundant, and the execution succeeds. If any node has no connection, it further checks if the node is online. If it is offline, it prompts the user to add a host connection after the node recovers. If all nodes are online, it prompts the user to add a path connection to a certain node. If the user chooses to add a connection, it checks again after adding it; if the user does not add another connection, the deployment can be set to fail or succeed based on the importance of the business.

[0124] To avoid misjudgment, as an optional implementation, the third determining unit mentioned above includes:

[0125] The fourth acquisition module is used to acquire the detection results of whether the paths between the storage system detection host and each controller node are available;

[0126] The first determining module is used to determine that the storage system is not in a single-controller path state and update the corresponding path identification information when the detection result shows that there are currently available paths between the host and each controller node.

[0127] The second determination module is used to determine the existence of a faulty path when the detection result shows that the number of controller nodes that are currently available and online between the host and a single controller node is greater than 1.

[0128] In the above embodiments, such as Figure 6As shown, after receiving a single-controller path status alarm from the multipath software, the storage system first checks its current status. If it is in a single-controller-node-on-line state (meaning other controller nodes are temporarily unavailable due to power failure or restart), no alarm needs to be reported to notify the user. If the check finds that all controller nodes are online and multiple controller node paths are available, then the host multipath software's judgment is incorrect or there is a delay in the host multipath software's synchronization information; updating the corresponding path's identifier information is sufficient. If the number of online controller nodes is greater than one and there is indeed a case where only a single controller node has a available path, then an alarm needs to be reported to the user to remind them to repair the faulty path in a timely manner.

[0129] To facilitate the transmission of alarm information, as an optional implementation, the first sending unit mentioned above includes:

[0130] The third sending module is used to send in-band management commands to the storage system. The in-band management commands include single-controller path status alarm information.

[0131] In the above implementation, the host sends an in-band management command to the storage system, using the SCSI Mode Select (10) command format to define the SCSI in-band management command. According to the SPC4 protocol document, the opcode of the SCSI Mode Select (10) command is 55h. According to the protocol document, mode page 00h is a vendor-defined field; mode page 20h to 3eh and all subpage codes below it are vendor-defined fields. The in-band management mode page is defined as 0x20, and the subpage code is 0x00. After the storage successfully executes the command, it returns a success message to the host. First, the SCSI in-band management command CDB uses a byte-10 length CDB structure, as shown in Table 1, where PARAMETER LIST LENGTH = 128+4=132B. The mode page header is 4B, and the Mode parameters (multipath information) are 128B. If the value is not specified, the PARAM_LIST_LEGNTH_ERR error code is returned. Secondly, because the subpage field is used, according to the SPC4 protocol specification, the sub_page mode page format is used as the data out buffer format, as shown in Table 2. PARAMETER DATALENGTH indicates the length of the multipath information, which is fixed at 128 bytes. Otherwise, the PARAM_LIST_LEGNTH_ERR error code is returned. The Mode Parameters are vendor-defined data, as shown in Table 3. Among them, multipath information: multipath information is set independently by the self-developed multipath software; path alarm information: indicates that a volume is in a single-controller path state. This part can be flexibly set; if there are many volumes, it can be omitted here. Reserved: filled with 0s. If it is not 0, the command terminates and returns a parametervalue invalid error to the host.

[0132] For a description of the features in the embodiment corresponding to the fault detection device of the storage system, please refer to the relevant description of the embodiment corresponding to the fault detection method of the storage system, which will not be repeated here.

[0133] Embodiments of this application also provide a data management platform, including:

[0134] The storage system includes multiple controller nodes;

[0135] There are data transmission paths between the host and multiple controller nodes of the storage system, with each controller node corresponding to at least one path;

[0136] Memory, used to store computer programs;

[0137] A processor is used to implement a fault detection method for any type of storage system when executing a computer program.

[0138] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described embodiments of the fault detection method for storage systems when it is run.

[0139] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0140] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described embodiments of the fault detection method for a storage system.

[0141] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described embodiments of the fault detection method for a storage system.

[0142] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0143] The above provides a detailed description of a fault detection method for a storage system provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to help understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method of failure detection of a storage system, characterized by, There are paths of data transmission between a host and a plurality of controller nodes of a storage system, each of the controller nodes corresponds to at least one of the paths, and the method comprises: The host obtains identification information of a current available path, the identification information of the path of the same controller node is the same, the identification information of the path of different controller nodes is different, and the current available path is the path on which the host can read and write the storage system; In the case that there is only one kind of identification information of the current available path, the host determines that the storage system is in a single-control path state; In the case that the storage system is in the single-control path state and there are a plurality of current available paths, the host determines that there is a non-redundant node failure in the storage system; In the case that the storage system is in the single-control path state and there is only one current available path, the host determines that there is a non-redundant path failure in the storage system.

2. The failure detection method of a storage system according to claim 1, wherein, The host obtains identification information of a current available path, comprising: In the case that the port group identifiers of the paths of different controller nodes are different, the host obtains the port group identifiers of the current available path to obtain the identification information of the current available path; In the case that the port group identifiers of the paths of different controller nodes are the same, the host obtains a target global port name or a target device name of the current available path to obtain the identification information of the current available path.

3. The failure detection method of a storage system according to claim 1, wherein After the host determines that the storage system is in a single-control path state, the method further comprises: The host sends single-control path state alarm information to the storage system; The host obtains the states of the controller nodes of the storage system, and determines that there is a failure path in the case that the number of the controller nodes in an online state is greater than 1.

4. The failure detection method of a storage system according to claim 1, wherein Before the host obtains the identification information of a current available path, the method further comprises: The host obtains state information of each of the paths; In the case that the state information of the path is an available state, the host determines that the path is the current available path; In the case that the state information of the path is an unavailable state or an offline state, the host determines that the path is not the current available path.

5. The failure detection method of a storage system according to claim 1, wherein, Before the host obtains the identification information of a current available path, the method further comprises: In the case that the host and the storage system complete mapping, the host detects whether the paths exist between the host and each of the controller nodes; In the case that the paths exist between the host and each of the controller nodes, the host determines that the deployment is successful; In the case that the paths do not exist between the host and a target controller node, the host sends prompt information, the prompt information is information of adding path connection between the host and the target controller node, and the target controller node is any one of the controller nodes.

6. The failure detection method of a storage system according to claim 5, wherein The host sends prompt information, comprising: The host obtains the state of the target controller node; In the case that the target controller node is online, the host sends first prompt information, which is information for adding path connection between the host and the target controller node; In the case that the target controller node is offline, the host sends second prompt information, which is information for adding path connection between the host and the target controller node after the target controller node is recovered.

7. The failure detection method of a storage system according to claim 5, wherein After the host sends the prompt information, the method further comprises: In the case that the path connection between the host and the target controller node is added, the host re-detects whether the path exists between the host and each of the controller nodes, and in the case that the path exists between the host and each of the controller nodes, the host determines that the deployment is successful; In the case that the path connection between the host and the target controller node is not added, the host determines that the deployment fails.

8. The failure detection method of a storage system according to claim 3, wherein, In the case that the number of the controller nodes in the online state is greater than 1, the host determines that there is a fault path, comprising: The host acquires detection results of the storage system detecting whether the path between the host and each of the controller nodes is available; In the case that the detection result is that the current available path exists between the host and each of the controller nodes, the host determines that the storage system is not in the single-control path state and updates the identification information of the corresponding path; In the case that the detection result is that the current available path exists only between the host and a single controller node and the number of the controller nodes in the online state is greater than 1, the host determines that there is a fault path.

9. The failure detection method of a storage system according to claim 3, wherein, The host sends single-control path state alarm information to the storage system, comprising: The host sends an in-band management command to the storage system, and the in-band management command comprises the single-control path state alarm information.

10. A data management platform, characterized by, Comprising: A storage system comprising a plurality of controller nodes; A host, wherein a path for data transmission exists between the host and the plurality of controller nodes of the storage system, and each of the controller nodes corresponds to at least one path; A memory for storing a computer program; A processor for executing the computer program to realize the steps of the fault detection method of the storage system according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method and system for automatically collecting and analyzing computer cluster node information

    CN105681070A

  • Test method of storage link, electronic equipment, storage medium and program product

    CN119847846A