Host active path switching method and apparatus

Through cross-controller reset mechanism and asynchronous event notification, the host can quickly identify and switch to a healthy path, solving the problems of path fault detection delay and misjudgment in the existing technology, and realizing efficient fault recovery and business continuity of the storage system.

CN121326645BActive Publication Date: 2026-02-17INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511871104.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-11
Publication Date
2026-02-17
Estimated Expiration
2045-12-11

AI Technical Summary

Technical Problem

In existing technologies, the waiting time for host confirmation of path failure is relatively long, and the requirements for KATO or CQT settings are relatively high, which can easily lead to service interruption and misjudgment of path failure. Furthermore, the failure recovery delay of passive detection mechanisms is relatively high, making it impossible to switch to backup paths in a timely manner, thus affecting service continuity.

Method used

Through the Cross Controller Reset (CCR) mechanism and Asynchronous Event Notification (AEN), the host can quickly reset the faulty controller, proactively report path faults, shorten fault detection time, and achieve proactive path switching.

Benefits of technology

By quickly resetting the fault controller through a healthy path, long waiting times for KATO or CQT timeouts are avoided, ensuring business continuity and improving the communication reliability and fault recovery efficiency of the storage system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121326645B_ABST
    Figure CN121326645B_ABST
Patent Text Reader

Abstract

The application discloses a host active path switching method and device, and relates to the technical field of path switching, and comprises the following steps: through a cross-controller reset mechanism, the host can quickly reset a faulty controller through a healthy controller, and asynchronous event notification can be performed, so that the storage controller can actively report to the host that controller failure has completed restart; in addition, through the controller instance random number and unique identifier mechanism, the recovered controller is prevented from being reset by mistake, the problems that in the prior art, the waiting time for the host to confirm path failure is relatively long, the KATO or CQT setting requirement is relatively high, service interruption and path failure misjudgment are easily caused, and the service continuity is greatly affected are solved, and in addition, the delay of the passive detection mechanism is relatively high, so that the host cannot be switched to the standby path in time, and the technical problems are solved, so that the technical effects of shortening the fault detection time, guaranteeing the service continuity and the data availability are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of path switching, and particularly relates to a host active path switching method and device. BACKGROUND

[0002] NVMF (Non-Volatile Memory Express over Fabrics, a non-volatile memory host controller interface through a structure network protocol) host multipath technology establishes multiple physical or logical paths conforming to the interaction specification of the NVMF protocol between a host and a storage device, when a certain path or a certain controller fails, the system can automatically switch to a healthy path or a healthy controller to guarantee business continuity and data availability.

[0003] At present, in the related art, the host periodically sends a keep alive command based on a passive detection mechanism, and the storage controller returns a corresponding response after receiving the keep alive command; if the controller does not receive the command within a preset timeout time, the host considers that the communication is interrupted, and a controller reset or path switching operation can be triggered.

[0004] However, in the related art, the waiting time for the host to confirm the path failure is long, and the KATO (Keep Alive Time Out, keep alive timeout) or CQT (Command Quiesce Time Out, command quiesce timeout) setting requirement is high, which easily leads to business interruption and path failure misjudgment, greatly affecting business continuity; in addition, the passive detection mechanism has a high fault recovery delay, so that the host cannot timely switch to a backup path, which needs to be urgently solved. SUMMARY

[0005] The present application provides a host active path switching method and device to at least solve the technical problems in the related art that the waiting time for the host to confirm the path failure is long, and the KATO or CQT setting requirement is high, which easily leads to business interruption and path failure misjudgment, greatly affecting business continuity; in addition, the passive detection mechanism has a high fault recovery delay, so that the host cannot timely switch to a backup path.

[0006] The application provides a host active path switching method, comprising the following steps: determining an active path between a target host and a first storage controller in a preset storage cluster system, controlling the target host to communicate with the first storage controller through the active path, and detecting whether a first communication loss occurs between the target host and the first storage controller, wherein when the first communication loss occurs, a first cross-controller reset instruction is sent from the target host to a second storage controller in the storage cluster system; based on the first cross-controller reset instruction, a plurality of controller field information of the first storage controller is inquired to determine whether the first storage controller meets a preset cross-controller reset trigger requirement according to the plurality of controller field information; if the first storage controller meets the cross-controller reset trigger requirement, a second cross-controller reset instruction is sent from the second storage controller to the first storage controller to restart the first storage controller based on the second cross-controller reset instruction, and an asynchronous event notification containing first storage controller restart information is sent from the second storage controller to the target host, otherwise an asynchronous event notification of log page change is sent from the second storage controller to the target host; based on the asynchronous event notification, the second storage controller inquires a target cross-controller reset log page to obtain corresponding log information, and determines whether the first storage controller meets a preset path failure requirement according to the log information, wherein when the preset path failure requirement is met, a target active path of the target host is switched to the second storage controller.

[0007] The application further provides a host active path switching device, comprising: a communication detection module, configured to determine an active path between a target host and a first storage controller in a preset storage cluster system, control the target host to communicate with the first storage controller through the active path, and detect whether a first communication loss occurs between the target host and the first storage controller, wherein, when the first communication loss occurs, a first cross-controller reset instruction is sent from the target host to a second storage controller in the storage cluster system; a query module, configured to query a plurality of controller field information of the first storage controller based on the first cross-controller reset instruction, and determine whether the first storage controller meets a preset cross-controller reset trigger requirement according to the plurality of controller field information; a reset module, configured to, if the first storage controller meets the cross-controller reset trigger requirement, send a second cross-controller reset instruction from the second storage controller to the first storage controller, restart the first storage controller based on the second cross-controller reset instruction, and send an asynchronous event notification containing first storage controller restart information from the second storage controller to the target host, or send an asynchronous event notification of log page change from the second storage controller to the target host; and a switching module, configured to, based on the asynchronous event notification, control the second storage controller to query a target cross-controller reset log page to obtain corresponding log information, and determine whether the first storage controller meets a preset path failure requirement according to the log information, wherein, when the preset path failure requirement is met, a target active path of the target host is switched to the second storage controller.

[0008] The application further provides an electronic device, comprising: a memory configured to store a computer program; and a processor configured to execute the computer program to implement the steps of any of the host active path switching methods.

[0009] The application further provides a nonvolatile computer readable storage medium, wherein the nonvolatile computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of any of the host active path switching methods.

[0010] The application further provides a computer program product, comprising a computer program, and the computer program is executed by a processor to implement the steps of any of the host active path switching methods.

[0011] By the present application, the active path between the target host and the first storage controller in the preset storage cluster system can be determined, and the target host is controlled to communicate with the first storage controller through the active path, and whether the first communication loss occurs between the target host and the first storage controller is detected, wherein, when the first communication loss occurs, the first cross-controller reset instruction is sent to the second storage controller in the storage cluster system through the target host; based on the first cross-controller reset instruction, the controller field information of the first storage controller is queried to determine whether the first storage controller meets the preset cross-controller reset trigger requirement according to the controller field information; if the first storage controller meets the cross-controller reset trigger requirement, the second cross-controller reset instruction is sent to the first storage controller by the second storage controller, so that the first storage controller is restarted based on the second cross-controller reset instruction, and the asynchronous event notification containing the first storage controller restart information is sent to the target host by the second storage controller, otherwise the asynchronous event notification of the log page change is sent to the target host by the second storage controller; based on the asynchronous event notification, the second storage controller queries the target cross-controller reset log page to obtain the corresponding log information, and determines whether the first storage controller meets the preset path failure requirement according to the log information, wherein, in the case of meeting the preset path failure requirement, the target active path of the target host is switched to the second storage controller, so that the technical problems that the waiting time of host path fault confirmation is long and the KATO or CQT setting requirement is high in the related art, which easily leads to business interruption and path fault misjudgment, greatly affects the business continuity, can be solved; in addition, the fault recovery delay of the passive detection mechanism is high, so that the host cannot be switched to the standby path in time, which achieves the technical effects that the faulty controller is quickly reset through the healthy path, the long-time waiting for KATO or CQT timeout is avoided, the storage subsystem can actively report the path fault to the host, and the fault detection time is shortened. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0013] Figure 1 The flow chart of the host active path switching method provided by the embodiments of the present application is shown in the figure.

[0014] Figure 2 The path topology diagram between the host and the storage controller provided by an embodiment of the present application is shown in the figure.

[0015] Figure 3A schematic diagram illustrating the execution logic of host active path switching, provided as an embodiment of this application;

[0016] Figure 4 This is an example diagram of a host activity path switching device according to an embodiment of this application.

[0017] Among them, 10 is the host active path switching device, 100 is the communication detection module, 200 is the query module, 300 is the reset module, and 400 is the switching module. Detailed Implementation

[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0019] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0020] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0021] The specific application environment architecture or specific hardware architecture on which the execution of the host active path switching method depends is described here.

[0022] Embodiments of this application provide a method for switching the active path of a host.

[0023] like Figure 1 The diagram shown is a flowchart of a host active path switching method according to an embodiment of this application. The host active path switching method includes the following steps:

[0024] In step S101, an active path between the target host and a first storage controller in a preset storage cluster system is determined, and the target host is controlled to communicate with the first storage controller through the active path, and it is detected whether a first communication loss occurs between the target host and the first storage controller, wherein when the first communication loss occurs, a first cross-controller reset instruction is sent to a second storage controller in the storage cluster system through the target host.

[0025] Those skilled in the art should appreciate that in the related art, the host can periodically send a Keep Alive command, and the storage controller returns a corresponding response after receiving the Keep Alive command; if the storage controller does not receive the Keep Alive command within a preset timeout period, the host considers that the communication is interrupted, so as to trigger the storage controller reset or path switching operation.

[0026] However, in the related art, the waiting time for the host to confirm the path failure and perform the recovery operation is long, which easily leads to service interruption and affects the availability of the storage system. If the KATO or CQT is set too short, the host may misjudge the path failure and perform the recovery operation when the controller is still processing I / O, thereby causing data damage. If the KATO or CQT is set too long, the fault recovery time is further prolonged, which affects the service continuity. In addition, the related art mostly relies on the host to periodically send a Keep Alive signal to detect the path state, which cannot actively notify the host when the controller detects a fault, thereby increasing the delay of fault recovery and possibly causing the host to fail to switch to a backup path in time.

[0027] Therefore, the embodiments of the present application can make the host quickly reset the faulty controller through a healthy active path through a Cross Controller Reset (CCR) mechanism, avoid long waiting time, and cause KATO / CQT timeout; in addition, the embodiments of the present application can also introduce CCR to complete Async Event notify (AEN), so that the storage subsystem can actively report the path failure to the host, thereby shortening the fault detection time.

[0028] Specifically, in actual implementation, the embodiment of the present application can first assume that the storage cluster system (i.e., a system of jointly multiple storage controllers to provide storage services externally uniformly, in the embodiment of the present application, the storage cluster system can also be equivalent to an NVM (Non-Volatile Memory) subsystem) includes a first storage controller and a second storage controller (the priority of the first storage controller is greater than the priority of the second storage controller); the embodiment of the present application can first establish an active path between the host and multiple storage controllers in the storage cluster system to determine the active path between the host and the first storage controller, so that the host communicates with the first storage controller through the active path.

[0029] Secondly, during the business communication, the embodiment of the present application can detect the first communication loss between the host and the first storage controller through the heartbeat signal detection or the command timeout (such as Keep Alive timeout) analysis strategy, and when the first communication loss occurs, the corresponding cross-controller reset instruction (i.e., the first cross-controller reset instruction) is sent from the host to the second storage controller.

[0030] For example, as shown in FIG. 1, in the embodiment of the present application, the host can be connected to the corresponding network and the storage controller through the first active path and the second active path, i.e., the host, the network and the first storage controller in the storage cluster system are connected through the first active path, and the host, the network and the second storage controller are connected through the second active path, so that when the first communication loss between the host and the first storage controller is detected, the embodiment of the present application controls the host and the second storage controller to communicate through the second active path. Figure 2

[0031] Thus, the embodiment of the present application can perform cross-controller reset operation when the first communication loss between the host and the first storage controller is detected, so as to quickly switch the communication object, reduce the business interruption, and improve the communication reliability and business continuity of the storage cluster system.

[0032] ​Optionally, in an embodiment of the present application, determining the active path between the target host and the first storage controller in the preset storage cluster system comprises: controlling the target host to send a probe instruction to a plurality of storage controllers in the storage cluster system through a preset underlying communication protocol, and after the plurality of storage controllers receive the probe instruction, generating corresponding response data and sending the response data to the target host; extracting a plurality of device information in the response data by the target host, and determining whether the plurality of storage controllers meet a preset availability requirement according to the plurality of device information, wherein in the case that the plurality of storage controllers all meet the preset availability requirement or the first storage controller in the plurality of storage controllers meets the preset availability requirement, the active path between the target host and the first storage controller is determined.

[0033] It should be noted that the embodiment of the present application can first deploy a multi-path software on the host side as a key component of the control host, and the host synchronously sends a probe instruction to a plurality of storage controllers in the storage cluster system relying on a preset underlying communication protocol, such as a fiber channel protocol.

[0034] The above-mentioned probe instruction contains core contents such as path connectivity detection, performance parameter query, and state diagnosis. After the plurality of storage controllers receive the instruction, the running state, port occupancy rate, data transmission bandwidth, link stability and other information of the plurality of storage controllers are quickly collected, corresponding response data is generated and real-time feedback is fed back to the host.

[0035] Secondly, the host can extract a plurality of device information such as device model, port state, link delay, fault tolerance capability in the response data through the multi-path software, and comprehensively determine according to a preset availability requirement (such as link connectivity, performance threshold, fault recovery capability, etc.).

[0036] When all the storage controllers meet the availability requirement, or the first storage controller (the controller with the highest priority or the best performance) meets the requirement, the multi-path software will automatically determine the active path between the host and the first storage controller, and simultaneously perform redundant backup on other paths that meet the requirement.

[0037] Therefore, the embodiment of the present application can realize intelligent optimization and redundant backup of the path, improve the stability of storage access, optimize the load distribution of the link, reduce the risk of single point failure, and guarantee the efficiency and reliability of data transmission.

[0038] Optionally, in an embodiment of the present application, the control target host is in business communication with the first storage controller through an active path, and whether a first communication loss occurs between the target host and the first storage controller is detected, including: sending a periodic interaction signal to the first storage controller through the target host, and obtaining an interaction response signal corresponding to the periodic interaction signal generated by the first storage controller, so as to obtain connection state data between the target host and the first storage controller according to the interaction response signal; obtaining a plurality of business commands sent by the target host to the first storage controller in the process of business communication, and generating command response data corresponding to the plurality of business commands; based on the connection state data and the command response data, it is judged whether the periodic interaction signal satisfies a preset continuous non-response condition, and whether the business command satisfies a preset response range out-of-limit requirement; if the periodic interaction signal satisfies the preset continuous non-response condition, and the business command satisfies the preset response range out-of-limit requirement, it is determined that the business communication between the target host and the first storage controller occurs the first communication loss, otherwise it is determined that the business communication between the target host and the first storage controller does not occur the communication loss.

[0039] As an implementable way, after the host and the first storage controller complete device discovery and establish business communication, the multi-path software on the host side can trigger a dual monitoring mechanism; on the one hand, the embodiment of the present application can send a periodic interaction signal in the form of a heartbeat to the first storage controller according to a preset period (such as 5 seconds / time), and capture the interaction response signal returned by the first storage controller in real time, so as to extract connection state data such as link delay and signal strength.

[0040] On the other hand, the embodiment of the present application can record a plurality of business commands such as read / write instructions and state queries sent by the host in the process of business communication, and synchronously generate command response data containing response time and execution result, and obtain a continuous non-response threshold (such as 3 times of continuous non-response) and a business command response range (such as a normal response time of 0.1-1 second) preset by the system; then, the multi-path software can perform correlation analysis on the connection state data and the command response data: when the periodic interaction signal continuously fails to reach the threshold, and the response time of the business command exceeds the preset range, it is determined that the business communication between the host and the first storage controller occurs the first communication loss; otherwise, it is determined that the communication is normal.

[0041] In addition, as an implementable way, when the number of continuous non-responses corresponding to the periodic interaction signal reaches the corresponding threshold, or the number of response range out-of-limit times corresponding to the business command reaches the corresponding threshold, it can also be determined that the communication between the host and the first storage controller is abnormally interrupted.

[0042] Therefore, the embodiment of the present application accurately identifies the first communication loss through dual monitoring, reduces the communication misjudgment, thereby providing reliable data guidance and technical support for subsequent path switching, and ensuring the continuity of business execution.

[0043] Optionally, in one embodiment of the present application, when the first communication loss occurs, the target host sends a first cross-controller reset instruction to the second storage controller in the storage cluster system, comprising: the target host sends an identification controller command to each of the plurality of storage controllers, and obtains the message information corresponding to the identification controller command generated by the storage controller; parsing the message information to obtain the corresponding message parsing data, and extracting a plurality of target field information in the message parsing data that meet the preset criticality requirement, and constructing the first cross-controller reset instruction according to the plurality of target field information, and sending the first cross-controller reset instruction to the second storage controller through the target host.

[0044] It should be noted that in the process of device discovery of the storage controller by the host multipath software, the identification controller command is issued to all storage controllers, and the return message of the command is parsed, and the CID (Controller Identity, controller identifier), CIU (Controller Instance Uniquifier, controller instance unique identifier) and CIRN (Controller Instance Random Number, controller instance random number) and other key field values (i.e. a plurality of target field information) are recorded.

[0045] In an embodiment of the present application, the returned plurality of key field values are as shown in Table 1:

[0046] Table 1

[0047]

[0048] Then, the embodiment of the present application can encapsulate the plurality of key field values into a cross-controller reset instruction (i.e. first cross-controller reset instruction opcode: 0x6F, wherein opcode: 0x6F represents that the instruction opcode is 0x6F) according to the preset protocol format, and clearly mark the session information, fault path identifier and switching priority to be synchronized, and send the first cross-controller reset instruction to the second storage controller to initiate a cross-controller reset or synchronization request, thereby instructing the storage controller to prepare for state synchronization or failover.

[0049] In the first cross-controller reset instruction opcode: 0x6F, the instruction word setting of command dword10 is as shown in Tables 2, 3 and 4:

[0050] Table 2

[0051]

[0052] Table 3

[0053]

[0054] Table 4

[0055]

[0056] Therefore, the embodiment of the present application can quickly trigger the cross-controller state synchronization operation, thereby laying a solid foundation for failover, shortening the path switching delay, and guaranteeing the continuity and consistency of service data.

[0057] In step S102, based on the first cross-controller reset instruction, the multiple controller field information of the first storage controller is queried to determine whether the first storage controller meets the preset cross-controller reset trigger requirement according to the multiple controller field information.

[0058] Further, after the second storage controller receives the first cross-controller reset instruction sent by the host, the embodiment of the present application can query the multiple controller field information of the first storage controller according to the first cross-controller reset instruction, and determine whether the first storage controller meets the cross-controller reset trigger requirement by using the multiple controller field information.

[0059] Therefore, the embodiment of the present application can accurately determine whether the cross-controller reset requirement is met by querying the controller instance random number and unique identifier and other field information of the first storage controller, so as to accurately trigger the reset operation, avoid misoperation, and guarantee the accuracy and rationality of the cross-controller switching.

[0060] Alternatively, in an embodiment of the present application, based on the first cross-controller reset instruction, the multiple controller field information of the first storage controller is queried to determine whether the first storage controller meets the preset cross-controller reset trigger requirement according to the multiple controller field information, including: based on the first cross-controller reset instruction, the second storage controller queries the multiple controller field information of the first storage controller, wherein the multiple controller field information includes host identity information, controller instance random number and controller instance unique identifier; the second storage controller compares the multiple target field information with the multiple controller field information to obtain a corresponding comparison result, and determines whether the multiple target field information and the multiple controller field information are the same according to the comparison result; if the multiple target field information and the multiple controller field information are the same, it is determined that the first storage controller meets the cross-controller reset trigger requirement, otherwise it is determined that the first storage controller does not meet the cross-controller reset trigger requirement.

[0061] Specifically, the embodiment of the present application can query the HNQN (Host NVMe Qualified Name), CIRN, CIU and other controller field information of the first storage controller in the same NVM subsystem as the second storage controller through the second storage controller.

[0062] It should be noted that the embodiment of the present application can confirm whether the storage controller is still the same instance as before the communication loss by comparing the CIRN values returned by the two identification controller commands, and if the CIRN does not match, it indicates that the storage controller has been reinitialized, and the cross-controller reset operation should be skipped to avoid resetting the new instance.

[0063] Secondly, in an NVMe (Non-Volatile Memory Express) storage cluster, the storage controller identifier may be reused after multiple restarts of the storage controller, but the incremental nature of CIU ensures that even if the controller identifier CID is the same, the CIU of the new and old controller instances is different.

[0064] The second storage controller judges whether the first storage controller is a completely new instance (i.e., whether the plurality of target field information and the plurality of controller field information are the same) by comparing the HNQN, CIU and CIRN, and if the match is successful (i.e., the plurality of target field information and the plurality of controller field information are the same), it indicates that the first storage controller is still the previous instance, and the cross-controller reset operation needs to be triggered for subsequent operations; if any of the values do not match, the verification fails, indicating that the first storage controller is no longer the instance before the host communication was disconnected, and its state may have been invalidated.

[0065] Thus, the embodiment of the present application can accurately identify whether the first storage controller is the original instance by comparing the HNQN, CIRN and CIU field information, thereby avoiding misoperation on the new instance, ensuring reasonable triggering of the cross-controller reset, and ensuring the stability of the storage cluster system.

[0066] In step S103, if the first storage controller meets the cross-controller reset trigger requirement, a second cross-controller reset instruction is sent from the second storage controller to the first storage controller, the first storage controller is restarted based on the second cross-controller reset instruction, and an asynchronous event notification containing the first storage controller restart information is sent from the second storage controller to the target host, otherwise a log page change asynchronous event notification is sent from the second storage controller to the target host.

[0067] Afterwards, the embodiment of the present application can control the second storage controller to send a second cross-controller reset instruction to the first storage controller to perform a restart operation, and send an asynchronous event notification containing the restart information to the target host, when the first storage controller meets the cross-controller reset trigger requirement; otherwise, the embodiment of the present application can send an asynchronous event notification of log page change to the host through the second storage controller.

[0068] Therefore, the embodiment of the present application judges whether the first storage controller meets the cross-controller reset trigger requirement, so as to timely feedback the state information of the first storage controller according to the corresponding judgment result, and guarantee the reasonable execution of the cross-controller reset operation.

[0069] Optionally, in an embodiment of the present application, if the first storage controller meets the cross-controller reset trigger requirement, a second cross-controller reset instruction is sent to the first storage controller by using the second storage controller, so as to restart the first storage controller based on the second cross-controller reset instruction, and send an asynchronous event notification containing the first storage controller restart information to the target host through the second storage controller, including: sending the second cross-controller reset instruction to the first storage controller through the second storage controller, and controlling the first storage controller to perform a restart operation according to the second cross-controller reset instruction; sending an asynchronous event request command to the target host by the target host, and subscribing to the asynchronous event notification of the corresponding storage controller based on the asynchronous event request command; judging whether the first storage controller restarts successfully after performing the restart operation, wherein, in the case that the first storage controller restarts successfully, the asynchronous event notification containing the first storage controller restart information is sent to the target host through the second storage controller.

[0070] In the actual execution process, the embodiment of the present application can send an AERC (Async Event Request Command, asynchronous event request command) to subscribe to the CCR asynchronous event notification to each storage controller in the NVMe subsystem through the self-developed multi-path plug-in of the host side; the host side submits the asynchronous event request command of NVMe to send the asynchronous notification of the specified type event to the NVM subsystem, and when the storage side occurs the internal state change caused by the event of this type, the host side is actively informed.

[0071] In the embodiment of the present application, a self-defined event subscription type can be added: the storage side occupies the event subscription type reserved in the NVMe standard protocol opcode 0x0C, and adds type 0x7F to represent that the cross-controller reset completion log page has been changed; the setting instruction word dword10 submitted by the CCR asynchronous event command opcode: 0x0C is as shown in Table 5:

[0072] Table 5

[0073]

[0074] After the host sends the asynchronous event request command subscription CCR asynchronous event notification to each storage controller, the embodiment of the application can send a CCR instruction (i.e., a second cross-controller reset instruction) to the first storage controller through the second storage controller to trigger the first storage controller to perform a restart operation.

[0075] After the first storage controller restarts successfully, the embodiment of the application can add entry information corresponding to the asynchronous event notification containing the first storage controller restart information to the preset cross-storage controller reset completion page, and send an asynchronous event notification containing the first storage controller restart information to the host through the second storage controller.

[0076] Thus, the embodiment of the application can quickly reset the faulty storage controller through the host by constructing a custom log page, an asynchronous event type, and a CCR command, and can actively report to the host that the storage controller has been restarted, greatly shortening the detection time of the storage controller fault.

[0077] In step S104, based on the asynchronous event notification, the second storage controller is controlled to query the target cross-controller reset log page to obtain corresponding log information, and whether the first storage controller meets the preset path failure requirement is determined according to the log information, wherein in the case of meeting the preset path failure requirement, the target active path of the target host is switched to the second storage controller.

[0078] After that, the embodiment of the application can control the second storage controller to query the target cross-controller reset log page according to the asynchronous event notification to obtain corresponding log information; and then, the embodiment of the application can determine whether the first storage controller meets the preset path failure requirement, and if so, the target active path of the host is switched to the second storage controller.

[0079] Thus, the embodiment of the application can accurately determine the path state, thereby timely completing the path switching, avoiding the interruption of business communication, and improving the reliability and fault tolerance capability of the storage system.

[0080] Optionally, in an embodiment of the present application, before the second storage controller queries the target cross-controller reset log page, the method further comprises: occupying a preset custom log page in a preset standard protocol by the storage controller, and determining log page basic configuration information of the preset custom log page, wherein the log page basic configuration information comprises a storage capacity and a data structure specification; determining a special identifier and a record format of the preset custom log page based on the log page basic configuration information, to determine a write rule of the preset custom log page according to the special identifier and the record format; in a case where the storage controller satisfies a preset write trigger condition, obtaining storage controller reset information, and writing the storage controller reset information into the preset custom log page based on the write rule, to generate the target cross-controller reset log page.

[0081] In the implementation process, the embodiment of the present application first occupies a reserved log page in the NVMe standard protocol by the storage controller, and defines it as a special custom log page, that is, a cross-controller reset completion log page, the special identifier of which is set to 0x7F.

[0082] Secondly, the embodiment of the present application can explicitly configure the basic configuration of the log page. Specifically, the embodiment of the present application can set the fixed storage capacity of the log page to meet the recording requirement, and formulate a data structure specification containing fields such as reset time, controller identifier, operation result, etc. Based on the above configuration, the embodiment of the present application can determine the log record format (such as field order, data type), and formulate a special write rule of “triggering write only after CCR operation is completed”.

[0083] When the storage controller completes the cross-controller reset (satisfies the preset write trigger condition), the embodiment of the present application can automatically collect the reset information of the storage controller, such as the restart time, reset result, fault reason, etc. of the first storage controller, and write it into the custom log page (that is, 0x7F log page adds a record) according to the preset write rule, to generate a complete target cross-controller reset log page.

[0084] Therefore, the embodiment of the present application can construct a special log page by means of the NVMe protocol reserved page, to standardize the reset information record, so as to provide reliable data basis for path switching, and effectively improve the efficiency of fault diagnosis and the maintainability of the system.

[0085] Optionally, in an embodiment of the present application, based on the asynchronous event notification, the second storage controller is controlled to query a target cross-controller reset log page to obtain corresponding log information, and whether the first storage controller meets the preset path failure requirement is determined according to the log information, including: the target host parses the received asynchronous event notification to obtain corresponding notification analysis data, and based on the notification analysis data, sends a target log page acquisition command to the second storage controller; the second storage controller queries the target cross-controller reset log page according to the target log page acquisition command to obtain corresponding log information, and sends the log information to the target host; the log entries in the log information are parsed to obtain corresponding log analysis data, and the valid bit information in the log analysis data is extracted, and whether the valid bit is a preset valid value is determined according to the valid bit information, wherein in the case that the valid bit is the preset valid value, it is determined that the first storage controller meets the preset path failure requirement.

[0086] In actual execution process, when the host receives the asynchronous event notification of cross-controller reset, the embodiment of the present application can parse the received asynchronous event notification to obtain corresponding notification analysis data, and based on the notification analysis data, determine that the log information of cross-controller reset has been updated; further, the embodiment of the present application can send a target log page acquisition command (i.e. get log page command) to the second storage controller to query a special CCR log page (i.e. target cross-controller reset log page), and the opcode of the target log page acquisition command is 0x02 (i.e. get log page command opcode: 0x02); wherein the key field (i.e. log information) of the 10th double word (i.e. commandDword10) of the target log page acquisition command is shown in Table 6:

[0087] Table 6

[0088]

[0089] After that, the embodiment of the present application can return a success state and a log entry by the second storage controller, to indicate that the entry is associated with the first storage controller, and its valid bit (V) is set to 0. Wherein, the meaning of V=0 is: informing the host that the storage controller instance has been invalid, the host should stop sending I / O request to it, and should regard it as reset.

[0090] It should be noted that in the embodiment of the present application, the descriptor format of the target log page acquisition command is shown in Table 7, and the host can parse the corresponding entry according to the format:

[0091] Table 7

[0092]

[0093] Secondly, in the embodiments of the present application, the ICID (ImpactedController ID, identifier of the restarted storage controller) and other fields in the cross-controller reset entry can be parsed according to Table 8:

[0094] Table 8

[0095]

[0096] Therefore, the embodiments of the present application can trigger the log query operation by parsing the asynchronous notification, and accurately determine whether the path is invalid based on the valid bit information, thereby improving the accuracy and efficiency of path switching.

[0097] Optionally, in an embodiment of the present application, when the preset path invalidation requirement is met, the target active path of the target host is switched to the second storage controller, comprising: obtaining a plurality of potential connection paths corresponding to the second storage controller, and performing real-time link detection operation on the plurality of potential connection paths to generate corresponding probe results; based on the probe results, extracting at least one target candidate path corresponding to the second storage controller from the plurality of potential connection paths, and performing quantitative rating on the at least one target candidate path to obtain corresponding rating results, and filtering out the target active path with the highest rating from the at least one target candidate path according to the rating results; suspending the uncompleted service request on the active path corresponding to the first storage controller, marking the request state of the uncompleted service request, and migrating the service request after marking the request state to the target active path, and associating the second storage controller with the target active path, so as to switch the target active path of the target host to the second storage controller.

[0098] It should be noted that after the host confirms that the first storage controller path is invalid, the embodiments of the present application can start the switching process through the multi-path software on the host side. Specifically, the embodiments of the present application can first obtain the potential connection paths (such as different ports, link combinations) of the second storage controller, generate probe results by real-time detection of link bandwidth, delay and packet loss rate, and extract available target candidate paths therefrom.

[0099] Secondly, the embodiments of the present application can perform quantitative rating based on path stability, transmission rate and load condition, and filter out the path with the highest rating as the target active path; subsequently, the embodiments of the present application can suspend the uncompleted service request of the original path, mark its state as "to be retried", and migrate the request to the new path, and associate the second storage controller with the new path, so as to complete the active path switching.

[0100] Therefore, the embodiments of the present application realize fast failover, guarantee the continuity of services, and improve the reliability of storage access through intelligent routing and request migration operations.

[0101] In addition, as an implementable manner, after the active path is switched to the second storage controller, the embodiments of the present application can start periodic automatic detection of the path between the first storage controller and the host, and the periodic automatic detection includes link connectivity detection and data transmission stability detection; a path fault type identifier is generated based on the detection result, and a corresponding maintenance instruction is triggered according to the fault type identifier, and the maintenance instruction includes a link reset instruction and a port restart instruction; during the maintenance process, the path maintenance state is monitored in real time, and when the path maintenance state meets a preset repair completion condition, the repaired path is verified for availability, and the availability verification includes response speed testing and data consistency checking; if the verification is passed, the path is marked as a standby path, and when a preset path switching condition is met, the active path is switched back to the first storage controller from the second storage controller.

[0102] Specifically, after the active path is switched to the second storage controller, the specific process of starting the intelligent repair process of the path between the first storage controller and the host by the embodiments of the present application is as follows:

[0103] 1. A multi-level periodic detection mechanism is constructed to perform link connectivity pulse detection and data packet round-trip delay monitoring within a basic period (such as 10 seconds / time), and to carry out large-flow data transmission stress testing within an extended period (such as 5 minutes / time), and to synchronously collect link error rate, port temperature and other environmental parameters, thereby forming a multi-dimensional detection matrix;

[0104] 2. A fault type identification model is constructed based on the multi-dimensional detection matrix, and a fault type identifier containing fault levels (fatal / non-fatal) and influence ranges is automatically generated by comparing a preset fault feature library (such as 3 consecutive connectivity interruptions corresponding to loose links, transmission delay increasing by more than 3 times corresponding to bandwidth congestion); and a link reset instruction is triggered for fatal faults, and a hardware management module is restarted for physical ports in linkage, and a protocol layer session reconstruction instruction is executed for non-fatal faults;

[0105] 3. A repair progress dynamic tracking mechanism is introduced, and information such as link reset progress and port restart stage is obtained in real time by embedding a state feedback field in the repair instruction; when it is detected that the link is re-established and there is no abnormality for 5 consecutive basic periods, it is determined that the repair completion condition is met, and the availability verification process is automatically started, that is, the response speed stability is verified through 1000 times of cyclic read-write test of the preset verification data packet, the consistency is ensured by comparing the data check values before and after the repair, and the performance fluctuation coefficient in the verification process is recorded synchronously;

[0106] 4、After verification, the embodiment of the application can mark the path as "hot standby" state and include it in the path resource pool dynamic management, and the priority can be dynamically adjusted according to real-time performance monitoring data (such as average response time, bandwidth utilization); when it is monitored that the load rate of the second storage controller exceeds the threshold (such as 80%) or the standby path performance advantage lasts for 3 extension periods, the smooth switching mechanism is triggered, that is, the new service request is first shunted to the standby path, and after the original path is completed, the active path is switched back to the first storage controller, and the second storage controller is used as the standby path.

[0107] Therefore, the embodiment of the application can realize automatic maintenance and verification of the path, and flexibly switch back to the original path after the path is repaired, thereby improving the path resource utilization and system redundancy capability, and guaranteeing the continuity of storage access.

[0108] The execution logic of the host active path switching method of the application is described below by combining the accompanying drawings.

[0109] Figure 3 The execution logic of the host active path switching method of the application is described below by combining the accompanying drawings. Figure 3 As shown in the figure, the execution process of the host active path switching method of the application is as follows:

[0110] S301: subscribing to an asynchronous event of cross-controller reset completion log page update from the host to the first storage controller;

[0111] S302: subscribing to an asynchronous event of cross-controller reset completion log page update from the host to the second storage controller;

[0112] S303: sending an identification controller command from the host to the first storage controller to query the controller identifier, the controller instance unique identifier and the controller instance random number of the first storage controller;

[0113] S304: sending an identification controller command from the host to the second storage controller to query the controller identifier, the controller instance unique identifier and the controller instance random number of the second storage controller;

[0114] S305: selecting the path of the first storage controller as the corresponding active path by the host, and issuing service data;

[0115] S306: the host detects the loss of communication with the first storage controller for the first time through a heartbeat signal detection or a command timeout analysis strategy;

[0116] S307: issuing a cross-controller reset instruction to the second storage controller by the host;

[0117] S308: The controller identifier, controller instance unique identifier, and controller instance random number sent by the host to reset the target storage controller are parsed through the second storage controller;

[0118] S309: Query the controller identifier, controller instance unique identifier, and controller instance random number of the first storage controller in the same storage cluster system through the second storage controller;

[0119] S3010: Match the corresponding controller identifier, controller instance unique identifier, and controller instance random number through the second storage controller;

[0120] S3011: Determine if the match is successful. If the match is successful, proceed to S3012; otherwise, proceed to S3013.

[0121] S3012: Trigger the reset operation of the first storage controller through the second storage controller, add a first storage controller cross-controller reset completion entry to the cross-controller reset completion log page, and proceed to S3013;

[0122] S3013: Asynchronous event notification of log page changes completed by returning a host cross-controller reset via the second storage controller;

[0123] S3014: Send a log page retrieval command to the second storage controller via the host to query a dedicated cross-controller log page;

[0124] S3015: After the host learns that the first storage controller has been restarted, it confirms that the original first storage controller instance has failed and switches the corresponding active path to the second storage controller.

[0125] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0126] Embodiments of this application also provide a host active path switching device.

[0127] like Figure 4 As shown, the host active path switching device 10 includes: a communication detection module 100, a query module 200, a reset module 300, and a switching module 400.

[0128] The communication detection module 100 is configured to determine an active path between the target host and a first storage controller in the preset storage cluster system, control the target host to communicate with the first storage controller through the active path, and detect whether a first communication loss occurs between the target host and the first storage controller. When the first communication loss occurs, the target host sends a first cross-controller reset instruction to a second storage controller in the storage cluster system.

[0129] The query module 200 is configured to query a plurality of controller field information of the first storage controller based on the first cross-controller reset instruction, and determine whether the first storage controller meets a preset cross-controller reset trigger requirement according to the plurality of controller field information.

[0130] The reset module 300 is configured to send a second cross-controller reset instruction to the first storage controller by using the second storage controller if the first storage controller meets the cross-controller reset trigger requirement, restart the first storage controller based on the second cross-controller reset instruction, and send an asynchronous event notification containing first storage controller restart information to the target host by using the second storage controller, or send an asynchronous event notification of a log page change to the target host by using the second storage controller if the first storage controller does not meet the cross-controller reset trigger requirement.

[0131] The switching module 400 is configured to query a target cross-controller reset log page by using the second storage controller based on the asynchronous event notification, obtain corresponding log information, and determine whether the first storage controller meets a preset path failure requirement according to the log information. If the first storage controller meets the preset path failure requirement, the target active path of the target host is switched to the second storage controller.

[0132] Optionally, in an embodiment of the present application, the communication detection module 100 includes a first control unit and a first determination unit.

[0133] The first control unit is configured to control the target host to send a detection instruction to a plurality of storage controllers in the storage cluster system through a preset underlying communication protocol, and generate corresponding response data after the plurality of storage controllers receive the detection instruction, and send the response data to the target host.

[0134] The first determination unit is configured to extract a plurality of device information in the response data by using the target host, and determine whether the plurality of storage controllers meet a preset availability requirement according to the plurality of device information. If the plurality of storage controllers meet the preset availability requirement or the first storage controller in the plurality of storage controllers meets the preset availability requirement, the active path between the target host and the first storage controller is determined.

[0135] Optionally, in an embodiment of the present application, the communication detection module 100 further comprises an interaction unit, a generation unit, a second judgment unit and a first determination unit.

[0136] The interaction unit is configured to send a periodic interaction signal to the first storage controller through the target host, and obtain an interaction response signal corresponding to the periodic interaction signal generated by the first storage controller, so as to obtain connection state data between the target host and the first storage controller according to the interaction response signal.

[0137] The generation unit is configured to obtain a plurality of service commands sent by the target host to the first storage controller in a service communication process, and generate command response data corresponding to the plurality of service commands.

[0138] The second judgment unit is configured to judge whether the periodic interaction signal meets a preset continuous non-response condition and whether the service command meets a preset response range overrun requirement based on the connection state data and the command response data.

[0139] The first determination unit is configured to determine that the service communication between the target host and the first storage controller has a first communication loss if the periodic interaction signal meets the preset continuous non-response condition and the service command meets the preset response range overrun requirement, and otherwise determine that the service communication between the target host and the first storage controller has no communication loss.

[0140] Optionally, in an embodiment of the present application, the communication detection module 100 further comprises an identification unit and a first analysis unit.

[0141] The identification unit is configured to send an identification controller command to a plurality of storage controllers through the target host respectively, and obtain packet information corresponding to the identification controller command generated by the storage controller.

[0142] The first analysis unit is configured to analyze the packet information to obtain corresponding packet analysis data, extract a plurality of target field information meeting a preset key requirement in the packet analysis data, construct a first cross-controller reset instruction according to the plurality of target field information, and send the first cross-controller reset instruction to the second storage controller through the target host.

[0143] Optionally, in an embodiment of the present application, the query module 200 comprises a second control unit, a comparison unit and a second determination unit.

[0144] The second control unit is configured to control the second storage controller to query a plurality of controller field information of the first storage controller based on the first cross-controller reset instruction, wherein the plurality of controller field information comprises host identity information, controller instance random number and controller instance unique identifier.

[0145] The comparison unit is configured to compare the plurality of target field information with the plurality of controller field information through the second storage controller to obtain a corresponding comparison result, and determine whether the plurality of target field information and the plurality of controller field information are the same according to the comparison result.

[0146] The second determination unit is configured to determine that the first storage controller meets the cross-controller reset trigger requirement if the plurality of target field information and the plurality of controller field information are the same, and determine that the first storage controller does not meet the cross-controller reset trigger requirement otherwise.

[0147] Optionally, in an embodiment of the present application, the reset module 300 comprises a first sending unit, a subscription unit and a third determination unit.

[0148] The first sending unit is configured to send a second cross-controller reset instruction to the first storage controller through the second storage controller, and control the first storage controller to perform a restart operation according to the second cross-controller reset instruction.

[0149] The subscription unit is configured to send an asynchronous event request command to the plurality of storage controllers through the target host respectively, and subscribe to corresponding asynchronous event notifications of the plurality of storage controllers based on the asynchronous event request command.

[0150] The third determination unit is configured to determine whether the first storage controller restarts successfully after performing the restart operation, wherein in the case that the first storage controller restarts successfully, the second storage controller sends an asynchronous event notification containing first storage controller restart information to the target host.

[0151] Optionally, in an embodiment of the present application, the switching module 400 comprises a second analysis unit, a second sending unit and an extraction unit.

[0152] The second analysis unit is configured to analyze the received asynchronous event notification through the target host to obtain corresponding notification analysis data, and send a target log page acquisition command to the second storage controller based on the notification analysis data.

[0153] The second sending unit is configured to query a target cross-controller reset log page according to the target log page acquisition command by using the second storage controller to obtain corresponding log information, and send the log information to the target host.

[0154] The extraction unit is configured to analyze log entries in the log information to obtain corresponding log analysis data, extract valid bit information in the log analysis data, and determine whether the valid bit is a preset valid value according to the valid bit information, wherein in the case that the valid bit is the preset valid value, it is determined that the first storage controller meets a preset path invalidation requirement.

[0155] Optionally, in an embodiment of the present application, the switching module 400 further comprises a link detection unit, a rating unit and a marking unit.

[0156] The link detection unit is configured to acquire a plurality of potential connection paths corresponding to the second storage controller, and perform real-time link detection on the plurality of potential connection paths to generate a corresponding probe result.

[0157] The rating unit is configured to extract at least one target candidate path corresponding to the second storage controller from the plurality of potential connection paths based on the probe result, and perform quantitative rating on the at least one target candidate path to obtain a corresponding rating result, and filter out a target active path with the highest rating from the at least one target candidate path according to the rating result.

[0158] The marking unit is configured to suspend an unfinished service request on an active path corresponding to the first storage controller, mark a request state of the unfinished service request, migrate the service request after marking the request state to the target active path, and associate the second storage controller with the target active path, so as to switch the target active path of the target host to the second storage controller.

[0159] Optionally, in an embodiment of the present application, the host active path switching device 10 further comprises a first determination module, a second determination module and a writing module.

[0160] The first determination module is configured to occupy a preset custom log page in a preset standard protocol by the storage controller before controlling the second storage controller to query the target cross-controller reset log page, and determine log page basic configuration information of the preset custom log page, wherein the log page basic configuration information includes storage capacity and data structure specification.

[0161] The second determination module is configured to determine an exclusive identifier and a record format of the preset custom log page based on the log page basic configuration information, so as to determine a writing rule of the preset custom log page according to the exclusive identifier and the record format.

[0162] The writing module is configured to acquire storage controller reset information in the case that the storage controller meets a preset writing trigger condition, and write the storage controller reset information into the preset custom log page based on the writing rule, so as to generate the target cross-controller reset log page.

[0163] The features of the embodiments of the host active path switching device can be referred to the related descriptions of the embodiments of the host active path switching method, which will not be repeated here.

[0164] The embodiment of the present application further provides an electronic device, comprising a memory and a processor, the memory stores a computer program, and the processor is arranged to run the computer program to execute the steps in any of the above host active path switching method embodiments.

[0165] The embodiment of the present application further provides a non-volatile computer readable storage medium, which stores a computer program, and the computer program is arranged to execute the steps in any of the above host active path switching method embodiments when running.

[0166] In an example embodiment, the above non-volatile computer readable storage medium can include but is not limited to: a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0167] The embodiment of the present application further provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps in any of the above host active path switching method embodiments.

[0168] The embodiment of the present application further provides another computer program product, which comprises a non-volatile computer readable storage medium, and the non-volatile computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps in any of the above host active path switching method embodiments.

[0169] The skilled person can further realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in general terms in the above description. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0170] The above describes in detail a host active path switching method, device, equipment and medium provided by the present application. The principles and implementation modes of the present application are described by applying specific examples, and the above description of the embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that, for those skilled in the art, without departing from the principles of the present application, the present application can be improved and modified in several ways, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A host active path switching method, characterized by, The method comprises the following steps: determining an active path between a target host and a first storage controller in a preset storage cluster system, controlling the target host to communicate with the first storage controller through the active path, and detecting whether a first communication loss occurs between the target host and the first storage controller, wherein, when the first communication loss occurs, a first cross-controller reset instruction is sent from the target host to a second storage controller in the storage cluster system; based on the first cross-controller reset instruction, querying a plurality of controller field information of the first storage controller to determine whether the first storage controller meets a preset cross-controller reset trigger requirement according to the plurality of controller field information; if the first storage controller meets the cross-controller reset trigger requirement, sending a second cross-controller reset instruction from the second storage controller to the first storage controller to restart the first storage controller based on the second cross-controller reset instruction, and sending an asynchronous event notification containing the first storage controller restart information from the second storage controller to the target host, otherwise sending an asynchronous event notification of a log page change from the second storage controller to the target host; based on the asynchronous event notification, controlling the second storage controller to query a target cross-controller reset log page to obtain corresponding log information, and determining whether the first storage controller meets a preset path failure requirement according to the log information, wherein, when the preset path failure requirement is met, switching a target active path of the target host to the second storage controller.

2. The host active path switching method of claim 1, wherein, The method comprises the following steps: controlling the target host to send a probe instruction to a plurality of storage controllers in the storage cluster system through a preset underlying communication protocol, and generating corresponding response data after the plurality of storage controllers receive the probe instruction, and sending the response data to the target host; extracting a plurality of device information in the response data through the target host, and determining whether the plurality of storage controllers meet a preset availability requirement according to the plurality of device information, wherein, when the plurality of storage controllers all meet the preset availability requirement or the first storage controller in the plurality of storage controllers meets the preset availability requirement, an active path between the target host and the first storage controller is determined.

3. The host active path switching method of claim 1, wherein, The method comprises the following steps: sending a periodic interaction signal from the target host to the first storage controller, and obtaining an interaction response signal corresponding to the periodic interaction signal generated by the first storage controller to obtain connection state data between the target host and the first storage controller according to the interaction response signal; acquire a plurality of service commands sent by the target host to the first storage controller in a service communication process, and generate command response data corresponding to the plurality of service commands; determine whether the periodic interaction signal meets a preset continuous non-response condition and whether the service command meets a preset response range overrun requirement based on the connection state data and the command response data; if the periodic interaction signal meets the preset continuous non-response condition and the service command meets the preset response range overrun requirement, determine that the service communication between the target host and the first storage controller has a first communication loss, otherwise, determine that the service communication between the target host and the first storage controller has no communication loss.

4. The host active path switching method of claim 1, wherein, when the first communication loss occurs, the target host sends a first cross-controller reset instruction to a second storage controller in the storage cluster system, comprising: The target host sends an identification controller command to a plurality of storage controllers respectively, and acquires message information corresponding to the identification controller command generated by the storage controller; parse the message information to obtain corresponding message parsing data, extract a plurality of target field information meeting a preset key requirement in the message parsing data, and construct the first cross-controller reset instruction according to the plurality of target field information, and send the first cross-controller reset instruction to the second storage controller through the target host.

5. The host active path switching method of claim 4, wherein, based on the first cross-controller reset instruction, query a plurality of controller field information of the first storage controller to determine whether the first storage controller meets a preset cross-controller reset trigger requirement according to the plurality of controller field information, comprising: based on the first cross-controller reset instruction, control the second storage controller to query a plurality of controller field information of the first storage controller, wherein the plurality of controller field information includes host identity information, controller instance random number and controller instance unique identifier; compare the plurality of target field information with the plurality of controller field information through the second storage controller to obtain a corresponding comparison result, and determine whether the plurality of target field information and the plurality of controller field information are the same according to the comparison result; if the plurality of target field information and the plurality of controller field information are the same, it is determined that the first storage controller meets the cross-controller reset trigger requirement, otherwise it is determined that the first storage controller does not meet the cross-controller reset trigger requirement.

6. The host active path switching method of claim 5, wherein, if the first storage controller meets the cross-controller reset trigger requirement, send a second cross-controller reset instruction to the first storage controller using the second storage controller, restart the first storage controller based on the second cross-controller reset instruction, and send an asynchronous event notification containing the first storage controller restart information to the target host through the second storage controller, comprising: sending the second cross-controller reset instruction to the first storage controller through the second storage controller, and controlling the first storage controller to perform a restart operation according to the second cross-controller reset instruction; sending asynchronous event request commands to the plurality of storage controllers through the target host respectively, and subscribing to asynchronous event notifications of the plurality of storage controllers based on the asynchronous event request commands; determining whether the first storage controller restarts successfully after performing the restart operation, wherein, in the case that the first storage controller restarts successfully, sending an asynchronous event notification containing first storage controller restart information to the target host through the second storage controller.

7. The host active path switching method of claim 6, wherein, controlling the second storage controller to query a target cross-controller reset log page to obtain corresponding log information based on the asynchronous event notification, and determining whether the first storage controller meets a preset path failure requirement according to the log information, including: parsing the received asynchronous event notification through the target host to obtain corresponding notification parsing data, and sending a target log page acquisition command to the second storage controller based on the notification parsing data; querying the target cross-controller reset log page through the second storage controller according to the target log page acquisition command to obtain corresponding log information, and sending the log information to the target host; parsing log entries in the log information to obtain corresponding log parsing data, and extracting valid bit information in the log parsing data, and determining whether the valid bit is a preset valid value according to the valid bit information, wherein, in the case that the valid bit is the preset valid value, it is determined that the first storage controller meets the preset path failure requirement.

8. The host active path switching method of claim 7, wherein, in the case that the preset path failure requirement is met, switching a target active path of the target host to the second storage controller, including: obtaining a plurality of potential connection paths corresponding to the second storage controller, and performing real-time link detection operations on the plurality of potential connection paths to generate corresponding probe results; based on the probe results, extracting at least one target candidate path corresponding to the second storage controller from the plurality of potential connection paths, and performing quantitative rating on the at least one target candidate path to obtain corresponding rating results, and selecting a target active path with the highest rating from the at least one target candidate path according to the rating results; suspending uncompleted service requests on an active path corresponding to the first storage controller, marking the request state of the uncompleted service requests, migrating the service requests with the marked request state to the target active path, and associating the second storage controller with the target active path to switch the target active path of the target host to the second storage controller.

9. The host active path switching method of claim 1, wherein, before controlling the second storage controller to query the target cross-controller reset log page, further including: The preset custom log page in the preset standard protocol is occupied by a storage controller, and log page basic configuration information of the preset custom log page is determined, wherein the log page basic configuration information comprises a storage capacity and a data structure specification; Based on the log page basic configuration information, a unique identifier and a record format of the preset custom log page are determined, so as to determine a write rule of the preset custom log page according to the unique identifier and the record format; In a case where the storage controller meets a preset write trigger condition, storage controller reset information is obtained, and the storage controller reset information is written into the preset custom log page based on the write rule, so as to generate the target cross-controller reset log page.

10. An electronic device, comprising: Comprise: a memory for storing a computer program; a processor for executing the computer program to implement the steps of the host active path switching method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Fault switching method, device and equipment for multiple storage controllers and storage medium

    CN116340040A

  • Server cooperative control method, storage medium and electronic equipment

    CN119917350A