Multipath fault recovery method and device, electronic device and storage medium

By performing multi-level verification and repair operations on multi-path fault paths of a computer storage system, the problem of data inconsistency after multi-path fault recovery is solved, and the stability and reliability of the storage system are improved.

CN120407263BActive Publication Date: 2025-09-12INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510889393.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-09-12
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

In the existing computer storage system, data consistency and service reliability cannot be guaranteed after multi-path failure recovery, especially in large-scale enterprise storage environments, where it is difficult to solve problems such as drive letter drift.

Method used

By obtaining the status data of multiple faulty paths in the target storage system, multi-level verification processing is performed, including device attribute verification, shared consistency verification, and scheduling metadata verification. Path repair operations are performed on paths that fail verification, and the repaired paths are registered in the active path pool.

Benefits of technology

It improves the accuracy and reliability of multi-path fault recovery, reduces the risk of data inconsistency and system failure, reduces system maintenance costs and downtime, and ensures the stability and efficiency of the storage system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407263B_ABST
    Figure CN120407263B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and apparatus for multi-path fault recovery, an electronic device, and a storage medium. The method obtains status data of storage devices associated with multiple fault paths in a target storage system, and detects whether the multiple fault paths have recovered to a normal state based on the status data. When it is detected that the path status of the target fault path has recovered, a multi-level verification process is performed on the target fault path. If the target fault path fails the multi-level verification, a path repair operation is performed on the target fault path that failed the multi-level verification based on the verification type that failed. The repaired target fault path and the target fault path that passed the multi-level verification are registered in an active path pool. Compared with related technologies, when it is detected that the path status of the target fault path has recovered, a multi-level verification process is performed on the target fault path, which can comprehensively and carefully check whether there are potential problems in the recovered path.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of storage technology, and in particular to a method and device for multi-path fault recovery, an electronic device, and a storage medium. Background Art

[0002] In computer storage systems, multipathing software is a core component for ensuring high storage availability. The reliability of its path failure recovery mechanism directly impacts business continuity. Currently, an event-driven mechanism is commonly used for path status management: when a storage controller link recovers from a fault, the multipathing service directly activates the path and adds it to the pool of available paths. While these solutions can implement basic path switching, they struggle to ensure data consistency and service reliability in large-scale enterprise storage environments due to issues such as drive letter drift during the recovery phase. Summary of the Invention

[0003] The present disclosure provides a multi-path fault recovery method and device, electronic device, and storage medium, which are mainly intended to solve the problem that data consistency cannot be guaranteed after a multi-path fault in a computer storage system.

[0004] According to a first aspect of the present disclosure, a method for multipath failure recovery is provided, comprising:

[0005] Acquire status data of storage devices associated with multiple faulty paths in the target storage system, and detect whether the multiple faulty paths have recovered to a normal state based on the status data;

[0006] When the path status of the target fault path is detected to be restored, a multi-level verification process is performed on the target fault path. The verification types of the multi-level verification process include: device attribute verification, shared consistency verification, and scheduling metadata verification;

[0007] If the target fault path fails the multi-level verification, a path repair operation is performed on the target fault path that fails the multi-level verification based on the verification type that fails;

[0008] Register the repaired target fault path and the target fault path that passes multi-level verification into the active path pool.

[0009] Optionally, before performing the multi-level verification process on the target fault path, the multi-path fault recovery method further includes:

[0010] Extract the historical path status of the target fault path from the status data and determine whether the current path status of the target fault path is consistent with the historical path status;

[0011] When the current path status is inconsistent with the historical path status, determine whether the target fault path is in an active state based on the current path status;

[0012] If it is in the active state, device attribute verification and shared consistency verification are performed on the target fault path.

[0013] Optionally, perform device attribute verification and shared consistency verification on the target fault path, including:

[0014] Perform device capacity verification on the storage device capacity associated with the target failure path and perform trustworthy verification on the device identifier;

[0015] Check whether the registration key value of the target fault path recorded on the storage controller side matches the stored target key value.

[0016] Optionally, perform device capacity verification on the storage device associated with the target failure path and perform trustworthy verification on the device identifier, including:

[0017] Verify whether the device capacity data is empty or inconsistent with the device capacity of the target storage system to perform device capacity verification;

[0018] If the device capacity verification passes, the device identifier is authenticated.

[0019] Optionally, perform trustworthy verification on the device identifier, including:

[0020] Verify whether there is any anomaly in the device identifier. After confirming that there is no anomaly, verify whether the device identifier is consistent with the cached device identifier;

[0021] If they are inconsistent, it is determined that the device identifier trust verification has failed; if they are consistent, it is determined that the device identifier trust verification has passed.

[0022] Optionally, when the device capacity data verification fails and / or the device identifier verification fails, a path repair operation is performed on the faulty path that fails the multi-level verification based on the failed verification type, including:

[0023] Remove the device binding relationship of the target fault path and set it to the failed state. Scan and add the target fault path in the failed state.

[0024] Optionally, after the device capacity check passes and the device identifier verification passes, the multipath failure recovery method further includes:

[0025] Determine whether the historical path status of the target fault path is in an inactive state. If it is in an inactive state, increase the number of active paths by one.

[0026] Optionally, checking whether the registration key value of the target fault path recorded on the storage controller matches the stored target key value includes:

[0027] Verify whether the registration key value is empty;

[0028] If the registration key value is not empty, verify whether the registration key value is consistent with the target key value.

[0029] Optionally, based on the failed verification type, performing a path repair operation on the target fault path that failed the multi-level verification further includes:

[0030] When it is determined that the registration key value is consistent with the target key value, the target fault path is re-registered based on the target key value.

[0031] Optionally, the multipath failure recovery method further includes:

[0032] After performing a shared consistency check on the target failure path, a scheduling metadata check is performed on the target failure path.

[0033] Optionally, the multipath failure recovery method further includes:

[0034] If the current path status is consistent with the historical path status, the scheduling metadata verification is performed on the target fault path.

[0035] Optionally, perform scheduling metadata verification on the target failure path, including:

[0036] Determine the ownership type of the volume associated with the target failure path;

[0037] Based on the attribution type, scheduling metadata verification and maintenance are performed on the target fault path.

[0038] Optionally, based on the attribute type, scheduling metadata verification and maintenance are performed on the target fault path, including:

[0039] If the volume is unowned, the block priority and port group information of the target failure path are compared with the kernel-state saved values; any inconsistent block priority and / or port group information is updated according to the kernel-state saved values;

[0040] If the volume has an ownership, the block priority of the target failure path is verified with the value saved in the kernel state; the block priority that is inconsistent is updated according to the value saved in the kernel state.

[0041] Optionally, before performing scheduling metadata verification on the target failed path, the multipath failure recovery method further includes:

[0042] Determine whether the current path state of the target fault path is active;

[0043] When the state is determined to be active, the detection interval of the target fault path is incremented by one;

[0044] It is determined whether the target fault path is currently in the early stage of path recovery. If it is in the early stage of path recovery, scheduling metadata verification is performed on the target fault path.

[0045] Optionally, determining whether the target fault path is currently in the early stage of path recovery includes:

[0046] Compare the detection time of the target fault path with the initial recovery threshold, which is the product of the detection interval and the proportional coefficient.

[0047] If the detection duration does not reach the initial recovery threshold, the path is determined to be in the initial recovery stage.

[0048] When the detection duration reaches the initial recovery threshold, it is determined that the path is not in the initial recovery stage.

[0049] Optionally, after determining whether the current path state of the target faulty path is consistent with the historical path state, the multipath fault recovery method further includes:

[0050] Set the detection interval of the target fault path to the preset detection interval.

[0051] According to a second aspect of the present disclosure, a multipath fault recovery apparatus is provided, comprising:

[0052] a detection unit, configured to obtain status data of storage devices associated with multiple faulty paths in a target storage system, and detect whether the multiple faulty paths have recovered to a normal state based on the status data;

[0053] A verification unit, configured to perform multi-level verification processing on the target fault path when detecting that the path state of the target fault path has recovered, wherein the verification types of the multi-level verification processing include: device attribute verification, shared consistency verification, and scheduling metadata verification;

[0054] a repair unit, configured to, if the target faulty path fails the multi-level verification, perform a path repair operation on the target faulty path that fails the multi-level verification based on a verification type that fails;

[0055] The registration unit is used to register the repaired target fault path and the target fault path that has passed the multi-level verification into the active path pool.

[0056] Optionally, the multipath failure recovery device further includes:

[0057] a first judgment unit configured to extract a historical path state of the target faulty path from the state data before performing multi-stage verification processing on the target faulty path, and to judge whether a current path state of the target faulty path is consistent with the historical path state; and if the current path state is inconsistent with the historical path state, to judge whether the target faulty path is in an active state based on the current path state;

[0058] The first execution unit is configured to perform a device attribute check and a shared consistency check on the target fault path if the state is activated.

[0059] Optionally, the first execution unit includes:

[0060] A first verification module is configured to perform device capacity verification on the capacity of the storage device associated with the target fault path and to perform credibility verification on the device identifier;

[0061] The second verification module is used to detect whether the registration key value of the target fault path recorded on the storage controller side matches the stored target key value.

[0062] Optionally, the first verification module is further configured to:

[0063] Verify whether the device capacity data is empty or inconsistent with the device capacity of the target storage system to perform device capacity verification;

[0064] If the device capacity verification passes, the device identifier is authenticated.

[0065] Optionally, the second verification module is further configured to:

[0066] Verify whether there is any anomaly in the device identifier. After confirming that there is no anomaly, verify whether the device identifier is consistent with the cached device identifier;

[0067] If they are inconsistent, it is determined that the device identifier trust verification has failed; if they are consistent, it is determined that the device identifier trust verification has passed.

[0068] Optionally, the repair unit includes:

[0069] Add a module for releasing the device binding relationship of the target fault path and setting it to a failed state when the device capacity data verification fails and / or the device identifier verification fails, and scanning and adding the target fault path in the failed state.

[0070] Optionally, the multipath failure recovery device further includes:

[0071] The second judgment unit is configured to judge whether the historical path status of the target fault path is inactive after the device capacity check and the device identifier verification are passed, and if so, increase the number of active paths by one.

[0072] Optionally, the second verification module is further configured to:

[0073] Verify whether the registration key value is empty;

[0074] If the registration key value is not empty, verify whether the registration key value is consistent with the target key value.

[0075] Optionally, the repair unit also includes:

[0076] The registration module is used to re-register the target fault path based on the target key value when it is determined that the registration key value is consistent with the target key value.

[0077] Optionally, the multipath failure recovery device further includes:

[0078] The second execution unit is configured to perform a scheduling metadata check on the target fault path after performing a shared consistency check on the target fault path.

[0079] Optionally, the second execution unit is further configured to perform scheduling metadata verification on the target fault path if the current path state is consistent with the historical path state.

[0080] Optionally, the second execution unit includes:

[0081] A judgment module, used to judge the type of the volume associated with the target failure path;

[0082] The maintenance module is used to perform scheduling metadata verification and maintenance on the target fault path based on the attribution type.

[0083] Optionally, the maintenance module is also used to:

[0084] If the volume is unowned, the block priority and port group information of the target failure path are compared with the kernel-state saved values; any inconsistent block priority and / or port group information is updated according to the kernel-state saved values;

[0085] If the volume has an ownership, the block priority of the target failure path is verified with the value saved in the kernel state; the block priority that is inconsistent is updated according to the value saved in the kernel state.

[0086] Optionally, the multipath failure recovery device further includes:

[0087] a third judgment unit, configured to judge whether the current path state of the target fault path is an active state before performing scheduling metadata verification on the target fault path; and if the current path state is determined to be an active state, incrementing the detection interval of the target fault path by one;

[0088] The fourth judgment unit is configured to judge whether the target fault path is currently in an early stage of path recovery. If the target fault path is in the early stage of path recovery, a scheduling metadata check is performed on the target fault path.

[0089] Optionally, the fourth judgment unit includes:

[0090] A comparison module is used to compare the detection time of the target fault path with the initial recovery threshold, where the initial recovery threshold is the product of the detection interval and the proportional coefficient;

[0091] A first determination module is configured to determine that the path is in the initial stage of path recovery when the detection duration does not reach the initial stage of recovery threshold;

[0092] The second determination module is configured to determine that the path is not in the initial stage of path recovery when the detection duration reaches an initial stage of recovery threshold.

[0093] Optionally, the multipath failure recovery device further includes:

[0094] The setting unit is configured to set the detection interval of the target fault path to a preset detection interval after determining whether the current path state of the target fault path is consistent with the historical path state.

[0095] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0096] at least one processor; and

[0097] a memory communicatively connected to the at least one processor; wherein,

[0098] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the multipath failure recovery method described in the first aspect.

[0099] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the multipath failure recovery method described in the first aspect.

[0100] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the multipath failure recovery method as described in the first aspect.

[0101] The present disclosure provides a multipath fault recovery method and apparatus, an electronic device, and a storage medium, relating to the field of storage technology. By detecting multiple paths, the present disclosure enables real-time monitoring and detection of the status of faulty paths, thereby enabling timely response when the paths are restored. Upon detecting that the path status of a target faulty path has recovered, a multi-level verification process is performed on the target faulty path, including device attribute verification, shared consistency verification, and scheduling metadata verification. This multi-level verification process comprehensively and meticulously checks the restored path for potential issues. Device attribute verification ensures that the basic attributes and configurations of path devices meet requirements; shared consistency verification avoids issues with shared access by multiple hosts; and scheduling metadata verification ensures the correctness of path load balancing and scheduling policies. This multi-level verification mechanism improves the accuracy and reliability of path recovery and reduces the risk of data inconsistency or system failures caused by improper path recovery. Path repair operations are performed on target faulty paths that fail the multi-level verification process. This allows the system to promptly implement repair measures when potential issues are detected after path recovery, preventing further failures or data inconsistencies caused by these issues. The path repair operation can provide targeted processing for different types of verification failures, improving repair efficiency and accuracy and reducing system maintenance costs and downtime. Through real-time monitoring and detection of faulty path status, multi-level verification processing, path repair operations, and registration with the active path pool, the accuracy, reliability, and availability of multipath-driven path recovery methods can be improved, providing a more stable and efficient path recovery solution for storage systems.

[0102] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0103] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0104] Figure 1 A flowchart of a multipath fault recovery method provided by an embodiment of the present disclosure;

[0105] Figure 2 A flowchart of another multipath fault recovery method provided by an embodiment of the present disclosure;

[0106] Figure 3 A schematic structural diagram of a multipath fault recovery device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0107] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0108] The following describes a method and apparatus for multipath failure recovery, an electronic device, and a storage medium according to embodiments of the present disclosure with reference to the accompanying drawings.

[0109] Figure 1 A flowchart of a multipath fault recovery method provided by an embodiment of the present disclosure is provided.

[0110] like Figure 1 As shown, the method comprises the following steps:

[0111] Step 101: Acquire status data of storage devices associated with multiple faulty paths in a target storage system, and detect whether the multiple faulty paths have recovered to a normal state based on the status data.

[0112] In an embodiment of the present disclosure, status data for storage devices associated with multiple failed paths in a target storage system is obtained, and based on this obtained status data, the system detects whether these failed paths have returned to normal. Specifically, the system communicates with each storage device in the target storage system via a pre-defined interface or protocol to collect status information directly related to the failed paths. This status information includes, but is not limited to, the device's online / offline status, device capacity, device identifier (e.g., WWID), and device response status. After collecting the status data, the system executes a series of detection logic to determine whether the failed path has returned to normal. This detection process may involve comparing and analyzing the status data, and evaluating it based on pre-defined rules or thresholds. For example, the system may check whether the device has transitioned from offline to online, whether the device capacity matches expectations, and whether the device identifier remains consistent. If the detection results indicate that the failed path has returned to normal, the system will proceed with subsequent path recovery or verification processes. If the detection results indicate that the failed path has not yet returned to normal, the system may continue to monitor the path's status or perform appropriate fault handling operations based on pre-defined policies.

[0113] Through the above steps, the system can accurately and in real time obtain status information about the faulty path and make appropriate judgments and actions based on this information. This not only improves the system's ability to monitor faulty paths but also provides reliable data support for subsequent path recovery or fault handling. Furthermore, this step, as the starting point of the entire path recovery process, lays the foundation for the smooth execution of subsequent steps, thereby ensuring the effectiveness and reliability of the entire multipath-driven path recovery method.

[0114] Step 102 : When it is detected that the path status of the target fault path has recovered, a multi-level verification process is performed on the target fault path. The verification types of the multi-level verification process include: device attribute verification, shared consistency verification, and scheduling metadata verification.

[0115] In the embodiment of the present disclosure, if the system detects that the path status of the target faulty path has returned to normal, the system proceeds to step 102 to perform a multi-level verification process on the target faulty path. The multi-level verification process is intended to ensure that the restored path meets system requirements in all aspects, avoiding subsequent failures or data inconsistencies caused by potential problems. The verification types of the multi-level verification process mainly include but are not limited to the following three aspects:

[0116] Device attribute verification primarily verifies the basic properties of the storage device associated with the restored path. Verification includes, but is not limited to, ensuring that the device's unique identifier (such as WWID), device capacity, and device model are consistent with the system's preset or recorded information. Device attribute verification ensures that the device associated with the restored path meets system requirements both physically and at the configuration level. Device attribute verification helps prevent path failures caused by device misconfiguration or improper device replacement, ensuring that the storage device used by the system has the correct properties and configuration.

[0117] Shared consistency checks are primarily performed on storage volumes accessed by multiple hosts. These checks aim to ensure that the restored path maintains data consistency and access synchronization in a multi-host environment. This may involve verifying the storage volume's registration keys, access permissions, and locking mechanisms to prevent data conflicts or inconsistencies caused by simultaneous access by multiple hosts. Shared consistency checks help avoid data inconsistencies in multi-host shared access scenarios, ensuring system data integrity and consistency.

[0118] Scheduling metadata verification primarily verifies the load balancing and I / O scheduling metadata of the restored path. Verification includes, but is not limited to, ensuring that path priority information, block allocation information, port group information, and other information are consistent with system presets or recorded information. Scheduling metadata verification ensures that the restored path functions properly in terms of load balancing and I / O scheduling according to system requirements. Scheduling metadata verification helps prevent load imbalances or I / O scheduling anomalies caused by scheduling metadata errors, ensuring the proper allocation and efficient utilization of system resources.

[0119] In summary, the multi-level verification process in step 102 comprehensively verifies the target fault path after recovery through three aspects: device attribute verification, shared consistency verification, and scheduling metadata verification, ensuring that the path meets system requirements in all aspects, thereby improving the stability and reliability of the system.

[0120] Step 103 : If the target faulty path fails the multi-level verification, a path repair operation is performed on the target faulty path that fails the multi-level verification based on the verification type that fails.

[0121] In an embodiment of the present disclosure, if the multi-level verification processing (including device attribute verification, shared consistency verification, and scheduling metadata verification) performed by the system on the target fault path shows that the path fails one or more verification types, the system will perform corresponding path repair operations on the target fault path that fails the multi-level verification based on the verification type that fails. Specifically, the path repair operation will be customized according to the verification type that fails. For the case where the device attribute verification fails: If the device attribute verification fails, it indicates that the storage device associated with the restored path does not meet the system requirements at the physical or configuration level. At this time, the system will check the device identifier, capacity, model and other attributes, and try to reconfigure or replace the device to ensure that the device attributes meet the system preset or recorded information. By repairing the device attributes, the system can ensure that the storage device used meets the requirements at the physical and configuration levels, avoiding path failures caused by device configuration errors.

[0122] If the shared consistency check fails, this indicates data inconsistencies or access synchronization issues on the restored path in a multi-host shared access scenario. The system will then check the volume's registration key, access permissions, and locking mechanisms, and take appropriate measures (such as resynchronizing data and adjusting access permissions) to restore data consistency and access synchronization. By repairing shared consistency, the system avoids data conflicts and inconsistencies in multi-host shared access scenarios, ensuring system data integrity and consistency.

[0123] Scheduling metadata verification failure: If scheduling metadata verification fails, this indicates anomalies in load balancing and I / O scheduling on the restored path. The system then checks the path's priority information, block allocation information, port group information, and other information, and adjusts or reconfigures this metadata to ensure that the path functions properly in terms of load balancing and I / O scheduling according to system requirements. By repairing scheduling metadata, the system prevents load imbalances or I / O scheduling anomalies caused by scheduling metadata errors, ensuring the proper allocation and efficient utilization of system resources.

[0124] In summary, the path repair operation in step 103 is customized based on the failed verification type, aiming to resolve any issues with the device attributes, shared consistency, or scheduling metadata of the restored path, thereby ensuring that the path can function normally according to system requirements and improving system stability and reliability.

[0125] Step 104 : Register the repaired target fault path and the target fault path that has passed the multi-level verification into the active path pool.

[0126] In the disclosed embodiments, after completing repair operations on target faulty paths (for paths that failed multi-level verification) and confirming that the target faulty paths have passed multi-level verification, these repaired or confirmed normal target faulty paths are registered in the active path pool. The system adds path information (including but not limited to path identifiers, associated device information, verification pass status, etc.) that meets active status requirements to the active path pool through internal management interfaces or protocols. The active path pool is a logical collection used by the system to manage and schedule available storage paths. It maintains information about all currently functioning storage paths in the system.

[0127] By registering repaired or confirmed normal paths into the active path pool, the system ensures that these paths are quickly and accurately allocated and used when needed, thereby improving the overall availability and responsiveness of the storage system.

[0128] The present disclosure provides a multipath fault recovery method. By performing multipath detection, the present disclosure enables real-time monitoring and detection of faulty path status, enabling timely response upon path recovery. Upon detecting that the target faulty path has recovered, a multi-level verification process is performed on the target faulty path, including device attribute verification, shared consistency verification, and scheduling metadata verification. This multi-level verification mechanism comprehensively and meticulously checks the restored path for potential issues. Device attribute verification ensures that the basic attributes and configurations of path devices meet requirements; shared consistency verification prevents multi-host shared access issues; and scheduling metadata verification ensures the correctness of path load balancing and scheduling policies. This multi-level verification mechanism improves the accuracy and reliability of path recovery and reduces the risk of data inconsistencies or system failures caused by improper path recovery. Path repair operations are performed on target faulty paths that fail the multi-level verification process. This allows the system to promptly implement remediation measures when potential issues are detected after path recovery, preventing further failures or data inconsistencies caused by these issues. Path repair operations can provide targeted processing for different types of verification failures, improving repair efficiency and accuracy and reducing system maintenance costs and downtime. Through real-time monitoring and detection of faulty path status, multi-level verification processing, path repair operations, and registration with the active path pool, the accuracy, reliability, and availability of multipath-driven path recovery methods can be improved, providing a more stable and efficient path recovery solution for storage systems.

[0129] In order to clearly illustrate the embodiment of the present disclosure, this embodiment provides a flowchart of another multipath failure recovery method.

[0130] like Figure 2 As shown, the method comprises the following steps:

[0131] Step 201: Obtain status data of storage devices associated with multiple faulty paths in a target storage system. Extract the historical path status of the target faulty path from the status data and determine whether the current path status of the target faulty path is consistent with the historical path status.

[0132] When the current path state is inconsistent with the historical path state, step 202 is executed; when the current path state is consistent with the historical path state, step 208 is executed.

[0133] Specifically, in step 201, status data for storage devices associated with multiple faulty paths in the target storage system is obtained. The historical path status of the target faulty path is further extracted from this status data to determine whether the current path status is consistent with the historical path status. The system communicates with each storage device in the target storage system via a pre-defined interface or protocol to collect status information directly related to the faulty path. This status information includes, but is not limited to, the device's online / offline status, device capacity, device identifier (e.g., WWID), device response status, and path connection status.

[0134] The system extracts the historical path status of the target fault path from the collected status data. This historical path status may be stored in the system's log files, database, or dedicated path status management module. This extracted historical path status includes, but is not limited to, the path's last successful connection time, the last failure time, and path connection / disconnection records.

[0135] Consistency check between current and historical path states: The system compares the extracted historical path states with the currently collected path states to determine whether the current path state is consistent with the historical path state. This comparison may include key information such as path connection status, device response status, and device identifiers. If the current and historical path states are consistent in all key information, the system considers them consistent. If any inconsistency exists, the system considers them inconsistent.

[0136] Step 202: Based on the current path status, determine whether the target fault path is in an active state.

[0137] If it is in the active state, execute step 203.

[0138] Specifically, in step 202, the current status data of the target fault path has been obtained through the aforementioned steps. This data may include key information such as the path's connection status, device response status, and I / O operation status. The system has a preset set of criteria for determining the path's activation status. These criteria may be based on the path's connection status (e.g., whether the path has successfully established a connection), device response status (e.g., whether the device can normally respond to I / O requests), and possibly business logic status (e.g., whether the path has been allocated to a business process).

[0139] Specifically, the system can check the "connection status" field in the path status data. If the field indicates that the path has been successfully connected and the device response status is normal (such as no error I / O operation report), and it may be combined with the business logic status (such as the path has been used by the business process), the system determines that the target fault path is currently in an active state.

[0140] Step 203: Perform device capacity verification on the storage device capacity associated with the target fault path, and perform credibility verification on the device identifier.

[0141] As a specific implementation method of this embodiment, "performing a device capacity check on the associated storage device capacity of the target fault path and performing a trustworthy verification on the device identifier" includes but is not limited to: verifying whether the device capacity data is empty or inconsistent with the device capacity of the target storage system to perform a device capacity check; if the device capacity check passes, then performing a trustworthy verification on the device identifier.

[0142] Furthermore, when performing "trusted verification of the device identifier", the following methods may be used, but are not limited to: verifying whether there is an abnormality in the device identifier, and after determining that there is no abnormality, verifying whether the device identifier is consistent with the cached device identifier; if they are inconsistent, determining that the trusted verification of the device identifier has failed; if they are consistent, determining that the trusted verification of the device identifier has passed.

[0143] Specifically, in step 203, a capacity verification operation is performed on the storage device associated with the target fault path. During the verification process, the system obtains the actual capacity data of the device and compares it with the pre-recorded or expected device capacity in the target storage system. Specific verification procedures include, but are not limited to, checking whether the device capacity data is empty (i.e., no valid capacity information is obtained) and whether the device capacity is inconsistent with the device capacity record in the target storage system. If the device capacity data is empty or inconsistent with the record, the device capacity verification is determined to have failed.

[0144] Device Identifier Trust Verification: If the device capacity verification passes, the system will further verify the device identifier. A device identifier (such as the WWID) uniquely identifies a storage device in the system and is used to distinguish between different storage devices. During verification, the system first checks the device identifier for anomalies (such as format errors or invalid characters). If any anomalies are found, the device identifier trust verification will be deemed to have failed.

[0145] If the device identifier is normal, the system will further verify whether it is consistent with the device identifier stored in the cache. If it is inconsistent, it means that the device identifier may have been tampered with or there is another problem, and the device identifier trust verification will be determined to have failed. If it is consistent, the device identifier trust verification will be determined to have passed.

[0146] The system can integrate a device capacity check and device identifier trusted verification module into the path recovery process. The module receives capacity data and identifier information from the storage device or target storage system and performs the above check and verification operations.

[0147] To improve the efficiency and accuracy of checksum verification, the system can use hashing algorithms, checksums, or other data comparison techniques to quickly determine the consistency of device capacity and identifiers. Furthermore, the system can record the checksum verification results in a log file or database for subsequent auditing and troubleshooting.

[0148] When the device capacity data verification fails and / or the device identifier verification fails, step 204 is executed.

[0149] After the device capacity check passes and the device identifier verification passes, step 205 is executed.

[0150] Step 204 : Release the device binding relationship of the target fault path and set it to a failed state, and scan and add the target fault path in the failed state.

[0151] Specifically, in step 204, the storage device to which the target fault path is bound is identified and determined. By calling the corresponding management interface or sending a control command, the system releases the binding relationship between the target fault path and the storage device. This operation ensures that the path is no longer associated with a specific device, preparing for subsequent path recovery or reconfiguration. After the binding relationship is released, the system marks the status of the target fault path as "failed". This marking operation can be achieved by updating the path status database, modifying the path configuration file, or sending a status change notification. The marking of the failed status enables the system to identify and handle abnormal conditions of the path, avoiding attempts to use the path before it is fully restored, thereby preventing potential data errors or system failures.

[0152] The system initiates a path scanning mechanism, performing periodic or triggered scans of all paths in the storage system. During this scan, the system detects and identifies failed paths. Once a failed path is detected, the system triggers the path addition process, re-adding the path to the path management module. During this addition process, the system reconfigures path parameters, binds storage devices (if applicable), and updates the path status to "available" or another appropriate state.

[0153] Step 205 : Determine whether the historical path status of the target fault path is in an inactive state. If it is in an inactive state, increase the number of active paths by one.

[0154] Specifically, in step 205, the system retrieves the historical path status information of the target fault path from the path status management module or the relevant database. The historical path status information generally records the status of the path at a certain point in the past, including whether it is activated, the connection status, the time when the fault occurred, etc. The system parses the retrieved historical path status information to determine whether the target fault path is in an inactive state in the historical records. The inactive state may mean that the path has historically experienced a fault, has been manually disabled, or has not been used for other reasons. The judgment logic can be based on a status flag, a status code, or a specific status description text. For example, if the historical status information contains "inactive", "disabled", or similar flags, the system determines that the path has been inactive historically.

[0155] If the target fault path is determined to have a historically inactive path status, the system increments the number of active paths. Specifically, the system accesses a counter or variable that counts the number of active paths and increments it by one. This operation reflects the successful restoration of the previously inactive path, thereby increasing the number of currently available active paths.

[0156] Step 206: Check whether the registration key value of the target fault path recorded on the storage controller side matches the stored target key value.

[0157] As a specific implementation method of this embodiment, "detecting whether the registration key value of the target fault path recorded on the storage controller side matches the stored target key value" can be verified and matched in the following ways, but not limited to: verifying whether the registration key value is empty; if the registration key value is not empty, verifying whether the registration key value is consistent with the target key value.

[0158] When it is determined that the registration key value is inconsistent with the target key value, step 207 is executed.

[0159] Specifically, in step 206, a specific query command is sent through the communication interface (such as SCSI, iSCSI, FC, or other protocol interfaces) of the storage controller to obtain the registration key value recorded on the storage controller side for the target fault path. The registration key value is usually a unique identifier used to distinguish different paths or devices in the storage system. The system reads the pre-stored target key value from the local storage or configuration file. This target key value is set during the path configuration or initialization process and is used to compare with the registration key value on the storage controller side. The system first verifies whether the registration key value obtained from the storage controller side is empty. If the registration key value is empty, it may mean that the storage controller side has not correctly recorded the registration information of the path, or the path information has been deleted or damaged. If the registration key value is not empty, the system further verifies whether the registration key value is consistent with the pre-stored target key value. This verification process can be implemented through string comparison, hash value comparison, or other data comparison technologies.

[0160] Step 207: re-register the target fault path based on the target key value.

[0161] Specifically, in step 207, the validity of the obtained target key value is confirmed. The target key value is determined in step 206 by verifying the matching of the registration key value recorded by the storage controller end with the pre-stored target key value. It serves as the unique identifier of the path in the storage system and is used for subsequent path registration operations. Based on the target key value, the system prepares the information required to re-register the target fault path. This information may include but is not limited to the path identifier, associated storage device information, path configuration parameters (such as priority, bandwidth allocation, etc.), and any necessary authentication or authorization information. The system initiates a re-registration request for the target fault path by communicating with the management interface or protocol of the storage system. In the request, the system will include the target key value and the above-prepared path registration information.

[0162] After receiving the re-registration request, the storage system verifies the validity of the target key value and updates or creates the corresponding path record based on the path registration information in the request. If the verification passes and the information is complete, the storage system successfully registers the target failed path and returns a registration success response.

[0163] Step 208: Determine the ownership type of the volume associated with the target failure path; and perform scheduling metadata verification and maintenance on the target failure path based on the ownership type.

[0164] As a specific implementation method of this embodiment, "scheduling metadata verification and maintenance of the target fault path based on the ownership type" includes but is not limited to: if there is no owned volume, the block priority and port group information of the target fault path are verified with the kernel state saved value; the inconsistent block priority and / or port group information are updated according to the kernel state saved value; if there is an owned volume, the block priority of the target fault path is verified with the kernel state saved value; the inconsistent block priority is updated according to the kernel state saved value.

[0165] Specifically, in step 208, the ownership type of the volume associated with the target fault path is determined by querying the relevant information in the storage system or the path management module. The ownership type is generally divided into two types: "unowned volume" and "owned volume". Unowned volume means that the volume is not exclusively used by a specific application or service, while owned volume means that the volume has been explicitly bound to a certain application or service. Scheduling metadata verification and maintenance in the case of unowned volume: If the judgment result is that the volume associated with the target fault path is an unowned volume, the system will further verify the block priority and port group information of the target fault path. The verification process involves comparing the block priority and port group information in the current path configuration with the corresponding values ​​saved in the kernel state. The kernel state saved value is usually the latest and most accurate path configuration information recorded and maintained by the system kernel when the system is initialized or the path configuration is updated. If the verification finds that the block priority and / or port group information are inconsistent, the system will update this information according to the kernel state saved value to ensure the consistency and accuracy of the path configuration.

[0166] Scheduling metadata verification and maintenance with owned volumes: If the volume associated with the target failure path is determined to be owned, the system will only verify the block priorities of the target failure path. This verification method is similar to the case without owned volumes, comparing the block priorities in the current path configuration with the values ​​saved in kernel state. If the verification finds a discrepancy in the block priorities, the system will update them according to the values ​​saved in kernel state to ensure that the path configuration meets the latest kernel state requirements.

[0167] Step 209 : Register the repaired target fault path and the target fault path that has passed the multi-level verification into the active path pool.

[0168] Specifically, in step 209, it is confirmed that the target faulty path has been repaired. Repair operations may include, but are not limited to, hardware replacement, software configuration adjustment, and network connection restoration, and are performed based on the fault type and repair plan detected in steps 201 to 208. Confirmation of repair completion can be based on automatically detected path status changes, manual confirmation from an administrator, or a completion flag set during the repair process.

[0169] Multi-level verification confirmation: The system verifies that the target fault path has passed all pre-defined multi-level verification processes. These verification processes may include, but are not limited to, path status verification (as described in step 201), device capacity verification and identifier verification (as described in step 203), path binding removal and re-registration (as described in steps 204 and 207), and scheduling metadata verification and maintenance (as described in step 208). Passing multi-level verification means that the path meets the system's required availability and stability standards across multiple dimensions.

[0170] Registering with the active path pool: After confirming that the target faulty path has been repaired and passed multiple levels of verification, the system registers the path with the active path pool. The active path pool is a collection of currently available path information in the system for use in subsequent I / O operation scheduling. Registration operations may involve updating the path status database, modifying the path configuration file, or sending a registration request to the path management module through a specific management interface. After successful registration, the path is marked as "active" and becomes available for system scheduling.

[0171] As an implementable manner of this embodiment, before performing scheduling metadata verification on the target fault path, the multipath fault recovery method further includes: determining whether the current path state of the target fault path is an active state;

[0172] When the target fault path is determined to be in the activated state, the detection interval of the target fault path is incremented by one; it is determined whether the target fault path is currently in the early stage of path recovery. If so, a scheduling metadata check is performed on the target fault path.

[0173] Specifically, in the multipath fault recovery method, to ensure the rationality and efficiency of scheduling metadata verification, a series of pre-judgments and operations are required before executing scheduling metadata verification. The specific contents are as follows:

[0174] Step 1: Determine whether the current path status of the target fault path is active.

[0175] The system obtains the current path status information of the target fault path by accessing the path status management module or calling related interfaces. Path status information is usually stored in the form of a specific status identifier or status code, such as "1" for active status and "0" for inactive status.

[0176] The system compares the acquired path status information with the preset activation status identifier to determine whether the target fault path is currently in the activation state.

[0177] Step 2: When the state is determined to be active, the detection interval of the target fault path is increased by one.

[0178] If the target fault path is determined to be active in step 1, the system increments the detection interval. The detection interval is the time interval between regular system checks on the target fault path, typically measured in seconds, minutes, or other time units. The system accesses the variable or configuration item storing the target fault path detection interval and increments its value by one. For example, if the current detection interval is 5 seconds, it will increase to 6 seconds after the increment.

[0179] Step 3: Determine whether the target fault path is currently in the early stage of path recovery. If it is in the early stage of path recovery, perform scheduling metadata verification on the target fault path.

[0180] The system determines whether the target faulty path is currently in the early stages of path recovery based on preset criteria. This criteria can be time-based, such as determining a path is in the early stages of recovery within a certain time threshold (e.g., 30 seconds) from the time the path failure occurs. It can also be event-based, such as determining a path is in the early stages of recovery within a certain period of time after a specific repair operation (such as hardware replacement or network connection restoration). If the target faulty path is determined to be in the early stages of path recovery, the system triggers the scheduling metadata verification process to verify the scheduling metadata for the target faulty path to ensure its accuracy and consistency.

[0181] As an implementable method of this embodiment, determining whether the target fault path is currently in the early stage of path recovery includes: comparing the detection time of the target fault path with an early stage recovery threshold, where the early stage recovery threshold is the product of the detection interval and the proportional coefficient; when the detection time does not reach the early stage recovery threshold, determining that it is in the early stage of path recovery; when the detection time reaches the early stage recovery threshold, determining that it is not in the early stage of path recovery.

[0182] Specifically, the detection interval of the target fault path is obtained. The detection interval is a time interval preset by the system for periodically detecting the status of the target fault path, usually expressed in time units such as seconds and milliseconds. For example, the detection interval may be set to 5 seconds. The system also obtains a preset proportional coefficient. The proportional coefficient is an empirical value or a value set based on system performance requirements, which is used to adjust the size of the initial recovery threshold. The value range of the proportional coefficient can be determined based on actual conditions, for example, between 0.1 and 1, or it may be greater than 1.

[0183] The system calculates the initial recovery threshold by multiplying the detection interval by the proportional coefficient. The calculation formula is: Initial recovery threshold = detection interval × proportional coefficient. For example, if the detection interval is 5 seconds and the proportional coefficient is 4, the initial recovery threshold is 20 seconds.

[0184] The system records the time when the target fault path begins detection, typically when a fault is detected and the recovery process is initiated. During subsequent determinations, the system obtains the current time and calculates the difference between the current time and the time when the target fault path begins detection. This difference is the detection duration. For example, if the target fault path begins detection at 10:00:00 and the current time is 10:00:02, the detection duration is 2 seconds.

[0185] As an implementable manner of this embodiment, after determining whether the current path state of the target faulty path is consistent with the historical path state, the multipath fault recovery method further includes: setting the detection interval of the target faulty path to a preset detection interval.

[0186] Specifically, in the multi-path fault recovery method process, determining whether the current path state of the target fault path is consistent with the historical path state is a key link. After this, setting the detection interval of the target fault path to the preset detection interval is executed, which is of great significance to the stable operation and efficient recovery of the system.

[0187] The system obtains the current path status information of the target fault path by accessing the path status management module or related data storage areas. Path status information typically includes the path's connection status (e.g., connected, disconnected), activation status (e.g., activated, inactivated), and fault status (e.g., normal, faulty). At the same time, the system retrieves historical path status information of the target fault path from historical records. Historical path status information records the path's status at a certain point in the past and can be used to compare it with the current status. The system compares the current path status information with the historical path status information item by item to determine whether the two are consistent. If all key status items are the same, the status is determined to be consistent; if any key status item is different, the status is determined to be inconsistent.

[0188] The preset detection interval is a system-defined time value that specifies the interval for periodic detection of the target faulty path. This preset detection interval can be set based on factors such as system performance requirements, storage device characteristics, and business needs. For example, for services with high real-time requirements, the preset detection interval can be set to a shorter value, such as 1 second; for services with relatively low real-time requirements, the preset detection interval can be set to a longer value, such as 10 seconds. After the system verifies the consistency of the current path status with the historical path status, regardless of whether the result is consistent or inconsistent, it sets the detection interval of the target faulty path to the preset detection interval. Specifically, the system accesses the variable or configuration item storing the target faulty path detection interval and modifies its value to the preset detection interval value. For example, if the preset detection interval is 5 seconds, the system modifies the value of this variable or configuration item from the original value (such as 3 seconds or 8 seconds) to 5 seconds.

[0189] It should be noted that the embodiments of the present disclosure may include multiple steps. For the convenience of description, these steps are numbered, but these numbers do not limit the execution time slots or execution order between the steps; these steps can be implemented in any order, and the embodiments of the present disclosure do not limit this.

[0190] Corresponding to the above-mentioned multipath fault recovery method, the present disclosure also provides a multipath fault recovery device. Since the device embodiment of the present disclosure corresponds to the above-mentioned method embodiment, details not disclosed in the device embodiment can be referred to the above-mentioned method embodiment and will not be repeated in this disclosure.

[0191] Figure 3 A schematic diagram of a multi-path fault recovery device provided by an embodiment of the present disclosure is shown in FIG. Figure 3 Shown, including:

[0192] The detection unit 31 is configured to obtain status data of storage devices associated with multiple faulty paths in the target storage system, and detect whether the multiple faulty paths have recovered to a normal state based on the status data;

[0193] A verification unit 32 is configured to perform a multi-level verification process on the target fault path when detecting that the path state of the target fault path has recovered. The verification types of the multi-level verification process include: device attribute verification, shared consistency verification, and scheduling metadata verification;

[0194] A repair unit 33 is configured to perform a path repair operation on the target faulty path that fails the multi-level verification based on the verification type that fails if the target faulty path fails the multi-level verification;

[0195] The registration unit 34 is configured to register the repaired target fault path and the target fault path that has passed the multi-level verification into the active path pool.

[0196] The present disclosure provides a multipath fault recovery device. By performing multipath detection, the device monitors and detects the status of faulty paths in real time, enabling timely response when the paths are restored. Upon detecting that the path status of a target faulty path has recovered, the device performs a multi-level verification process on the target faulty path, including device attribute verification, shared consistency verification, and scheduling metadata verification. This multi-level verification mechanism comprehensively and meticulously checks the restored path for potential issues. Device attribute verification ensures that the basic attributes and configurations of path devices meet requirements; shared consistency verification prevents multi-host shared access issues; and scheduling metadata verification ensures the correctness of path load balancing and scheduling policies. This multi-level verification mechanism improves the accuracy and reliability of path recovery and reduces the risk of data inconsistencies or system failures caused by improper path recovery. Path repair operations are performed on target faulty paths that fail the multi-level verification process. This allows the system to promptly implement remedial measures when potential issues are detected after path recovery, preventing further failures or data inconsistencies caused by these issues. Path repair operations can provide targeted processing for different types of verification failures, improving repair efficiency and accuracy and reducing system maintenance costs and downtime. Through real-time monitoring and detection of faulty path status, multi-level verification processing, path repair operations, and registration with the active path pool, the accuracy, reliability, and availability of multipath-driven path recovery methods can be improved, providing a more stable and efficient path recovery solution for storage systems.

[0197] It should be noted that the above explanation of the method embodiment is also applicable to the device of this embodiment, and the principles are the same, which is not limited in this embodiment.

[0198] For the description of the features in the embodiment corresponding to the apparatus for multipath failure recovery, reference may be made to the relevant description of the embodiment corresponding to the method for multipath failure recovery, which will not be described in detail here.

[0199] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned multipath fault recovery method embodiments.

[0200] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned multipath failure recovery method embodiments when running.

[0201] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0202] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned multipath failure recovery method embodiments are implemented.

[0203] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned multi-path fault recovery method embodiments are implemented.

[0204] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0205] The above describes in detail a multi-path fault recovery method and device, electronic device, and storage medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only intended to help understand the method and core concept of the present application. It should be noted that, for those skilled in the art, without departing from the principles of the present application, several improvements and modifications may be made to the present application, and such improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A multipath fault recovery method, characterized in that: include: Acquire status data of storage devices associated with multiple faulty paths in the target storage system, and detect whether the multiple faulty paths have recovered to a normal state based on the status data; When the path status of the target failed path is detected to have recovered, a multi-level verification process is performed on the target failed path. The verification types of the multi-level verification process include: device attribute verification, shared consistency verification, and scheduling metadata verification. The shared consistency verification performs consistency verification on storage volumes shared by multiple hosts, and the scheduling metadata verifies the metadata of the restored path in load balancing and I / O scheduling. If the target fault path fails the multi-level verification, performing a path repair operation on the target fault path that fails the multi-level verification based on the verification type that fails; Register the repaired target fault path and the target fault path that passes multi-level verification into the active path pool.

2. The method for multipath failure recovery according to claim 1, wherein: Before performing multi-level verification processing on the target fault path, the method further includes: Extracting a historical path state of the target fault path from the state data, and determining whether the current path state of the target fault path is consistent with the historical path state; When the current path state is inconsistent with the historical path state, judging whether the target fault path is in an active state based on the current path state; If it is in the activated state, a device attribute check and a shared consistency check are performed on the target fault path.

3. The multipath failure recovery method according to claim 2, wherein: The performing of device attribute verification and shared consistency verification on the target fault path includes: Performing device capacity verification on the capacity of the storage device associated with the target fault path and performing credibility verification on the device identifier; Detect whether the registration key value of the target fault path recorded by the storage controller matches the stored target key value.

4. The multipath failure recovery method according to claim 3, wherein: The performing device capacity verification on the storage device capacity associated with the target fault path and performing trustworthy verification on the device identifier includes: Verifying whether the device capacity data is empty or inconsistent with the device capacity of the target storage system, thereby performing device capacity verification; If the device capacity verification passes, the device identifier is authenticated.

5. The method for multipath failure recovery according to claim 4, wherein: The performing trustworthy verification on the device identifier includes: Verifying whether there is an anomaly in the device identifier, and after determining that there is no anomaly, verifying whether the device identifier is consistent with a cached device identifier; If they are inconsistent, it is determined that the device identifier trust verification has failed; if they are consistent, it is determined that the device identifier trust verification has passed.

6. The multipath failure recovery method according to claim 4, characterized in that: When the device capacity data verification fails and / or the device identifier verification fails, performing a path repair operation on the faulty path that fails the multi-level verification based on the failed verification type includes: The device binding relationship of the target fault path is released and the target fault path is set to a failed state, and the target fault path in the failed state is scanned and added.

7. The method for multipath failure recovery according to claim 4, wherein: After the device capacity check passes and the device identifier verification passes, the method further includes: It is determined whether the historical path state of the target fault path is in an inactive state. If it is in an inactive state, the number of active paths is increased by one.

8. The multipath failure recovery method according to claim 3, wherein: The detecting whether the registration key value of the target fault path recorded by the storage controller matches the stored target key value includes: Verify whether the registration key value is empty; If the registration key value is not empty, verify whether the registration key value is consistent with the target key value.

9. The multipath failure recovery method according to claim 8, characterized in that: The performing of a path repair operation on a target faulty path that fails the multi-level verification based on the failed verification type further includes: When it is determined that the registration key value is consistent with the target key value, the target fault path is re-registered based on the target key value.

10. The multipath failure recovery method according to claim 9, characterized in that: The method further comprises: After performing a shared consistency check on the target failure path, a scheduling metadata check is performed on the target failure path.

11. The multipath failure recovery method according to claim 2, wherein: The method further comprises: If the current path state is consistent with the historical path state, a scheduling metadata check is performed on the target fault path.

12. The method for multipath failure recovery according to any one of claims 10 or 11, characterized in that: The performing scheduling metadata verification on the target fault path includes: Determining the ownership type of the volume associated with the target failure path; Based on the attribution type, scheduling metadata verification and maintenance are performed on the target fault path.

13. The multipath failure recovery method according to claim 12, wherein: The performing scheduling metadata verification and maintenance on the target fault path based on the belonging type includes: If the volume is unowned, the block priority and port group information of the target failure path are compared with the kernel state stored values; the block priority and / or port group information that are inconsistent are updated according to the kernel state stored values; If there is an owned volume, the block priority of the target failure path is verified with the kernel state saved value; the block priority that is inconsistent is updated according to the kernel state saved value.

14. The method for multipath failure recovery according to any one of claims 10 or 11, characterized in that: Before performing scheduling metadata verification on the target fault path, the method further includes: Determine whether the current path state of the target fault path is an activated state; When the state is determined to be active, the detection interval of the target fault path is incremented by one; It is determined whether the target fault path is currently in an early stage of path recovery. If it is in the early stage of path recovery, scheduling metadata verification is performed on the target fault path.

15. The method for multipath failure recovery according to claim 14, wherein: The determining whether the target fault path is currently in an early stage of path recovery includes: Comparing the detection duration of the target fault path with an initial recovery threshold, where the initial recovery threshold is the product of the detection interval and a proportional coefficient; When the detection duration does not reach the initial recovery threshold, it is determined that the path is in the initial recovery stage; When the detection duration reaches the initial recovery threshold, it is determined that the path is not in the initial recovery stage.

16. The method for multipath failure recovery according to claim 2, wherein: After determining whether the current path state of the target fault path is consistent with the historical path state, the method further includes: The detection interval of the target fault path is set to a preset detection interval.

17. A multi-path fault recovery device, characterized in that: include: a detection unit, configured to obtain status data of storage devices associated with multiple faulty paths in a target storage system, and detect whether the multiple faulty paths have recovered to a normal state based on the status data; A verification unit is configured to perform a multi-level verification process on the target failed path upon detecting that the path state of the target failed path has recovered. The verification types of the multi-level verification process include: device attribute verification, shared consistency verification, and scheduling metadata verification; the shared consistency verification performs consistency verification on storage volumes shared by multiple hosts, and the scheduling metadata verifies metadata related to load balancing and I / O scheduling of the restored path; a repair unit, configured to, if the target faulty path fails the multi-level verification, perform a path repair operation on the target faulty path that fails the multi-level verification based on a verification type that fails; The registration unit is used to register the repaired target fault path and the target fault path that has passed the multi-level verification into the active path pool.

18. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the multipath failure recovery method according to any one of claims 1 to 16.

19. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the multipath failure recovery method according to any one of claims 1 to 16.

20. A computer program product, characterized in that The invention comprises a computer program, which implements the multipath failure recovery method according to any one of claims 1 to 16 when being executed by a processor.

Citation Information

Patent Citations

  • Multi-path anomaly detection and repair method, device, equipment and medium

    CN115686921A

  • Multi-path exception processing method and device, computer equipment and storage medium

    CN117573405A