Multi-path fault recovery method and device, electronic equipment and storage medium

Through the multi-path failure recovery method of the computer storage system, multi-level checksum path repair operations are performed, data consistency and reliability problems after multi-path failure are solved, and more stable and efficient path recovery is achieved.

CN120407263AActive Publication Date: 2025-08-01INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510889393.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-08-01
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

In the computer storage system, data consistency and service reliability cannot be guaranteed after multi-path failure recovery, especially in large-scale enterprise storage environments, path switching is difficult to ensure data consistency and service reliability.

Method used

By obtaining the fault path status data of the target storage system, performing multi-level verification processing, including device attribute verification, shared consistency verification and scheduling metadata verification, repair operations are performed on paths that fail to pass the verification, and registering the repaired path to the active path pool.

Benefits of technology

Improve the accuracy and reliability of path recovery, reduce the risk of data inconsistency or system failure caused by improper path recovery, improve the stability and availability of the system, and reduce system maintenance costs and downtime.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407263A_ABST
    Figure CN120407263A_ABST
Patent Text Reader

Abstract

The invention provides a multi-path fault recovery method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining state data of storage equipment associated with a plurality of fault paths in a target storage system, and detecting whether the plurality of fault paths are recovered to a normal state or not based on the state data; when it is detected that the path state of the target fault path is recovered, performing multi-stage verification processing on the target fault path; if the target fault path does not pass the multi-stage verification, performing path repair operation on the target fault path which does not pass the multi-stage verification based on the non-passed verification type; and registering the repaired target fault path and the target fault path passing the multi-stage verification to an active path pool. Compared with the prior art, when it is detected that the path state of the target fault path is recovered, multi-stage verification processing is performed on the target fault path, and whether the recovered path has potential problems or not can be comprehensively and meticulously checked.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of storage technologies, and in particular, to a method and apparatus for multi-path fault recovery, an electronic device, and a storage medium. Background Art

[0002] In the field of computer storage systems, multi-path software, as a core component to ensure high storage availability, the reliability of its path fault recovery mechanism directly affects business continuity. Currently, an event-driven mechanism is generally adopted for path status management: when a storage controller link recovers from a fault state to a normal state, the multi-path service will directly activate the path and add it to the available path pool. Although the related solutions can implement the basic path switching function, due to reasons such as drive letter drift during its recovery stage, it is difficult to ensure data consistency and service reliability in a large-scale enterprise storage environment. Summary of the Invention

[0003] The present disclosure provides a method and apparatus for multi-path fault recovery, an electronic device, and a storage medium. Its main purpose is to solve the problem that data consistency cannot be guaranteed after a multi-path fault in a computer storage system.

[0004] According to a first aspect of the present disclosure, there is provided a method for multi-path fault recovery, including: Obtaining status data of multiple fault path-associated storage devices in a target storage system, and detecting whether the multiple fault paths have recovered to a normal state based on the status data; When it is detected that the path status of a target fault path has recovered, performing multi-level verification processing on the target fault path, and the verification types of the multi-level verification processing include: device attribute verification, shared consistency verification, and scheduling metadata verification; If the target fault path fails the multi-level verification, based on the unpassed verification type, performing a path repair operation on the target fault path that fails the multi-level verification; Registering the repaired target fault path and the target fault path that passes the multi-level verification to the active path pool.

[0005] Optionally, before performing the multi-level verification processing on the target fault path, the method for multi-path fault recovery further includes: Extracting the historical path status of the target fault path from the status data, and determining whether the current path status of the target fault path is consistent with the historical path status; When the current path status is inconsistent with the historical path status, determining whether the target fault path is in an active state based on the current path status; If it is in an active state, performing device attribute verification and shared consistency verification on the target fault path.

[0006] Optionally, perform device attribute verification and shared consistency verification on the target fault path, including: Perform device capacity verification on the storage device capacity associated with the target failure path and perform trustworthy verification on the device identifier; Check whether the registration key value of the target fault path recorded on the storage controller side matches the stored target key value.

[0007] Optionally, perform device capacity verification on the storage device associated with the target failure path and perform trustworthy verification on the device identifier, including: Verify whether the device capacity data is empty or inconsistent with the device capacity of the target storage system to perform device capacity verification; If the device capacity verification passes, the device identifier is authenticated.

[0008] Optionally, perform trustworthy verification on the device identifier, including: Verify whether there is any anomaly in the device identifier. After confirming that there is no anomaly, verify whether the device identifier is consistent with the cached device identifier; If they are inconsistent, it is determined that the device identifier trust verification has failed; if they are consistent, it is determined that the device identifier trust verification has passed.

[0009] Optionally, when the device capacity data verification fails and / or the device identifier verification fails, a path repair operation is performed on the faulty path that fails the multi-level verification based on the failed verification type, including: Remove the device binding relationship of the target fault path and set it to the failed state. Scan and add the target fault path in the failed state.

[0010] Optionally, after the device capacity check passes and the device identifier verification passes, the multipath failure recovery method further includes: Determine whether the historical path status of the target fault path is in an inactive state. If it is in an inactive state, increase the number of active paths by one.

[0011] Optionally, checking whether the registration key value of the target fault path recorded on the storage controller matches the stored target key value includes: Verify whether the registration key value is empty; If the registration key value is not empty, verify whether the registration key value is consistent with the target key value.

[0012] Optionally, based on the failed verification type, performing a path repair operation on the target faulty path that failed the multi-level verification further includes: When it is determined that the registration key value is consistent with the target key value, the target fault path is re-registered based on the target key value.

[0013] Optionally, the multi-path fault recovery method further includes: After performing a shared consistency check on the target fault path, perform a scheduling metadata check on the target fault path.

[0014] Optionally, the multi-path fault recovery method further includes: If the current path state is the same as the historical path state, perform a scheduling metadata check on the target fault path.

[0015] Optionally, performing a scheduling metadata check on the target fault path includes: Determine the ownership type of the volume associated with the target fault path; [[ID= 13]] Based on the ownership type, perform scheduling metadata check and maintenance on the target fault path.

[0016] Optionally, based on the ownership type, performing scheduling metadata check and maintenance on the target fault path includes: If it is a volume without ownership, verify the block priority and port group information of the target fault path with the kernel state saved value; update the inconsistent block priority and / or port group information according to the kernel state saved value; If it is a volume with ownership, verify the block priority of the target fault path with the kernel state saved value; update the inconsistent block priority according to the kernel state saved value.

[0017] Optionally, before performing a scheduling metadata check on the target fault path, the multi-path fault recovery method further includes: Determine whether the current path state of the target fault path is the active state; When it is determined to be the active state, increment the detection interval of the target fault path by one; Determine whether the target fault path is currently in the initial stage of path recovery. If it is in the initial stage of path recovery, perform a scheduling metadata check on the target fault path.

[0018] Optionally, determining whether the target fault path is currently in the initial stage of path recovery includes: Compare the detection duration of the target fault path with the initial recovery threshold. The initial recovery threshold is the product of the detection interval and the proportionality coefficient; When the detection duration does not reach the initial recovery threshold, it is determined to be in the initial stage of path recovery; When the detection duration reaches the initial recovery threshold, it is determined not to be in the initial stage of path recovery.

[0019] Optionally, after determining whether the current path state of the target fault path is the same as the historical path state, the multi-path fault recovery method further includes: Set the detection interval of the target fault path to the preset detection interval.

[0020] According to a second aspect of the present disclosure, there is provided an apparatus for multi-path fault recovery, including: A detection unit, configured to obtain status data of a plurality of fault-path associated storage devices in a target storage system, and detect whether the plurality of fault paths are restored to a normal state based on the status data; A verification unit, configured to perform multi-level verification processing on a target fault path when it is detected that the path status of the target fault path is restored, and the verification types of the multi-level verification processing include: device attribute verification, shared consistency verification, and scheduling metadata verification; A repair unit, configured to, if the target fault path fails the multi-level verification, perform a path repair operation on the target fault path that fails the multi-level verification based on the verification type that fails; A registration unit, configured to register the repaired target fault path and the target fault path that passes the multi-level verification to an active path pool.

[0021] Optionally, the apparatus for multi-path fault recovery further includes: A first determination unit, configured to, before performing multi-level verification processing on a target fault path, extract the historical path status of the target fault path from the status data, and determine whether the current path status of the target fault path is consistent with the historical path status; when the current path status is inconsistent with the historical path status, determine whether the target fault path is in an active state based on the current path status; A first execution unit, configured to, if it is in an active state, perform device attribute verification and shared consistency verification on the target fault path.

[0022] Optionally, the first execution unit includes: A first verification module, configured to perform device capacity verification on the capacity of the associated storage device of the target fault path, and perform a credibility verification on the device identifier; A second verification module, configured to detect whether the registered key value of the target fault path recorded at the storage controller end matches the stored target key value.

[0023] Optionally, the first verification module is further configured to: Verify whether the device capacity data is empty or inconsistent with the device capacity of the target storage system to perform device capacity verification; If the device capacity verification passes, perform a credibility verification on the device identifier.

[0024] Optionally, the second verification module is further configured to: Verify whether the device identifier is abnormal, and after determining that there is no abnormality, verify whether the device identifier is consistent with the cached device identifier; If they are inconsistent, it is determined that the device identifier trust verification has failed; if they are consistent, it is determined that the device identifier trust verification has passed.

[0025] Optionally, the repair unit includes: Add a module for releasing the device binding relationship of the target fault path and setting it to a failed state when the device capacity data verification fails and / or the device identifier verification fails, and scanning and adding the target fault path in the failed state.

[0026] Optionally, the multipath failure recovery device further includes: The second judgment unit is configured to judge whether the historical path status of the target fault path is inactive after the device capacity check and the device identifier verification are passed, and if so, increase the number of active paths by one.

[0027] Optionally, the second verification module is further configured to: Verify whether the registration key value is empty; If the registration key value is not empty, verify whether the registration key value is consistent with the target key value.

[0028] Optionally, the repair unit also includes: The registration module is used to re-register the target fault path based on the target key value when it is determined that the registration key value is consistent with the target key value.

[0029] Optionally, the multipath failure recovery device further includes: The second execution unit is configured to perform a scheduling metadata check on the target fault path after performing a shared consistency check on the target fault path.

[0030] Optionally, the second execution unit is further configured to perform scheduling metadata verification on the target fault path if the current path state is consistent with the historical path state.

[0031] Optionally, the second execution unit includes: A judgment module, used to judge the type of the volume associated with the target failure path; The maintenance module is used to perform scheduling metadata verification and maintenance on the target fault path based on the attribution type.

[0032] Optionally, the maintenance module is also used to: If the volume is unowned, the block priority and port group information of the target failure path are compared with the kernel-state saved values; any inconsistent block priority and / or port group information is updated according to the kernel-state saved values; If the volume has an ownership, the block priority of the target failure path is verified with the value saved in the kernel state; the block priority that is inconsistent is updated according to the value saved in the kernel state.

[0033] Optionally, the multi-path fault recovery device further includes: A third determination unit, configured to determine whether the current path state of the target fault path is an active state before performing scheduling metadata verification on the target fault path; when it is determined to be in an active state, increment the detection interval of the target fault path by one; A fourth determination unit, configured to determine whether the target fault path is currently in the initial stage of path recovery, and if so, perform scheduling metadata verification on the target fault path.

[0034] Optionally, the fourth determination unit includes: A comparison module, configured to compare the detection duration of the target fault path with a recovery initial threshold, where the recovery initial threshold is the product of the detection interval and a proportionality coefficient; A first determination module, configured to determine that it is in the initial stage of path recovery when the detection duration does not reach the recovery initial threshold; A second determination module, configured to determine that it is not in the initial stage of path recovery when the detection duration reaches the recovery initial threshold.

[0035] Optionally, the multi-path fault recovery device further includes: A setting unit, configured to set the detection interval of the target fault path to a preset detection interval after determining whether the current path state of the target fault path is consistent with the historical path state.

[0036] According to a third aspect of the present disclosure, there is provided an electronic device, including: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the multi-path fault recovery method described in the foregoing first aspect.

[0037] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the multi-path fault recovery method described in the foregoing first aspect.

[0038] According to a fifth aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program, when executed by a processor, implements the multi-path fault recovery method described in the foregoing first aspect.

[0039] The present disclosure provides a method and apparatus for multi-path fault recovery, an electronic device, and a storage medium, relating to the field of storage technology. By detecting multi-paths, the present disclosure can monitor and detect the status of faulty paths in real time, so as to respond in a timely manner when the paths are restored. When it is detected that the path status of a target faulty path is restored, multi-level verification processing is performed on the target faulty path, including device attribute verification, shared consistency verification, and scheduling metadata verification. The multi-level verification processing mechanism can comprehensively and meticulously check whether there are potential problems in the restored path. Device attribute verification ensures that the basic attributes and configurations of path devices meet the requirements; shared consistency verification avoids multi-host shared access problems; and scheduling metadata verification ensures the correctness of path load balancing and scheduling policies. This multi-level verification mechanism can improve the accuracy and reliability of path restoration, and reduce the risk of data inconsistency or system failures caused by improper path restoration. For a target faulty path that fails to pass the multi-level verification, a path repair operation is performed. This enables the system to take timely repair measures when potential problems are found after path restoration, avoiding further failures or data inconsistencies caused by potential problems. The path repair operation can handle different types of verification failures in a targeted manner, improving the repair efficiency and accuracy, and reducing system maintenance costs and downtime. Through mechanisms such as real-time monitoring and detection of faulty path status, multi-level verification processing, path repair operations, and registration to the active path pool, the accuracy, reliability, and availability of the multi-path driver path restoration method can be improved, providing a more stable and efficient path restoration solution for the storage system.

[0040] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them: Figure 1 It is a schematic flow chart of a method for multi-path fault recovery provided by an embodiment of the present disclosure; Figure 2 It is a schematic flow chart of another method for multi-path fault recovery provided by an embodiment of the present disclosure; Figure 3 It is a schematic structural diagram of an apparatus for multi-path fault recovery provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The exemplary embodiments of the present disclosure will be described below in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.

[0043] The method, apparatus, electronic device, and storage medium for multi-path fault recovery according to the embodiments of the present disclosure will be described below with reference to the accompanying drawings.

[0044] Figure 1 It is a schematic flowchart of a method for multi-path fault recovery provided by an embodiment of the present disclosure.

[0045] As Figure 1 shown, the method includes the following steps: Step 101, obtain the status data of multiple storage devices associated with the fault paths in the target storage system, and detect whether the multiple fault paths have been restored to the normal state based on the status data.

[0046] In the embodiments of the present disclosure, the status data of the storage devices associated with multiple fault paths in the target storage system is obtained, and whether these fault paths have been restored to the normal state is detected based on the obtained status data. Specifically, the system communicates with each storage device in the target storage system through a preset interface or protocol to collect the status information directly related to the fault paths. These status information includes but is not limited to the online / offline status of the device, device capacity, device identifier (such as WWID), and device response status, etc. After collecting the status data, the system will execute a series of detection logics to determine whether the fault paths have returned to normal. This detection process may involve comparing and analyzing the status data and evaluating according to preset rules or thresholds. For example, the system may check whether the device has changed from the offline state to the online state, whether the device capacity is consistent with the expectation, and whether the device identifier remains the same, etc. If the detection result shows that the fault path has been restored to the normal state, the system will further execute the subsequent path recovery or verification process; if the detection result shows that the fault path has not been restored to the normal state, the system may continue to monitor the status of the path or execute corresponding fault handling operations according to the preset policy.

[0047] Through the above steps, the system can obtain the status information of the fault path in real time and accurately, and make corresponding judgments and processing based on this information. This not only improves the system's monitoring ability for the fault path, but also provides reliable data support for subsequent path recovery or fault handling. At the same time, as the starting point of the entire path recovery process, this step lays a foundation for the smooth execution of subsequent steps, thus ensuring the effectiveness and reliability of the entire multi-path drive path recovery method.

[0048] Step 102: When it is detected that the path status of the target fault path has been restored, perform multi-level verification processing on the target fault path. The verification types of the multi-level verification processing include: device attribute verification, shared consistency verification, and scheduling metadata verification.

[0049] In the embodiments of the present disclosure, if the system detects that the path status of the target fault path has been restored to the normal state, it enters Step 102 to perform multi-level verification processing on the target fault path. The multi-level verification processing aims to ensure that the restored path meets the system requirements in all aspects and avoid subsequent faults or data inconsistencies caused by potential problems. The verification types of the multi-level verification processing mainly include but are not limited to the following three aspects: The device attribute verification mainly performs basic attribute verification on the storage device associated with the restored path. The verification content includes but is not limited to whether the unique identifier of the device (such as WWID), device capacity, device model, etc. are consistent with the information preset or recorded in the system. Through the device attribute verification, it can be ensured that the device associated with the restored path meets the system requirements both physically and configurationally. The device attribute verification helps prevent path faults caused by incorrect device configuration or improper device replacement, and ensures that the storage devices used by the system have the correct attributes and configurations.

[0050] The shared consistency verification is mainly carried out for the storage volume involving multi-host shared access. The verification content aims to ensure that the restored path can maintain data consistency and access synchronization in a multi-host environment. This may involve verifying the registration key values, access permissions, locking mechanisms, etc. of the storage volume to prevent data conflicts or inconsistencies caused by simultaneous access by multiple hosts. The shared consistency verification helps avoid data inconsistency problems in the multi-host shared access scenario and ensures the integrity and consistency of the system data.

[0051] The scheduling metadata verification mainly verifies the metadata of the restored path in terms of load balancing and I / O scheduling. The verification content includes, but is not limited to, whether the priority information, block allocation information, port group information, etc. of the path are consistent with the information preset or recorded in the system. Through the scheduling metadata verification, it can be ensured that the restored path can work properly according to the system requirements in terms of load balancing and I / O scheduling. The scheduling metadata verification helps to prevent load imbalance or I / O scheduling anomalies caused by incorrect scheduling metadata, and ensures the reasonable allocation and efficient utilization of system resources.

[0052] In summary, the multi-level verification process in step 102 comprehensively verifies the restored target failed path through three aspects: device attribute verification, shared consistency verification, and scheduling metadata verification, ensuring that the path meets the system requirements in all aspects, thereby improving the stability and reliability of the system.

[0053] Step 103, if the target failed path fails the multi-level verification, perform a path repair operation on the target failed path that fails the multi-level verification based on the failed verification type.

[0054] In the embodiments of the present disclosure, if the multi-level verification process (including device attribute verification, shared consistency verification, and scheduling metadata verification) performed by the system on the target failed path shows that the path fails one or some verification types, the system will perform corresponding path repair operations on the target failed path that fails the multi-level verification based on the failed verification type. Specifically, the path repair operation will be customized according to the failed verification type. For the case where the device attribute verification fails: If the device attribute verification fails, it indicates that the storage device associated with the restored path does not meet the system requirements at the physical or configuration level. At this time, the system will check the attributes such as device identifier, capacity, model, etc., and try to reconfigure or replace the device to ensure that the device attributes conform to the information preset or recorded in the system. By repairing the device attributes, the system can ensure that the storage devices used meet the requirements at the physical and configuration levels, and avoid path failures caused by incorrect device configurations.

[0055] For the case where the shared consistency verification fails: If the shared consistency verification fails, it indicates that there are data inconsistency or access synchronization problems in the restored path in the multi-host shared access scenario. At this time, the system will check the registration key values, access permissions, locking mechanisms, etc. of the storage volume, and take corresponding measures (such as resynchronizing data, adjusting access permissions) to restore data consistency and access synchronization. By repairing the shared consistency, the system can avoid data conflicts or inconsistencies in the multi-host shared access scenario and ensure the integrity and consistency of system data.

[0056] For the case where the scheduling metadata verification fails: If the scheduling metadata verification fails, it indicates that there are abnormalities in the restored path in terms of load balancing and I / O scheduling. At this time, the system will check the priority information, block allocation information, port group information, etc. of the path, and adjust or reconfigure these metadata to ensure that the path can work properly in terms of load balancing and I / O scheduling according to the system requirements. By repairing the scheduling metadata, the system can prevent problems such as load imbalance or I / O scheduling abnormalities caused by scheduling metadata errors, and ensure the reasonable allocation and efficient utilization of system resources.

[0057] In summary, the path repair operation in step 103 is customized based on the failed verification type, aiming to solve the problems existing in the restored path in terms of device attributes, shared consistency, or scheduling metadata, so as to ensure that the path can work properly according to the system requirements and improve the stability and reliability of the system.

[0058] Step 104, register the repaired target failed path and the target failed path that has passed multi-level verification to the active path pool.

[0059] In the embodiments of the present disclosure, after completing the repair operation of the target failed path (for the path that fails to pass multi-level verification) and confirming that the target failed path has passed multi-level verification, these repaired or confirmed normal target failed paths are registered to the active path pool. The system adds the path information that meets the active state requirements (including but not limited to path identifiers, associated device information, verification passed status, etc.) to the active path pool through an internal management interface or protocol. The active path pool is a logical set used by the system to manage and schedule available storage paths, and it maintains the information of all storage paths that can be used normally in the current system.

[0060] By registering the repaired or confirmed normal paths to the active path pool, the system can ensure that these paths can be quickly and accurately allocated and used when needed, thereby improving the overall availability and response speed of the storage system.

[0061] The present disclosure provides a method for multi-path fault recovery. By detecting multi-paths, the present disclosure can monitor and detect the status of faulty paths in real time, so as to make timely responses when the paths are restored. When the path status of the target faulty path is restored, multi-level verification processing is performed on the target faulty path, including device attribute verification, shared consistency verification, and scheduling metadata verification. The multi-level verification processing mechanism can comprehensively and meticulously check whether there are potential problems in the restored path. Device attribute verification ensures that the basic attributes and configurations of path devices meet the requirements; shared consistency verification avoids multi-host shared access problems; and scheduling metadata verification ensures the correctness of path load balancing and scheduling policies. This multi-level verification mechanism can improve the accuracy and reliability of path restoration, and reduce the risk of data inconsistency or system failures caused by improper path restoration. For the target faulty path that fails to pass the multi-level verification, path repair operations are performed. This enables the system to take timely repair measures when potential problems are found after path restoration, and avoid further failures or data inconsistencies caused by potential problems. Path repair operations can handle different types of verification failures in a targeted manner, improve repair efficiency and accuracy, and reduce system maintenance costs and downtime. Through mechanisms such as real-time monitoring and detection of faulty path status, multi-level verification processing, path repair operations, and registration in the active path pool, the accuracy, reliability, and availability of the multi-path drive path recovery method can be improved, providing a more stable and efficient path recovery solution for the storage system.

[0062] To clearly illustrate the embodiments of the present disclosure, this embodiment provides a flowchart of another method for multi-path fault recovery.

[0063] As Figure 2 shown, the method includes the following steps: Step 201, obtain the status data of multiple storage devices associated with faulty paths in the target storage system. From the status data, extract the historical path status of the target faulty path, and determine whether the current path status of the target faulty path is consistent with the historical path status.

[0064] When the current path status is inconsistent with the historical path status, execute step 202; when the current path status is consistent with the historical path status, execute step 208.

[0065] Specifically in step 201, obtain the status data of the storage devices associated with multiple failed paths in the target storage system, and further extract the historical path status of the target failed path from these status data to determine whether the current path status is consistent with the historical path status. The system communicates with each storage device in the target storage system through a preset interface or protocol to collect the status information directly related to the failed path. These status information includes, but is not limited to, the online / offline status of the device, device capacity, device identifier (such as WWID), device response status, and path connection status, etc.

[0066] The system extracts the historical path status of the target failed path from the collected status data. The historical path status may be stored in the system log file, database, or a dedicated path status management module. The extracted historical path status includes, but is not limited to, the last successful connection time of the path, the last failure time, connection / disconnection records of the path, etc.

[0067] Judgment of the consistency between the current path status and the historical path status: The system compares the extracted historical path status with the currently collected path status to determine whether the current path status is consistent with the historical path status. The content of the comparison may include key information such as path connection status, device response status, and device identifier. If the current path status is consistent with the historical path status in all key information, the system considers that the current path status is consistent with the historical path status; if there is any inconsistency, the system considers that the current path status is inconsistent with the historical path status.

[0068] In step 202, based on the current path status, determine whether the target failed path is in an active state.

[0069] If it is in an active state, execute step 203.

[0070] Specifically in step 202, the current status data of the target failed path has been obtained through the foregoing steps. These data may include key information such as path connection status, device response status, and I / O operation status. The system presets a set of criteria for judging the path active state. These criteria may be based on the path connection status (such as whether the path has been successfully established), device response status (such as whether the device can normally respond to I / O requests), and possible business logic status (such as whether the path has been assigned to a certain business process for use).

[0071] Specifically, the system can check the "connection status" field in the path status data. If this field indicates that the path has been successfully established and the device response status is normal (such as no error I / O operation report), and may be combined with the business logic status (such as the path has been used by the business process), the system determines that the target failed path is currently in an active state.

[0072] Step 203: Perform device capacity verification on the associated storage device capacity of the target failure path, and perform trusted verification on the device identifier.

[0073] As a specific implementation manner of this embodiment, "performing device capacity verification on the associated storage device capacity of the target failure path and performing trusted verification on the device identifier" includes, but is not limited to: verifying whether the device capacity data is empty or inconsistent with the device capacity of the target storage system to perform device capacity verification; if the device capacity verification passes, then perform trusted verification on the device identifier.

[0074] Furthermore, when performing "trusted verification on the device identifier", it can be adopted, but is not limited to: verifying whether the device identifier is abnormal, and after determining that there is no abnormality, verifying whether the device identifier is consistent with the cached device identifier; if they are inconsistent, then determine that the trusted verification of the device identifier fails; if they are consistent, then determine that the trusted verification of the device identifier passes.

[0075] Specifically, in step 203, a verification operation is performed on the associated storage device capacity of the target failure path. During the verification process, the system will obtain the actual capacity data of the device and compare it with the device capacity previously recorded or expected in the target storage system. The specific verification content includes, but is not limited to: checking whether the device capacity data is empty (that is, no valid capacity information is obtained), and whether the device capacity is inconsistent with the device capacity record in the target storage system. If the device capacity data is empty or inconsistent with the record, it is determined that the device capacity verification fails.

[0076] Trusted verification of the device identifier: If the device capacity verification passes, the system will further perform trusted verification on the device identifier. The device identifier (such as WWID) is the unique identifier of the storage device in the system and is used to distinguish different storage devices. During the verification process, the system will first check whether the device identifier is abnormal (such as incorrect format, invalid characters, etc.). If there is an abnormality, it is directly determined that the trusted verification of the device identifier fails.

[0077] If the device identifier is normal, the system will further verify whether it is consistent with the device identifier saved in the cache. If they are inconsistent, it means that the device identifier may have been tampered with or there are other problems, and it is determined that the trusted verification of the device identifier fails; if they are consistent, it is determined that the trusted verification of the device identifier passes.

[0078] The system can integrate the device capacity verification and device identifier trusted verification module into the path recovery process. This module receives the capacity data and identifier information from the storage device or the target storage system and performs the above verification and validation operations.

[0079] To improve the efficiency and accuracy of checksum verification, the system can use hashing algorithms, checksums, or other data comparison techniques to quickly determine the consistency of device capacity and identifiers. Furthermore, the system can record the checksum verification results in a log file or database for subsequent auditing and troubleshooting.

[0080] When the device capacity data verification fails and / or the device identifier verification fails, step 204 is executed.

[0081] After the device capacity check passes and the device identifier verification passes, step 205 is executed.

[0082] Step 204 : Release the device binding relationship of the target fault path and set it to a failed state, and scan and add the target fault path in the failed state.

[0083] Specifically, in step 204, the storage device to which the target fault path is bound is identified and determined. By calling the corresponding management interface or sending a control command, the system releases the binding relationship between the target fault path and the storage device. This operation ensures that the path is no longer associated with a specific device, preparing for subsequent path recovery or reconfiguration. After the binding relationship is released, the system marks the status of the target fault path as "failed". This marking operation can be achieved by updating the path status database, modifying the path configuration file, or sending a status change notification. The marking of the failed status enables the system to identify and handle abnormal conditions of the path, avoiding attempts to use the path before it is fully restored, thereby preventing potential data errors or system failures.

[0084] The system initiates a path scanning mechanism, performing periodic or triggered scans of all paths in the storage system. During this scan, the system detects and identifies failed paths. Once a failed path is detected, the system triggers the path addition process, re-adding the path to the path management module. During this addition process, the system reconfigures path parameters, binds storage devices (if applicable), and updates the path status to "available" or another appropriate state.

[0085] Step 205 : Determine whether the historical path status of the target fault path is in an inactive state. If it is in an inactive state, increase the number of active paths by one.

[0086] Specifically in step 205, the system retrieves the historical path status information of the target fault path from the path status management module or the relevant database. The historical path status information usually records the status of the path at a certain past time point, including whether it is active, connection status, fault occurrence time, etc. The system parses the retrieved historical path status information to determine whether the target fault path was in an inactive state in the historical record. An inactive state may mean that the path has had a fault in the past, was manually disabled, or was not used for other reasons. The judgment logic can be based on status flag bits, status codes, or specific status description texts. For example, if the historical status information contains flags such as "inactive", "disabled", or similar, the system determines that the path was in an inactive state in the past.

[0087] If the judgment result is that the historical path status of the target fault path is in an inactive state, the system will perform an increment operation on the number of active paths. Specifically, the system will access a counter or variable used to count the number of active paths and increment its value by one. This operation reflects that the system has successfully restored a path that was previously in an inactive state, thus increasing the current number of available active paths.

[0088] In step 206, it is detected whether the registration key value of the target fault path recorded on the storage controller side matches the stored target key value.

[0089] As a specific implementation manner of this embodiment, "detecting whether the registration key value of the target fault path recorded on the storage controller side matches the stored target key value" can be verified and matched in the following ways but is not limited to them: verifying whether the registration key value is empty; in the case where the registration key value is not empty, verifying whether the registration key value is consistent with the target key value.

[0090] When it is determined that the registration key value is inconsistent with the target key value, step 207 is executed.

[0091] Specifically in step 206, a specific query command is sent through the communication interface with the storage controller (such as protocol interfaces like SCSI, iSCSI, FC, etc.) to obtain the registration key value recorded at the storage controller end for the target failed path. The registration key value is usually a unique identifier used to distinguish different paths or devices in the storage system. The system reads the pre-stored target key value from local storage or a configuration file. This target key value is set during path configuration or initialization and is used to compare with the registration key value at the storage controller end. The system first verifies whether the registration key value obtained from the storage controller end is empty. If the registration key value is empty, it may mean that the storage controller end has not correctly recorded the registration information for this path, or the path information has been deleted or damaged. If the registration key value is not empty, the system further verifies whether this registration key value is consistent with the pre-stored target key value. This verification process can be achieved through techniques such as string comparison, hash value comparison, or other data comparison techniques.

[0092] Step 207: Re-register the target failed path based on the target key value.

[0093] Specifically in step 207, the validity of the obtained target key value is confirmed. The target key value is determined in step 206 by verifying the matching of the registration key value recorded at the storage controller end with the pre-stored target key value. It serves as the unique identifier of the path in the storage system and is used for subsequent path registration operations. The system prepares the information required to re-register the target failed path based on the target key value. This information may include, but is not limited to, path identifiers, associated storage device information, path configuration parameters (such as priority, bandwidth allocation, etc.), and any necessary authentication or authorization information. The system initiates a re-registration request for the target failed path by communicating with the management interface or protocol of the storage system. In the request, the system will include the target key value and the above-prepared path registration information.

[0094] After receiving the re-registration request, the storage system will verify the validity of the target key value and update or create the corresponding path record according to the path registration information in the request. If the verification passes and the information is complete, the storage system will successfully register the target failed path and return a registration success response.

[0095] Step 208: Determine the ownership type of the volume associated with the target failed path; based on the ownership type, perform scheduling metadata verification and maintenance on the target failed path.

[0096] As a specific implementation manner of this embodiment, "scheduling metadata verification and maintenance for the target failure path based on the attribution type" includes, but is not limited to: if it is a volume without attribution, verify the block priority and port group information of the target failure path against the values saved in the kernel state; update the inconsistent block priority and / or port group information according to the values saved in the kernel state; if it is a volume with attribution, verify the block priority of the target failure path against the values saved in the kernel state; update the inconsistent block priority according to the values saved in the kernel state.

[0097] Specifically in step 208, by querying the relevant information in the storage system or the path management module, determine the attribution type of the volume associated with the target failure path. The attribution type is generally divided into two types: "volume without attribution" and "volume with attribution". A volume without attribution means that the volume is not exclusively occupied by a specific application or service, while a volume with attribution means that the volume has been clearly bound to a certain application or service. Scheduling metadata verification and maintenance in the case of a volume without attribution: If the judgment result is that the volume associated with the target failure path is a volume without attribution, the system will further verify the block priority and port group information of the target failure path. The verification process involves comparing the block priority and port group information in the current path configuration with the corresponding values saved in the kernel state. The values saved in the kernel state are usually the latest and most accurate path configuration information recorded and maintained by the system kernel during system initialization or path configuration update. If it is found that the block priority and / or port group information is inconsistent during verification, the system will update this information according to the values saved in the kernel state to ensure the consistency and accuracy of the path configuration.

[0098] Scheduling metadata verification and maintenance in the case of a volume with attribution: If the judgment result is that the volume associated with the target failure path is a volume with attribution, the system will only verify the block priority of the target failure path. The verification method is similar to that in the case of a volume without attribution, that is, comparing the block priority in the current path configuration with the values saved in the kernel state. If it is found that the block priority is inconsistent during verification, the system will update it according to the values saved in the kernel state to ensure that the path configuration meets the latest requirements of the kernel state.

[0099] Step 209, register the repaired target failure path and the target failure path that has passed multi-level verification into the active path pool.

[0100] Specifically in step 209, confirm that the repair operation of the target failure path has been completed. The repair operation may include, but is not limited to: hardware failure replacement, software configuration adjustment, network connection restoration, etc., and is specifically executed according to the failure type and repair plan detected in steps 201 to 208. The basis for confirming the completion of the repair can be the path status change automatically detected by the system, the manual confirmation instruction of the administrator, or the preset completion flag in the repair process.

[0101] Multi-level verification confirmation: The system verifies that the target fault path has passed all pre-defined multi-level verification processes. These verification processes may include, but are not limited to, path status verification (as described in step 201), device capacity verification and identifier verification (as described in step 203), path binding removal and re-registration (as described in steps 204 and 207), and scheduling metadata verification and maintenance (as described in step 208). Passing multi-level verification means that the path meets the system's required availability and stability standards across multiple dimensions.

[0102] Registering with the active path pool: After confirming that the target faulty path has been repaired and passed multiple levels of verification, the system registers the path with the active path pool. The active path pool is a collection of currently available path information in the system for use in subsequent I / O operation scheduling. Registration operations may involve updating the path status database, modifying the path configuration file, or sending a registration request to the path management module through a specific management interface. After successful registration, the path is marked as "active" and becomes available for system scheduling.

[0103] As an implementable manner of this embodiment, before performing scheduling metadata verification on the target fault path, the multipath fault recovery method further includes: determining whether the current path state of the target fault path is an active state; When the target fault path is determined to be in the activated state, the detection interval of the target fault path is incremented by one; it is determined whether the target fault path is currently in the early stage of path recovery. If it is in the early stage of path recovery, a scheduling metadata check is performed on the target fault path.

[0104] Specifically, in the multipath fault recovery method, to ensure the rationality and efficiency of scheduling metadata verification, a series of pre-judgments and operations are required before executing scheduling metadata verification. The specific contents are as follows: Step 1: Determine whether the current path status of the target fault path is active.

[0105] The system obtains the current path status information of the target fault path by accessing the path status management module or calling related interfaces. Path status information is usually stored in the form of a specific status identifier or status code, such as "1" for active status and "0" for inactive status.

[0106] The system compares the acquired path status information with the preset activation status identifier to determine whether the target fault path is currently in the activation state.

[0107] Step 2: When the state is determined to be active, the detection interval of the target fault path is increased by one.

[0108] If, after the judgment in Step 1, it is determined that the current path state of the target fault path is the active state, the system will perform an operation to increase the detection interval. The detection interval refers to the time interval at which the system periodically detects the target fault path, usually measured in time units such as seconds and minutes. The system accesses the variable or configuration item storing the detection interval of the target fault path and increments its value by one. For example, if the current detection interval is 5 seconds, it becomes 6 seconds after incrementing.

[0109] Step 3: Determine whether the target fault path is currently in the initial stage of path recovery. If it is in the initial stage of path recovery, perform scheduling metadata verification on the target fault path.

[0110] The system determines whether the target fault path is currently in the initial stage of path recovery according to the preset initial path recovery judgment criteria. The initial path recovery judgment criteria can be based on time factors. For example, starting from the time when the path fails, it is determined to be in the initial stage of path recovery within a certain time threshold (such as 30 seconds); it can also be based on event factors. For example, within a certain period of time after completing specific repair operations (such as hardware replacement, network connection recovery, etc.) it is determined to be in the initial stage of path recovery. If it is determined that the target fault path is currently in the initial stage of path recovery, the system will trigger the scheduling metadata verification process to verify the scheduling metadata of the target fault path to ensure the accuracy and consistency of the scheduling metadata.

[0111] As an implementable manner of this embodiment, determining whether the target fault path is currently in the initial stage of path recovery includes: comparing the detection duration of the target fault path with the initial recovery threshold, where the initial recovery threshold is the product of the detection interval and the proportionality coefficient; when the detection duration does not reach the initial recovery threshold, it is determined to be in the initial stage of path recovery; when the detection duration reaches the initial recovery threshold, it is determined not to be in the initial stage of path recovery.

[0112] Specifically, obtain the detection interval of the target fault path. The detection interval is the time interval preset by the system for periodically detecting the status of the target fault path, usually expressed in time units such as seconds and milliseconds. For example, the detection interval may be set to 5 seconds. The system also obtains a preset proportionality coefficient. The proportionality coefficient is an empirical value or a value set according to the system performance requirements, used to adjust the size of the initial recovery threshold. The value range of the proportionality coefficient can be determined according to the actual situation. For example, it is between 0.1 and 1, or may be greater than 1.

[0113] The system calculates the initial recovery threshold by multiplying the detection interval by the proportionality coefficient. The calculation formula is: initial recovery threshold = detection interval × proportionality coefficient. For example, if the detection interval is 5 seconds and the proportionality coefficient is 4, the initial recovery threshold is 20 seconds.

[0114] The system records the time point when the target fault path starts to be detected, usually when a fault occurs in the target fault path and the recovery process is initiated. In the subsequent judgment process, the system obtains the current time and calculates the time difference between the current time and the time point when the target fault path starts to be detected. This time difference is the detection duration. For example, if the time point when the target fault path starts to be detected is 10:00:00 and the current time is 10:00:02, the detection duration is 2 seconds.

[0115] As an implementable way of this embodiment, after determining whether the current path state of the target fault path is consistent with the historical path state, the multi-path fault recovery method further includes: setting the detection interval of the target fault path to a preset detection interval.

[0116] Specifically, in the process flow of the multi-path fault recovery method, determining whether the current path state of the target fault path is consistent with the historical path state is a key link. After that, performing the operation of setting the detection interval of the target fault path to a preset detection interval is of great significance for the stable operation and efficient recovery of the system.

[0117] The system obtains the current path state information of the target fault path by accessing the path state management module or the relevant data storage area. The path state information usually includes the connection state of the path (such as connected, not connected), the activation state (such as activated, not activated), the fault state (such as normal, faulty), etc. At the same time, the system retrieves the historical path state information of the target fault path from the historical records. The historical path state information records the state of the path at a certain time in the past and can be used for comparison with the current state. The system compares the current path state information with the historical path state information item by item to determine whether they are consistent. If all the key state items are the same, it is determined that the states are consistent; if there is any key state item that is different, it is determined that the states are inconsistent.

[0118] The preset detection interval is a time value preset by the system, which is used to specify the time interval for regularly detecting the target fault path. This preset detection interval can be set according to factors such as the system's performance requirements, the characteristics of the storage device, and business requirements. For example, for services with high real-time requirements, the preset detection interval can be set shorter, such as 1 second; while for services with relatively low real-time requirements, the preset detection interval can be set longer, such as 10 seconds. After completing the consistency judgment between the current path state and the historical path state, regardless of whether the judgment result is consistent or inconsistent, the system sets the detection interval of the target fault path to the preset detection interval. Specifically, the system will access the variable or configuration item storing the detection interval of the target fault path and modify its value to the value of the preset detection interval. For example, if the preset detection interval is 5 seconds, the system will modify the value of this variable or configuration item from the original value (such as 3 seconds or 8 seconds) to 5 seconds.

[0119] It should be noted that the embodiments of the present disclosure may include multiple steps. For the convenience of description, these steps are numbered, but these labels are not intended to limit the execution time slots and execution orders between the steps; these steps can be implemented in any order, and the embodiments of the present disclosure do not make any limitations in this regard.

[0120] Corresponding to the above multi-path fault recovery method, the present disclosure also proposes a multi-path fault recovery device. Since the device embodiments of the present disclosure correspond to the above method embodiments, the details not disclosed in the device embodiments can be referred to the above method embodiments, and will not be elaborated in the present disclosure.

[0121] Figure 3 FIG. is a schematic structural diagram of a multi-path fault recovery device provided by an embodiment of the present disclosure, as Figure 3 shown, including: A detection unit 31, configured to obtain status data of multiple fault path associated storage devices in the target storage system, and detect whether the multiple fault paths are restored to a normal state based on the status data; A verification unit 32, configured to perform multi-level verification processing on the target fault path when it is detected that the path state of the target fault path is restored. The verification types of the multi-level verification processing include: device attribute verification, shared consistency verification, and scheduling metadata verification; A repair unit 33, configured to perform a path repair operation on the target fault path that fails to pass the multi-level verification based on the verification type that fails to pass if the target fault path fails to pass the multi-level verification; A registration unit 34, configured to register the repaired target fault path and the target fault path that passes the multi-level verification into the active path pool.

[0122] The present disclosure provides an apparatus for multi - path fault recovery. By detecting multi - paths, the present disclosure can monitor and detect the status of faulty paths in real time, so as to make a timely response when the paths are restored. When the path status of the target faulty path is detected to be restored, multi - level verification processing is performed on the target faulty path, including device attribute verification, shared consistency verification, and scheduling metadata verification. The multi - level verification processing mechanism can comprehensively and meticulously check whether there are potential problems in the restored path. Device attribute verification ensures that the basic attributes and configurations of the path devices meet the requirements; shared consistency verification avoids multi - host shared access problems; and scheduling metadata verification ensures the correctness of the path load balancing and scheduling policies. This multi - level verification mechanism can improve the accuracy and reliability of path restoration, and reduce the risk of data inconsistency or system failures caused by improper path restoration. For the target faulty path that fails to pass the multi - level verification, path repair operations are performed. This enables the system to take timely repair measures when potential problems are found after path restoration, avoiding further failures or data inconsistencies caused by potential problems. Path repair operations can handle different types of verification failures in a targeted manner, improving the repair efficiency and accuracy, and reducing system maintenance costs and downtime. Through mechanisms such as real - time monitoring and detection of faulty path status, multi - level verification processing, path repair operations, and registration to the active path pool, the accuracy, reliability, and availability of the multi - path driver path restoration method can be improved, providing a more stable and efficient path restoration solution for the storage system.

[0123] It should be noted that the foregoing explanation of the method embodiments also applies to the apparatus of this embodiment, with the same principle, and will not be further limited in this embodiment.

[0124] For the description of the features in the corresponding embodiment of the multi - path fault recovery apparatus, reference can be made to the relevant description in the corresponding embodiment of the multi - path fault recovery method, which will not be elaborated here one by one.

[0125] An embodiment of the present application also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above - mentioned method embodiments of multi - path fault recovery.

[0126] An embodiment of the present application also provides a computer - readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above - mentioned method embodiments of multi - path fault recovery when running.

[0127] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disks, magnetic disks, or optical discs that can store computer programs.

[0128] An embodiment of the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the method embodiments of the above multi-path fault recovery.

[0129] Another embodiment of the present application also provides a computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the method embodiments of the above multi-path fault recovery.

[0130] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0131] The above has introduced in detail a multi-path fault recovery method, apparatus, electronic device, and storage medium provided by the present application. Specific examples are used herein to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for multi-path fault recovery, characterized in that, include: Acquire status data of storage devices associated with multiple faulty paths in the target storage system, and detect whether the multiple faulty paths have recovered to a normal state based on the status data; When detecting that the path state of the target fault path has recovered, performing a multi-level verification process on the target fault path, wherein the verification types of the multi-level verification process include: device attribute verification, shared consistency verification, and scheduling metadata verification; If the target fault path fails the multi-level verification, performing a path repair operation on the target fault path that fails the multi-level verification based on the verification type that fails; Register the repaired target fault path and the target fault path that passes multi-level verification into the active path pool.

2. The method for multipath fault recovery according to claim 1, characterized in that Before performing multi-level verification processing on the target fault path, the method further includes: Extracting a historical path state of the target fault path from the state data, and determining whether the current path state of the target fault path is consistent with the historical path state; When the current path state is inconsistent with the historical path state, judging whether the target fault path is in an active state based on the current path state; If it is in the activated state, a device attribute check and a shared consistency check are performed on the target fault path.

3. The method for multi-path fault recovery according to claim 2, wherein The performing of device attribute verification and shared consistency verification on the target fault path includes: Performing device capacity verification on the capacity of the storage device associated with the target fault path and performing credibility verification on the device identifier; Detect whether the registration key value of the target fault path recorded by the storage controller matches the stored target key value.

4. The method for multi-path fault recovery according to claim 3, wherein The performing device capacity verification on the storage device capacity associated with the target fault path and performing trustworthy verification on the device identifier includes: Verify whether the device capacity data is empty or inconsistent with the device capacity of the target storage system, so as to perform device capacity verification; If the device capacity verification passes, the device identifier is authenticated.

5. The method for multi-path fault recovery according to claim 4, wherein The performing trustworthy verification on the device identifier includes: Verifying whether there is an anomaly in the device identifier, and after determining that there is no anomaly, verifying whether the device identifier is consistent with a cached device identifier; If they are inconsistent, it is determined that the device identifier trust verification has failed; if they are consistent, it is determined that the device identifier trust verification has passed.

6. The method for multi-path fault recovery according to claim 4, wherein When the device capacity data verification fails and / or the device identifier verification fails, performing a path repair operation on the faulty path that fails the multi-level verification based on the failed verification type includes: The device binding relationship of the target fault path is released and the target fault path is set to a failed state, and the target fault path in the failed state is scanned and added.

7. The method for multi-path fault recovery according to claim 4, wherein After the device capacity check passes and the device identifier verification passes, the method further includes: It is determined whether the historical path state of the target fault path is in an inactive state. If it is in an inactive state, the number of active paths is increased by one.

8. The method for multipath fault recovery according to claim 3, wherein The detecting whether the registration key value of the target fault path recorded by the storage controller matches the stored target key value includes: Verify whether the registration key value is empty; When the registered key value is not empty, verify whether the registered key value is consistent with the target key value.

9. The method for multi-path fault recovery according to claim 8, wherein The performing a path repair operation on a target failure path that fails a multi-level check based on the failed check type further includes: When it is determined that the registered key value is consistent with the target key value, re-register the target failure path based on the target key value.

10. The method for multi-path fault recovery according to claim 9, wherein, The method further includes: After performing a shared consistency check on the target failure path, perform a scheduling metadata check on the target failure path.

11. The method for multipath fault recovery according to claim 2, wherein The method further includes: If the current path state is consistent with the historical path state, perform a scheduling metadata check on the target failure path.

12. The method for multipath fault recovery according to any one of claims 10 or 11, characterized in that, The performing a scheduling metadata check on the target failure path includes: Judging the ownership type of the volume associated with the target failure path; Based on the ownership type, perform scheduling metadata check maintenance on the target failure path.

13. The method for multi-path fault recovery according to claim 12, characterized in that, The performing scheduling metadata check maintenance on the target failure path based on the ownership type includes: If it is a volume without ownership, verify the block priority and port group information of the target failure path with the kernel state saved value; update the inconsistent block priority and / or port group information according to the kernel state saved value; If it is a volume with ownership, verify the block priority of the target failure path with the kernel state saved value; update the inconsistent block priority according to the kernel state saved value.

14. The method for multi-path fault recovery according to any one of claims 10 or 11, characterized in that, Before performing a scheduling metadata check on the target failure path, the method further includes: Judging whether the current path state of the target failure path is an active state; When it is determined to be in an active state, increment the detection interval of the target failure path by one; Judge whether the target failure path is currently in the initial stage of path recovery. If it is in the initial stage of path recovery, perform a scheduling metadata check on the target failure path.

15. The method for multi-path fault recovery according to claim 14, characterized in that, The judging whether the target failure path is currently in the initial stage of path recovery includes: Compare the detection duration of the target failure path with a recovery initial threshold, where the recovery initial threshold is the product of the detection interval and a proportionality coefficient; When the detection duration does not reach the recovery initial threshold, it is determined to be in the initial stage of path recovery; When the detection duration reaches the recovery initial threshold, it is determined not to be in the initial stage of path recovery.

16. The method for multipath fault recovery according to claim 2, wherein After judging whether the current path state of the target failure path is consistent with the historical path state, the method further includes: Set the detection interval of the target failure path to a preset detection interval.

17. A device for multi-path fault recovery, characterized in that, Includes: A detection unit, configured to obtain status data of multiple failure path associated storage devices in a target storage system, and detect whether the multiple failure paths are restored to a normal state based on the status data; A verification unit, configured to perform a multi-level verification process on the target failure path when it is detected that the path state of the target failure path is restored, and the verification types of the multi-level verification process include: device attribute verification, shared consistency verification, scheduling metadata verification; A repair unit, configured to, if the target fault path fails to pass multi-level verification, perform a path repair operation on the target fault path that fails to pass multi-level verification based on the verification type that fails to pass. A registration unit, configured to register the repaired target fault path and the target fault path that passes multi-level verification into the active path pool.

18. An electronic device, characterized in that, Comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the multi-path fault recovery method according to any one of claims 1-16.

19. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the multi-path fault recovery method according to any one of claims 1-16.

20. A computer program product, characterized in that, Comprising a computer program, which when executed by a processor implements the multi-path fault recovery method according to any one of claims 1-16.

Citation Information

Patent Citations

  • Method and device for processing multipath storage faults

    CN106775487A

  • Multi-path anomaly detection and repair method, device, equipment and medium

    CN115686921A

  • Multi-path exception processing method and device, computer equipment and storage medium

    CN117573405A

Cited By

  • System task exception processing method and device

    CN121326531A