Abnormality processing method and device, storage medium and electronic equipment

By automating the acquisition of storage system path status information, identifying anomaly types, and matching repair strategies, the problem of time-consuming, labor-intensive, and misjudgment-prone storage system anomaly handling has been solved, achieving efficient and accurate automatic repair and improving system stability and user experience.

CN120407262BActive Publication Date: 2026-03-03INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510855487.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2026-03-03
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

In existing technologies, storage system anomaly handling relies on manual diagnosis and repair, which is time-consuming, labor-intensive, and prone to misjudgment and omission, affecting the continuity and reliability of user business.

Method used

By acquiring the path status information of multiple storage volumes in the storage system, the system automatically identifies the anomaly type and matches it with a preset repair strategy library to achieve automated repair processing.

Benefits of technology

It improves the efficiency and accuracy of anomaly handling, reduces manual intervention, enhances the stability and operational efficiency of the storage system, and ensures the continuity and security of data access.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407262B_ABST
    Figure CN120407262B_ABST
Patent Text Reader

Abstract

The application discloses an exception processing method and device, a storage medium and an electronic device, and relates to the technical field of data storage. The method comprises the following steps: firstly, acquiring path state information corresponding to a plurality of storage volumes in a storage system; secondly, determining an exception type of the storage system according to the path state information; thirdly, matching the exception type with a repair strategy in a preset repair strategy library corresponding to the storage system; and fourthly, performing repair processing on the storage system based on the matched repair strategy. Compared with the prior art, the application effectively reduces manual intervention, improves the efficiency and accuracy of exception processing, avoids the business impact caused by misjudgment and missed judgment, has good user interaction and system adaptability, and significantly improves the stability and operation and maintenance efficiency of the storage system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data storage technology, and in particular to an anomaly handling method, apparatus, storage medium, and electronic device. Background Technology

[0002] With the rapid development of informatization and digitalization, storage systems have become the core medium for data storage in various industries. Storage multipathing technology is widely used to improve the reliability and performance of storage systems. By configuring multiple paths for storage devices, the system can automatically switch to another path when one path fails, thereby ensuring the continuity and reliability of data access.

[0003] Currently, storage systems may encounter various anomalies during operation, and these problems are typically diagnosed and repaired manually. However, this approach is not only time-consuming and labor-intensive, but also prone to misdiagnosis and missed diagnoses, impacting user business operations. Summary of the Invention

[0004] This disclosure provides an anomaly handling method, apparatus, storage medium, and electronic device. Its main purpose is to address the problems in related technologies where manual repair of system anomalies is not only time-consuming and labor-intensive, but also prone to misjudgments and omissions, thus impacting user services.

[0005] Firstly, this application provides an exception handling method, including:

[0006] Obtain path status information corresponding to multiple storage volumes in the storage system;

[0007] Based on the path status information, determine the anomaly type of the storage system;

[0008] Match the anomaly type with the repair strategies in the preset repair strategy library corresponding to the storage system;

[0009] The storage system is repaired based on the matched repair strategy.

[0010] Secondly, this application provides an anomaly handling apparatus, comprising:

[0011] The acquisition module is configured to acquire path status information corresponding to multiple storage volumes in the storage system.

[0012] The determination module is configured to determine the anomaly type of the storage system based on the path status information;

[0013] The matching module is configured to match the exception type with the repair strategies in the preset repair strategy library corresponding to the storage system;

[0014] The processing module is configured to perform repair processing on the storage system based on the matched repair strategy.

[0015] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method of the first aspect.

[0016] Fourthly, this application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the computer program to implement the method of the first aspect.

[0017] Fifthly, this application provides a computer program product having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method of the first aspect.

[0018] The present disclosure provides an anomaly handling method, apparatus, storage medium, and electronic device. The method includes: first, acquiring path status information corresponding to multiple storage volumes in a storage system; then, determining the anomaly type of the storage system based on the path status information; next, matching the anomaly type with repair strategies in a preset repair strategy library corresponding to the storage system; and finally, repairing the storage system based on the matched repair strategy. Compared with current related technologies, this application automates the acquisition of path status information of each storage volume in the storage system, determines and accurately locates the anomaly type, and combines it with a preset repair strategy library to achieve intelligent matching and automatic repair. This effectively reduces manual intervention, improves the efficiency and accuracy of anomaly handling, avoids business impacts caused by misjudgments and omissions, and also possesses good user interactivity and system adaptability, significantly improving the stability and operational efficiency of the storage system.

[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0020] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 A flowchart illustrating an exception handling method provided in an embodiment of this application is shown.

[0022] Figure 2 A schematic diagram illustrating an example provided in an embodiment of this application is shown;

[0023] Figure 3 A flowchart illustrating another exception handling method provided in an embodiment of this application is shown;

[0024] Figure 4 A flowchart illustrating an example provided in an embodiment of this application is shown;

[0025] Figure 5 A flowchart illustrating an example provided in an embodiment of this application is shown;

[0026] Figure 6 A schematic diagram of an exception handling device provided in an embodiment of this application is shown. Detailed Implementation

[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0028] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0029] In today's global information and digitalization context, human society's data is growing ever larger, giving rise to storage systems that have become the primary medium for backend data storage across various industries. As the demand for 24 / 7 data access becomes the norm, storage systems have gradually evolved through stages of demand for capacity, performance, and reliability, enabling them to store user data sufficiently (adequate space), access it quickly (sufficient performance), and maintain stable access (no data loss, reliable storage process capable of handling general failures). Storage multipathing technology is widely used to improve the reliability and performance of storage systems. By configuring multiple paths for storage devices, the system can automatically switch to another path when one path fails, thereby ensuring the continuity and reliability of data access.

[0030] When business users discover anomalies in their applications, such as performance degradation or read / write errors, they typically contact the administrator for troubleshooting. After logging into the environment, the administrator diagnoses the problem by checking the storage system's multipath environment and log information. However, this process often begins only after the business has already been affected, resulting in a time lag between problem discovery and resolution, during which the business remains in an abnormal state. For example, first, a business user discovers the anomaly and contacts the administrator; then, the administrator logs into the system, checks the storage multipath environment and log information to determine the cause of the problem; finally, based on the diagnostic results, repairs are performed until the problem is resolved and the business returns to normal operation. While this process ultimately resolves the issue, its reactive nature means that the business continues to be affected until the problem is fixed, impacting user experience and business continuity.

[0031] Based on the various anomalies that may be encountered during the operation of the aforementioned multi-path storage system, and to address the technical problems of manual system anomaly repair methods in related technologies being time-consuming, labor-intensive, and prone to misjudgments and omissions, thus impacting user services, this embodiment provides an anomaly handling method, such as... Figure 1 As shown, the method includes the following steps:

[0032] Step 101: Obtain the path status information corresponding to multiple storage volumes in the storage system.

[0033] In some examples, multiple storage volumes are created by logically partitioning the physical storage resources of a storage system into multiple independent storage units. Each storage volume can be configured, accessed, and managed independently. These storage volumes can provide customized storage solutions for different applications or services, support data read and write operations, and can perform management activities such as capacity adjustment, performance optimization, and fault recovery as needed to meet the data storage requirements of different business scenarios.

[0034] For example, path status information may include key indicators such as the operating status, connectivity, I / O response latency, error frequency, and path switching records of each data path in the storage system. These indicators are used to reflect the health status and performance of the data access path in real time, providing an important basis for anomaly detection and fault diagnosis.

[0035] For example, business application software services on host A access data through volumes provided by the storage system, utilizing the storage vendor's proprietary multipathing software to achieve load balancing and fault redundancy. Figure 2As shown, host A connects to switches A and B via a Fibre Channel Host Bus Adapter (FC HBA) (containing Port 1 and Port 2), and then connects to multiple ports (Port A, Port B, Port C, Port D) of the storage system, ultimately accessing volumes 1 to N in the storage system. Storage multipathing serves as a crucial bridge connecting front-end user software and back-end storage volumes. Failures in path status or on the link devices directly impact the input / output (I / O) read / write performance of the front-end host software, leading to service interruptions or performance degradation. Therefore, timely detection, diagnosis, and repair of multipath anomalies are essential to ensure the stable operation and high availability of user services. Figure 2 It can be seen that multipath software plays a core role in monitoring and managing the status of these paths. By detecting and automatically switching paths in real time, it effectively avoids the impact of single path failures on business, thereby ensuring business continuity and data security.

[0036] Step 102: Determine the exception type of the storage system based on the path status information.

[0037] In some examples, based on key indicators such as the operating status, connectivity, I / O response latency, error frequency, and path switching records of each data path in the storage system, the system can determine in real time whether there are any abnormalities in each storage path in the storage system, such as path interruption, excessive I / O latency, frequent errors, or abnormal path switching, thereby promptly identifying potential faults and providing a basis for subsequent diagnosis and repair.

[0038] For example, the types of anomalies in a storage system include, but are not limited to, hard disk failure, path failure, excessively high data access latency, frequent I / O errors, insufficient storage capacity, data corruption or loss, performance bottlenecks, and configuration errors. These anomalies may affect the normal operation of the system and the security of the data, either individually or in combination.

[0039] Step 103: Match the exception type with the repair strategy in the preset repair strategy library corresponding to the storage system.

[0040] For example, the repair strategy library may include standard handling solutions for various storage path anomalies, such as automatically switching to a backup path, reactivating a failed path, adjusting the I / O queue depth, or triggering alarms to notify operations and maintenance personnel to intervene.

[0041] In some examples, by precisely matching the exception type with the corresponding repair strategy, the system can quickly execute the corresponding repair actions, thereby effectively restoring the normal operation of the storage path and improving the stability and availability of the storage system.

[0042] Step 104: Based on the matched repair strategy, perform repair processing on the storage system.

[0043] In some examples, the system can perform repair operations according to preset steps in the repair strategy, which may include, but are not limited to, replacing the faulty hard drive, reconfiguring the storage path, optimizing I / O scheduling to reduce latency, and correcting configuration errors.

[0044] For example, during the repair process, a detailed processing strategy report can also be generated, recording all operational details and changes. This report will then be submitted to the management platform for administrator review and follow-up, ensuring the stable operation of the storage system and data security.

[0045] Compared with related technologies, this embodiment first obtains the path status information corresponding to multiple storage volumes in the storage system; then, based on the path status information, it determines the anomaly type of the storage system; next, it matches the anomaly type with the repair strategies in the preset repair strategy library corresponding to the storage system; and finally, it repairs the storage system based on the matched repair strategy. Compared with current related technologies, this embodiment, based on the path status information, can automatically determine the anomaly type of the storage system and accurately identify the specific anomaly type based on anomaly characteristics, such as hard disk failure, path interruption, or performance bottleneck, and intelligently match the identified anomaly type with the repair solutions in the preset repair strategy library, improving the accuracy and efficiency of fault response. Then, based on the matched repair strategy, it performs automated repair processing on the storage system, which not only effectively shortens the fault response time but also reduces the need for manual intervention and maintenance costs. It achieves closed-loop management of anomaly detection, diagnosis, matching, and repair, significantly improving the stability, reliability, and intelligent operation and maintenance level of the storage system, and ensuring the continuity and security of data access.

[0046] As a refinement of this embodiment, the storage system can be repaired using, but is not limited to, the following methods: Figure 3 As shown, Figure 3 This is a flowchart illustrating an exception handling method provided in an embodiment of the present disclosure, including:

[0047] Step 201: Obtain the path status information corresponding to multiple storage volumes in the storage system.

[0048] The path status information includes multipath status, number of paths, device status corresponding to multiple storage volumes, and input / output performance metrics.

[0049] In some examples, multipathing software can monitor the multipath status of each storage volume in a storage system in real time. This not only tracks the status of each path (e.g., whether it's active, failed, or recovering), but also monitors changes in the number of paths to ensure data transfer redundancy and reliability. Simultaneously, the software checks the overall health of each storage device, including but not limited to its online status and fault conditions. Furthermore, multipathing software collects and analyzes I / O performance metrics such as response time, throughput, and input / output operations per second (IOPS) to promptly identify potential bottlenecks or abnormal behavior and make corresponding adjustments, thereby optimizing the overall performance and stability of the storage system.

[0050] Step 202: Based on the path status information, determine whether there are any abnormalities in the storage path in the storage system.

[0051] For example, multipath software can determine in real time whether there are any anomalies in the storage paths of the storage system, such as path disconnection, excessive latency, or communication failure, based on path status information. When an anomaly is detected in the multipath, the system immediately captures the relevant anomaly information and performs in-depth analysis and diagnosis. This process includes a comprehensive evaluation of key indicators such as path switching records, I / O error logs, and device response time to quickly locate the cause of the fault and determine whether it is due to a physical connection problem, storage device failure, or misconfiguration.

[0052] Optionally, step 202 may specifically include: detecting the storage system based on the multipath status, the number of paths, the device status corresponding to multiple storage volumes, and input / output performance indicators; and determining the anomaly type of the storage system based on the detection results.

[0053] In some examples, multipathing software, as a key component in managing the multipathing devices corresponding to storage volumes, has the ability to monitor the underlying storage paths in real time. When the status of each path (such as online, offline, fault recovery, etc.) or the number of paths changes, the multipathing software can capture these changes immediately, ensuring that the system responds promptly and adjusts the path usage strategy accordingly.

[0054] For example, in the event of I / O performance degradation (such as increased latency) or I / O errors, the multipath software can quickly detect and record relevant anomaly information, including path-level I / O latency metrics and error logs. This efficient monitoring and capture mechanism provides solid data support for subsequent fault diagnosis, performance optimization, and automatic repair, ensuring the stable operation and high availability of the storage system in complex environments.

[0055] Optionally, the above-mentioned determination of the anomaly type of the storage system based on the detection results may specifically include: determining whether there is a fault in the storage path of the storage system, whether there is a fault in the device corresponding to the storage volume, and whether there is an error in the configuration of the storage path.

[0056] In some examples, storage system failures can be mainly categorized into three types: path failures, device failures, and configuration errors. Path failures typically manifest as anomalies in one or more storage paths, such as degraded path I / O performance (including increased I / O latency), path I / O processing efficiency significantly lower than other paths (e.g., significantly lower throughput), and the "path oscillation" phenomenon caused by unstable path states, where I / O fails intermittently and recovers periodically. Device failures refer to the unavailability of all paths associated with a particular storage volume; for example, the volume cannot perform normal read / write operations, but other volume devices in the same storage system can still be accessed normally. This is common when the underlying storage device crashes or the connection is completely interrupted. Configuration errors are mostly caused by improper policy settings, such as data I / O being assigned to a non-optimal path, resulting in overall I / O performance not reaching the expected level and affecting business operation efficiency.

[0057] For example, based on the detection results of the multipath software, it can be determined whether there is a fault in the storage path of the storage system. If it is found that the I / O performance of a certain path has dropped significantly, the processing efficiency is lower than that of other paths, or there are periodic I / O failures and recovery phenomena, then it can be diagnosed as a path fault.

[0058] For example, it can also be used to determine whether the device corresponding to the storage volume is faulty. If all paths of a certain volume become unavailable, causing read and write operations to be unable to proceed normally, while other volumes in the same storage system can still work normally, it indicates that the device may have failed.

[0059] For example, one can also determine whether there is a configuration error by analyzing whether the selection of I / O paths is reasonable, such as if the I / O distribution path is not optimal, resulting in overall performance being lower than expected.

[0060] Step 203: If there is an anomaly in the storage path, the anomaly diagnosis model is used to diagnose and analyze the path status information to identify the anomaly type of the storage system.

[0061] In some examples, the abnormal information obtained from multipath software analysis and diagnosis can be classified into different types of faults based on the nature and scope of the abnormality, such as path interruption, frequent path switching, I / O timeout, or device unreachability.

[0062] Optionally, the method in this embodiment may further include: training the historical abnormal data of the storage system based on machine learning algorithms to construct an abnormality diagnosis model.

[0063] For example, a pre-built system model can be used to simulate the normal operation of a multipath storage system, and the differences can be located and the root cause of the fault can be found by comparing the current abnormal state. Alternatively, "diagnosis based on machine learning algorithms" can be introduced, and a predictive model can be built by training on historical data. When a new anomaly occurs, algorithms such as decision trees, support vector machines (SVM) or neural networks can be used to quickly identify the anomaly type and cause, thereby achieving more intelligent and efficient fault diagnosis and repair capabilities.

[0064] Step 204: Match the exception type with the repair strategy in the preset repair strategy library corresponding to the storage system.

[0065] For example, during read / write operations on storage volume multipath devices by business software, the multipath software captures and monitors key indicators such as I / O latency, throughput, and I / O success / failure rate for each path in real time. If the I / O latency of a certain path is found to be significantly higher than other paths, reaching a preset threshold, the path is considered abnormal. Further analysis of host-side logs (such as message logs) near this point in time is then performed to check for driver-layer timeout records (e.g., log information containing keywords such as LPFC and timeout). Based on the detected high I / O latency and host-side driver-layer anomalies, the system matches the information with existing "repair strategies." For this type of path failure, an exemplary repair strategy is as follows: when increased I / O latency is confirmed to be accompanied by host-side driver anomalies (such as timeouts), a preliminary judgment is made that it is a link problem. At this time, measures are automatically taken to isolate the affected path to prevent excessive I / O latency from impacting user services, and the administrator is simultaneously notified to investigate the specific cause of the link failure so that further action can be taken in a timely manner.

[0066] Optionally, the method in this embodiment may further include: maintaining a preset repair strategy library using keyword rules based on the anomaly type and cause diagnosed by the anomaly diagnosis model.

[0067] In some examples, this method of maintaining a pre-defined repair strategy library allows for the continuous accumulation of multipath anomalies and their handling and repair strategies based on current human experience into the multipath software's automatic anomaly repair strategy. This enables the software to identify and judge all past anomalies based on human experience, freeing administrators from repetitive problem analysis. The initial collection and judgment of anomaly information is automated by the multipath software, while administrators are primarily responsible for reviewing and confirming the software's preliminary analysis results and timely reports of anomalies and diagnostic findings. This method can be summarized as a rule-based diagnostic and repair approach.

[0068] Step 205: Based on the matched repair strategy, perform repair processing on the storage system.

[0069] For example, such as Figure 4 As shown, when the system successfully matches a repair strategy corresponding to the current anomaly type in the repair strategy library, it will automatically generate a complete report containing anomaly information and suggested handling strategies. If the current configuration is an automatic handling strategy, the system will immediately execute the corresponding repair operation according to the matched repair strategy. After the repair is completed, the system will automatically archive the anomaly information and handling strategy report and notify the administrator via SMS or email to ensure that they are aware of the processing results. If the current configuration is an interactive confirmation handling strategy, the system will not automatically execute the repair operation, but will notify the administrator via SMS or email, requesting them to review the anomaly information and suggested handling strategies. After the administrator confirms that everything is correct and provides feedback of approval, the system will automatically process according to the repair strategy. If the administrator determines that there is a risk or the problem is not yet clear and the review is not passed, the system will pause the processing flow and wait for further manual intervention and guidance from the administrator. This mechanism effectively balances the needs of automated response and manual control, ensuring the security and reliability of the storage system during anomaly handling.

[0070] Optionally, step 205 may specifically include: generating a processing strategy report containing anomaly information and processing suggestions for the storage path based on the matched repair strategy; repairing the storage system according to the repair strategy; and submitting the processing strategy report to the management platform.

[0071] For example, based on the matched repair strategy, the system will automatically generate a processing strategy report. The report includes anomaly information of the storage path and corresponding processing suggestions, providing clear guidance for fault repair. The system can then automatically repair the storage system according to the repair strategy to restore it to normal operation.

[0072] For example, after the repair is completed, the handling strategy report will be submitted to the management platform for administrators to review and analyze, ensuring that the problem handling process is traceable and providing data support for subsequent optimization of diagnostic strategies.

[0073] Optionally, the method in this embodiment may further include: if no abnormality is detected in the storage path or no repair strategy is matched, the abnormal information of the storage system is reported to the administrator and awaits manual intervention.

[0074] In some examples, if no remediation strategy is found, the administrator is notified via SMS or email, and the administrator will intervene manually as soon as possible after receiving the notification.

[0075] In some embodiments, such as Figure 5As shown, the multipath anomaly handling process consists of four core modules, forming a closed-loop monitoring, analysis, repair, and notification system. First, the multipath anomaly monitoring and capture module monitors key indicators during I / O operations in real time, including I / O latency, throughput, and success / failure rates, to promptly detect potential anomalies. Once an anomaly is detected, the data is transmitted to the multipath anomaly analysis module, which performs preliminary diagnosis and analysis of the captured anomalies, such as querying host-side logs or environmental information at relevant time points to further confirm the specific cause of the anomaly. Then, based on the analysis results, the multipath anomaly repair strategy management module matches corresponding repair strategies according to keyword rules. These strategies are based on accumulated experience and are continuously updated and maintained, aiming to provide effective solutions for different types of anomalies.

[0076] In some examples, regardless of whether a suitable handling strategy is matched, the email / SMS notification module can inform the administrator of the details of the abnormal issue and the relevant handling strategy as soon as possible, ensuring that the administrator can intervene quickly and take appropriate measures.

[0077] Compared to current related technologies, this embodiment provides an intelligent automatic repair method for storage multipath anomalies. By automatically monitoring metrics such as latency, throughput, and I / O success rate during the I / O process, it can automatically detect and repair anomalies in storage multipath systems, reducing manual intervention, mitigating the risk of data access interruptions due to anomalies, and improving system reliability and availability. The automatic repair module can quickly execute repair operations according to preset repair strategies, significantly shortening the processing time for anomalies and improving system maintenance efficiency. It reduces reliance on system administrators, lowers the risk of misjudgments and omissions due to manual intervention, thereby reducing system maintenance costs. Through a flexible repair strategy library and automatic repair mechanism, it can easily adapt to different types of storage multipath systems and anomalies, enhancing system scalability.

[0078] Embodiments of this application also provide an exception handling device, as... Figure 1 The specific implementation of the method shown is as follows: Figure 6 As shown, the device includes: an acquisition module 31, a determination module 32, a matching module 33, and a processing module 34.

[0079] The acquisition module 31 is configured to acquire path status information corresponding to multiple storage volumes in the storage system;

[0080] The determination module 32 is configured to determine the anomaly type of the storage system based on the path status information;

[0081] Matching module 33 is configured to match the exception type with the repair strategy in the preset repair strategy library corresponding to the storage system;

[0082] The processing module 34 is configured to perform repair processing on the storage system based on the matched repair strategy.

[0083] In some examples of this embodiment, the path status information includes multi-path status, number of paths, device status corresponding to the multiple storage volumes, and input / output performance indicators; accordingly, the determining module 32 is specifically configured to detect the storage system based on the multi-path status, number of paths, device status corresponding to the multiple storage volumes, and input / output performance indicators; and determine the anomaly type of the storage system based on the detection results.

[0084] In some examples of this embodiment, the determining module 32 is further configured to determine, based on the detection results, whether there is a fault in the storage path of the storage system, whether there is a fault in the device corresponding to the storage volume, and whether there is an error in the configuration of the storage path.

[0085] In some examples of this embodiment, the determining module 32 is further configured to train the historical abnormal data of the storage system based on a machine learning algorithm to construct an anomaly diagnosis model; and to use the anomaly diagnosis model to diagnose and analyze the path status information to identify the anomaly type of the storage system.

[0086] In some examples of this embodiment, the matching module 33 is further configured to maintain the preset repair strategy library using keyword rules based on the anomaly type and cause diagnosed by the anomaly diagnosis model.

[0087] In some examples of this embodiment, the processing module 34 is further configured to generate a processing strategy report containing the abnormal information and processing suggestions of the storage path based on the matched repair strategy; repair the storage system according to the repair strategy; and submit the processing strategy report to the management platform.

[0088] In some examples of this embodiment, the processing module 34 is further configured to report the abnormal information of the storage system to the administrator and wait for manual intervention if no abnormality is detected in the storage path or no repair strategy is matched.

[0089] It should be noted that other corresponding descriptions of the functional units involved in the exception handling device provided in this embodiment can be found in [reference]. Figure 1 The corresponding descriptions in [the document] will not be repeated here.

[0090] Based on the above, Figure 1and Figure 3 Accordingly, this embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method. Figure 1 and Figure 3 The method shown.

[0091] Based on the above, Figure 1 and Figure 3 Accordingly, this embodiment also provides a computer program product having a computer program stored thereon, which, when executed by a processor, implements the above-described method. Figure 1 and Figure 3 The method shown.

[0092] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.

[0093] Based on the above, Figure 1 and Figure 3 The method shown, and Figure 6 To achieve the above objectives, the present application also provides an electronic device, such as a personal computer or a server, in the illustrated virtual device embodiment. This device includes a storage medium and a processor; the storage medium stores a computer program; the processor executes the computer program to implement the above-described virtual device. Figure 1 and Figure 3 The method shown.

[0094] In some embodiments, the aforementioned physical device may further include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, an input unit such as a keyboard, etc., and optionally, a USB interface, a card reader interface, etc. In some embodiments, the network interface may include a standard wired interface, a wireless interface (such as a Wi-Fi interface), etc.

[0095] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.

[0096] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.

[0097] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms, or it can be implemented by hardware. By applying the solution of this embodiment, compared with the current related technologies, this embodiment introduces an anomaly monitoring and capture module, which enables the system to capture information at the first moment when multi-path anomalies occur, ensuring that no problem is missed; subsequently, the captured anomaly information is preliminarily judged using an automatic diagnosis and analysis mechanism to form a clear diagnostic conclusion; on this basis, combined with the existing repair strategy library, the system can automatically match the applicable processing solution to achieve rapid response and repair of anomalies; to further improve the security and flexibility of the processing process, an email or SMS notification mechanism is added on top of the automatic processing mechanism, allowing administrators to intervene and perform manual interactive confirmation at key stages, thereby constructing a multi-path anomaly handling process that combines automation efficiency with manual control capabilities.

[0098] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0099] The above are merely specific embodiments of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to these embodiments, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. An exception handling method, characterized in that, include: Obtain path status information corresponding to multiple storage volumes in the storage system. The path status information includes multi-path status, number of paths, device status corresponding to the multiple storage volumes, and input / output performance indicators. Based on the path status information, the anomaly type of the storage system is determined, including: The storage system is tested based on the multipath status, the number of paths, the device status corresponding to the multiple storage volumes, and the input / output performance indicators. Based on the detection results, determine the anomaly type of the storage system, including: based on the detection results, determine whether there is a fault in the storage path of the storage system, whether there is a fault in the device corresponding to the storage volume, and whether there is an error in the configuration of the storage path; Match the anomaly type with the repair strategies in the preset repair strategy library corresponding to the storage system; Based on the matched repair strategy, the storage system is repaired, including: generating a processing strategy report containing anomaly information and processing suggestions for the storage path; repairing the storage system according to the repair strategy; and submitting the processing strategy report to the management platform.

2. The method according to claim 1, characterized in that, Before determining the anomaly type of the storage system based on the path status information, the method further includes: An anomaly diagnosis model is constructed by training the historical anomaly data of the storage system based on machine learning algorithms; Determining the anomaly type of the storage system based on the path status information includes: The anomaly diagnosis model is used to diagnose and analyze the path status information to identify the anomaly type of the storage system.

3. The method according to claim 2, characterized in that, The method further includes: Based on the anomaly types and causes diagnosed by the anomaly diagnosis model, the preset repair strategy library is maintained using keyword rules.

4. The method according to claim 1, characterized in that, The method further includes: If no abnormality is detected in the storage path or no repair strategy is matched, the abnormal information of the storage system will be reported to the administrator and await manual intervention.

5. An anomaly handling device, characterized in that, include: The acquisition module is configured to acquire path status information corresponding to multiple storage volumes in the storage system. The path status information includes multi-path status, number of paths, device status corresponding to the multiple storage volumes, and input / output performance indicators. The determination module is configured to determine the anomaly type of the storage system based on the path status information. Specifically, the determination module is configured to detect the storage system based on the multi-path status, the number of paths, the device status corresponding to the multiple storage volumes, and the input / output performance indicators; and to determine the anomaly type of the storage system based on the detection results. The determination module is further configured to determine, based on the detection results, whether there is a fault in the storage path of the storage system, whether there is a fault in the device corresponding to the storage volume, and whether there is an error in the configuration of the storage path. The matching module is configured to match the exception type with the repair strategies in the preset repair strategy library corresponding to the storage system; The processing module is configured to perform repair processing on the storage system based on the matched repair strategy, including: generating a processing strategy report containing anomaly information and processing suggestions for the storage path based on the matched repair strategy; The storage system is repaired according to the repair strategy, and the processing strategy report is submitted to the management platform.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 4.

7. An electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 4.

Citation Information

Patent Citations

  • Fault disposal method and device for storage system

    CN106874136A

  • Multi-path exception processing method and device, computer equipment and storage medium

    CN117573405A