Abnormality processing method and device, storage medium and electronic equipment
By automatically obtaining the storage system path status information, identifying exception types and matching repair strategies, the problem of manual repair of the storage system is solved, and the problem of time-consuming and labor-intensive and easy to misjudgment is easily achieved, efficient and accurate exception handling is achieved, and system stability and user experience are improved.
Patent Information
- Application Number
- CN202510855487.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-06-24
AI Technical Summary
In the prior art, storage system exception handling relies on manual diagnosis and repair, which is time-consuming and labor-intensive and prone to misjudgment and misjudgment, affecting user business continuity and reliability.
By obtaining the path status information of multiple storage volumes in the storage system, automatically identifying the exception type and matching it with the preset repair policy library, automated repair processing is achieved.
It improves the efficiency and accuracy of exception handling, reduces manual intervention, reduces the risks of misjudgment and misjudgment, and improves the stability and operation and maintenance efficiency of the storage system.
Smart Images

Figure CN120407262A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data storage, and in particular, to an exception handling method, apparatus, storage medium, and electronic device. Background Art
[0002] With the rapid development of informatization and digitalization, storage systems have become the core medium for data storage in various industries. Storage multipath technology is widely used to improve the reliability and performance of storage systems. By configuring multiple paths for storage devices, the system can automatically switch to another path when one path fails, thus ensuring the continuity and reliability of data access.
[0003] Currently, for various abnormal problems that may occur during the operation of storage systems, it usually relies on manual diagnosis and repair of problems. However, this method not only takes time and effort but also is prone to misjudgment and missed judgment, affecting user services. Summary of the Invention
[0004] The present disclosure provides an exception handling method, apparatus, storage medium, and electronic device. Its main purpose is to solve the problem that the method of manually repairing system exceptions in the related art not only takes time and effort but also is prone to misjudgment and missed judgment, affecting user services.
[0005] In a first aspect, the present application provides an exception handling method, including: Obtaining path status information corresponding to multiple storage volumes in a storage system; Determining the exception type of the storage system according to the path status information; Matching the exception type with repair strategies in a preset repair strategy library corresponding to the storage system; Performing a repair process on the storage system based on the matched repair strategy.
[0006] In a second aspect, the present application provides an exception handling apparatus, including: An obtaining module configured to obtain path status information corresponding to multiple storage volumes in a storage system; A determining module configured to determine the exception type of the storage system according to the path status information; A matching module configured to match the exception type with repair strategies in a preset repair strategy library corresponding to the storage system; A processing module configured to perform a repair process on the storage system based on the matched repair strategy.
[0007] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method of the first aspect is implemented.
[0008] In a fourth aspect, the present application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor. When the processor executes the computer program, the method of the first aspect is implemented.
[0009] In a fifth aspect, the present application provides a computer program product, on which a computer program is stored. When the computer program is executed by a processor, the method of the first aspect is implemented.
[0010] The present disclosure provides an exception handling method, apparatus, storage medium, and electronic device. The method includes: first, obtaining path status information corresponding to multiple storage volumes in a storage system; then, determining the exception type of the storage system according to the path status information; then, matching the exception type with a repair policy in a preset repair policy library corresponding to the storage system; and based on the matched repair policy, performing a repair process on the storage system. Compared with the current related technologies, the present application automatically obtains the path status information of each storage volume in the storage system, judges and accurately locates the exception type, and realizes intelligent matching and automatic repair in combination with the preset repair policy library, effectively reducing manual intervention, improving the efficiency and accuracy of exception handling, avoiding the business impact caused by misjudgment and missed judgment, and at the same time having good user interactivity and system adaptability, significantly improving the stability and operation and maintenance efficiency of the storage system.
[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0013] Figure 1 Shows a schematic flowchart of an exception handling method provided by an embodiment of the present application; Figure 2 Shows a schematic diagram of an example provided by an embodiment of the present application; Figure 3 Shows a schematic flowchart of another exception handling method provided by an embodiment of the present application; Figure 4Shows a schematic flowchart of an example provided by an embodiment of the present application; Figure 5 Shows a schematic flowchart of an example provided by an embodiment of the present application; Figure 6 Shows a schematic structural diagram of an exception handling device provided by an embodiment of the present application. Detailed implementation manners
[0014] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0015] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0016] In the context of today's global informatization and digitalization, the data information of human society is becoming increasingly huge, and storage systems have emerged as the main medium for backend data storage in all walks of life. With the all-weather data access demand becoming the norm, storage systems have gradually gone through the development stages of capacity, performance, and reliability requirements, enabling them to store user data (enough space), access it quickly (sufficient performance), and access it stably (without loss, reliable storage process, and can handle general failures). Storage multi-path technology is widely used to improve the reliability and performance of storage systems. By configuring multiple paths for storage devices, the system can automatically switch to another path when one path fails, thus ensuring the continuity and reliability of data access.
[0017] When business users find abnormalities in business applications, such as performance degradation or business read / write errors, they usually contact the administrator for troubleshooting. After logging in to the environment, the administrator diagnoses problems by checking the multipath environment and log information of the storage system. However, this process often starts after the business has been affected, resulting in a certain cycle between problem discovery and repair completion, during which the business is in an abnormal operating state. For example, first, the business user discovers an abnormality and contacts the administrator; then, the administrator logs in to the system, troubleshoots the storage multipath environment and log information to determine the cause of the problem; finally, repairs are made according to the diagnostic results until the problem is solved and the business resumes normal operation. Although this process can ultimately solve the problem, its passive response characteristic causes the business to be continuously affected before the problem is repaired, affecting user experience and business continuity.
[0018] Based on the various abnormal problems that may occur during the operation of the above storage multipath system, in order to improve the technical problem that the method of manually repairing system abnormalities in related technologies is not only time-consuming and laborious, but also prone to misjudgment and missed judgment, which affects the user's business. This embodiment provides an abnormal handling method, as Figure 1 shown, the method includes the following steps: Step 101, obtain the path status information corresponding to multiple storage volumes in the storage system.
[0019] In some examples, multiple storage volumes are multiple independent storage units created by logically partitioning the physical storage resources of the storage system. Each storage volume can be configured, accessed, and managed separately. These storage volumes can provide customized storage solutions for different applications or services, support data read / write operations, and can be managed for capacity adjustment, performance optimization, and fault recovery according to requirements to meet the data storage needs in different business scenarios.
[0020] For example, the path status information may include key indicators such as the running status, connectivity, I / O response latency, error occurrence frequency, and path switching records of each data path in the storage system, which are used to reflect the health status and performance of the data access path in real time and provide an important basis for abnormal detection and fault diagnosis.
[0021] Exemplarily, the business application software service on host A accesses and stores data through the volumes provided by the storage system, and uses the multipath software developed by the storage manufacturer to achieve load balancing and fault redundancy. As Figure 2As shown in the figure, host A is connected to switches A and B through Fibre Channel Host Bus Adapters (FC HBAs) (including Port 1 and Port 2), and then connected to multiple ports (Port A, Port B, Port C, Port D) of the storage system, and finally accesses volumes 1 to N in the storage system. As a key bridge connecting the front-end user business software and the back-end storage volumes, once a path state or a device on the link fails, it will directly affect the input / output (I / O) read and write performance of the front-end host business software, resulting in service interruption or performance degradation. Therefore, it is crucial to detect and diagnose and repair multi-path anomalies in a timely manner to ensure the stable operation and high availability of user services. Combining Figure 2 it can be seen that the multi-path software plays a core role in monitoring and managing these path states. By detecting and automatically switching paths in real time, it effectively avoids the impact of a single path failure on the service, thus ensuring service continuity and data security.
[0022] Step 102: Determine the type of anomaly of the storage system according to the path state information.
[0023] In some examples, according to key metrics such as the running state, connectivity, I / O response latency, error occurrence frequency, and path switching records of each data path in the storage system, the system can judge in real time whether there are anomalies in each storage path in the storage system, such as path interruption, excessive I / O latency, frequent errors, or abnormal path switching, etc., so as to discover potential faults in a timely manner and provide a basis for subsequent diagnosis and repair.
[0024] Exemplarily, the types of anomalies of the storage system include but are not limited to hard disk failures, path failures, excessive data access latency, frequent I / O errors, insufficient storage capacity, data corruption or loss, performance bottlenecks, and configuration errors, etc. These anomaly problems may affect the normal operation of the system and data security alone or in combination.
[0025] Step 103: Match the type of anomaly with the repair strategy in the preset repair strategy library corresponding to the storage system.
[0026] Exemplarily, the repair strategy library may contain standard disposal schemes for various storage path anomalies, such as automatically switching to an alternate path, reactivating a failed path, adjusting the I / O queue depth, or triggering an alarm to notify the operation and maintenance personnel to intervene and handle, etc.
[0027] In some examples, by accurately matching the type of anomaly with the corresponding repair strategy, the system can quickly execute the corresponding repair actions, thereby effectively restoring the normal running state of the storage path and improving the stability and availability of the storage system.
[0028] Step 104: Based on the matched repair strategy, perform repair processing on the storage system.
[0029] In some examples, the system can execute repair operations according to the preset steps in the repair strategy, which may include but are not limited to replacing faulty hard disks, reconfiguring storage paths, optimizing I / O scheduling to reduce latency, correcting configuration errors, etc.
[0030] Exemplarily, during the repair process, a detailed processing strategy report can also be generated simultaneously, recording all operation details and change contents, and finally submitting this report to the management platform for administrators to review and subsequent tracking to ensure the stable operation and data security of the storage system.
[0031] Compared with the related technology, in this embodiment, first, the path status information corresponding to multiple storage volumes in the storage system is obtained; then, according to the path status information, the abnormal type of the storage system is determined; then, the abnormal type is matched with the repair strategies in the preset repair strategy library corresponding to the storage system; based on the matched repair strategy, repair processing is performed on the storage system. Compared with the current related technology, in this embodiment, based on mastering the path status information, the system can automatically determine the abnormal type of the storage system, accurately identify the specific abnormal type according to the abnormal characteristics, such as hard disk failure, path interruption or performance bottleneck, etc., and intelligently match the identified abnormal type with the repair solutions in the preset repair strategy library, improving the accuracy and efficiency of fault response. Then, based on the matched repair strategy, automated repair processing is performed on the storage system, which not only effectively shortens the fault response time but also reduces the need for manual intervention and operation and maintenance costs. It realizes the closed-loop management of abnormal detection, diagnosis, matching, and repair, significantly improves the stability, reliability, and operation and maintenance intelligence level of the storage system, and ensures the continuity and security of data access.
[0032] As a refinement of this embodiment, the storage system can be repaired in the following ways, but not limited to, such as Figure 3 as shown Figure 3 is a schematic flowchart of an abnormal processing method provided by an embodiment of the present disclosure, including: Step 201: Obtain the path status information corresponding to multiple storage volumes in the storage system.
[0033] Among them, the path status information includes multi-path status, number of paths, device status corresponding to multiple storage volumes, and input / output performance metrics.
[0034] In some examples, the multipath status of each storage volume in the storage system can be monitored in real time through multipath software, not only tracking the status of each path (such as whether it is active, failed, or recovering), but also monitoring changes in the number of paths to ensure the redundancy and reliability of data transmission. At the same time, the software checks the overall health of each storage device, including but not limited to the online status and fault conditions of the device. In addition, the multipath software also collects and analyzes I / O performance metrics, such as response time, throughput, and input / output operations per second (IOPS), in order to detect potential bottlenecks or abnormal behaviors in a timely manner and make corresponding adjustments, thereby optimizing the performance and stability of the overall storage system.
[0035] Step 202: Determine whether there is an abnormality in the storage paths in the storage system according to the path status information.
[0036] For example, the multipath software can determine in real time whether there is an abnormality in the storage paths in the storage system according to the path status information, such as problems like path disconnection, excessive delay, or communication failure. When an abnormal state of the multipath is detected, the system will immediately capture the relevant abnormal information and conduct in-depth analysis and diagnosis. This process includes a comprehensive evaluation of key metrics such as path switching records, I / O error logs, and device response time to quickly locate the cause of the failure and determine whether it is due to a physical connection problem, a storage device failure, or improper configuration.
[0037] Optionally, step 202 may specifically include: detecting the storage system based on the multipath status, the number of paths, the device status corresponding to multiple storage volumes, and the input / output performance metrics; determining the type of abnormality in the storage system according to the detection results.
[0038] In some examples, as a key component for managing the multipath devices corresponding to the storage volumes, the multipath software has the ability to monitor the underlying storage paths in real time. When the status of each path (such as going online, going offline, fault recovery, etc.) or the number of paths changes, the multipath software can capture these changes in a timely manner to ensure that the system responds in a timely manner and adjusts the path usage strategy.
[0039] Exemplarily, in the case of a decrease in I / O performance (such as an increase in delay) or the occurrence of I / O errors, the multipath software can also quickly sense and record the relevant abnormal information, including I / O delay metrics and error logs at the path level. This efficient monitoring and capture mechanism provides solid data support for subsequent fault diagnosis, performance optimization, and automatic repair, ensuring the stable operation and high availability of the storage system in a complex environment.
[0040] Optionally, determining the abnormal type of the storage system according to the detection result may specifically include: judging whether there is a fault in the storage path of the storage system, whether there is a fault in the device corresponding to the storage volume, and whether there is a configuration error in the storage path according to the detection result.
[0041] In some examples, the fault types in the storage system can be mainly divided into three categories: path fault, device fault, and configuration error. Among them, path faults usually manifest as abnormalities in one or more storage paths, such as a decrease in path I / O performance (including an increase in I / O latency), a significantly lower path I / O processing efficiency than other paths (such as a significantly lower throughput), and the "path oscillation" phenomenon caused by an unstable path state, that is, I / O fails and then recovers normally from time to time, showing periodic fluctuations. A device fault means that all paths associated with a certain storage volume are unavailable. For example, normal read and write operations cannot be performed on this volume, but other volume devices in the same storage system can still be accessed normally, which is common in the case of the underlying storage device crashing or the connection being completely interrupted. Configuration errors are mostly caused by improper policy settings. For example, data I / O is sent to a non-optimal path, resulting in the overall I / O performance not reaching the expected level and affecting the business operation efficiency.
[0042] Exemplarily, according to the detection result of the multipath software, it can be judged whether there is a fault in the storage path of the storage system. If it is found that the I / O performance of a certain path has decreased significantly, the processing efficiency is lower than other paths, or there is a periodic I / O failure and recovery phenomenon, it can be diagnosed as a path fault.
[0043] For example, it is also possible to judge whether there is a fault in the device corresponding to the storage volume. If all paths of a certain volume become unavailable, resulting in the inability to perform normal read and write operations, while other volumes in the same storage system can still work normally, it indicates that the device may have failed.
[0044] For example, it is also possible to determine whether there is a configuration error by analyzing whether the selection of the I / O path is reasonable. For example, the overall performance is lower than expected because the I / O distribution path is not optimal.
[0045] Step 203: If there is an abnormality in the storage path, use the abnormality diagnosis model to diagnose and analyze the path status information to identify the abnormal type of the storage system.
[0046] In some examples, the abnormal information obtained by analyzing and diagnosing the multipath software can be classified, and different types of faults are classified according to the nature and influence range of the abnormality, such as path interruption, frequent path switching, I / O timeout, or device unreachability.
[0047] Optionally, the method of this embodiment may specifically further include: training the historical abnormal data of the storage system based on a machine learning algorithm to construct an abnormality diagnosis model.
[0048] Exemplarily, a pre - built system model can be utilized to simulate the normal operating state of the multi - path storage system. By comparing with the current abnormal state, the difference points can be located and the root cause of the fault can be found. Additionally, "diagnosis based on machine learning algorithms" can be introduced. By training on historical data, a prediction model is established. When a new anomaly appears, algorithms such as decision trees, support vector machines (SVM), or neural networks are used to quickly identify the type and cause of the anomaly, thereby achieving more intelligent and efficient fault diagnosis and repair capabilities.
[0049] Step 204: Match the anomaly type with the repair strategies in the preset repair strategy library corresponding to the storage system.
[0050] For example, during the read - write operation of the storage volume multi - path device by the business software, the multi - path software will capture and monitor key metrics such as I / O latency, throughput, and I / O success - failure rate of each path in real - time. Once it is found that the I / O latency of a certain path is significantly higher than other paths and reaches the preset threshold, it is determined that this path is abnormal, and the relevant host - side logs (such as message logs) near this time point are further analyzed to check whether there is a record of driver - layer timeout (for example, log information containing keywords such as lpfc and timeout). Based on the detected large I / O latency and the information of host - side driver - layer anomaly, the system will match with the existing "repair strategies". For such path - fault problems, an exemplary repair strategy is: when it is confirmed that the I / O latency increases and there is an accompanying host - side driver anomaly (such as timeout), it is initially judged as a link problem; at this time, measures are automatically taken to isolate the affected path to prevent its excessive I / O latency from affecting the user's business, and the administrator is notified synchronously to troubleshoot the specific cause of the link failure for timely further action.
[0051] Optionally, the method of this embodiment may further specifically include: maintaining the preset repair strategy library using keyword rules according to the anomaly type and cause diagnosed by the anomaly diagnosis model.
[0052] In some examples, through this way of maintaining the preset repair strategy library, the multi - path anomaly phenomena, processing, and repair strategies currently handled by manual experience can be continuously accumulated and added to the automatic repair strategies for multi - path software anomalies, enabling it to identify and judge all the occurred anomaly problems based on manual experience, liberating the administrator from repetitive problem analysis. The collection and judgment of primary problem anomaly information are automatically completed by the multi - path software, and the administrator is mainly responsible for reviewing and confirming the preliminary analysis results of the multi - path software and the anomaly information and diagnosis reports reported in the first time. This method can be summarized as a rule - based diagnosis and repair method.
[0053] Step 205: Based on the matched repair strategy, perform repair processing on the storage system.
[0054] Exemplarily, as Figure 4 shown, when the system successfully matches a repair strategy corresponding to the current exception type in the repair strategy library, a complete report containing the exception information and the recommended handling strategy will be automatically generated. If the current configuration is the automatic handling strategy, the system will immediately execute the corresponding repair operation according to the matched repair strategy. After the repair is completed, the system will automatically archive the exception information and the handling strategy report for this time, and notify the administrator by text message or email to ensure that they understand the processing result. If the current configuration is the interactive confirmation handling strategy, the system will not automatically execute the repair operation, but will notify the administrator by text message or email, asking them to review the exception information and the recommended handling strategy. After the administrator confirms and gives feedback to agree, the system will perform automatic processing according to the repair strategy. If the administrator determines that there are risks or the problems are not clear and does not pass the review, the system will suspend the processing flow and wait for further manual intervention and guidance from the administrator. This mechanism effectively balances the needs of automated response and manual control, ensuring the security and reliability of the storage system during the exception handling process.
[0055] Optionally, step 205 may specifically include: generating a handling strategy report containing the exception information of the storage path and the handling suggestions based on the matched repair strategy; repairing the storage system according to the repair strategy, and submitting the handling strategy report to the management platform.
[0056] For example, based on the matched repair strategy, the system will automatically generate a handling strategy report, which contains the exception information of the storage path and the corresponding handling suggestions, providing clear guidance for fault repair. Furthermore, the system can perform automatic repair operations on the storage system according to the repair strategy to restore the normal operating state.
[0057] Exemplarily, after the repair is completed, the handling strategy report will be submitted to the management platform for the administrator to view and subsequent analysis, ensuring that the problem handling process is traceable and providing data support for subsequent optimization of the diagnostic strategy.
[0058] Optionally, the method of this embodiment may specifically further include: if no exception in the storage path is detected or no repair strategy is matched, report the exception information of the storage system to the administrator and wait for manual intervention for processing.
[0059] In some examples, if no repair strategy is matched, the administrator will be notified by text message or email, and the administrator will perform manual intervention for processing as soon as possible after receiving the notice.
[0060] In some embodiments, as Figure 5As shown in the figure, the multi - path exception problem handling process consists of four core modules, forming a closed - loop monitoring, analysis, repair, and notification system. First, the multi - path exception problem monitoring capture module monitors key indicators during the I / O operation in real - time, including I / O latency time, throughput, success and failure rates, etc., to detect potential abnormal situations in a timely manner. Once abnormal information is detected, the data will be transmitted to the multi - path exception problem analysis module, which is responsible for the preliminary diagnosis and analysis of the captured exception. For example, it queries the host - side logs or environmental information at relevant time points to further confirm the specific cause of the exception. Then, based on the analysis results, the multi - path exception problem repair strategy management module matches corresponding repair strategies according to keyword rules. These strategies are accumulated based on past experience and continuously updated and maintained, aiming to provide effective solutions for different types of exceptions.
[0061] In some examples, whether or not a suitable processing strategy is matched, the email / sms notification module can inform the administrator of the details of the exception problem and relevant processing strategies in the first time, ensuring that they can intervene quickly and take corresponding measures.
[0062] Compared with the current related technologies, an intelligent automatic repair method for storage multi - path exception problems provided by this embodiment can automatically detect and repair abnormal problems in the storage multi - path system by automatically monitoring index factors such as latency, throughput, and I / O success rate during the I / O process, reducing manual intervention, reducing the risk of data access interruption caused by abnormal problems, and improving the reliability and availability of the system. Among them, the automatic repair module can quickly execute repair operations according to the preset repair strategies, greatly shortening the processing time of abnormal problems and improving the maintenance efficiency of the system. It reduces the dependence on system administrators and reduces the risks of misjudgment and missed judgment caused by manual intervention, thus reducing the maintenance cost of the system. Through a flexible repair strategy library and an automatic repair mechanism, it can easily adapt to different types of storage multi - path systems and abnormal problems, enhancing the scalability of the system.
[0063] The embodiment of the present application also provides an exception handling device, as Figure 1 a specific implementation of the Figure 6 shown method. As shown in the figure, the device includes: an acquisition module 31, a determination module 32, a matching module 33, and a processing module 34.
[0064] The acquisition module 31 is configured to acquire path status information corresponding to multiple storage volumes in the storage system; The determination module 32 is configured to determine the exception type of the storage system according to the path status information; The matching module 33 is configured to match the exception type with the repair strategies in the preset repair strategy library corresponding to the storage system; The processing module 34 is configured to perform a repair process on the storage system based on the matched repair policy.
[0065] In some examples of this embodiment, the path status information includes the multi-path status, the number of paths, the device status corresponding to the multiple storage volumes, and the input / output performance metrics; correspondingly, the determination module 32 is specifically configured to detect the storage system based on the multi-path status, the number of paths, the device status corresponding to the multiple storage volumes, and the input / output performance metrics; and determine the abnormal type of the storage system according to the detection result.
[0066] In some examples of this embodiment, the determination module 32 is further specifically configured to determine whether there is a fault in the storage path of the storage system, whether there is a fault in the device corresponding to the storage volume, and whether there is an error in the configuration of the storage path according to the detection result.
[0067] In some examples of this embodiment, the determination module 32 is further specifically configured to train the historical abnormal data of the storage system based on a machine learning algorithm to construct an abnormal diagnosis model; and use the abnormal diagnosis model to diagnose and analyze the path status information to identify the abnormal type of the storage system.
[0068] In some examples of this embodiment, the matching module 33 is further specifically configured to maintain the preset repair policy library according to the abnormal type and cause diagnosed by the abnormal diagnosis model by using keyword rules.
[0069] In some examples of this embodiment, the processing module 34 is further specifically configured to generate a processing policy report including the abnormal information and processing suggestions of the storage path based on the matched repair policy; repair the storage system according to the repair policy, and submit the processing policy report to the management platform.
[0070] In some examples of this embodiment, the processing module 34 is further specifically configured to report the abnormal information of the storage system to the administrator and wait for manual intervention if no abnormality is detected in the storage path or no repair policy is matched.
[0071] It should be noted that other corresponding descriptions of each functional unit involved in an abnormal processing device provided in this embodiment can be referred to the corresponding description in Figure 1 and will not be elaborated here.
[0072] Based on the above as Figure 1 and Figure 3Accordingly, in the method described above, this embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method as described above, such as Figure 1 and Figure 3 .
[0073] Based on the method as described above, such as Figure 1 and Figure 3 , accordingly, this embodiment also provides a computer program product, on which a computer program is stored. When the computer program is executed by a processor, it implements the method as described above, such as Figure 1 and Figure 3 .
[0074] Based on such an understanding, the technical solution of this application can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of various implementation scenarios of this application.
[0075] Based on the method as described above, such as Figure 1 and Figure 3 , and Figure 6 shown in the virtual device embodiment, to achieve the above object, this embodiment of the application also provides an electronic device, such as a personal computer or a server. The device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the method as described above, such as Figure 1 and Figure 3 .
[0076] In some embodiments, the above-mentioned physical device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, and so on. The user interface may include a display screen (Display), an input unit such as a keyboard (Keyboard), etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may include a standard wired interface, a wireless interface (such as a WI-FI interface), etc. in some embodiments.
[0077] Those skilled in the art can understand that the above-mentioned physical device structure provided in this embodiment does not limit the physical device, and it may include more or fewer components, or combine certain components, or arrange different components.
[0078] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing the hardware and software resources of the above-mentioned physical devices, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement communication between components inside the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0079] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform, or can also be implemented by hardware. By applying the solution of this embodiment, compared with the current related technologies, in this embodiment, by introducing an abnormal problem monitoring and capturing module, the system can capture information at the first time when a multi-path abnormality occurs, ensuring that the problem will not be missed; subsequently, using the automatic diagnosis and analysis mechanism to make a preliminary judgment on the captured abnormal information to form a clear diagnosis conclusion; on this basis, combined with the existing repair strategy library, the system can automatically match applicable processing solutions to achieve fast response and repair of abnormalities; to further improve the security and flexibility of the processing process, a mail or SMS notification mechanism is added on top of the automatic processing mechanism, and the administrator can intervene at key links and perform manual interaction confirmation, thereby constructing a multi-path abnormal processing process with both automation efficiency and manual control capabilities.
[0080] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0081] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments herein, but will conform to the widest scope consistent with the principles and novel features claimed herein.
Claims
1. An exception handling method, characterized in that, Including: Obtain the path status information corresponding to multiple storage volumes in the storage system; Determine the exception type of the storage system according to the path status information; Match the exception type with the repair strategies in the preset repair strategy library corresponding to the storage system; Based on the matched repair strategy, perform a repair process on the storage system.
2. The method according to claim 1, wherein The path status information includes multi-path status, path quantity, device status corresponding to the multiple storage volumes, and input / output performance metrics; The determining the exception type of the storage system according to the path status information includes: Detect the storage system based on the multi-path status, the path quantity, the device status corresponding to the multiple storage volumes, and the input / output performance metrics; Determine the exception type of the storage system according to the detection result.
3. The method according to claim 2, wherein The determining the exception type of the storage system according to the detection result includes: According to the detection result, determine whether there is a fault in the storage path of the storage system, whether there is a fault in the device corresponding to the storage volume, and whether there is an error in the configuration of the storage path.
4. The method according to claim 1, wherein Before the determining the exception type of the storage system according to the path status information, the method further includes: Train the historical exception data of the storage system based on a machine learning algorithm to construct an exception diagnosis model; The determining the exception type of the storage system according to the path status information includes: Use the exception diagnosis model to diagnose and analyze the path status information to identify the exception type of the storage system.
5. The method according to claim 4, wherein The method further includes: Maintain the preset repair strategy library according to the exception type and exception cause diagnosed by the exception diagnosis model using keyword rules.
6. The method according to claim 3, characterized in that, The performing a repair process on the storage system based on the matched repair strategy includes: Generate a processing strategy report including the exception information and processing suggestions of the storage path based on the matched repair strategy; Repair the storage system according to the repair strategy and submit the processing strategy report to the management platform.
7. The method according to claim 3, characterized in that, The method further includes: If no exception in the storage path is detected or no repair strategy is matched, report the exception information of the storage system to the administrator and wait for manual intervention for processing.
8. An exception handling device, characterized in that, Including: An acquisition module configured to obtain the path status information corresponding to multiple storage volumes in the storage system; A determination module configured to determine the exception type of the storage system according to the path status information; A matching module configured to match the exception type with the repair strategies in the preset repair strategy library corresponding to the storage system; A processing module configured to perform a repair process on the storage system based on the matched repair strategy.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 7.
10. An electronic device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Fault disposal method and device for storage system
CN106874136A
Multi-path anomaly detection and repair method, device, equipment and medium
CN115686921A
Multi-path exception processing method and device, computer equipment and storage medium
CN117573405A