Disk exception processing method, device, equipment and storage medium

By monitoring disk status data to predict anomalies and dynamically determining the timing and method of processing, the lag problem of post-processing of disk anomalies is solved, and timely prevention of disk anomalies is achieved, reducing data loss and interruption.

CN120523689BActive Publication Date: 2025-10-10INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511006444.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-10-10
Estimated Expiration
2045-07-22

AI Technical Summary

Technical Problem

In the existing technology, disk anomalies can only be handled manually after the disk anomaly occurs, and cannot be prevented in time, resulting in data loss or service interruption.

Method used

By monitoring the status data of the target disk, predicting the probability and type of anomalies, dynamically determining the probability threshold, and when the probability of anomalies is greater than the threshold, taking preventive measures by using a processing method that matches the anomaly type.

Benefits of technology

Timely processing before disk abnormalities occur can reduce data loss and service interruptions, and improve the pertinence and effectiveness of processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523689B_ABST
    Figure CN120523689B_ABST
Patent Text Reader

Abstract

The application provides a disk abnormality processing method, device and equipment and a storage medium, and can be applied to the technical field of storage. The disk abnormality processing method comprises the following steps: determining an abnormality probability and a predicted abnormality type of a target disk according to state data of the target disk; in the case that the abnormality probability indicates that the target disk has a risk of abnormality, determining a probability threshold for the target disk according to a disk type of the target disk and an importance level of the data stored in the target disk; and in the case that the abnormality probability is greater than the probability threshold, performing abnormality processing on the target disk by using an abnormality processing method matched with the predicted abnormality type according to the predicted abnormality type.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of storage technology, and more particularly to a disk exception processing method, device, equipment, medium and program product. Background Art

[0002] With the rapid development of cloud computing and big data technologies, data storage demands are growing exponentially. Disks, as core storage media, store critical data such as business systems, databases, and user information. Disk failures can lead to data loss or service interruptions. Therefore, timely disk maintenance is essential.

[0003] However, during disk operation and maintenance, maintenance personnel can usually only manually handle disk anomalies after disk anomalies have occurred, and cannot perform maintenance in time before disk anomalies occur, resulting in a lag. Summary of the Invention

[0004] In view of the above problems, the present invention provides a disk exception processing method, apparatus, device, medium and program product.

[0005] According to a first aspect of the present invention, a disk exception handling method is provided, comprising: determining an exception probability and a predicted exception type of the target disk based on status data of the target disk; in a case where the exception probability indicates that there is a risk of an exception occurring on the target disk, determining a probability threshold for the target disk based on the disk type of the target disk and the importance level of data stored in the target disk; and in a case where the exception probability is greater than the probability threshold, performing exception handling on the target disk based on the predicted exception type using an exception handling method that matches the predicted exception type.

[0006] The second aspect of the present invention provides a disk exception handling device, including: an exception determination module, used to determine the exception probability and predicted exception type of the above-mentioned target disk based on the status data of the target disk; a threshold determination module, used to determine the probability threshold for the above-mentioned target disk according to the disk type of the target disk and the importance level of the data stored in the above-mentioned target disk when the above-mentioned exception probability indicates that the above-mentioned target disk has an exception risk; and an exception handling module, used to perform exception handling on the above-mentioned target disk according to the above-mentioned predicted exception type, using an exception handling method that matches the above-mentioned predicted exception type when the above-mentioned exception probability is greater than the above-mentioned probability threshold.

[0007] A third aspect of the present invention provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0008] The fourth aspect of the present invention further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.

[0009] The fifth aspect of the present invention further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor.

[0010] According to an embodiment of the present invention, by utilizing the target disk's disk type and the importance level of the stored data to dynamically determine the probability threshold for executing exception handling operations on the target disk, the timing for implementing exception handling methods can be accurately determined, allowing for preemptive exception handling of the target disk before an anomaly occurs, thereby reducing data loss or service interruptions caused by target disk anomalies. Furthermore, if the target disk's anomaly probability exceeds the probability threshold, an exception handling method matching the predicted anomaly type is employed. This allows for the combined use of information from both the fault type and the fault probability dimensions to determine a targeted exception handling method, thereby improving the effectiveness of exception handling. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above contents and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings.

[0012] Figure 1 A diagram illustrating an application scenario of a disk exception handling method, apparatus, device, medium, and program product according to an embodiment of the present invention is shown.

[0013] Figure 2 A flowchart of a disk exception processing method according to an embodiment of the present invention is shown.

[0014] Figure 3 A flow chart of dynamic multi-dimensional monitoring according to an embodiment of the present invention is shown.

[0015] Figure 4 A schematic diagram of an abnormality prediction model according to an embodiment of the present invention is shown.

[0016] Figure 5 A flowchart of an exception handling method according to an embodiment of the present invention is shown.

[0017] Figure 6 A structural block diagram of a disk exception processing device according to an embodiment of the present invention is shown.

[0018] Figure 7 A block diagram of an electronic device suitable for implementing a disk exception processing method according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0019] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concept of the present invention.

[0020] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise", "include", etc. used herein indicate the presence of the features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.

[0021] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0022] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0023] Traditional disk operation and maintenance usually involves manual log analysis, hardware replacement, or data migration by operation and maintenance personnel after a disk anomaly occurs. In distributed scenarios such as edge computing, manual response to anomalies leads to long service interruptions.

[0024] Therefore, in order to provide early warning and response in the early stage of disk storage performance degradation, an embodiment of the present invention provides a disk exception handling method, which is characterized in that the method includes: determining the abnormality probability and predicted abnormality type of the target disk based on the status data of the target disk; in a case where the abnormality probability indicates that there is a risk of abnormality occurring on the target disk, determining the probability threshold for the target disk based on the disk type of the target disk and the importance level of the data stored in the target disk; and in a case where the abnormality probability is greater than the probability threshold, using an exception handling method that matches the predicted abnormality type according to the predicted abnormality type to perform abnormality handling on the target disk.

[0025] Embodiments of the present invention utilize the target disk's disk type and the importance level of the stored data to dynamically determine the probability threshold for executing exception handling operations on the target disk. This allows accurate determination of the timing for implementing exception handling methods, enabling preventive measures to be taken before a disk anomaly occurs. Furthermore, if the target disk's anomaly probability exceeds the probability threshold, an exception handling method matching the predicted anomaly type is employed. This approach combines information from both the fault type and the fault probability to determine a targeted exception handling method, improving the effectiveness of exception handling.

[0026] Figure 1 A diagram illustrating an application scenario of a disk exception handling method, apparatus, device, medium, and program product according to an embodiment of the present invention is shown.

[0027] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a first node 110 and a second node 120 , wherein the first node 110 includes a target disk 111 .

[0028] The first node 110 may be a server that provides various services, such as a backend management server that supports websites browsed by a terminal device (for example only). The backend management server may analyze and process received data such as user requests, and feed back the processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal device.

[0029] The second node 120 may be a server for providing exception handling services, for example, performing exception handling on the target disk 111 in the first node 110 .

[0030] It should be noted that the disk exception handling method provided in the embodiment of the present invention can generally be executed by the second node 120. Accordingly, the disk exception handling device provided in the embodiment of the present invention can generally be set in the second node 120. The disk exception handling method provided in the embodiment of the present invention can also be executed by a server or server cluster that is different from the second node 120 and can communicate with the first node 110 and / or the second node 120. Accordingly, the disk exception handling device provided in the embodiment of the present invention can also be set in a server or server cluster that is different from the second node 120 and can communicate with the first node 110 and / or the second node 120.

[0031] For example, the second node 120 determines the abnormality probability and predicted abnormality type of the target disk 111 based on the status data of the target disk 111; when the abnormality probability indicates that the target disk 111 has an abnormality risk, the probability threshold for the target disk 111 is determined according to the disk type of the target disk 111 and the importance level of the data stored in the target disk 111; and when the abnormality probability is greater than the probability threshold, according to the predicted abnormality type, an abnormality handling method matching the predicted abnormality type is adopted to perform abnormality handling on the target disk 111, such as backing up the stored data in the target disk 111 to other nodes.

[0032] It should be understood that Figure 1 The number of the first node, the second node and the target disk in is only . According to the implementation requirements, there can be any number of the first node, the second node and the target disk.

[0033] The following will be based on Figure 1 The scene described by Figures 2 to 5 The disk exception handling method according to the embodiment of the present invention is described in detail.

[0034] Figure 2 A flowchart of a disk exception processing method according to an embodiment of the present invention is shown.

[0035] like Figure 2 As shown, the disk exception handling method of this embodiment includes operations S210 to S230.

[0036] In operation S210 , an abnormality probability and a predicted abnormality type of the target disk are determined based on the status data of the target disk.

[0037] In operation S220 , when the abnormality probability indicates that the target disk has an abnormality risk, a probability threshold for the target disk is determined according to the disk type of the target disk and the importance level of data stored in the target disk.

[0038] In operation S230 , when the abnormality probability is greater than the probability threshold, an abnormality handling method matching the predicted abnormality type is adopted according to the predicted abnormality type to perform abnormality handling on the target disk.

[0039] In an embodiment of the present invention, in order to promptly handle the abnormality of the target disk, the target disk may be dynamically monitored to predict the abnormal state of the target disk based on the real-time state data of the target disk.

[0040] When dynamically monitoring the target disk, the target disk's status data in different dimensions can be obtained from multiple dimensions such as the hardware layer, driver layer, and application layer. This makes the obtained status data more comprehensive than the disk's Self-Monitoring Analysis and Reporting Technology (SMART) data, thereby improving the accuracy of exception prediction and exception handling.

[0041] In embodiments of the present invention, when predicting the abnormal state of a target disk, the probability and type of abnormality that the target disk will experience an abnormality can be predicted. For example, the probability and type of abnormality that the target disk will experience an abnormality can be determined based on a mapping relationship between the state data, the probability of abnormality, and the type of abnormality. Furthermore, the state data of the target disk in a future time period can be predicted to obtain predicted state data, which can then be used to characterize the abnormal state of the target disk in the future time period.

[0042] After obtaining the anomaly probability and the predicted anomaly type, the risk of the target disk experiencing an anomaly can be determined based on the anomaly probability. In some embodiments, the anomaly probability range is divided into intervals corresponding to low risk, medium risk, high risk, and emergency, respectively. Based on the interval in which the anomaly probability falls, whether the target disk has an anomaly risk is determined.

[0043] In other embodiments, the abnormality probability may be compared with a risk probability threshold. If the abnormality probability is greater than the risk probability threshold, it is determined that the target disk has an abnormality risk. Otherwise, it is determined that the target disk does not have an abnormality risk. The risk probability threshold may be set to 50%, for example.

[0044] When it is determined that the abnormality probability indicates that the target disk has an abnormality risk, the abnormality probability may be compared with a probability threshold of the target disk to determine whether abnormality processing needs to be performed on the target disk.

[0045] In an embodiment of the present invention, to improve the pertinence of the probability threshold, the probability threshold can be set to dynamically change according to the disk type of the target disk and the importance level of the data stored on the target disk. For example, a lower probability threshold can be set for disk types with poor risk resistance and higher importance levels.

[0046] When the abnormality probability is greater than the probability threshold, the target disk is more likely to have an abnormality and needs to be processed for the abnormality. At this time, the target disk can be processed for the abnormality according to the abnormality processing method that matches the predicted abnormality type to perform the abnormality processing on the target disk in a targeted manner.

[0047] For example, you can use other disks to back up the data stored on the target disk, or use other nodes to back up the data stored on the node where the target disk is located to avoid data loss. For another example, you can use other nodes to take over the tasks performed by the node where the target disk is located to avoid service interruption.

[0048] According to an embodiment of the present invention, by utilizing the target disk's disk type and the importance level of the stored data to dynamically determine the probability threshold for executing exception handling operations on the target disk, the timing for implementing exception handling methods can be accurately determined, allowing for preemptive exception handling of the target disk before an anomaly occurs, thereby reducing data loss or service interruptions caused by target disk anomalies. Furthermore, if the target disk's anomaly probability exceeds the probability threshold, an exception handling method matching the predicted anomaly type is employed. This allows for the combined use of information from both the fault type and the fault probability dimensions to determine a targeted exception handling method, thereby improving the effectiveness of exception handling.

[0049] According to an embodiment of the present invention, based on the predicted exception type, an exception handling method matching the predicted exception type is adopted to perform exception handling on the target disk, including: in a case where the predicted exception type is a disk failure type or a performance bottleneck type, determining a backup node from multiple nodes based on the predicted exception type and the node status data of each of the multiple nodes in the server cluster; backing up the storage data of the target node where the target disk is located to the backup node, and switching the task currently executed by the target node to be executed by the backup node; and in a case where the predicted exception type is a disk array degradation type, backing up the storage data of the target disk, and switching the task currently executed by the target disk to be executed by a backup disk in the same disk array as the target disk.

[0050] In an embodiment of the present invention, the predicted anomaly type may include a disk failure type, a performance bottleneck type, and a disk array degradation type. The disk failure type may indicate that the target disk is likely to fail, the performance bottleneck type may indicate that the target disk has reached a performance bottleneck, and the disk array degradation type may indicate that other disks in the disk array where the target disk resides have failed, causing the disk array to degrade.

[0051] Because abnormalities such as disk failure, performance bottleneck, and disk array degradation may cause the target disk's read and write performance to degrade, and may even result in data loss and service interruption, when the probability of the target disk abnormality is high, you can back up data in a timely manner or use a neighboring node to take over the service.

[0052] When the predicted exception type is a disk failure type or a performance bottleneck type, since the target disk may store data of higher importance, in order to avoid the target disk exception causing node exception, which in turn causes data loss or service interruption, the backup node can be used to back up the data of the target node where the target disk is located, and the backup node can be used to take over the service of the target node.

[0053] When using the backup node to back up the data of the target node where the target disk is located, since the target node may include multiple disks, the storage data of these multiple disks can be backed up to the backup node, and the sharding strategy of the distributed file system can be adjusted.

[0054] When the predicted exception type is disk array degradation, the target disk will be affected by the faulty disk in the same disk array, resulting in decreased read and write speeds. At this time, the faulty disk needs to be processed to restore the disk array level in a timely manner to reduce the impact on the target disk.

[0055] Since a backup disk is usually provided in a disk array, the backup disk can be used to back up the data in the failed disk and perform the task that the failed disk was performing.

[0056] In some embodiments, the backup disk may be a hot spare disk in a disk array. When the backup disk is used to back up data in a failed disk, all stored data in the failed disk may be backed up to the backup disk.

[0057] In other embodiments, the backup disk may be a high-performance disk in a disk array, and hotspot data in a failed disk may be migrated to the backup disk, thereby triggering redundancy checking.

[0058] According to the embodiment of the present invention, different exception handling methods are used for different predicted exception types, so that exception handling is targeted, thereby improving the effectiveness of exception handling for the target disk.

[0059] According to an embodiment of the present invention, a backup node is determined from a plurality of nodes based on a predicted anomaly type and respective node status data of a plurality of nodes in a server cluster, including: determining the disk status score, load status score, geographic location score and disk capacity score of the plurality of nodes based on the respective disk status sub-data, load status sub-data, geographic location sub-data and disk capacity sub-data of the plurality of nodes; determining the node score of the plurality of nodes based on the respective disk status score, load status score, geographic location score and disk capacity score of the plurality of nodes; determining a backup node set from the plurality of nodes based on the respective node scores of the plurality of nodes; in a case where the predicted anomaly type is a disk failure type, determining the node in the backup node set whose disk status score meets a first predetermined condition as a backup node; and in a case where the predicted anomaly type is a performance bottleneck type, determining the node in the backup node set whose load status score meets a second predetermined condition as a backup node.

[0060] According to an embodiment of the present invention, the node status data includes disk status sub-data, load status sub-data, geographic location sub-data, and disk capacity sub-data.

[0061] In some embodiments, the disk status sub-data may represent the health status of multiple disks in a node. For example, a health score for each disk in a node may be determined based on the status data of each disk, and then the node's disk status score may be determined based on the average of the health scores of the multiple disks.

[0062] When determining a disk's health score, you can set an initial health score for the disk. Using rule matching, the initial health score is updated based on the disk's status data to determine the disk's real-time health score. For example, if the disk's remaining life is less than 10% and the temperature is greater than 70°C, the health score is reduced by 30%.

[0063] When determining the load status score, the load status score can be determined based on the read and write request processing volume of the node's central processing unit and disk. For example, the load status score can be determined based on the read and write request volume interval in which the read and write request volume falls.

[0064] When determining the geographic location score, the geographic location sub-data may be the server cabinet where the node is located, and the geographic location score may be determined based on the distance between the server cabinet where the node is located and the server cabinet where the target node is located.

[0065] When determining the disk capacity score, the total remaining disk capacity of the node can be determined based on the remaining capacity of each of the multiple disks in the node, and the node can be scored based on the total remaining disk capacity.

[0066] When determining the node score, different weights can be set for the disk status score, load status score, geographic location score, and disk capacity score, and the node score is determined as the score obtained by weighted summation of the disk status score, load status score, geographic location score, and disk capacity score.

[0067] In some embodiments, before determining the set of backup nodes, a preliminary screening of multiple nodes may be performed. For example, nodes with a disk status score less than 80, nodes with a load status score less than 70, nodes that are not in the same cabinet as the target node, and nodes with low disk capacity scores may be excluded.

[0068] After preliminary screening of multiple nodes, an initial node set is obtained, and then multiple nodes with node scores greater than a preset score threshold can be determined from the initial node set as a backup node set.

[0069] In the embodiment of the present invention, when selecting a backup node from the backup node set, the backup nodes may be screened according to different predicted anomaly types and using different preset conditions.

[0070] When the predicted abnormality type is a disk failure, more attention may be paid to the node's disk health status, and thus a node whose disk health score meets a first predetermined condition may be selected as a backup node. The first predetermined condition may be, for example, the highest disk health score.

[0071] When the predicted abnormality type is a performance bottleneck type, more attention may be paid to the performance of the node, and thus a node whose load status score meets a second predetermined condition may be selected as a backup node. The second predetermined condition may be, for example, the highest load status score.

[0072] According to an embodiment of the present invention, a backup node that matches the predicted exception type is determined based on the node's disk status sub-data, load status sub-data, geographic location sub-data, and disk capacity sub-data, so that the backup node has a higher matching degree with the target node, thereby improving the exception handling efficiency.

[0073] According to an embodiment of the present invention, a backup node set is determined from a plurality of nodes based on respective node scores of the plurality of nodes, including: for each node, determining the abnormality probability of each disk based on status data of each disk in the node; determining the disk abnormality score of each node based on the abnormality probability of each disk; determining a plurality of initial backup nodes from the plurality of nodes based on the disk abnormality scores; and determining a backup node set from the plurality of nodes based on the respective node scores of the plurality of initial backup nodes.

[0074] Since the data stored on the target disk may be of high importance, in order to avoid data loss, when selecting a backup node from multiple nodes, it is also necessary to determine the disk anomaly score of each node based on the status data of each disk in each node, so as to avoid transferring the data stored on the target disk to a node or disk with a higher anomaly risk.

[0075] In some embodiments, for each node, the maximum value of the abnormality probabilities of multiple disks in the node can be used as the disk abnormality probability of the node, and the disk abnormality score of the node can be determined based on the disk abnormality probability, so that the disk with the highest risk in the node can be considered when selecting the node, thereby reducing the data loss caused by backup node abnormality after data backup and improving the reliability of the data backup process.

[0076] According to an embodiment of the present invention, the abnormality of the disk in each node is taken into consideration when selecting a backup node, so that nodes with lower abnormality risks can be screened out, thereby improving the reliability of the data backup process.

[0077] According to an embodiment of the present invention, the abnormality probability and predicted abnormality type of the target disk are determined based on the status data of the target disk, including: using an abnormality prediction model to process the status data to obtain the abnormality probability, predicted abnormality type and predicted status data.

[0078] In an embodiment of the present invention, the anomaly prediction model includes an encoder and an output layer connected in series with the encoder. The encoder can be obtained by adopting a deep learning model (Transformer) based on the attention mechanism combined with a temporal convolutional network (TCN), and is used to encode the state data to obtain state features. Accordingly, the state data can be the time series state data of the target disk within 1 minute before the current moment, and the predicted state data can be obtained.

[0079] In some embodiments, the time series state data can be stored in a time series database. After obtaining state data of different dimensions from the hardware layer, the driver layer, and the application layer, the state data can be indexed according to the node identifier of the target node where the target disk is located, the timestamp, and the type of the state data, and the indexed state data can be stored in the time series database.

[0080] When obtaining status data from the hardware layer, the disk's interface protocol is used to obtain status data such as disk wear, remaining life, temperature, and error rate, and the disk's health status data is obtained through SMART data monitoring.

[0081] When obtaining status data from the driver layer, a probe based on the extended Berkeley packet filter can be deployed in the operating system kernel storage driver stack of the target node where the target disk is located to capture status data such as input and output queue depth anomalies, direct memory access transmission timeouts, and driver layer error codes in real time.

[0082] When obtaining data from the application layer, parameters such as the number of read and write requests, latency, and cache hit rate of the node file system can be obtained as status data, and can be combined with business load characteristics such as read and write ratio for correlation analysis.

[0083] Optionally, the status data in the time series database can be visualized, such as using visualization tools to display the topology of the server cluster, health score heat map, and performance trends in real time.

[0084] Optionally, for the time series status data in the time series database, a data aggregation strategy can be adopted to count the average value, maximum value, and standard deviation of the time series status data every 5 minutes, store the average value, maximum value, and standard deviation of the time series status data, and delete the original time series status data to retain the status data change characteristics while reducing the data volume and saving storage space.

[0085] Since the anomaly prediction model in the present invention needs to simultaneously output the anomaly probability, predicted anomaly type and predicted state data, a multi-task learning branch can be added to the output layer. Specifically, the output layer is configured to include a first sub-output layer, a second sub-output layer and a third sub-output layer connected in parallel. The first sub-output layer is used to generate the anomaly probability based on the state characteristics, the second sub-output layer is used to generate the predicted anomaly type based on the state characteristics, and the third sub-output layer is used to generate the predicted state data based on the state characteristics.

[0086] In some embodiments, the first sub-output layer may include an activation function, such as a sigmoid function, for outputting anomaly probability, the second sub-output layer may include an activation function, such as a softmax function, for outputting predicted anomaly type, and the third sub-output layer may output predicted status data through regression, such as the read and write speed of the target disk in the next 5 minutes.

[0087] When training the anomaly prediction model, a different loss function can be used for each sub-output layer to specifically improve the accuracy of each sub-output layer, and the loss function of each sub-output layer can be further added to obtain the joint loss function of the entire anomaly prediction model to improve the overall accuracy of the anomaly prediction model.

[0088] For example, the first sample state data of the sample disk can be input into the initial abnormality prediction model to obtain the initial abnormality probability, initial predicted abnormality type and initial predicted state data output by the initial abnormality prediction model; the sample abnormality state, sample abnormality type and second sample state data of the sample disk are obtained as labels, wherein the sample abnormality state includes an abnormality probability of 100% or an abnormality probability of 0%, and the time period corresponding to the second sample state data is after the first sample state data; using the respective loss functions of each sub-output layer, the first loss value between the initial abnormality probability and the sample abnormality state, the second loss value between the initial predicted abnormality type and the sample abnormality type, the third loss value between the initial predicted state data and the second sample state data, and the joint loss value are determined; the parameters of the first initial sub-output layer are adjusted according to the first loss value, the parameters of the second initial sub-output layer are adjusted using the second loss value, the parameters of the third initial sub-output layer are adjusted using the third loss value, and the parameters of the initial abnormality prediction model as a whole are adjusted using the joint loss value until the first loss value, the second loss value, the third loss value and the joint loss value meet the preset conditions to obtain the abnormality prediction model.

[0089] According to an embodiment of the present invention, by predicting the abnormal state of the target disk from three dimensions: abnormality probability, predicted abnormality type and predicted status data, the status of the target disk can be more comprehensively characterized from the predicted data of multiple dimensions, thereby improving the accuracy of determining the abnormality handling method based on the abnormality probability, predicted abnormality type and predicted status data.

[0090] Figure 3 A flow chart of dynamic multi-dimensional monitoring according to an embodiment of the present invention is shown.

[0091] like Figure 3 As shown, dynamic multi-dimensional monitoring of a disk includes operations S310 to S360.

[0092] In operation S310 , status data of the disk at the hardware layer is collected.

[0093] In operation S320 , state data of the disk at the driver layer is collected.

[0094] In operation S330 , state data of the disk at the application layer is collected.

[0095] In operation S340 , indexes are created for the status data of the disk at the hardware layer, the status data of the driver layer, and the status data of the application layer, and the indexes are stored in a time series database.

[0096] In operation S350 , data aggregation is performed on the historical data in the time series database.

[0097] In operation S360 , the state data in the time series database is visualized.

[0098] Figure 4 A schematic diagram of an abnormality prediction model according to an embodiment of the present invention is shown.

[0099] like Figure 4 As shown, the abnormality prediction model M400 includes an encoder M410 and an output layer M420 connected in series with the encoder M410. The output layer M420 includes a first sub-output layer M421, a second sub-output layer M422, and a third sub-output layer M423 connected in parallel.

[0100] like Figure 4 As shown, when the abnormality prediction model M400 is used to process the state data 401, the state data 401 can be input into the encoder M410, which outputs the state feature 402. After the state feature 402 is obtained, the state feature 402 can be input into the first sub-output layer M421, the second sub-output layer M422, and the third sub-output layer M423, respectively, to obtain the abnormality probability 403 output by the first sub-output layer M421, the predicted abnormality type 404 output by the second sub-output layer M422, and the predicted state data 405 output by the third sub-output layer M423.

[0101] According to an embodiment of the present invention, the disk exception handling method also includes: when the exception probability is less than or equal to a probability threshold, analyzing the predicted state data of the target disk in a future time period to obtain state change information of the target disk; and when the state change information represents that the state change degree of the target disk meets a preset condition, using an exception handling method that matches the predicted exception type to perform exception handling on the target disk.

[0102] After the predicted status data is generated by the abnormality prediction model, the abnormality handling method for the target disk may be determined by further combining the predicted status data with the abnormality probability and the predicted abnormality type.

[0103] In an embodiment of the present invention, in order to avoid failure to timely perform exception processing on the target disk due to inaccurate prediction of the exception probability, the abnormal state of the target disk can be further determined in combination with the predicted state data when the exception probability is less than or equal to the probability threshold, and the target disk can still be processed for exception when the state change information indicates that the state change degree of the target disk meets the preset conditions.

[0104] When determining the state change information of the target disk, the state change information may be determined based on the maximum value and the minimum value of the predicted state data in the predicted state data for the future period, for example, the read / write speed decreases by 60%.

[0105] After determining the state change information, it can be determined whether the state change information meets a preset condition. The preset condition can, for example, indicate that the state change degree of the target disk is greater than a preset threshold. For example, the state change information is determined to meet the preset condition when the read / write speed change degree is greater than 50%.

[0106] In an embodiment of the present invention, when an exception handling method matching the predicted exception type is used to handle the exception of the target disk, the specific handling method is similar to the above method for handling the exception of the target disk, and will not be repeated here.

[0107] According to an embodiment of the present invention, by performing exception processing on the target disk when the state change information represents that the state change degree of the target disk meets the preset conditions, the timing of performing exception processing is determined by combining data from two dimensions, namely, the abnormality probability and the predicted state data. This can reduce the failure to perform exception processing in a timely manner due to misjudgment and reduce the occurrence of abnormalities in the target disk.

[0108] According to an embodiment of the present invention, the disk abnormality processing method further includes: generating alarm information for the target disk when the state change information indicates that the state change degree of the target disk does not meet a preset condition.

[0109] When the abnormal probability of the target disk is less than or equal to the preset threshold and the state change degree of the target disk does not meet the preset conditions, the probability of the target disk being abnormal is small, so there is no need to handle the abnormality of the target disk to avoid waste of resources.

[0110] Since the target disk still has the risk of abnormality, although the target disk is not handled abnormally, an alarm message for the target disk can be generated to prompt the operation and maintenance personnel, so that the operation and maintenance personnel can manually determine whether the target disk needs to be handled abnormally, thereby reducing service interruption or data loss caused by target disk abnormality.

[0111] According to an embodiment of the present invention, when the degree of state change does not meet the preset conditions and the abnormality probability is low, an alarm message is directly generated to reduce resource waste.

[0112] Figure 5 A flowchart of an exception handling method according to an embodiment of the present invention is shown.

[0113] like Figure 5 As shown, the exception handling method includes operations S510 to S560.

[0114] In operation S510 , an abnormality probability, a predicted abnormality type, and predicted status data of the target disk are determined based on status data of the target disk.

[0115] In operation S520, it is determined whether the abnormal probability is greater than a probability threshold. If the abnormal probability is greater than the probability threshold, operation S550 is performed, otherwise operation S530 is performed.

[0116] In operation S530 , the predicted state data is analyzed to obtain state change information representing a degree of state change of the target disk in a future period.

[0117] In operation S540, it is determined whether the degree of change in the state of the target disk satisfies a preset condition. If the degree of change in the state of the target disk satisfies the preset condition, operation S550 is performed, otherwise operation S560 is performed.

[0118] In operation S550 , an exception handling method matching the predicted exception type is used to perform exception handling on the target disk.

[0119] In operation S560 , warning information for the target disk is generated.

[0120] According to an embodiment of the present invention, a probability threshold for a target disk is determined based on the disk type of the target disk and the importance level of data stored in the target disk, including: determining a first probability threshold for the disk type based on the disk type; determining a second probability threshold for the importance level based on the importance level; and determining a probability threshold based on the first probability threshold and the second probability threshold.

[0121] In an embodiment of the present invention, different probability thresholds may be set for different disk types and importance levels, so that when determining the probability threshold for a target disk, the probability threshold can be directly determined based on the disk type and importance level.

[0122] Optionally, a mapping relationship between disk types and probability thresholds may be pre-stored in a first mapping table. When determining the first probability threshold for a disk type, a match may be performed in the first mapping table based on the disk type, and the probability threshold matching the disk type may be determined as the first probability threshold.

[0123] When setting the mapping relationship between disk types and probability thresholds, you can set a lower probability threshold for disks with poor risk resistance to prevent abnormalities in these disks in advance.

[0124] Optionally, a mapping relationship between importance levels and probability thresholds may be pre-stored in the second mapping table. When determining the second probability threshold for the importance level, a match may be performed in the second mapping table according to the importance level, and the probability threshold matching the importance level may be determined as the second probability threshold.

[0125] When setting the mapping relationship between importance levels and probability thresholds, a lower probability threshold can be set for a higher importance level to ensure the security of important data.

[0126] According to an embodiment of the present invention, by combining data in two dimensions, namely, disk type and importance level, and dynamically determining a probability threshold, the timing for handling an exception on a target disk can be determined more accurately.

[0127] Based on the above disk exception handling method, the present invention also provides a disk exception handling device. Figure 6 The device is described in detail.

[0128] Figure 6 A structural block diagram of a disk exception processing device according to an embodiment of the present invention is shown.

[0129] like Figure 6 As shown, the disk exception handling device 600 of this embodiment includes an exception determination module 610 , a threshold determination module 620 and an exception handling module 630 .

[0130] The abnormality determination module 610 is used to determine the abnormality probability and predict the abnormality type of the target disk according to the status data of the target disk. In one embodiment, the abnormality determination module 610 can be used to perform the operation S210 described above, which will not be repeated here.

[0131] Threshold determination module 620 is configured to determine a probability threshold for the target disk based on the disk type of the target disk and the importance level of the data stored on the target disk, when the abnormality probability indicates that the target disk has an abnormality risk. In one embodiment, threshold determination module 620 can be configured to perform operation S220 described above, and will not be further described here.

[0132] The exception handling module 630 is used to perform exception handling on the target disk based on the predicted exception type when the exception probability is greater than the probability threshold. In one embodiment, the exception handling module 630 can be used to perform operation S230 described above, which will not be repeated here.

[0133] According to an embodiment of the present invention, the exception handling module 630 includes a first processing submodule and a second processing submodule.

[0134] The first processing sub-module is used to determine a backup node from multiple nodes based on the predicted exception type and the node status data of multiple nodes in the server cluster when the predicted exception type is a disk failure type or a performance bottleneck type; back up the storage data of the target node where the target disk is located to the backup node, and switch the task currently executed by the target node to be executed by the backup node.

[0135] The second processing submodule is used to back up the storage data of the target disk when the predicted abnormality type is the disk array degradation type, and switch the task currently executed by the target disk to a backup disk in the same disk array as the target disk.

[0136] According to an embodiment of the present invention, the node status data includes disk status sub-data, load status sub-data, geographic location sub-data, and disk capacity sub-data.

[0137] According to an embodiment of the present invention, the first processing submodule includes a first scoring unit, a second scoring unit, a set determination unit, a fault determination unit, and a performance determination unit.

[0138] The first scoring unit is used to determine the disk status score, load status score, geographic location score and disk capacity score of each of the multiple nodes based on the disk status sub-data, load status sub-data, geographic location sub-data and disk capacity sub-data of each of the multiple nodes.

[0139] The second scoring unit is configured to determine a node score for each of the multiple nodes based on the disk status score, load status score, geographic location score, and disk capacity score of each of the multiple nodes.

[0140] The set determining unit is used to determine a backup node set from a plurality of nodes according to respective node scores of the plurality of nodes.

[0141] The fault determination unit is configured to, when the predicted abnormality type is a disk fault type, determine a node in the backup node set whose disk status score meets a first predetermined condition as a backup node.

[0142] The performance determination unit is configured to, when the predicted abnormality type is a performance bottleneck type, determine a node in the backup node set whose load status score meets a second predetermined condition as a backup node.

[0143] According to an embodiment of the present invention, the abnormality determination module 610 includes an abnormality determination submodule.

[0144] The anomaly determination submodule is used to process the state data using the anomaly prediction model to obtain anomaly probability, predicted anomaly type and predicted state data; wherein the anomaly prediction model includes an encoder and an output layer connected in series with the encoder, the output layer includes a first sub-output layer, a second sub-output layer and a third sub-output layer connected in parallel, the encoder is used to encode the state data to obtain state features, the first sub-output layer is used to generate anomaly probability based on the state features, the second sub-output layer is used to generate predicted anomaly type based on the state features, and the third sub-output layer is used to generate predicted state data based on the state features.

[0145] According to an embodiment of the present invention, the threshold determination module 620 includes a first determination submodule, a second determination submodule, and a third determination submodule.

[0146] The first determination submodule is configured to determine a first probability threshold for the disk type according to the disk type.

[0147] The second determination submodule is configured to determine a second probability threshold for the importance level according to the importance level.

[0148] The third determination submodule is configured to determine a probability threshold according to the first probability threshold and the second probability threshold.

[0149] According to an embodiment of the present invention, the disk exception handling device 600 further includes a change determination module and a change processing module.

[0150] The change determination module is used to analyze the predicted state data of the target disk in the future period when the abnormal probability is less than or equal to the probability threshold, and obtain the state change information of the target disk.

[0151] The change processing module is used to perform exception processing on the target disk using an exception processing method that matches the predicted exception type when the state change information indicates that the state change degree of the target disk meets a preset condition.

[0152] According to an embodiment of the present invention, the disk exception processing device 600 further includes an alarm generation module.

[0153] The alarm generating module is used for generating alarm information for the target disk when the state change information indicates that the state change degree of the target disk does not meet a preset condition.

[0154] According to embodiments of the present invention, any multiple modules among the anomaly determination module 610, the threshold determination module 620, and the anomaly handling module 630 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present invention, at least one of the anomaly determination module 610, the threshold determination module 620, and the anomaly handling module 630 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of these. Alternatively, at least one of the anomaly determination module 610, the threshold determination module 620, and the anomaly handling module 630 may be at least partially implemented as a computer program module that, when executed, performs the corresponding functionality.

[0155] Figure 7 A block diagram of an electronic device suitable for implementing a disk exception processing method according to an embodiment of the present invention is shown.

[0156] like Figure 7 As shown, an electronic device 700 according to an embodiment of the present invention includes a processor 701, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 702 or programs loaded from a storage unit 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or related chipsets and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include onboard memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0157] The RAM 703 stores various programs and data required for the operation of the electronic device 700. The processor 701, ROM 702, and RAM 703 are connected to each other via a bus 704. The processor 701 executes the programs in the ROM 702 and / or RAM 703 to perform the various operations of the method flow according to the embodiment of the present invention. It should be noted that the programs may also be stored in one or more memories other than the ROM 702 and RAM 703. The processor 701 may also execute the programs stored in one or more memories to perform the various operations of the method flow according to the embodiment of the present invention.

[0158] According to an embodiment of the present invention, electronic device 700 may further include an input / output (I / O) interface 705, which is also connected to bus 704. Electronic device 700 may also include one or more of the following components connected to I / O interface 705: an input section 706 including a keyboard, mouse, etc.; an output section 707 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 708 including a hard disk; and a communication section 709 including a network interface card such as a LAN card or modem. Communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to I / O interface 705 as needed. Removable media 711, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 710 as needed, so that computer programs read from the removable media can be installed into storage section 708 as needed.

[0159] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the present invention.

[0160] According to an embodiment of the present invention, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, a computer-readable storage medium may include the ROM 702 and / or RAM 703 described above, and / or one or more memories other than ROM 702 and RAM 703.

[0161] An embodiment of the present invention further includes a computer program product comprising a computer program containing program code for executing the method shown in the flowchart. When the computer program product is executed in a computer system, the program code is used to cause the computer system to implement the disk exception handling method provided in an embodiment of the present invention.

[0162] The computer program executes the above functions defined in the system / device of the embodiment of the present invention when the computer program is executed by the processor 701. According to the embodiment of the present invention, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0163] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 709, and / or installed from a removable medium 711. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0164] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709 and / or installed from the removable medium 711. When the computer program is executed by the processor 701, the above-described functions defined in the system of the embodiment of the present invention are performed. According to the embodiment of the present invention, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.

[0165] According to an embodiment of the present invention, the program code for executing the computer program provided by the embodiment of the present invention can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).

[0166] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0167] It will be understood by those skilled in the art that the features described in the various embodiments of the present invention may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention may be combined and / or coupled in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or couplings fall within the scope of the present invention.

[0168] The above describes embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.

Claims

1. A disk exception handling method, characterized in that: The method comprises: Determining, based on the state data of the target disk, an abnormality probability and a predicted abnormality type of the target disk; In a case where the abnormality probability indicates that the target disk has an abnormality risk, determining a probability threshold for the target disk according to a disk type of the target disk and an importance level of data stored in the target disk; and When the abnormality probability is greater than the probability threshold, according to the predicted abnormality type, an abnormality handling method matching the predicted abnormality type is adopted, so as to perform abnormality handling on the target disk by using the abnormality handling method determined by combining information of the predicted abnormality type and the abnormality probability, including: In a case where the predicted abnormality type is a disk failure type or a performance bottleneck type, determining a backup node from the multiple nodes based on the predicted abnormality type and node status data of each of the multiple nodes in the server cluster; backing up the storage data of the target node where the target disk is located to the backup node, and switching the task currently executed by the target node to be executed by the backup node; and When the predicted abnormality type is a disk array degradation type, a backup operation is performed on the storage data of the target disk, and the task currently executed by the target disk is switched to a backup disk in the same disk array as the target disk.

2. The method according to claim 1, characterized in that The node status data includes disk status sub-data, load status sub-data, geographic location sub-data and disk capacity sub-data; The determining a backup node from the multiple nodes according to the predicted abnormality type and the node status data of each of the multiple nodes in the server cluster includes: Determine, based on the disk status sub-data, load status sub-data, geographic location sub-data, and disk capacity sub-data of each of the plurality of nodes, a disk status score, a load status score, a geographic location score, and a disk capacity score of each of the plurality of nodes; Determining a node score for each of the plurality of nodes based on the disk status score, load status score, geographic location score, and disk capacity score of each of the plurality of nodes; Determining a backup node set from the plurality of nodes according to respective node scores of the plurality of nodes; In a case where the predicted abnormality type is a disk failure type, determining a node in the backup node set whose disk status score meets a first predetermined condition as the backup node; and In a case where the predicted abnormality type is a performance bottleneck type, a node in the backup node set whose load status score meets a second predetermined condition is determined as the backup node.

3. The method according to claim 1, characterized in that The determining, based on the status data of the target disk, the abnormality probability and the predicted abnormality type of the target disk, includes: Using an abnormality prediction model, the state data is processed to obtain the abnormality probability, the predicted abnormality type and the predicted state data; In which, the abnormality prediction model includes an encoder and an output layer connected in series with the encoder, the output layer includes a first sub-output layer, a second sub-output layer and a third sub-output layer connected in parallel, the encoder is used to encode the state data to obtain state features, the first sub-output layer is used to generate the abnormality probability according to the state features, the second sub-output layer is used to generate the predicted abnormality type according to the state features, and the third sub-output layer is used to generate the predicted state data according to the state features.

4. The method according to claim 1, wherein The determining of the probability threshold for the target disk according to the disk type of the target disk and the importance level of the data stored in the target disk includes: determining a first probability threshold for the disk type based on the disk type; determining a second probability threshold for the importance level based on the importance level; and The probability threshold is determined according to the first probability threshold and the second probability threshold.

5. The method according to claim 1, wherein The method further comprises: When the abnormal probability is less than or equal to the probability threshold, analyzing predicted state data of the target disk in a future period to obtain state change information of the target disk; and When the state change information indicates that the state change degree of the target disk meets a preset condition, an exception handling method matching the predicted exception type is adopted to perform exception handling on the target disk.

6. The method according to claim 5, characterized in that The method further comprises: When the status change information indicates that the status change degree of the target disk does not meet a preset condition, alarm information for the target disk is generated.

7. A disk abnormality processing device, characterized in that: The device comprises: An abnormality determination module, configured to determine an abnormality probability and predict an abnormality type of the target disk according to the status data of the target disk; a threshold determination module, configured to determine a probability threshold for the target disk based on a disk type of the target disk and an importance level of data stored in the target disk, when the abnormal probability indicates that the target disk has an abnormality risk; and an exception handling module, configured to, when the exception probability is greater than the probability threshold, adopt an exception handling method that matches the predicted exception type according to the predicted exception type, so as to perform exception handling on the target disk by utilizing the exception handling method determined by combining information of the predicted exception type and the exception probability; The exception handling module includes: A first processing submodule is configured to, when the predicted abnormality type is a disk failure type or a performance bottleneck type, determine a backup node from the multiple nodes based on the predicted abnormality type and node status data of each of the multiple nodes in the server cluster; back up the storage data of the target node where the target disk is located to the backup node, and switch the task currently executed by the target node to be executed by the backup node; The second processing submodule is used to perform a back-up operation on the storage data of the target disk when the predicted abnormality type is a disk array degradation type, and switch the task currently executed by the target disk to a backup disk in the same disk array as the target disk.

8. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Fault recovery method and device for hard disk in server, equipment and storage medium

    CN119621394A

  • Hard disk fault warning method and device

    CN120276944A