Disk exception determination method and device, program product and storage medium

By conducting multi-dimensional data analysis and scientific quantitative processing of the target disk, the problem of prolonged business interruption caused by disk anomalies, which cannot be identified in a timely manner in existing technologies, has been solved, enabling accurate identification of disk failures and reduction of business interruption time.

CN120994478APending Publication Date: 2025-11-21JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511082408.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies cannot promptly detect disk anomalies, leading to prolonged business interruptions.

Method used

By acquiring the target disk's performance data, transmission latency data, and abnormal disk data, performing mapping and normalization operations, calculating the disk abnormality probability, and determining the disk abnormality when the probability is greater than a preset threshold.

Benefits of technology

It improves the accuracy and reliability of disk anomaly detection, enables early prediction of disk failures, reduces the risk of business interruption, and achieves scientific allocation and dynamic scheduling of resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994478A_ABST
    Figure CN120994478A_ABST
Patent Text Reader

Abstract

The invention discloses a disk exception determination method and device, a program product and a storage medium, and relates to the field of computers.The method comprises the steps that first data, second data and third data of a target disk are obtained; respectively executing a first processing operation on the first data, the second data and the third data to obtain a first target value, a second target value and a third target value; determining the probability that the target disk is abnormal based on the first target value, the second target value and the third target value; and when the probability is greater than a preset threshold, determining that the target disk is abnormal. Through scientific quantitative processing of the data related to the target disk, the problem that the service is interrupted for a long time due to the fact that the life cycle of the disk cannot be predicted in time in related technologies is solved, and the technical effects of accurately judging the disk fault and reducing the service interruption time are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computers, and particularly relates to a method and device for determining disk abnormalities, a program product and a storage medium. BACKGROUND

[0002] In the global range, the size of enterprise-level storage systems is increasingly large, and disk failures have become a key factor affecting the reliability of storage systems. In the related art, the passive data reconstruction process not only consumes a large amount of resources, but also leads to a decline in service quality.

[0003] In the face of this challenge, the related art mainly focuses on discrete fault responses, but fails to determine disk abnormalities in a timely manner, thereby still having problems of long-term interruption of business and inability to handle business in a timely manner. SUMMARY

[0004] The present application provides a method and device for determining disk abnormalities, a program product and a storage medium to at least solve the problem that the related art cannot predict the life cycle of a disk in a timely manner, resulting in long-term interruption of business.

[0005] The present application provides a method for determining disk abnormalities, comprising: obtaining first data, second data and third data of a target disk, wherein the first data is used to represent the performance of the target disk, the second data is used to represent the transmission delay of the target disk, and the third data is used to represent the abnormality of an abnormal disk, and the abnormal disk is of the same type as the target disk; performing a first processing operation on the first data, the second data and the third data to obtain a first target value, a second target value and a third target value, wherein the first processing operation comprises the following steps: performing a mapping operation on the first data to obtain the first target value, performing a normalization operation on the second data to obtain the second target value, and performing the normalization operation on the third data to obtain the third target value; determining a probability of the target disk having an abnormality based on the first target value, the second target value and the third target value; and determining that the target disk has an abnormality when the probability is greater than a preset threshold.

[0006] The application further provides a disk abnormality determination device, comprising: an acquisition module, configured to acquire first data, second data and third data of a target disk, wherein the first data is used to represent performance of the target disk, the second data is used to represent transmission delay of the target disk, and the third data is used to represent an abnormality of an abnormal disk which is of the same type as the target disk; a processing module, configured to perform a first processing operation on the first data, the second data and the third data respectively to obtain a first target value, a second target value and a third target value, wherein the first processing operation comprises the following steps: performing a mapping operation on the first data to obtain the first target value, performing a normalization operation on the second data to obtain the second target value, and performing the normalization operation on the third data to obtain the third target value; a first determination module, configured to determine a probability of the target disk having an abnormality based on the first target value, the second target value and the third target value; and a second determination module, configured to determine that the target disk has an abnormality when the probability is greater than a preset threshold.

[0007] According to still another embodiment of the application, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the steps of any of the method embodiments.

[0008] According to still another embodiment of the application, a computer readable storage medium is provided, comprising a computer program stored therein, wherein the computer program is configured to execute the steps of any of the method embodiments when running.

[0009] According to still another embodiment of the application, an electronic device is provided, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the computer program to perform the steps of any of the method embodiments.

[0010] This application performs a mapping operation on the first data of the acquired target disk to obtain a first target value, performs a normalization operation on the second data to obtain a second target value, and performs a normalization operation on the third data to obtain a third target value. Then, when the probability of the target disk exhibiting an anomaly, determined based on the first, second, and third target values, exceeds a preset threshold, the target disk is determined to be an anomaly. Therefore, through comprehensive analysis of multi-dimensional data from the target disk and scientific quantification of data related to the target disk, the accuracy and reliability of disk anomaly detection are improved. This enables early prediction of disk failures, determination of resource usage, and facilitates subsequent dynamic resource allocation, reducing the risk of business interruption. It solves the problem in related technologies where timely prediction of disk lifecycle is impossible, leading to prolonged business interruptions, and achieves the technical effect of accurately identifying disk failures and reducing business interruption time. Attached Figure Description

[0011] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A schematic diagram of the hardware environment for a method for determining disk anomalies provided in an embodiment of this application;

[0013] Figure 2 A flowchart illustrating a method for determining disk anomalies provided in an embodiment of this application;

[0014] Figure 3 A flowchart illustrating the execution of a target operation on an abnormal disk, provided in an embodiment of this application;

[0015] Figure 4 A flowchart illustrating a method for determining disk anomalies in a distributed storage system, provided as an embodiment of this application;

[0016] Figure 5 This is a structural block diagram of a disk anomaly determination device provided in an embodiment of this application. Detailed Implementation

[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0018] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0019] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0020] The methods and embodiments provided in this application can be executed on a server device or a similar computing device. Taking running on a server device as an example, Figure 1 This is a schematic diagram of the hardware environment for a method for determining disk anomalies according to an embodiment of this application. Figure 1 As shown, the server device may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The server device may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the server equipment described above. For example, the server equipment may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0021] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to a disk anomaly determination method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thus implementing the aforementioned method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to server devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0022] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the server device. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0023] This embodiment provides a method for determining disk anomalies. Figure 2 This is a flowchart of a method for determining disk anomalies according to an embodiment of this application, as follows: Figure 2 As shown, the process includes the following steps:

[0024] Step S202: Obtain first data, second data, and third data of the target disk, wherein the first data is used to represent the performance of the target disk, the second data is used to represent the transmission latency of the target disk, and the third data is used to represent the abnormality of the abnormal disk, wherein the abnormal disk is of the same type as the target disk.

[0025] Optionally, the solution in this embodiment is applicable to storage systems using erasure coding technology, including but not limited to Ceph distributed storage systems. Application scenarios for the solution in this embodiment include, but are not limited to: storage systems in large data centers, financial industry data storage systems in the financial sector, cloud computing infrastructure storage systems, and IoT devices.

[0026] Optionally, the first data in this embodiment is used to represent the performance of the target disk, including but not limited to SMART parameters (such as the number of reallocated sectors, seek error rate, disk temperature, etc.). These parameters can be used to determine the health status of the disk and potential performance degradation trends.

[0027] Optionally, the second data in this embodiment is used to represent the transmission latency of the target disk, including but not limited to the response time of the target disk during read and write operations.

[0028] Optionally, the third data in this embodiment is used to represent information about known abnormal disks of the same type as the target disk. It mainly involves statistical analysis of the failure status of disks in the same batch or of the same model, and is used to assess batch risk caused by manufacturing defects or design problems.

[0029] Optionally, in this embodiment, the abnormal disk is the same model as the target disk. Specifically, the abnormal disk includes, but is not limited to, disks belonging to the same batch as the target disk or belonging to the same storage system as the target disk.

[0030] Step S204: Perform a first processing operation on the first data, the second data, and the third data respectively to obtain a first target value, a second target value, and a third target value. The first processing operation includes the following steps: perform a mapping operation on the first data to obtain the first target value; perform a normalization operation on the second data to obtain the second target value; and perform the normalization operation on the third data to obtain the third target value.

[0031] Optionally, the mapping operation in this embodiment is used to convert the first data into a value that satisfies the first numerical range, reflecting the degree of disk degradation, including but not limited to using the sigmoid function for mapping.

[0032] Optionally, the normalization operation in this embodiment is used to ensure that the second target value meets the second value range and the third target value meets the third value range.

[0033] Step S206: Determine the probability that the target disk is abnormal based on the first target value, the second target value, and the third target value.

[0034] Optionally, the probability of the target disk being abnormal determined based on the first target value, the second target value, and the third target value in this embodiment includes, but is not limited to, calculations obtained by weighted combination of S (corresponding to the first target value), L (corresponding to the second target value), and B (corresponding to the third target value).

[0035] Step S208: When the above probability is greater than a preset threshold, it is determined that the above target disk is abnormal.

[0036] Optionally, the preset threshold in this embodiment includes, but is not limited to, a threshold determined based on the impact of disk anomalies in the storage system where the target disk resides on the execution of the target service.

[0037] In this embodiment, the entity executing the above steps can be a management node, monitoring software, or a specially designed fault prediction module in the storage system. It can also be a specific processor or controller installed in the storage system where the target disk resides, or a processor or processing device set up relatively independently of the storage system where the target disk resides, but is not limited to these. For example, it could be a Ceph storage cluster management node, or software deployed in a data center or cloud service for disk health monitoring and fault prediction.

[0038] Through the above steps, a mapping operation is performed on the first data of the acquired target disk to obtain a first target value. A normalization operation is performed on the second data to obtain a second target value, and a normalization operation is performed on the third data to obtain a third target value. Then, if the probability of an anomaly occurring in the target disk, determined based on the first, second, and third target values, exceeds a preset threshold, the target disk is determined to be an anomaly. Therefore, by comprehensively analyzing multi-dimensional data from the target disk and scientifically quantifying data related to the target disk, the accuracy and reliability of disk anomaly detection are improved. This allows for early prediction of disk failures, determination of resource usage, and facilitates dynamic resource allocation, reducing the risk of business interruption. It solves the problem in related technologies where timely prediction of disk lifecycles is impossible, leading to prolonged business interruptions, and achieves the technical effect of accurately identifying disk failures and reducing business interruption time.

[0039] In one exemplary embodiment, performing a mapping operation on the first data to obtain the first target value includes: parsing the first data, determining the degree of degradation of the target disk, and obtaining a first value, wherein the degree of degradation is used to represent the performance degradation level of the target disk; calculating the difference between the first value and a first threshold to obtain a first difference, wherein the first difference is used to represent the anomaly level of the target disk; calculating the product between the first difference and a second value to obtain a first product, wherein the second value is used to represent the degree of influence of the first value on the anomaly of the target disk; and mapping the first product to a value that satisfies a first value range using an objective function to obtain the first target value.

[0040] Optionally, in this embodiment, the first value is used to represent the performance degradation level of the target disk. The first value is a quantitative indicator obtained after parsing the first data, and the first value usually increases as the usage time of the target disk increases. For example, if the number of relocation sectors in the first data increases significantly, it indicates that the physical layer of the target disk is degrading, thereby determining the performance degradation level of the target disk at this time based on a preset performance degradation table.

[0041] Optionally, in this embodiment, the first threshold is a baseline value determined by historical data or experience, used to distinguish between a healthy disk state and an abnormal state. When the first value exceeds the first threshold, it indicates that the target disk is in a medium-risk state and may begin to show abnormalities.

[0042] Optionally, the first difference in this embodiment reflects the degree of deviation between the degree of disk degradation and the normal healthy state, i.e., the abnormality level.

[0043] Optionally, in this embodiment, the second value represents the weight of the first difference in the impact of disk anomalies. The second value is a pre-set value based on disk characteristics and business requirements, used to quantify the degree of impact of changes in the first value on disk anomalies.

[0044] Optionally, the objective function in this embodiment is a mathematical transformation used to ensure that the first product maps to a predefined first numerical range (typically 0 to 1). In this way, regardless of the size of the original data, the first objective value can represent the disk's anomaly risk on a uniform scale, facilitating subsequent probability calculations and comparisons.

[0045] Optionally, in this embodiment, it includes, but is not limited to, using formulas The first target value is calculated, where S represents the performance degradation index of the target disk (i.e., the first target value mentioned above), and S is a value between 0 and 1. SMART represents the SMART parameter value of the target disk (corresponding to the first value mentioned above). θ (corresponding to the first threshold mentioned above) indicates that when the SMART parameter reaches this value, the target disk is in a medium-risk state. In the Sigmoid function (corresponding to the target function mentioned above), when SMART = θ, S = 0.5. k (corresponding to the second value mentioned above) is used to control the steepness of the Sigmoid function. The larger k is, the steeper the function changes around θ, which is used to reflect the degree of impact of the SMART parameter change on the abnormality of the target disk.

[0046] For example, in a Ceph distributed storage system, administrators want to predict disk anomalies by monitoring disk health in real time, so as to perform data migration and resource scheduling in advance, reducing business interruption and data reconstruction time. The Ceph monitoring module periodically retrieves SMART data from each disk. For example, if a disk has 100 relocation sectors, historical data analysis shows that a degradation level of 0.8 (corresponding to the first value mentioned above) is determined when the number of relocation sectors is 100. The health status threshold for this disk (corresponding to the first threshold mentioned above) is set to 0.5 to distinguish between normal and abnormal disk states. The difference between the current degradation level of 0.8 and the threshold of 0.5 is 0.3 (corresponding to the first difference mentioned above), indicating that the disk performance is starting to deviate from the normal level. Assuming the second value is set to 2, reflecting the strong correlation between the degradation level and the risk of disk anomalies, the product of 0.3 and 2 is 0.6 (corresponding to the first product mentioned above), meaning that the performance anomaly of this disk significantly contributes to the overall anomaly risk. Then, the sigmoid function is used to map 0.6 to a value between 0 and 1, such as 0.89, as the first target value, indicating that the disk has a high risk of performance anomalies. The Ceph system can calculate the probability of subsequent disk anomalies based on the performance degradation level of the disk, helping to automatically identify abnormal disks and trigger subsequent dynamic resource scheduling, such as allocating more resources for data reconstruction, while protecting normal business operations from being affected, and ensuring the high reliability and high performance of the storage system.

[0047] Through the above steps, the difference between the first value parsed from the first data and the first threshold is calculated to obtain the first difference value representing the anomaly level of the target disk. Then, the first product is calculated using the first difference and the second value representing the degree of influence of the first value on the anomaly of the target disk. Finally, the first product is converted into a first target value within the first value range using an objective function. By progressively quantifying and confirming the values, the performance degradation state of the target disk is more accurately assessed, improving the sensitivity of anomaly detection. At the same time, the mapping operation ensures the consistency of the value range, facilitating subsequent probability calculations.

[0048] In one exemplary embodiment, performing a normalization operation on the second data to obtain the second target value includes: extracting a third value and a fourth value from the second data, wherein the third value represents the average transmission delay of the target disk within a preset time period, and the fourth value represents the initial transmission delay of the target disk; calculating the difference between the third value and the fourth value to obtain a second difference, wherein the second difference represents the increment of the transmission delay of the target disk within the preset time period; and calculating the ratio between the second difference and a second threshold to obtain the second target value, wherein the second target value satisfies a second value range.

[0049] Optionally, the third value in this embodiment is used to represent the average transmission latency of the target disk within a preset time period (e.g., the past 5 minutes, 1 hour, etc.). This value is calculated by summarizing the latency of each input / output (I / O) operation within this time period and is used to measure the current performance status of the target disk.

[0050] Optionally, the fourth value in this embodiment is used to represent the initial transmission latency of the target disk, that is, the average latency recorded when the storage system starts up or when the target disk is first used, which can be regarded as the ideal or baseline performance level of the disk.

[0051] Optionally, the second difference in this embodiment reflects the trend of the target disk's transmission delay changing over time. If the second difference is large, it indicates that the performance of the target disk may have declined significantly.

[0052] Optionally, the second threshold in this embodiment is a standard value used to assess whether the disk latency change is abnormal. If the second difference exceeds this threshold, the increase in transmission latency is considered abnormal and may serve as a warning about the health status of the disk.

[0053] Optionally, in this embodiment, the second target value is obtained by calculating the ratio of the second difference to the second threshold, and is used to represent the contribution of the increase in disk transfer latency to disk anomalies. The second target value satisfies a second numerical range (e.g., 0 to 1).

[0054] Optionally, in this embodiment, it includes, but is not limited to, using formulas The second target value is calculated, where L represents the incremental indicator of the transmission delay of the target disk (i.e., the second target value mentioned above), ΔLatency represents the impact on the business caused by the deterioration of disk performance, ΔLatency represents the increment of disk I / O delay within a preset time (e.g., 5 minutes) (i.e., the second difference mentioned above), ΔLatency = average transmission delay (i.e., the third value mentioned above) - initial transmission delay (i.e., the fourth value mentioned above), and Latency_max (i.e., the second threshold mentioned above) is the preset maximum incremental threshold for transmission delay.

[0055] For example, in high-performance trading systems in the financial field, even slight changes in disk I / O latency can lead to a significant decrease in trading speed. Real-time monitoring of disk read / write latency and prediction of anomalies in target disks are necessary. The monitoring system continuously collects I / O latency data for a critical disk, recording its average latency over the past hour (the third value mentioned above) as 10ms, while the disk's initial latency (the fourth value mentioned above) is 3ms. Subtracting the initial latency of 3ms from the current average latency of 10ms yields a second difference of 7ms, indicating a recent decline in disk performance. Based on historical data and disk performance analysis, a second threshold is set at 5ms, meaning any latency increase exceeding 5ms is considered abnormal. Dividing the second difference of 7ms by the second threshold of 5ms yields a second target value of approximately 1.4. However, since the second target value must meet a certain range (e.g., 0 to 1), it needs to be normalized. Assuming the normalized second target value is 1, this indicates that the disk latency change has reached a warning level, posing a high risk of anomalies. By continuously monitoring disk I / O latency and applying the methods described above, financial trading systems can promptly identify signs of disk performance degradation. This facilitates subsequent probability calculations to determine disk anomalies, allowing the system to take timely preventative measures, such as data migration and resource reallocation, to reduce potential business interruption risks and ensure smooth trading processes and data security.

[0056] Through the above steps, the difference between the average transmission delay extracted from the second data and the initial transmission delay is calculated to obtain a second difference value representing the increment of transmission delay. Then, the ratio of this second difference to a second threshold is calculated to obtain a second target value representing the degree of delay anomaly. By analyzing the second data, sudden increases in transmission delay can be identified. Subsequently, through step-by-step calculations, the impact of transmission delay changes is standardized, further enhancing the accuracy of anomaly detection on the target disk and facilitating timely warnings of potential disk failures.

[0057] In an exemplary embodiment, performing the normalization operation on the third data to obtain the third target value includes: parsing the third data to obtain a fifth value and a sixth value, wherein the fifth value represents the number of disks of the abnormal disks and the sixth value represents the number of disks of the same type as the target disk; calculating the ratio between the fifth value and the sixth value to obtain the third target value, wherein the third target value satisfies the third value range.

[0058] Optionally, the fifth value in this embodiment is the number of disks of the same type as the target disk that have been confirmed to have anomalies. Specifically, in actual use, in order to more accurately analyze the probability of the target disk having anomalies, the fifth value may be the number of disks of the same model that have had anomalies and are located in the same storage system or the same storage cluster as the target disk.

[0059] Optionally, the sixth value in this embodiment is used to represent the total number of all disks of the same type as the target disk. Specifically, in actual use, in order to more accurately analyze the probability of the target disk malfunctioning, the sixth value can be the total number of disks of the same model located in the same storage system or the same storage cluster as the target disk.

[0060] Optionally, in this embodiment, the third target value is the ratio between the fifth value and the sixth value, and satisfies a predetermined third value range (such as between 0 and 1), which is used to quantify the risk level of the target disk due to the failure of a group of disks of the same model.

[0061] Optionally, the third target value in this embodiment is obtained by correlation analysis based on the abnormal historical data of disks of the same model. Since disks of the same batch or model may experience abnormalities within a similar time period due to manufacturing defects or firmware problems.

[0062] Optionally, the third target value in this embodiment can also be output by analyzing historical abnormal data and establishing an anomaly correlation model of disks of the same model. When a certain number of disks of the same model have anomalies, the anomaly correlation model can predict that other disks of the same model may also fail in the near future and output the corresponding third target value.

[0063] Optionally, the anomaly correlation model for disks of the same model established in this embodiment aims to predict and assess the current disk's anomaly risk by analyzing historical anomaly data of disks of the same model. This model typically employs a multi-layered architecture, using machine learning techniques (such as logistic regression, neural networks, random forests, etc.) to train the model, enabling it to identify disk failure patterns and predict the probability of failures in other disks of the same type. This allows for multi-level, multi-dimensional analysis of disk anomalies. The anomaly correlation model includes, but is not limited to, a data collection layer, a feature extraction layer, and a dynamic evaluation layer. The data collection layer is responsible for collecting and organizing raw data related to disk anomalies, including anomaly records and related performance data of disks of the same model as the target disk. The feature extraction layer extracts and transforms features from the data output by the data collection layer, converting the raw data into feature vectors that the model can process. The dynamic evaluation layer assesses the anomaly risk of the target disk in real time, dynamically adjusting the evaluation strategy based on the model's prediction results and the current system status (such as load and resource availability), and outputs a third target value for the target disk.

[0064] Optionally, in this embodiment, it includes, but is not limited to, using formulas The third target value is calculated, where B (the aforementioned third target value) reflects the risk of batch anomalies, N_failed represents the number of disks of the same model and production batch that have already failed (the aforementioned fifth value), and N_batch represents the total number of disks in the batch (the aforementioned sixth value). For example, if there are 100 disks in the same batch, and 10 have failed, then B = 0.1.

[0065] Through the above steps, the third data is parsed to obtain the fifth value representing the number of disks with abnormality, and the sixth value representing the number of disks of the same type as the target disk. Then, the ratio between the fifth value and the sixth value is calculated to obtain the third target value. By parsing the third data, the potential fault distribution of disks of the same model is revealed, which helps to avoid the risk of batch failures in advance. Through the correlation analysis of the third data, the comprehensiveness and foresight of anomaly prediction are enhanced.

[0066] In an exemplary embodiment, determining the probability of an anomaly in the target disk based on the first target value, the second target value, and the third target value includes: calculating the product between the first target value and a first weight to obtain a second product; calculating the product between the second target value and a second weight to obtain a third product; calculating the product between the third target value and a third weight to obtain a fourth product; and calculating the sum of the second product, the third product, and the fourth product to obtain the probability; wherein the first weight, the second weight, and the third weight are all greater than or equal to 0, the first weight, the second weight, and the third weight are all less than or equal to 1, and the sum of the first weight, the second weight, and the third weight is 1.

[0067] Optionally, in this embodiment, the first weight, the second weight, and the third weight are used to represent the degree of influence of each dimension on the final result in the calculation of the overall anomaly prediction probability.

[0068] Optionally, in this embodiment, the probability of the target disk exhibiting anomalies is calculated using the formula RiskScore = α·S + β·L + γ·B, where RiskScore represents the probability of the target disk exhibiting anomalies, α, β, and γ correspond to the first weight, second weight, and third weight, respectively, α + β + γ = 1, α represents the influence weight of the physical layer (i.e., the first target value) on the overall probability of anomalies in the target disk, β represents the influence weight of the performance layer (i.e., the second target value) on the overall probability of anomalies in the target disk, γ represents the influence weight of the batch layer (i.e., the third target value) on the overall probability of anomalies in the target disk, and S, L, and B correspond to the first target value, second target value, and third target value, respectively. S is a value between 0 and 1. SMART represents the SMART parameter value of the target disk (corresponding to the first value mentioned above). θ (corresponding to the first threshold mentioned above) indicates that when the SMART parameter reaches this value, the target disk is in a medium-risk state. k (corresponding to the second value mentioned above) controls the steepness of the Sigmoid function (corresponding to the objective function mentioned above). ΔLatency is used to represent the increment of disk I / O latency within a preset time period (e.g., 5 minutes) (i.e., the second difference mentioned above). ΔLatency = Average transmission latency (i.e., the third value mentioned above) - Initial transmission latency (i.e., the fourth value mentioned above). Latency_max (i.e., the second threshold mentioned above) is a preset maximum increment threshold for transmission latency. N_failed represents the number of disks of the same model and production batch that have failed (i.e., the fifth value mentioned above), and N_batch represents the total number of disks in the batch (i.e., the sixth value mentioned above). S, L, and B are all in the range [0, 1].

[0069] For example, in large data centers, operations and maintenance personnel need to perform real-time health monitoring on thousands or even tens of thousands of disks. Suppose that at a certain moment, the first target value obtained after parsing the SMART data of a certain disk is 0.8, indicating that the disk's physical condition is poor and it is close to failure. Simultaneously, it is detected that the disk's I / O latency has increased by twice the average threshold in the past half hour, resulting in a second target value of 0.6, showing a significant decline in disk performance. System records show that the failure rate of disks of the same model is 10%, resulting in a third target value of 0.1, suggesting possible batch defects or other common problems. Based on past experience and analysis, the weights assigned to the physical layer, performance layer, and batch layer are α = 0.5, β = 0.4, and γ = 0.1 respectively (corresponding to the weights of the first, second, and third target values ​​mentioned above), to reflect their importance and priority in failure prediction. The second product = first target value × first weight = 0.8 × 0.5 = 0.4, the third product = second target value × second weight = 0.6 × 0.4 = 0.24, the fourth product = third target value × third weight = 0.1 × 0.1 = 0.01, and the final probability of the target disk failing = second product + third product + fourth product = 0.4 + 0.24 + 0.01 = 0.65. This means that, based on the current physical degradation, performance decline, and batch correlation, the probability of this disk failing in the near future is assessed at 65%, which is much higher than the normal failure probability of a typical disk. Therefore, operations personnel should take immediate action, such as initiating data replication processes, adjusting disk read / write strategies, or preparing to replace the disk, to reduce potential business risks and the probability of data loss. In this way, the data center can not only monitor the health status of individual disks in real time, but also comprehensively consider the collective performance of disks of the same model, ensuring the scientific and timely nature of resource scheduling and fault prediction strategies, thereby effectively improving the stability and reliability of the storage system and reducing unplanned service interruptions.

[0070] By multiplying the first, second, and third target values ​​by their respective weights and then calculating the weighted sum, the probability of the target disk exhibiting anomalies is obtained. Through weight adjustment, the influence of different data dimensions on the anomaly judgment of the target disk is balanced, providing a comprehensive mathematical model for assessing anomaly probabilities. This enhances the scientific and rational nature of decision-making and allows for a more accurate determination of the probability of the target disk exhibiting anomalies.

[0071] In an exemplary embodiment, after determining that the target disk is abnormal when the probability exceeds a preset threshold, the method further includes: determining resources allocated to the storage system based on the resource usage information of the storage system, wherein the storage system includes the target disk, and the resources are used to perform target services or to perform target operations on the target disk, wherein the target operations include migrating data from the target disk to a first disk, and the first disk is a disk in a normal state; generating control instructions based on the resource allocation information, wherein the control instructions are used to instruct the storage system to perform a target partitioning operation, and the target partitioning operation is used to partition a service domain and a reconstruction domain based on the allocation information.

[0072] Optionally, the resource usage information in this embodiment is data about the current resource consumption of the storage system, including but not limited to the utilization rate of the Central Processing Unit (CPU), network bandwidth usage, and disk I / O operation frequency.

[0073] Optionally, in this embodiment, the first disk is a disk in a normal state and is used as the target for data migration. When the target disk is predicted to malfunction, the system will migrate the data on the target disk to the first disk to maintain data redundancy and availability.

[0074] Optionally, the resources in this embodiment include computing resources and network resources, such as CPU time, network bandwidth, I / O operation capabilities, etc., which are used to perform business operations and exception handling operations.

[0075] Optionally, the target service in this embodiment is an application or service that runs normally on the storage system and depends on the stability and high performance of the storage system. When allocating resources, the needs of the target service should be given priority.

[0076] Optionally, the target operation in this embodiment refers to fault response operations performed on the target disk, including but not limited to data migration, disk replacement, and data reconstruction. These operations aim to reduce or eliminate the impact of disk failures on business operations.

[0077] Optionally, the control instructions in this embodiment are generated by system management software or an intelligent scheduler to guide resource allocation and task scheduling within the storage system.

[0078] Optionally, the target partitioning operation in this embodiment is used to dynamically partition storage system resources. Through the execution of control instructions, the system resources are divided into two parts: a business domain and a reconstruction domain. The business domain focuses on maintaining normal business operations, while the reconstruction domain handles data recovery and resource reallocation related to disk failures.

[0079] Optionally, in actual use, resources allocated to the storage system may be determined based on the storage system's resource usage information through the following methods: Remote Direct Memory Access (RDMA) technology combined with Quality of Service (QoS) policies may be used to isolate and allocate network bandwidth (e.g., allocating specific bandwidth limits for reconstructed traffic to ensure business traffic is not affected); cgroup (control group) technology may be used to isolate and manage system resources (such as CPU) quotas, allocating specific CPU cores or setting CPU usage limits for the target operation to ensure sufficient CPU resources for business processing tasks; to prevent disk I / O operations during the execution of the target operation from excessively affecting normal business I / O, the token bucket algorithm may be used to limit the IOPS (i.e., I / O operations per second) during the execution of the target operation, and the rate of I / O operations may be limited by controlling the token generation rate.

[0080] Through the above steps, after confirming that a disk anomaly will occur, resources are allocated based on storage system resource usage information, and corresponding control commands are generated to instruct the storage system to perform the partitioning of the business domain and the reconstruction domain. By determining the probability of the target disk anomaly, dynamic resource scheduling is achieved in a timely manner, ensuring that business operations are not significantly affected when handling disk anomalies. By pre-planning resources, the data reconstruction process is accelerated, improving the overall availability and efficiency of the storage system. This resolves the coupling contradiction between resource contention and business interruption in related technologies, providing a collaborative solution that combines disk lifecycle prediction with dynamic resource partitioning.

[0081] In an exemplary embodiment, after generating control instructions based on the resource allocation information, the method further includes: sending the control instructions to the storage system to instruct the storage system to allocate the resources to the business domain or the reconstruction domain, wherein the business domain is used to perform the target business using the resources, and the reconstruction domain is used to perform the target operation on the target disk using the resources, and the business domain and the reconstruction domain are logical divisions of the storage system based on the purpose of using the resources.

[0082] Optionally, in this embodiment, the service domain is a set of resources allocated in the storage system to prioritize the processing of service requests. The scheduling and use of these resources aim to maintain business continuity and quality of service (such as low latency and high throughput).

[0083] Optionally, in this embodiment, the reconstruction domain is a resource area in the storage system specifically allocated for performing data reconstruction and disk anomaly recovery operations. These resources are rapidly mobilized before or during an anomaly to quickly restore data redundancy and reduce business interruption time.

[0084] Optionally, the target business in this embodiment is a critical application or service running on the system that needs to guarantee high performance and high availability, such as a financial transaction processing system or a real-time data analysis service.

[0085] Optionally, the logical partitioning in this embodiment refers to the flexible partitioning and management of resources at the software level according to their different uses and functions, so as to achieve independent operation and optimized scheduling of business domains and reconfiguration domains.

[0086] For example, in a large financial data center, the distributed storage system Ceph is providing data storage and read / write services for high-volume stock trading applications. The system detects an anomaly in the SMART data indicator of a certain disk (i.e., the target disk mentioned above), indicating a potential risk. To prevent the disk anomaly from impacting the trading application, the system needs to quickly handle the disk anomaly without affecting normal business operations. Based on anomaly prediction analysis, control instructions are generated. These instructions detail how to allocate a portion of the current resource pool (e.g., 20% CPU, 30% network bandwidth, and 10% IOPS) for anomaly handling, while ensuring that the resource requirements of the business domain are not affected. The control instructions are sent to the storage system, which responds by dynamically allocating resources. The business domain continues to have sufficient resources (e.g., 80% CPU, 70% network bandwidth, and 90% IOPS) to handle real-time requests from the stock trading application. The reconstruction domain is allocated the necessary resources for data migration and reconstruction tasks. Resources in the business domain are strictly protected to ensure that the stock trading application can continue to operate efficiently, providing customers with low-latency, highly reliable trading services and avoiding any business interruptions that may be caused by disk failure handling. The reconstruction domain uses its allocated resources to quickly initiate the data migration process, migrating data from the target disk to the first disk in a normal state, while simultaneously performing erasure coding calculations to restore data redundancy. Through this logical division, even during disk anomaly handling, the performance indicators of the stock trading application (such as transaction latency and throughput) remain stable, reducing business interruption time and even avoiding business interruption. At the same time, disk anomalies are handled in a timely and efficient manner, and data redundancy is quickly restored, ensuring the overall reliability of the system. This achieves the technical effect of efficiently responding to and resolving storage resource failures without affecting critical business operations, thus ensuring both business continuity and system stability.

[0087] Optionally, Figure 3This is a flowchart illustrating the execution of a target operation on an abnormal disk according to an embodiment of this application. In a large data center employing a Ceph distributed storage system, a multi-dimensional fault prediction module, a dynamic resource partitioning module, and an Object Storage Daemon Node (OSD) are configured. These modules and OSD nodes cooperate with each other, through methods such as... Figure 3 The steps shown perform the target operation on the abnormal disk:

[0088] Step S302: The multi-dimensional fault prediction module determines the probability of the target disk malfunctioning and determines the prediction result based on the probability. Specifically, when the probability is greater than a preset threshold, the target disk is determined to be malfunctioning, and the corresponding prediction result is obtained. The probability is calculated using the formula RiskScore = α·S + β·L + γ·B. RiskScore represents the probability of the target disk malfunctioning. α, β, and γ correspond to the first, second, and third weights, respectively. α + β + γ = 1, where α represents the influence weight of the physical layer (i.e., the first target value) on the overall malfunction probability of the target disk, β represents the influence weight of the performance layer (i.e., the second target value) on the overall malfunction probability of the target disk, and γ represents the influence weight of the batch layer (i.e., the third target value) on the overall malfunction probability of the target disk. S, L, and B correspond to the first, second, and third target values, respectively. S is a value between 0 and 1. SMART represents the SMART parameter value of the target disk (corresponding to the first value mentioned above). θ (corresponding to the first threshold mentioned above) indicates that when the SMART parameter reaches this value, the target disk is in a medium-risk state. k (corresponding to the second value mentioned above) controls the steepness of the Sigmoid function (corresponding to the objective function mentioned above). ΔLatency is used to represent the increment of disk I / O latency within a preset time period (e.g., 5 minutes) (i.e., the second difference mentioned above). ΔLatency = Average transmission latency (i.e., the third value mentioned above) - Initial transmission latency (i.e., the fourth value mentioned above). Latency_max (i.e., the second threshold mentioned above) is a preset maximum increment threshold for transmission latency. N_failed is used to represent the number of disks of the same model and the same production batch that have failed (i.e., the fifth value mentioned above), N_batch is used to represent the total number of disks in the batch (i.e., the sixth value mentioned above), and S, L, and B are all in the range [0, 1].

[0089] Step S304: When the prediction result indicates that the target disk will experience an anomaly, the dynamic resource partitioning module determines the allocated resources and generates control instructions. The dynamic resource partitioning module determines the resources allocated to the storage system based on the resource usage information of the storage system, and generates control instructions based on the allocation information of the above resources. The control instructions are used to instruct the storage system to perform target partitioning operations, and the target partitioning operations are used to partition the business domain and the reconstruction domain based on the allocation information.

[0090] Step S306: The target node divides resources into the reconstruction domain and the business domain based on control instructions. The business domain is a set of resources in the storage system that are allocated to prioritize the processing of business requests. The scheduling and use of these resources are aimed at maintaining business continuity and service quality (such as low latency and high throughput). The reconstruction domain is a resource area in the storage system that is specifically allocated to perform data reconstruction and disk anomaly recovery operations. These resources are quickly mobilized before or when an anomaly occurs to quickly restore data redundancy and reduce business interruption time.

[0091] In step S308, the target node migrates the data from the target disk to the first disk in the reconstruction domain and performs encoding calculations (such as recalculating erasure codes) to generate additional verification data in addition to the original data, so that even if some disks fail, the system can still recover the complete data.

[0092] Through the above steps, control commands are sent to the storage system to execute the logical division of business domains and refactoring domains, further ensuring that resources are allocated reasonably, optimizing the performance of business processing and exception handling, realizing precise control of resources, avoiding interference of target operations with business, reducing the interruption time of target business, and improving the response speed and resource utilization of the storage system.

[0093] The above method will be illustrated with a specific example below. Figure 4 This is a flowchart illustrating a method for determining disk anomalies in a distributed storage system according to an embodiment of this application. In a large data center employing a Ceph distributed storage system, disk anomalies can lead to passive reconstruction consuming significant resources and impacting service quality. The system administrator employs methods such as... Figure 4 The steps shown determine the probability of anomalies in the target disk and perform dynamic resource partitioning:

[0094] Step S402: Obtain first data, second data, and third data of the target disk. The first data represents the performance of the target disk, the second data represents the transmission latency of the target disk, and the third data represents the abnormality of the abnormal disk. The abnormal disk is of the same type as the target disk. The first data includes SMART data (e.g., number of relocation sectors, seek error rate, disk temperature). The second data includes the average transmission latency data of the target disk within a preset time and the initial latency baseline value of the target disk. The third data includes the number of disks of the same model as the target disk that have experienced abnormalities and the total number of disks of this model in this storage system.

[0095] Step S404: Perform a first processing operation on the first data, the second data, and the third data respectively to obtain a first target value, a second target value, and a third target value. The first processing operation specifically includes the following operations: S, L, and B correspond to the aforementioned first target value, second target value, and third target value, respectively. SMART represents the SMART parameter value of the target disk (corresponding to the first value mentioned above), θ (corresponding to the first threshold mentioned above) indicates that when the SMART parameter reaches this value, the target disk is in a medium-risk state, and k (corresponding to the second value mentioned above) controls the steepness of the Sigmoid function (corresponding to the objective function mentioned above).

[0096] ΔLatency is used to represent the increment of disk I / O latency within a preset time period (e.g., 5 minutes) (i.e., the second difference mentioned above). ΔLatency = Average transmission latency (i.e., the third value mentioned above) - Initial transmission latency (i.e., the fourth value mentioned above). Latency_max (i.e., the second threshold mentioned above) is a preset maximum increment threshold for transmission latency. N_failed is used to represent the number of disks of the same model and the same production batch that have failed (i.e., the fifth value mentioned above), N_batch is used to represent the total number of disks in the batch (i.e., the sixth value mentioned above), and S, L, and B are all in the range [0, 1].

[0097] Step S406: Using the first target value, the second target value, and the third target value, combined with the corresponding preset weight coefficients (first weight, second weight, and third weight), the probability of the target disk being abnormal is calculated by weighted summation, i.e., by the formula RiskScore=α·S+β·L+γ·B. RiskScore is used to represent the probability of the target disk being abnormal. α, β, and γ correspond to the first weight, second weight, and third weight, respectively. α+β+γ=1, where α represents the influence weight of the physical layer (i.e., the first target value) on the overall probability of the target disk being abnormal, β represents the influence weight of the performance layer (i.e., the second target value) on the overall probability of the target disk being abnormal, and γ represents the influence weight of the batch layer (i.e., the third target value) on the overall probability of the target disk being abnormal.

[0098] Step S408: When the calculated probability is greater than a preset threshold, it is determined that the target disk is abnormal;

[0099] Step S410: After determining that the target disk will experience an anomaly, analyze the resource usage information of the storage system (e.g., CPU utilization, network bandwidth, IOPS, etc.), and allocate resources and generate control instructions based on the resource usage information and business requirements. The control instructions are used to instruct the storage system to perform dynamic resource partitioning (i.e., the target partitioning operation mentioned above), and to partition the business domain and the reconstruction domain.

[0100] Step S412: The generated control command is sent to the storage system, instructing the storage system to allocate resources to the business domain or the reconstruction domain. The business domain is used to use resources to perform target business, and the reconstruction domain is used to use resources to perform target operations on the target disk. The business domain and the reconstruction domain are logical divisions of the storage system based on the purpose of resource use. The target operation includes the operation of migrating data from the target disk to the first disk, which is a disk in a normal state.

[0101] By following the steps above, data security is effectively protected by analyzing multi-dimensional data before obvious disk anomalies occur, thus providing early warnings and taking action. This solves the problem of the inability to predict disk lifecycles in a timely manner, which leads to prolonged business interruptions. It achieves the technical effect of accurately judging disk failures and reducing business interruption time. By combining disk lifecycle prediction with dynamic resource partitioning, the reliability and stability of the entire storage system are improved.

[0102] It should be noted that, through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0103] This embodiment also provides a disk anomaly determination device, which is used to implement the above embodiments and preferred embodiments, and will not be repeated as already described. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0104] Figure 5 This is a structural block diagram of a disk anomaly determination device according to an embodiment of this application, such as... Figure 5 As shown, the device includes:

[0105] The acquisition module 52 is used to acquire first data, second data and third data of the target disk, wherein the first data is used to represent the performance of the target disk, the second data is used to represent the transmission delay of the target disk, and the third data is used to represent the abnormality of the abnormal disk, wherein the abnormal disk is of the same type as the target disk.

[0106] Processing module 54 is used to perform a first processing operation on the first data, the second data, and the third data respectively to obtain a first target value, a second target value, and a third target value. The first processing operation includes the following steps: performing a mapping operation on the first data to obtain the first target value; performing a normalization operation on the second data to obtain the second target value; and performing the normalization operation on the third data to obtain the third target value.

[0107] The first determining module 56 is used to determine the probability that the target disk is abnormal based on the first target value, the second target value and the third target value.

[0108] The second determining module 58 is used to determine that the target disk is abnormal when the above probability is greater than a preset threshold.

[0109] In an exemplary embodiment, the processing module 54 is further configured to parse the first data, determine the degree of degradation of the target disk, and obtain a first value, wherein the degree of degradation is used to represent the performance degradation level of the target disk; calculate the difference between the first value and a first threshold to obtain a first difference, wherein the first difference is used to represent the abnormality level of the target disk; calculate the product between the first difference and a second value to obtain a first product, wherein the second value is used to represent the degree of influence of the first value on the abnormality of the target disk; and map the first product to a value that satisfies the first value range using an objective function to obtain the first target value.

[0110] In an exemplary embodiment, the processing module 54 is further configured to extract a third value and a fourth value from the second data, wherein the third value represents the average transmission delay of the target disk within a preset time period, and the fourth value represents the initial transmission delay of the target disk; calculate the difference between the third value and the fourth value to obtain a second difference, wherein the second difference represents the increment of the transmission delay of the target disk within the preset time period; and calculate the ratio between the second difference and a second threshold to obtain the second target value, wherein the second target value satisfies a second value range.

[0111] In an exemplary embodiment, the processing module 54 is further configured to parse the third data to obtain a fifth value and a sixth value, wherein the fifth value is used to represent the number of disks of the abnormal disk, and the sixth value is used to represent the number of disks of the same type as the target disk; calculate the ratio between the fifth value and the sixth value to obtain the third target value, wherein the third target value satisfies the third value range.

[0112] In an exemplary embodiment, the first determining module 56 is further configured to calculate the product between the first target value and the first weight to obtain a second product; calculate the product between the second target value and the second weight to obtain a third product; calculate the product between the third target value and the third weight to obtain a fourth product; and calculate the sum of the second product, the third product, and the fourth product to obtain the probability; wherein the first weight, the second weight, and the third weight are all greater than or equal to 0, the first weight, the second weight, and the third weight are all less than or equal to 1, and the sum of the first weight, the second weight, and the third weight is 1.

[0113] In an exemplary embodiment, the second determining module 58 is further configured to, after determining that the target disk is abnormal when the probability is greater than a preset threshold, determine the resources allocated to the storage system based on the resource usage information of the storage system, wherein the storage system includes the target disk, and the resources are used to perform target services or to perform target operations on the target disk, wherein the target operations include migrating data from the target disk to a first disk, and the first disk is a disk in a normal state; and generate control instructions based on the resource allocation information, wherein the control instructions are used to instruct the storage system to perform a target partitioning operation, and the target partitioning operation is used to partition a service domain and a reconstruction domain based on the allocation information.

[0114] In an exemplary embodiment, the second determining module 58 is further configured to generate a control instruction based on the resource allocation information and then send the control instruction to the storage system to instruct the storage system to allocate the resource to the business domain or the reconstruction domain. The business domain is used to perform the target business using the resource, and the reconstruction domain is used to perform the target operation on the target disk using the resource. The business domain and the reconstruction domain are logical divisions of the storage system based on the purpose of using the resource.

[0115] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.

[0116] For a description of the features in the embodiment of a disk anomaly determination device, please refer to the relevant description of the embodiment of a disk anomaly determination method, which will not be repeated here.

[0117] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0118] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when it is run.

[0119] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0120] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.

[0121] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.

[0122] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0123] The foregoing has provided a detailed description of the method, apparatus, program product, and storage medium for determining disk anomalies provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for determining disk anomalies, characterized in that, include: Acquire first data, second data, and third data of the target disk, wherein the first data is used to represent the performance of the target disk, the second data is used to represent the transmission latency of the target disk, and the third data is used to represent the abnormality of the abnormal disk, wherein the abnormal disk is of the same type as the target disk; A first processing operation is performed on the first data, the second data, and the third data respectively to obtain a first target value, a second target value, and a third target value. The first processing operation includes the following steps: performing a mapping operation on the first data to obtain the first target value; performing a normalization operation on the second data to obtain the second target value; and performing the normalization operation on the third data to obtain the third target value. The probability of the target disk exhibiting an anomaly is determined based on the first target value, the second target value, and the third target value; When the probability is greater than a preset threshold, it is determined that the target disk is abnormal.

2. The method according to claim 1, characterized in that, Performing a mapping operation on the first data to obtain the first target value includes: The first data is analyzed to determine the degree of degradation of the target disk and obtain a first value, wherein the degree of degradation is used to represent the performance degradation level of the target disk; Calculate the difference between the first value and the first threshold to obtain the first difference, wherein the first difference is used to represent the anomaly level of the target disk; Calculate the product between the first difference and the second value to obtain the first product, wherein the second value is used to represent the degree of influence of the first value on the anomaly of the target disk; The first product is mapped to a value that satisfies a first numerical range using an objective function, thus obtaining the first objective value.

3. The method according to claim 1, characterized in that, Performing a normalization operation on the second data to obtain the second target value includes: Extract the third and fourth values ​​from the second data, wherein the third value is used to represent the average transmission delay of the target disk within a preset time, and the fourth value is used to represent the initial transmission delay of the target disk; The difference between the third value and the fourth value is calculated to obtain a second difference, wherein the second difference is used to represent the increment of the transmission delay of the target disk within the preset time period; The ratio between the second difference and the second threshold is calculated to obtain the second target value, wherein the second target value satisfies the second value range.

4. The method according to claim 1, characterized in that, Performing the normalization operation on the third data to obtain the third target value includes: The third data is parsed to obtain a fifth value and a sixth value, wherein the fifth value is used to represent the number of disks of the abnormal disk, and the sixth value is used to represent the number of disks of the same type as the target disk; The ratio between the fifth value and the sixth value is calculated to obtain the third target value, wherein the third target value satisfies the third value range.

5. The method according to claim 1, characterized in that, Determining the probability of an anomaly in the target disk based on the first target value, the second target value, and the third target value includes: Calculate the product between the first target value and the first weight to obtain the second product; Calculate the product between the second target value and the second weight to obtain the third product; Calculate the product between the third target value and the third weight to obtain the fourth product; The probability is obtained by calculating the sum of the second product, the third product, and the fourth product. Wherein, the first weight, the second weight, and the third weight are all greater than or equal to 0, the first weight, the second weight, and the third weight are all less than or equal to 1, and the sum of the first weight, the second weight, and the third weight is 1.

6. The method according to claim 1, characterized in that, After determining that the target disk is abnormal when the probability is greater than a preset threshold, the method further includes: Based on the resource usage information of the storage system, the resources allocated to the storage system are determined, wherein the storage system includes the target disk, and the resources are used to execute target services or to perform target operations on the target disk. The target operations include the operation of migrating data in the target disk to a first disk, which is a disk in a normal state. Control instructions are generated based on the resource allocation information, wherein the control instructions are used to instruct the storage system to perform a target partitioning operation, and the target partitioning operation is used to partition a business domain and a refactoring domain based on the allocation information.

7. The method according to claim 6, characterized in that, After generating control instructions based on the resource allocation information, the method further includes: The control command is sent to the storage system to instruct the storage system to allocate the resource to the business domain or the reconstruction domain. The business domain is used to execute the target business using the resource, and the reconstruction domain is used to execute the target operation on the target disk using the resource. The business domain and the reconstruction domain are logical divisions of the storage system based on the purpose of using the resource.

8. A device for determining disk anomalies, characterized in that, include: The acquisition module is used to acquire first data, second data and third data of the target disk, wherein the first data is used to represent the performance of the target disk, the second data is used to represent the transmission latency of the target disk, and the third data is used to represent the abnormality of the abnormal disk, wherein the abnormal disk is of the same type as the target disk; The processing module is configured to perform a first processing operation on the first data, the second data, and the third data respectively to obtain a first target value, a second target value, and a third target value. The first processing operation includes the following steps: performing a mapping operation on the first data to obtain the first target value; performing a normalization operation on the second data to obtain the second target value; and performing the normalization operation on the third data to obtain the third target value. The first determining module is used to determine the probability that the target disk is abnormal based on the first target value, the second target value and the third target value; The second determining module is used to determine that the target disk is abnormal when the probability is greater than a preset threshold.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method described in any one of claims 1 to 7.