Control methods and devices for storage systems, electronic devices, and storage media

By identifying and replacing disks to be marked in high-risk disk groups using a pre-set risk prediction model, the problem of storage systems being unable to promptly handle disks at risk of failure is solved, ensuring the reliability of the storage system and data security.

CN119668527BActive Publication Date: 2025-10-31INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510190713.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-10-31
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

Existing storage systems are unable to handle disks at risk of failure in a timely manner, which can easily lead to the number of failed disks in the storage system exceeding the system's fault tolerance capacity, resulting in a high risk of storage system failure and data loss.

Method used

By calculating the disk failure risk index using a pre-set risk prediction model, high-risk disks and disk groups are identified. If no high-risk disk exists in a high-risk disk group, a disk to be marked is selected for replacement, thereby reducing the failure risk of the disk group and ensuring the reliability of the storage system.

Benefits of technology

It enables rapid identification of disks and disk groups with high failure risk, timely data backup and replacement, reducing the failure risk of the storage system and ensuring system reliability and data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119668527B_ABST
    Figure CN119668527B_ABST
Patent Text Reader

Abstract

This application discloses a control method and device, electronic device, and storage medium for a storage system, relating to the field of computer technology. By using a preset risk prediction model, the failure risk index of each disk in the storage system is calculated to identify high-risk disks. Then, based on the failure risk index of each disk, the failure risk index of disk groups is calculated to identify high-risk disk groups. From the high-risk disk groups where no high-risk disks exist, disks to be marked are selected. Subsequently, all high-risk disks and disks to be marked in the storage system are replaced. This allows for timely handling of disks and disk groups with failure risks in the storage system, achieving rapid identification of high-failure-risk disks and high-failure-risk disk groups, and timely data backup and replacement. This ensures that the number of failed disks in the storage system does not exceed the system's fault tolerance capacity, effectively reducing the failure risk of the storage system and ensuring system reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a control method and apparatus for a storage system, an electronic device, and a storage medium. Background Technology

[0002] With the development of computer and network technologies, network data is growing at an explosive rate. The massive amount of data poses a huge challenge to storage systems. Therefore, distributed storage systems built using servers have been widely used in data storage management. The increase in the scale of storage systems and the increase in the capacity of individual disks in storage systems have posed challenges to the reliability of storage systems. In order to improve the reliability of storage systems, fault-tolerant technologies are needed to improve the fault tolerance of storage systems.

[0003] In storage systems, there are two main types of common fault tolerance technologies: multi-replica fault tolerance and erasure coding fault tolerance. However, the fault tolerance capabilities of multi-replica fault tolerance and erasure coding fault tolerance are limited, and they cannot handle disks with failure risks in the storage system in a timely manner. This can easily lead to the number of failed disks in the storage system exceeding the system's fault tolerance capacity, causing the storage system to fail and resulting in low reliability. Summary of the Invention

[0004] This application provides a control method and apparatus, electronic device, and storage medium for a storage system, to at least solve the problem in related technologies that the inability to handle disks with failure risks in a storage system in a timely manner can easily lead to the number of failed disks in the storage system exceeding the system's fault tolerance capacity, causing the storage system to fail and resulting in a high risk of data loss.

[0005] This application provides a control method for a storage system, including:

[0006] Failure risk prediction is performed using a preset risk prediction model and disk information from the disks in the storage system to obtain the first failure risk index for each of the multiple disks. The storage system includes multiple disk groups, and each disk group includes multiple disks.

[0007] Based on the first failure risk index corresponding to each disk in each disk group and the average risk index corresponding to each disk group, the index calculation process is performed to obtain the second failure risk index corresponding to each disk group. The average risk index is calculated by averaging the first failure risk indices corresponding to each disk in the disk group.

[0008] High-risk disks and high-risk disk groups are identified, wherein a high-risk disk is a disk whose first failure risk index is greater than or equal to a first preset index threshold, and a high-risk disk group is a disk group whose second failure risk index is greater than or equal to a second preset index threshold.

[0009] If there are no high-risk disks in the high-risk disk group, the disks in the high-risk disk group are selected to obtain the disks to be marked, and the disks to be marked are marked as high-risk disks.

[0010] All high-risk disks in the storage system are replaced to obtain the replaced storage system.

[0011] This application also provides a control device for a storage system, including:

[0012] The prediction unit is used to perform failure risk prediction processing through a preset risk prediction model and disk information of the disks in the storage system to obtain the first failure risk index corresponding to each of the multiple disks. The storage system includes multiple disk groups, and each disk group includes multiple disks.

[0013] The calculation unit is used to perform index calculation processing based on the first failure risk index corresponding to each of the multiple disks included in each disk group and the average risk index corresponding to each disk group, to obtain the second failure risk index corresponding to each disk group. The average risk index is calculated by averaging the first failure risk indices corresponding to the multiple disks included in the disk group.

[0014] The identification unit is used to identify high-risk disks and high-risk disk groups, wherein a high-risk disk is a disk whose first failure risk index is greater than or equal to a first preset index threshold, and a high-risk disk group is a disk group whose second failure risk index is greater than or equal to a second preset index threshold.

[0015] The selection unit is used to select disks in the high-risk disk group when there are no high-risk disks in the high-risk disk group, to obtain disks to be marked, and to mark the disks to be marked as high-risk disks.

[0016] The processing unit is used to perform disk replacement processing on all high-risk disks in the storage system to obtain the replaced storage system.

[0017] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the control method of any of the above-described memory systems when executing the computer program.

[0018] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the control method of any of the above-described storage systems.

[0019] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described storage system control methods.

[0020] This application first uses a disk failure prediction model, i.e., a preset risk prediction model, to calculate the failure risk index of each disk in the storage system and identify disks with high failure risk. Since a disk group includes multiple disks, the failure risk index of the disk group can be calculated based on the failure risk index of each disk, thus identifying disk groups with high failure risk. However, because the failure risk index of a disk group is calculated from the failure risk indices of the disks included in the group, there may be situations where a high-risk disk group does not contain any high-risk disks. To further reduce the failure risk of the storage system, in addition to replacing high-risk disks, it is also necessary to reduce the failure risk of all high-risk disk groups. Therefore, it is necessary to select from the high-risk disk groups where no high-risk disks exist. A disk to be tagged is selected so that the failure risk of high-risk disk groups can be reduced by replacing the tagged disk. Then, all high-risk disks and disks to be tagged in the storage system are replaced. This can promptly handle disks and disk groups with failure risks in the storage system. This solves the technical problem in the existing technology that cannot handle disks with failure risks in the storage system in a timely manner, which can easily lead to the number of failed disks in the storage system exceeding the system's fault tolerance capacity, causing the storage system to fail and resulting in a high risk of data loss. The technology achieves the technical effect of quickly identifying disks and disk groups with high failure risks, and timely backing up and replacing data, ensuring that the number of failed disks in the storage system does not exceed the system's fault tolerance capacity, effectively reducing the failure risk of the storage system, and ensuring the reliability of the system. Attached Figure Description

[0021] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 A flowchart illustrating a control method for a storage system provided in an embodiment of this application;

[0023] Figure 2This is a schematic diagram of a disk and disk group provided in an embodiment of this application;

[0024] Figure 3 A flowchart illustrating a failure risk prediction process provided in this application embodiment;

[0025] Figure 4 A schematic diagram illustrating the calculation process of a second failure risk index provided in this application embodiment;

[0026] Figure 5 This is a schematic diagram of a disk selection process provided in an embodiment of this application;

[0027] Figure 6 A schematic diagram illustrating a disk failure risk level provided in an embodiment of this application;

[0028] Figure 7 This application provides a schematic diagram of the architecture of a distributed storage system.

[0029] Figure 8 This is a schematic diagram of the structure of a control device for a storage system provided in an embodiment of this application;

[0030] Figure 9 This is a schematic diagram of the structure of a control device for another storage system provided in an embodiment of this application. Detailed Implementation

[0031] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0032] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0033] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0034] Figure 1This document provides a flowchart illustrating a control method for a storage system, as provided in an embodiment of this application. The method will be described in detail below, taking into account the execution flow of the control method for the storage system.

[0035] like Figure 1 As shown, the control method of this storage system includes:

[0036] Step 101: Perform failure risk prediction processing by using a preset risk prediction model and disk information of the disks in the storage system to obtain the first failure risk index corresponding to each of the multiple disks. The storage system includes multiple disk groups, and each disk group includes multiple disks.

[0037] In the embodiments of this application, the preset risk prediction model is a pre-trained model that is selected by the user and can be used to predict the risk of disk failure. Regarding the type of preset risk prediction model, examples include: models designed based on random forest machine learning algorithms, models designed based on neural networks, models designed based on SVM, models designed based on k-NN, models designed based on decision trees, or models designed based on deep learning, etc. Specifically, this application does not limit the type of preset risk prediction model.

[0038] It should be noted that this preset risk prediction model can be deployed on a separate server or on the server node where the disk is located.

[0039] The First Failure Risk Index is a quantitative indicator output by a pre-defined risk prediction model, representing the probability of each disk failing within a certain timeframe. Types of the First Failure Risk Index include, but are not limited to, numerical values ​​between 0 and 1, percentages, etc. Generally, the higher the value of the First Failure Risk Index, the greater the risk of disk failure.

[0040] Step 102: Perform index calculation processing based on the first failure risk index corresponding to each disk in each disk group and the average risk index corresponding to each disk group to obtain the second failure risk index corresponding to each disk group. The average risk index is calculated by averaging the first failure risk indices corresponding to each disk in the disk group.

[0041] In the embodiments of this application, disk groups refer to the following: In a storage system, data is usually distributed across multiple disks, and the disks can be organized into multiple disk groups according to certain rules or strategies. Each disk group includes multiple disks, and the disks included in the disk group are jointly responsible for storing and accessing a portion of the data.

[0042] The second failure risk index is calculated at the disk group level, taking into account the failure risk of all disks within a disk group, thus providing a more comprehensive assessment that reflects the overall data loss or unavailability risk of the disk group.

[0043] By determining the second failure risk index of disk groups, a deeper understanding of storage system risks can be provided at the disk group level. Disk groups that may face higher risks due to the simultaneous failure of multiple disks can be identified, thereby enabling more effective risk mitigation measures to ensure data integrity and system reliability.

[0044] Step 103: Identify high-risk disks and high-risk disk groups, wherein a high-risk disk is a disk whose first failure risk index is greater than or equal to a first preset index threshold, and a high-risk disk group is a disk group whose second failure risk index is greater than or equal to a second preset index threshold.

[0045] In the embodiments of this application, the first preset index threshold is a pre-set numerical standard used to determine whether the failure risk of a single disk reaches a high risk level. When the first failure risk index of the disk is greater than or equal to the first preset index threshold, the disk is considered to have a high failure risk, i.e., a high-risk disk.

[0046] The second preset index threshold is also a pre-set numerical standard, which is a pre-set numerical standard used to judge the failure risk of a disk group. When the second failure risk index of a disk group is greater than or equal to the second preset index threshold, the disk group is considered to have a high failure risk as a whole.

[0047] The methods for marking high-risk disks include, but are not limited to: setting warning signs or color coding on the disk list in the storage management system, or recording the disk information (including disk identifier (ID), disk type, etc.) corresponding to high-risk disks in a preset storage location as a high-failure-risk disk group for further subsequent operations.

[0048] By identifying high-risk disks and high-risk disk groups, the management efficiency and data protection level of the storage system are improved, which helps to ensure the stable operation of the storage system and the security of data.

[0049] Step 104: If there are no high-risk disks in the high-risk disk group, select the disks in the high-risk disk group to obtain the disks to be marked, and mark the disks to be marked as high-risk disks.

[0050] In the embodiments of this application, when a disk group is a high-risk disk group, there may not necessarily be any high-risk disks in it. For example, if the first failure risk index of each disk in the disk group is relatively large, but all are less than the first preset threshold, then the second failure risk index calculated based on the first failure risk index may also be relatively large, and there is a possibility that it is greater than the second preset threshold. In this case, even if there are no high-risk disks in the disk group, this disk group is still a high-risk disk group.

[0051] It should be noted that although there are no high-risk disks in the high-risk disk group, because the disk group has a high risk of failure, it is necessary to select one disk in the high-risk disk group and mark it as a high-risk disk so that it can be replaced later to reduce the risk of disk group failure.

[0052] Specifically, to facilitate understanding of the embodiments of this disclosure, a schematic diagram of a disk and a disk group is provided, such as... Figure 2 As shown, although disk group 3 does not contain any high-risk disks, disk group 3 may be a high-risk disk group. Furthermore, disks are categorized as high-risk, medium-risk, and healthy. The disk classification includes, but is not limited to, the following methods: disks with a first failure risk index greater than or equal to a first preset index threshold are marked as high-risk disks; disks with a first failure risk index less than the first preset index threshold but greater than or equal to a third preset index threshold are marked as medium-risk disks; disks with a first failure risk index less than the third preset index threshold are marked as healthy disks. Alternatively,

[0053] Disks with a first failure risk index less than a first preset index threshold are marked as non-high-risk disks; disks with a first failure risk index greater than or equal to the first preset index threshold are marked as high-risk disks.

[0054] Specifically, regarding disk partitioning, it is only necessary to identify high-risk disks. The same applies to disk group partitioning. This application does not impose restrictions on other partitioning aspects (e.g., whether to partition into healthy or medium-risk areas).

[0055] Marking high-risk disks within high-risk disk groups helps improve the risk management capabilities of storage systems, ensuring that appropriate preventative measures can be taken to protect data security when the overall risk of a disk group is high.

[0056] Step 105: Perform disk replacement on all high-risk disks in the storage system to obtain the replaced storage system.

[0057] In the embodiments of this application, all high-risk disks are traversed and selected as disks to be replaced. Idle disks are selected from the storage system according to certain rules to replace the high-risk disks. The storage system uses a data copying method to copy the data from the high-risk disks to the selected idle disks. After the disk data copying is completed, the management service (Manager Daemon, MGRS) performs disk switching to obtain the storage system after replacement.

[0058] This application first uses a disk failure prediction model, i.e., a preset risk prediction model, to calculate the failure risk index of each disk in the storage system and identify disks with high failure risk, i.e., high-risk disks. Then, based on the failure risk index of each disk, it calculates the failure risk index of the disk group and identifies the disk group with high failure risk, i.e., the high-risk disk group. If there is no high-risk disk in the high-risk disk group, it selects one disk as the high-risk disk. Then, it adopts an active protection strategy to back up and replace the data on the high-risk disk. This achieves the technical effect of quickly identifying disks and disk groups with high failure risk, and timely backing up and replacing data, effectively reducing the failure risk of the storage system and ensuring the reliability of the system.

[0059] In one feasible embodiment of this application, regarding the failure risk prediction process for disks in a storage system, this application provides a schematic flowchart for failure risk prediction, as follows: Figure 3 As shown, it includes:

[0060] Step 301: Input the disk information of the disk in the storage system into the preset risk prediction model for failure risk prediction processing to obtain the initial failure risk index. The preset risk prediction model is a pre-configured model for predicting failure risk.

[0061] In the embodiments of this application, the disk information includes, but is not limited to: disk type, disk identifier (ID), and disk status data. After the disk information is input into the preset risk prediction model, the output is the disk failure risk index, i.e., the initial failure risk index.

[0062] The initial failure risk index is relatively messy data. For example, the initial failure risk index of one disk may be a value between 0 and 1, while the initial failure risk index of another disk may be a percentage. This is because the preset risk prediction model may include multiple sub-models, i.e., the first prediction model. The data output by different first prediction models may not have consistent units. Therefore, it is necessary to normalize the initial failure risk index to obtain the first failure risk index.

[0063] Step 302: Perform data normalization on the initial failure risk index to obtain the first failure risk index.

[0064] In the embodiments of this application, the initial failure risk index of different disks is normalized according to disk type, so that the failure risk index of different disks in the storage system can be compared. Methods for normalizing the initial failure risk data include, but are not limited to, Min-Max normalization, Z-Score normalization, or other nonlinear normalization methods. Specifically, this application does not limit the methods used to normalize the initial failure risk data.

[0065] Normalizing the initial failure risk index improves comparability: the initial failure risk index of all disks is on the same scale and can be directly compared; consistency: the normalized index is easier to compare and analyze across different storage environments and models; transparency: the normalization process increases the transparency of risk assessment, making the risk assessment results easier to understand.

[0066] It can improve the accuracy and practicality of disk failure risk assessment in storage systems, providing a more reliable basis for decision-making in the maintenance and management of storage systems.

[0067] In one feasible embodiment of this application, as a refinement of step 301 above, the initial failure risk index can be obtained in the following manner, but is not limited to: determining the first prediction model corresponding to each of multiple disks based on the disk type in the disk information, wherein the preset risk prediction model includes at least one first prediction model, and the first prediction model corresponds to one or more disks; inputting the disk information into the corresponding first prediction model for failure risk prediction processing to obtain the initial failure risk index, wherein the disk information includes at least disk type, disk status data, data acquisition time, and disk identifier.

[0068] In the embodiments of this application, the first prediction model is one or more sub-models in the preset risk prediction model, and each first prediction model is specifically designed or selected for a specific type of disk or disk group.

[0069] The disk type is used to determine the applicable first prediction model; disk status data includes at least self-monitoring, analysis, and reporting technology (SMART) data, error count data, and workload data, which are used to assess the health status of the disk; the data acquisition time records the time point at which the status data is collected, which is used to analyze the trend of disk status changes; the disk identifier is a unique identifier to track and identify the corresponding disk during the prediction process.

[0070] Determining the first prediction model means selecting a suitable first prediction model for each disk or group of similar disks based on the disk type; inputting disk information means inputting information such as disk type, status data, data acquisition time, and identification into the corresponding first prediction model; failure risk prediction processing means that the first prediction model calculates and analyzes based on the input disk information and outputs the initial failure risk index for each disk.

[0071] By providing customized failure risk predictions for different types of disks, the accuracy and efficiency of predictions are improved, which helps to more accurately identify and manage risks in storage systems, ensuring data security and stable system operation.

[0072] In one feasible embodiment of this application, before predicting the failure risk of a disk, it is necessary to obtain a preset risk prediction model in advance. When obtaining the preset risk prediction model, it can be achieved in the following ways, but not limited to: performing disk modeling processing based on the disk capacity, disk identifier, and disk type to obtain multiple disk models, wherein each disk model corresponds to at least one disk; and performing prediction model matching processing based on the disk models to obtain the preset risk prediction model.

[0073] In the embodiments of this application, each disk model corresponds to at least one disk. The disk model includes the key characteristics of the disk, such as capacity, identification and type, as well as other possible parameters, such as performance indicators, historical failure data, etc.

[0074] The matching process refers to selecting or configuring one or more suitable primary prediction models based on the characteristics of the disk model. These primary prediction models are pre-designed and capable of predicting failure risks for specific types of disk characteristics. Determining the preset risk prediction model involves identifying the most suitable primary prediction model for each disk model through the matching process. All primary prediction models together constitute the preset risk prediction model.

[0075] By providing a customized predictive model for each disk, the accuracy and applicability of failure risk prediction are improved, which helps to more effectively monitor and manage the health status of each disk in the storage system, identify potential failure risks in advance, and take corresponding preventive measures.

[0076] In one feasible embodiment of this application, before predicting the failure risk of a disk, it is necessary to obtain disk information in advance. When obtaining disk information, the following methods can be used, but are not limited to: collecting disk status data of multiple disks in the storage system according to a preset period; performing data set processing on the disk status data, disk capacity, disk identifier, disk type, and data collection time to obtain disk information, wherein the data collection time is the time of disk status data collection.

[0077] In the embodiments of this application, disk status data is collected at a preset period. The disk status data includes at least SMART data, error count data, and workload data. The preset period can be, for example, hour, day, week, etc. Specifically, this application does not limit the setting of the preset period.

[0078] Specifically, disk status data includes at least the following: number of media errors, number of pending and uncorrectable sector errors, number of reallocated sector errors, number of program / erase errors, cumulative disk power-on time, total number of NAND writes, total number of NAND reads and writes, and failure status.

[0079] Ensuring timely and comprehensive updates to disk information in the storage system provides accurate data input for risk prediction models, which helps improve the accuracy and real-time nature of disk failure risk prediction, enabling storage system administrators to take appropriate maintenance measures based on the latest disk information.

[0080] In one feasible embodiment of this application, as a refinement of step 102 above, regarding the calculation of the second failure risk index, this application provides a schematic diagram of the calculation process for the second failure risk index, as follows: Figure 4 As shown, it includes:

[0081] Step 401: In each disk group, the average risk index corresponding to each disk group is calculated by averaging the first failure risk index corresponding to each of the multiple disks included in the disk group.

[0082] In the embodiments of this application, it is necessary to calculate the failure risk index of all disk groups, i.e., the second failure risk index. Specifically, the calculation of the second failure risk index can be carried out in the following ways, but is not limited to: Suppose that a disk group occupies a total of N disks, and the first failure risk index of the j-th disk is... At this point, the calculation of the average risk index can be performed using, but is not limited to, formula (1):

[0083] Formula (1)

[0084] in, Here, N represents the average risk index, N is the number of disks in the disk group, and j is the j-th disk in the disk group. This is the first failure risk index for the j-th disk.

[0085] Step 402: In each disk group, data calculation and processing are performed based on the first failure risk index corresponding to each of the multiple disks included in the disk group and the average risk index corresponding to the disk group to obtain the discrete value corresponding to each disk group.

[0086] In the embodiments of this application, discrete values ​​can be represented by, but are not limited to, formula (2):

[0087] Formula (2)

[0088] in, This refers to discrete values, which are calculated based on the first failure risk index corresponding to each of the multiple disks and the average risk index corresponding to the disk group. They are intermediate data for calculating the second failure risk index and have no specific substantive meaning. However, they can be used to determine the data distribution characteristics of the first failure risk index corresponding to each of the multiple disks included in the disk group.

[0089] Step 403: In each disk group, perform index calculation processing based on the average risk index and discrete value corresponding to each disk group to obtain the second failure risk index corresponding to each disk group.

[0090] In the embodiments of this application, the calculation of the second failure risk index can be performed using, but is not limited to, formula (3):

[0091] Formula (3)

[0092] in, This is the second failure risk index.

[0093] By using the first and second failure risk indices, a more detailed and comprehensive risk assessment can be conducted, taking into account not only the risk of individual disks but also the overall risk status of disk groups. This helps to better identify and manage high-risk disk groups, thereby enabling more effective preventative measures to ensure data security and stable system operation.

[0094] In one feasible embodiment of this application, it is necessary to perform risk assessment on disks and disk groups based on a first failure risk index and a second failure risk index. Specifically, this can be achieved in the following ways, but not limited to: comparing the first failure risk index with a first preset index threshold; if the first failure risk index is less than the first preset index threshold, marking the disks with the first failure risk index less than the first preset index threshold as non-high-risk disks; if the first failure risk index is greater than or equal to the first preset index threshold, marking the disks with the first failure risk index greater than or equal to the first preset index threshold as high-risk disks; comparing the second failure risk index with a second preset index threshold; if the second failure risk index is less than the second preset index threshold, marking the disk groups with the second failure risk index less than the second preset index threshold as non-high-risk disk groups; if the second failure risk index is greater than or equal to the second preset index threshold, marking the disk groups with the second failure risk index greater than or equal to the second preset index threshold as high-risk disk groups.

[0095] In the embodiments of this application, a non-high-risk disk means that the disk has a low risk of failure during the current evaluation period. A high-risk disk indicates that the disk has a high risk of failure during the current evaluation period and may require additional monitoring or maintenance.

[0096] A non-high-risk disk group indicates that the disk group has a low overall risk of failure. A high-risk disk group indicates that the disk group has a high overall risk of failure and may require corresponding risk mitigation measures.

[0097] The risk level of disks and disk groups can be determined by the magnitude of the first and second failure risk indices. By classifying disks and disk groups according to their risk levels, resources can be allocated in a targeted manner, maintenance plans can be optimized, and the reliability and data protection capabilities of the storage system can be improved. This also helps to identify potential failure points in advance and take preventive measures to reduce data loss and system downtime.

[0098] In one feasible embodiment of this application, as a refinement of step 104 above, regarding disk selection in a high-risk disk group, this application provides a schematic flowchart for disk selection, as follows: Figure 5 As shown, it includes:

[0099] Step 501: Based on the magnitude of the first failure risk index corresponding to each of the multiple disks included in the high-risk disk group, perform risk index classification processing to obtain multiple levels of first failure risk index, wherein the first failure risk index of the higher level is greater than the first failure risk index of the lower level.

[0100] In the embodiments of this application, when the high-risk disk group does not contain a high-risk disk, a disk needs to be selected from the high-risk disk group as a high-risk disk, and this disk will be replaced later to reduce the risk index of the high-risk disk group.

[0101] Regarding the classification of the first failure risk index, the following methods can be used, but are not limited to: dividing the first failure risk index into K levels, and classifying each disk in the high-risk disk group into different levels based on the value of the first failure risk index. Specifically, a disk failure risk level matrix can be used to record this, such as... Figure 6 As shown, Figure 6 This is a schematic diagram illustrating a disk failure risk level provided in this application.

[0102] Step 502: Traverse the first failure risk indices of multiple levels to obtain the target failure risk index, and determine the first quantity of the target failure risk index, wherein the target failure risk index is the first failure risk index of the highest level among the first failure risk indices of multiple levels.

[0103] In the embodiments of this application, the process starts from the highest level and proceeds downwards to find the first failure risk index corresponding to one or more disks with the highest failure level in this high-risk disk group, which is the first failure risk index of the highest level.

[0104] Step 503: If the first number is not greater than the preset number threshold, the disk corresponding to the target failure risk index is identified as the disk to be marked.

[0105] In the embodiments of this application, if there is only one first failure risk index for the highest failure level of a high-risk disk group, then the disk corresponding to the first failure risk index is determined as a high failure risk disk.

[0106] Step 504: If the first quantity is greater than the preset quantity threshold, obtain the power-on time of the disk corresponding to each target failure risk index, and determine the disk with the longest power-on time among the disks corresponding to each target failure risk index as the disk to be marked.

[0107] In the embodiments of this application, if there are multiple first failure risk indices for the highest failure level of a high-risk disk group, the cumulative power-on time of the disks corresponding to the first failure risk indices for the highest failure level is compared, and the disk with the longest cumulative power-on time is determined as the disk to be marked.

[0108] By identifying disks to be tagged in high-risk disk groups, the risks of storage systems can be effectively managed. Even when no high-risk disks exist in the high-risk disk group, disks that need to be replaced can be identified to reduce the risk level of the disk group, which helps to improve the reliability and maintenance efficiency of the storage system.

[0109] In one feasible embodiment of this application, when performing disk replacement processing, the following methods may also be used, but are not limited to: performing disk selection processing in the storage system based on high-risk disks and preset selection conditions to obtain target free disks corresponding to each high-risk disk; copying the data in the high-risk disks to the corresponding target free disks to obtain copied target free disks; and replacing the high-risk disks with the copied target free disks to obtain the storage system after replacement processing.

[0110] In the embodiments of this application, the preset selection criteria are some custom-set conditions, such as disk performance, capacity, health status, location, etc. Specifically, this application does not impose any restrictions on the preset selection criteria.

[0111] For each high-risk disk, a corresponding target free disk is found. The target free disk will be used to replace the high-risk disk, therefore it needs to meet certain requirements and standards, namely preset selection criteria. Data replication refers to copying the data from the high-risk disk to the corresponding target free disk. This process includes, but is not limited to, data migration, synchronization, or mirroring operations to ensure data consistency and integrity. After replication is complete, the target free disk will contain a copy of the same data as the high-risk disk.

[0112] The process of replacing a high-risk disk with a copied target free disk includes, but is not limited to: updating the storage system configuration, disk mapping, Redundant Array of Independent Disks (RAID) configuration, file system, etc., to ensure that the system continues to operate normally. After the replacement process, the original high-risk disk can be removed, repaired, or used as a backup disk, while the storage system continues to be served by the new, lower-risk disk (i.e., the copied target free disk).

[0113] By replacing disks, potential points of failure can be removed from the storage system in a timely manner, thereby improving system stability and data reliability, helping to prevent data loss and system downtime, and ensuring the continuous availability of the storage system.

[0114] In one feasible embodiment of this application, before performing disk replacement, it is necessary to determine in advance whether there are enough free disks in the storage system. If the storage system does not have enough free disks, the operators need to be reminded to add new disks. Specifically, the following methods can be used, but are not limited to: determining whether there is a target free disk in the storage system based on high-risk disks and preset selection conditions; if there is no target free disk in the storage system, generating disk addition information, and performing alarm processing through preset alarm methods based on the disk addition information and preset alarm information.

[0115] In the embodiments of this application, the existence of the target free disk is first determined, that is, based on the high-risk disk and preset selection conditions, it is checked in the storage system whether there is a suitable free disk that can be used as a replacement.

[0116] If no suitable free disk is found in the storage system, the next step is to generate disk addition information. After confirming that there is no target free disk, disk addition information is generated. Disk addition information includes, but is not limited to, the type, quantity, and specifications of the disk to be added, so that the operations personnel can understand the specific actions to be taken.

[0117] Alarm handling refers to processing alarms using preset alarm methods based on generated disk information and preset alarm messages. Preset alarm messages are custom-defined, and preset alarm methods are custom-selected methods, including but not limited to: sending emails, SMS messages, system notifications, and sound alerts. The purpose is to notify storage system operators or relevant personnel of the current storage system status and the necessary actions to be taken. Specifically, this application does not impose any restrictions on preset alarm messages or preset alarm methods.

[0118] By implementing the above methods, it is ensured that relevant personnel can be notified in a timely manner when storage system resources are insufficient, so that necessary measures (such as adding new disks, optimizing storage resources, etc.) can be taken to maintain the normal operation of the system and the security of data. This helps to avoid potential storage resource bottlenecks and prevent system failures in advance.

[0119] In one feasible embodiment of this application, after disk replacement is performed, the operator is notified that the disk replacement has been completed. Specifically, but not limited to the following methods, processing log data of the storage system can be generated based on the data copying process in the high-risk disk; and a notification is given based on the processing log data and preset prompt information through a preset prompt method.

[0120] In the embodiments of this application, during the data copying process on high-risk disks, relevant operations and events are recorded to generate processing log data.

[0121] Processing log data includes, but is not limited to: the start and end times of the replication operation, the amount of data replicated, any errors or warnings during the replication process, the disk identifiers involved, operator information, etc. Processing log data is an important record of the operation history of the storage system, which is of great significance for tracking operation history, troubleshooting and system auditing.

[0122] The preset prompting method is a customizable selection, including but not limited to: pop-up windows, sound prompts, email notifications, and SMS notifications, with the aim of providing system administrators or operators with real-time operational feedback and system status updates. The preset prompting information is customizable and includes, but is not limited to: the progress of the copying operation, its success or failure status, and subsequent actions to be taken. Specifically, this application does not impose any restrictions on the preset prompting information or method.

[0123] By using a user feedback mechanism, the operation of the storage system is made transparent and traceable, which helps administrators to understand changes in the system's status in a timely manner and respond quickly to potential problems, thereby improving the management efficiency and reliability of the storage system.

[0124] In conjunction with the control method for the storage system, this application also provides an application environment architecture and a hardware architecture for executing the control method for the storage system, which are described herein.

[0125] like Figure 7 As shown, Figure 7 This application provides an architectural diagram of a distributed storage system, which mainly includes an application layer, a protocol layer, a management layer, and a hardware layer.

[0126] The application layer comprises programs that implement user requirements, including but not limited to: clients, applications, and virtual machines. The protocol layer primarily meets the access needs of different user protocols and consists of the Data Management and Protocol Fusion (LIBUDOSS) library. The LIBUDOSS library is a custom library developed for specific distributed storage functions or systems, providing different interface protocols that allow applications to interact with the distributed storage system.

[0127] The management layer is the core of the distributed storage system and is composed of the Unified Distributed Storage System (UDDOS). It mainly includes the following four types of services: (1) Monitoring Service (MONS): A daemon process used to maintain the cluster state mapping, including how many storage pools are in the cluster, how many disk groups (Placement Groups, PGs) are in each storage pool, and the mapping relationship between storage pools and PGs. (2) Management Service (MGRS): A daemon process responsible for managing cluster information and tracking runtime metrics and current status. (3) Object Storage Daemon (OSD): Each disk corresponds to one OSD daemon process, which is responsible for handling cluster data replication, recovery, and rebalancing, and provides information to the monitoring service and management service by checking the status of other OSD daemons. (4) Metadata Server (MDS): Manages the cluster's metadata. A storage pool is a logical storage unit in Ceph (Ceph Distributed File System) used to organize and manage data. There are two types of storage pools: multi-replica storage pools and erasure coding storage pools. A disk group (PG) is a physical data distribution unit in a distributed storage system, used to distribute data across different OSDs. A storage pool consists of multiple disk groups, each responsible for managing and storing a portion of the data.

[0128] The hardware layer consists of the physical disks that actually provide storage services.

[0129] It should be noted that this architecture is merely an example of a control method for executing a storage system and does not constitute a limitation of this application.

[0130] In summary, this application can achieve the following technical effects:

[0131] This application first uses a disk failure prediction model, i.e., a preset risk prediction model, to calculate the failure risk index of each disk in the storage system and identify disks with high failure risk. Since a disk group includes multiple disks, the failure risk index of the disk group can be calculated based on the failure risk index of each disk, thus identifying disk groups with high failure risk. However, because the failure risk index of a disk group is calculated from the failure risk indices of the disks included in the group, there may be situations where a high-risk disk group does not contain any high-risk disks. To further reduce the failure risk of the storage system, in addition to replacing high-risk disks, it is also necessary to reduce the failure risk of all high-risk disk groups. Therefore, it is necessary to select one disk from the high-risk disk groups where no high-risk disks exist. The process of marking disks to be marked allows for the reduction of the failure risk of high-risk disk groups by replacing these marked disks. Subsequently, all high-risk disks and marked disks in the storage system are replaced. This allows for the timely handling of disks and disk groups with failure risks in the storage system. It addresses the technical problem in existing technologies where the inability to promptly handle disks with failure risks leads to the number of failed disks exceeding the system's fault tolerance capacity, resulting in storage system failure and a high risk of data loss. The goal is to quickly identify high-risk disks and disk groups, perform timely data backup and replacement, ensure that the number of failed disks in the storage system does not exceed the system's fault tolerance capacity, effectively reduce the failure risk of the storage system, and guarantee system reliability.

[0132] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0133] Embodiments of this application also provide a control device for a storage system. Figure 8 A schematic diagram of the structure of a control device for a storage system provided in the application, such as... Figure 8 As shown, it includes:

[0134] The prediction unit 801 is used to perform failure risk prediction processing through a preset risk prediction model and disk information of the disks in the storage system to obtain the first failure risk index corresponding to each of the multiple disks. The storage system includes multiple disk groups, and each disk group includes multiple disks.

[0135] The calculation unit 802 is used to perform index calculation processing based on the first failure risk index corresponding to each of the multiple disks included in each disk group and the average risk index corresponding to each disk group to obtain the second failure risk index corresponding to each disk group. The average risk index is calculated by averaging the first failure risk indices corresponding to the multiple disks included in the disk group.

[0136] The identification unit 803 is used to identify high-risk disks and high-risk disk groups, wherein a high-risk disk is a disk whose first failure risk index is greater than or equal to a first preset index threshold, and a high-risk disk group is a disk group whose second failure risk index is greater than or equal to a second preset index threshold.

[0137] The selection unit 804 is used to select disks in the high-risk disk group when there are no high-risk disks in the high-risk disk group, to obtain disks to be marked, and to mark the disks to be marked as high-risk disks.

[0138] The processing unit 805 is used to perform disk replacement processing on all high-risk disks in the storage system to obtain a replaced storage system.

[0139] In one embodiment of this application, the prediction unit 801 is further configured to:

[0140] The disk information of the disk in the storage system is input into the preset risk prediction model for failure risk prediction processing to obtain the initial failure risk index. The preset risk prediction model is a pre-configured model for predicting failure risk.

[0141] The initial failure risk index is normalized to obtain the first failure risk index.

[0142] In one embodiment of this application, the prediction unit 801 is further configured to:

[0143] Based on the disk type in the disk information, a first prediction model is determined for each of the multiple disks. The preset risk prediction model includes at least one first prediction model, and the first prediction model corresponds to one or more disks.

[0144] The disk information is input into the corresponding first prediction model for failure risk prediction processing to obtain the initial failure risk index. The disk information includes at least the disk type, disk status data, data acquisition time, and disk identifier.

[0145] In one embodiment of this application, such as Figure 9 As shown, the control device for the storage system also includes:

[0146] Modeling unit 806 is used to perform disk modeling processing based on disk capacity, disk identifier, and disk type to obtain multiple disk models, wherein each disk model corresponds to at least one disk;

[0147] The matching unit 807 is used to perform prediction model matching processing based on the disk model to obtain a preset risk prediction model.

[0148] In one embodiment of this application, such as Figure 9 As shown, the control device for the storage system also includes:

[0149] The acquisition unit 808 is used to acquire disk status data of multiple disks in the storage system according to a preset period;

[0150] The collection unit 809 is used to perform data collection processing on disk status data, disk capacity, disk identifier, disk type, and data acquisition time to obtain disk information, wherein the data acquisition time is the acquisition time of disk status data.

[0151] In one embodiment of this application, the computing unit 802 is further configured to:

[0152] In each disk group, the average risk index corresponding to each disk group is calculated by averaging the first failure risk index corresponding to each of the multiple disks included in the disk group.

[0153] In each disk group, data is calculated and processed based on the first failure risk index corresponding to each of the multiple disks included in the disk group and the average risk index corresponding to the disk group, to obtain the discrete value corresponding to each disk group.

[0154] Within each disk group, an index calculation is performed based on the average risk index and discrete value corresponding to each disk group to obtain the second failure risk index corresponding to each disk group.

[0155] In one embodiment of this application, such as Figure 9 As shown, the control device for the storage system also includes:

[0156] The comparison unit 810 is used to compare the first failure risk index with the first preset index threshold. If the first failure risk index is less than the first preset index threshold, the disk with the first failure risk index less than the first preset index threshold is marked as a non-high-risk disk.

[0157] The comparison unit 810 is also used to mark disks whose first failure risk index is greater than or equal to the first preset index threshold as high-risk disks when the first failure risk index is greater than or equal to the first preset index threshold.

[0158] The comparison unit 810 is also used to compare the second failure risk index with the second preset index threshold. If the second failure risk index is less than the second preset index threshold, the disk group whose second failure risk index is less than the second preset index threshold is marked as a non-high-risk disk group.

[0159] The comparison unit 810 is also used to mark disk groups whose second failure risk index is greater than or equal to the second preset index threshold as high-risk disk groups when the second failure risk index is greater than or equal to the second preset index threshold.

[0160] In one embodiment of this application, the selection unit 804 is further configured to:

[0161] Based on the magnitude of the first failure risk index corresponding to each of the multiple disks included in the high-risk disk group, the risk index is graded to obtain multiple levels of first failure risk index, wherein the first failure risk index of the higher level is greater than the first failure risk index of the lower level.

[0162] Traverse the first failure risk indices of multiple levels to obtain the target failure risk index, and determine the first quantity of the target failure risk index, wherein the target failure risk index is the first failure risk index of the highest level among the first failure risk indices of multiple levels.

[0163] If the first number is not greater than the preset number threshold, the disk corresponding to the target failure risk index is identified as the disk to be marked.

[0164] If the first number is greater than the preset number threshold, obtain the power-on time of the disk corresponding to each target failure risk index, and determine the disk with the longest power-on time among the disks corresponding to each target failure risk index as the disk to be marked.

[0165] In one embodiment of this application, the processing unit 805 is further configured to:

[0166] Based on high-risk disks and preset selection criteria, disk selection is performed in the storage system to obtain the target free disks corresponding to each high-risk disk.

[0167] Copy the data from the high-risk disk to the corresponding target free disk to obtain the copied target free disk;

[0168] The high-risk disk is replaced by the copied target free disk, resulting in a replaced storage system.

[0169] In one embodiment of this application, such as Figure 9 As shown, the control device for the storage system also includes:

[0170] The determining unit 811 is used to determine whether there is a target free disk in the storage system based on high-risk disks and preset selection conditions;

[0171] The alarm unit 812 is used to generate disk addition information if there is no target free disk in the storage system, and to perform alarm processing according to the disk addition information and preset alarm information through preset alarm method.

[0172] In one embodiment of this application, such as Figure 9 As shown, the control device for the storage system also includes:

[0173] The generation unit 813 is used to generate storage system processing log data based on the data copying process in high-risk disks;

[0174] The prompting unit 814 is used to provide prompts based on the processing log data and preset prompt information, using a preset prompting method.

[0175] For a description of the features in the embodiment corresponding to the control device of the storage system, please refer to the relevant description of the embodiment corresponding to the control method of the storage system, which will not be repeated here.

[0176] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above-described control method embodiments of the storage system.

[0177] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described storage system control method embodiments when running.

[0178] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0179] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described storage system control method embodiments.

[0180] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described storage system control method embodiments.

[0181] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0182] The control method and apparatus, electronic device, and storage medium of a storage system provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. It should be noted that those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A control method for a storage system, characterized in that, include: Failure risk prediction is performed by using a preset risk prediction model and disk information of disks in the storage system to obtain a first failure risk index corresponding to each of the disks. Each storage pool in the storage system includes multiple disk groups, and each disk group includes multiple disks. The storage pool includes a multi-replica storage pool and an erasure coding storage pool. The disk group is the physical data distribution unit in the storage system, used to distribute data to different object storage guardian services and to manage and store the data. The discrete value corresponding to each disk group is calculated based on the first failure risk index corresponding to each of the disks included in each disk group and the average risk index corresponding to each disk group. The second failure risk index corresponding to each disk group is obtained by performing index calculation processing based on the discrete value and the average risk index. The average risk index is obtained by averaging the first failure risk indices corresponding to each of the disks included in the disk group. Identify high-risk disks and high-risk disk groups, wherein the high-risk disks are disks whose first failure risk index is greater than or equal to a first preset index threshold, the high-risk disk groups are disk groups whose second failure risk index is greater than or equal to a second preset index threshold, the high-risk disk groups are divided into high-risk disk groups including the high-risk disks and high-risk disk groups excluding the high-risk disks, and the high-risk disks are disks to be replaced; If the high-risk disk is not in the high-risk disk group, the disks in the high-risk disk group are selected to obtain disks to be marked, and the disks to be marked are marked as the high-risk disks; The disk selection process includes: classifying the first failure risk index into multiple levels of first failure risk index, determining the highest level of first failure risk index among the multiple levels of first failure risk index, and determining the disk with the longest cumulative power-on time among the highest level of first failure risk index as the disk to be marked. All the high-risk disks in the storage system are selected as disks to be replaced and a disk replacement process is performed to obtain a storage system after replacement. The second failure risk index can be calculated using the following formula: ; Where j is the j-th disk in the disk group, and the average risk index is The number of disks in the disk group is N, and the first failure risk index is The second failure risk index is .

2. The control method for the storage system according to claim 1, characterized in that, The step of performing failure risk prediction processing on the disks in the storage system using a preset risk prediction model to obtain a first failure risk index corresponding to each of the disks includes: The disk information of the disk in the storage system is input into the preset risk prediction model for failure risk prediction processing to obtain an initial failure risk index. The preset risk prediction model is a pre-configured model for predicting failure risk. The initial failure risk index is normalized to obtain the first failure risk index.

3. The control method for the storage system according to claim 2, characterized in that, The step of inputting the disk information of the disk in the storage system into the preset risk prediction model for failure risk prediction processing to obtain the initial failure risk index includes: Based on the disk type in the disk information, a first prediction model is determined for each of the multiple disks, wherein the preset risk prediction model includes at least one first prediction model, and the first prediction model corresponds to one or more of the disks; The disk information is input into the corresponding first prediction model for failure risk prediction processing to obtain the initial failure risk index. The disk information includes at least the disk type, disk status data, data acquisition time, and disk identifier.

4. The control method for the storage system according to claim 3, characterized in that, Before performing failure risk prediction processing using a preset risk prediction model and disk information from the disks in the storage system to obtain the first failure risk index corresponding to each of the multiple disks, the method further includes: Based on the disk capacity, disk identifier, and disk type, disk modeling is performed to obtain multiple disk models, wherein each disk model corresponds to at least one disk; Based on the disk model, a prediction model matching process is performed to obtain the preset risk prediction model.

5. The control method for the storage system according to claim 4, characterized in that, Before performing failure risk prediction processing using a preset risk prediction model and disk information from the disks in the storage system to obtain the first failure risk index corresponding to each of the multiple disks, the method further includes: According to a preset period, disk status data of multiple disks are collected in the storage system. The disk status data, disk capacity, disk identifier, disk type, and data acquisition time are processed into a data set to obtain the disk information, wherein the data acquisition time is the acquisition time of the disk status data.

6. The control method for the storage system according to claim 1, characterized in that, The step of calculating the second failure risk index for each disk group based on the first failure risk index corresponding to each of the multiple disks included in each disk group and the average risk index corresponding to each disk group includes: In each disk group, the average risk index corresponding to each disk group is calculated by averaging the first failure risk index corresponding to each of the multiple disks included in the disk group; In each disk group, data calculation and processing are performed based on the first failure risk index corresponding to each of the multiple disks included in the disk group and the average risk index corresponding to the disk group to obtain the discrete value corresponding to each disk group. In each disk group, an index calculation process is performed based on the average risk index and the discrete value corresponding to each disk group to obtain the second failure risk index corresponding to each disk group.

7. The control method for the storage system according to claim 1, characterized in that, After performing index calculation processing based on the first failure risk index corresponding to each of the plurality of disks included in each disk group and the average risk index corresponding to each disk group to obtain the second failure risk index corresponding to each disk group, the method further includes: The first failure risk index is compared with the first preset index threshold. If the first failure risk index is less than the first preset index threshold, the disk with the first failure risk index less than the first preset index threshold is marked as a non-high-risk disk. If the first failure risk index is greater than or equal to the first preset index threshold, the disks whose first failure risk index is greater than or equal to the first preset index threshold are marked as high-risk disks. The second failure risk index is compared with the second preset index threshold. If the second failure risk index is less than the second preset index threshold, the disk group whose second failure risk index is less than the second preset index threshold is marked as a non-high-risk disk group. If the second failure risk index is greater than or equal to the second preset index threshold, the disk group whose second failure risk index is greater than or equal to the second preset index threshold is marked as the high-risk disk group.

8. The control method for the storage system according to claim 1, characterized in that, If no high-risk disk is found in the high-risk disk group, the disks in the high-risk disk group are selected to obtain disks to be marked, including: Based on the magnitude of the first failure risk index corresponding to each of the multiple disks included in the high-risk disk group, risk index grading is performed to obtain the first failure risk index of the multiple levels, wherein the first failure risk index of the higher level is greater than the first failure risk index of the lower level. Traverse the first failure risk indices of the multiple levels to obtain the target failure risk index, and determine the first quantity of the target failure risk index, wherein the target failure risk index is the first failure risk index of the highest level among the first failure risk indices of the multiple levels; If the first quantity is not greater than a preset quantity threshold, the disk corresponding to the target failure risk index is determined as the disk to be marked. If the first quantity is greater than the preset quantity threshold, the cumulative power-on time of the disk corresponding to each target failure risk index is obtained, and the disk with the longest cumulative power-on time among the disks corresponding to each target failure risk index is determined as the disk to be marked.

9. The control method for the storage system according to claim 1, characterized in that, The step of performing disk replacement on all the high-risk disks in the storage system to obtain a replaced storage system includes: Based on the high-risk disks and preset selection criteria, disk selection processing is performed in the storage system to obtain the target free disks corresponding to each of the high-risk disks. The data in the high-risk disk is copied to the corresponding target free disk to obtain the copied target free disk; The high-risk disk is replaced by the copied target free disk to obtain the storage system after the replacement process.

10. The control method for the storage system according to claim 9, characterized in that, Before performing disk selection processing in the storage system based on the high-risk disks and preset selection criteria to obtain the free disks corresponding to each of the high-risk disks, the method further includes: Determine whether the target free disk exists in the storage system based on the high-risk disk and the preset selection criteria; If the target free disk does not exist in the storage system, disk addition information is generated, and alarm processing is performed according to the disk addition information and preset alarm information using a preset alarm method.

11. The control method for the storage system according to claim 9, characterized in that, After copying the data from the high-risk disk to the corresponding target free disk to obtain the copied target free disk, the method further includes: Based on the data copying process in the high-risk disk, the processing log data of the storage system is generated; Based on the processing log data and preset prompt information, prompts are processed using preset prompt methods.

12. A control device for a storage system, characterized in that, include: The prediction unit is used to perform failure risk prediction processing through a preset risk prediction model and disk information of the disks in the storage system to obtain a first failure risk index corresponding to each of the disks. Each storage pool in the storage system includes multiple disk groups, and each disk group includes multiple disks. The storage pool includes a multi-replica storage pool and an erasure coding storage pool. The disk group is the physical data distribution unit in the storage system, used to distribute data to different object storage guardian services and to manage and store the data. The calculation unit is configured to calculate a discrete value corresponding to the disk group based on the first failure risk index corresponding to each of the plurality of disks included in each disk group and the average risk index corresponding to each disk group, and to perform index calculation processing based on the discrete value and the average risk index to obtain a second failure risk index corresponding to each disk group, wherein the average risk index is obtained by averaging the first failure risk indices corresponding to the plurality of disks included in the disk group. An identification unit is used to identify high-risk disks and high-risk disk groups, wherein the high-risk disks are disks whose first failure risk index is greater than or equal to a first preset index threshold, the high-risk disk groups are disk groups whose second failure risk index is greater than or equal to a second preset index threshold, the high-risk disk groups are divided into high-risk disk groups including the high-risk disks and high-risk disk groups excluding the high-risk disks, and the high-risk disks are disks to be replaced. The selection unit is used to select disks in the high-risk disk group when the high-risk disk does not exist in the high-risk disk group, to obtain disks to be marked, and to mark the disks to be marked as the high-risk disks; The disk selection process includes: classifying the first failure risk index into multiple levels of first failure risk index, determining the highest level of first failure risk index among the multiple levels of first failure risk index, and determining the disk with the longest cumulative power-on time among the highest level of first failure risk index as the disk to be marked. The processing unit is used to perform disk replacement processing on all the high-risk disks in the storage system as the disks to be replaced, so as to obtain the storage system after replacement processing; The second failure risk index can be calculated using the following formula: ; Where j is the j-th disk in the disk group, and the average risk index is The number of disks in the disk group is N, and the first failure risk index is The second failure risk index is .

13. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the control method for the storage system as described in any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the control method for the storage system as described in any one of claims 1 to 11.

15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the control method for the storage system as described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Disk fault prediction method, device and equipment and computer storage medium

    CN113127274A