Disk management and control method and device and electronic equipment
By periodically collecting disk hardware status information, identifying abnormal values and calculating health scores, proactive data migration is achieved, solving the problem of data loss when the disk is damaged and improving data security.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-04-03
AI Technical Summary
In existing technologies, when a disk is damaged, there is a high possibility of data loss. If a data migration process is performed after a disk is damaged, data security is insufficient.
By periodically collecting hardware status information of the disk, identifying outliers, and calculating a health score based on the current weight of the target features, a data migration process is triggered when the health score falls below a predetermined threshold, thus achieving proactive data migration.
This reduces the likelihood of insufficient data security during data migration after actual disk failure, and improves the reliability of data stored on the disk.
Smart Images

Figure CN121785531A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a disk management method, apparatus and electronic device. Background Technology
[0002] In the field of data storage, disks may fail due to physical degradation or excessive usage. As the core asset of an information system, the security of data stored on disks is paramount, and data migration is a crucial means to ensure this.
[0003] In existing technologies, data migration is usually triggered by actual disk failure. That is, when a disk is damaged and identified as a bad disk, the management device will initiate a data migration process for that disk to transfer the data in that disk to other disks.
[0004] However, when a disk fails, there is a high probability of data loss. If a data migration process is performed after disk failure, the migrated data will also be lost. Therefore, existing technologies undoubtedly have insufficient data security. Summary of the Invention
[0005] The purpose of this application is to provide a disk management method, apparatus, and electronic device to improve the data reliability of data stored on the disk. The specific technical solution is as follows:
[0006] In a first aspect provided by the embodiments of this application, a disk management method is provided, the method comprising:
[0007] According to a predetermined collection cycle, the hardware status information of the target disk to be analyzed is collected; wherein, the hardware status information includes feature values of multiple target features of the target disk;
[0008] In response to the acquired hardware status information, identify whether there are any feature values that are outliers in the currently acquired hardware status information;
[0009] If present, obtain the current weight corresponding to each target feature; wherein, the current weight of each target feature is the weight set for that target feature within the current predetermined period; the period length of the predetermined period is greater than the period length of the predetermined collection period;
[0010] Based on the feature value of each target feature in the currently collected hardware status information and the obtained current weight, a health score is calculated to characterize the health status of the target disk; wherein, the health score is used to characterize the probability that the target disk has the risk of failure in the future time period;
[0011] When the obtained health score meets the predetermined triggering conditions, the data migration process for the target disk is triggered; wherein, the predetermined triggering conditions include the health score being lower than a predetermined score threshold.
[0012] Optionally, in one implementation, identifying whether there are any feature values belonging to outliers in the currently acquired hardware status information includes:
[0013] Obtain the time series features corresponding to each target feature; wherein, the time series features corresponding to each target feature are used to characterize the feature distribution of the target feature within the target time period;
[0014] For each target feature, based on the time series features corresponding to that target feature, analyze whether the feature value of that target feature in the currently acquired hardware status information is an outlier, and obtain the analysis result corresponding to that target feature;
[0015] Based on the analysis results corresponding to each target feature, determine whether there are any feature values that are outliers in the currently collected hardware status information.
[0016] Optionally, in one implementation, calculating a health score characterizing the health status of the target disk based on the feature value of each target feature in the currently acquired hardware status information and the obtained current weight includes:
[0017] For each target feature, the historical mean and historical standard deviation of the target feature are obtained, and the target index value corresponding to the target feature is calculated based on the feature value of the target feature in the currently collected hardware status information, the historical mean, and the historical standard deviation; wherein, the historical standard deviation of the target feature is the standard deviation corresponding to each historical feature value of the target feature, and the target index value corresponding to the target feature is used to characterize: the degree of impact of the target feature on disk damage when the target feature has the feature value in the currently collected hardware status information;
[0018] Based on the target index value corresponding to each target feature and the current weight corresponding to that target feature, a health score is calculated to characterize the health status of the target disk.
[0019] Optionally, in one implementation, calculating a health score to characterize the health status of the target disk based on the target indicator value corresponding to each target feature and the current weight corresponding to that target feature includes:
[0020] Using the first formula, a health score representing the health status of the target disk is calculated based on the target index value corresponding to each target feature and the current weight corresponding to that target feature; wherein, the first formula includes:
[0021] ;
[0022] Wherein, Score represents the health score of the target disk, and i represents the i-th target feature of the target disk. F represents the current weight corresponding to the i-th target feature. i The target index value that represents the i-th target feature.
[0023] Optionally, in one implementation, the calculation method for the target index value corresponding to each target feature includes:
[0024] Calculate the target index value corresponding to each target feature according to the second formula;
[0025] The second formula includes:
[0026] ;
[0027] F i The target index value that represents the i-th target feature. The feature value representing the i-th target feature in the currently acquired hardware status information. The historical mean representing the feature of the i-th target. The historical standard deviation characterizing the i-th target feature.
[0028] Optionally, in one implementation, each target feature corresponds to an initial weight; the method for determining the weight of each target feature within a predetermined period includes:
[0029] When the predetermined period is entered, the feature value sequence of the target feature of the target disk is obtained; wherein, the feature value sequence of the target feature includes the feature values of the target feature that have been collected;
[0030] Based on the feature value sequence of the target feature, the degree of impact of the target feature on the health status of the disk in the current predetermined period is estimated.
[0031] Based on the estimated degree of impact, the current weight of the target feature is adjusted according to a predetermined adjustment method, and the adjusted weight is used as the current weight set for the target feature within the current predetermined period. The predetermined adjustment method includes: if the estimated degree of impact within the current predetermined period does not match the estimated degree of impact in the previous predetermined period, the weight is increased in response to the increase in the estimated degree of impact relative to the previous predetermined period; otherwise, the weight is decreased.
[0032] Optionally, in one implementation, the data migration process for the target disk includes:
[0033] According to the predetermined disk selection criteria, select the disk to which the data stored in the target disk is to be migrated, obtain the disk to be used, and migrate the data stored in the target disk to the disk to be used according to the predetermined concurrency.
[0034] The predetermined disk selection criteria include: the most recent health score is higher than the target score, where the target score represents the health score with the highest probability of no disk failure risk within a specified time period; and the predetermined concurrency level represents the amount of data that can be migrated in parallel for each data migration process.
[0035] Optionally, in one implementation, after migrating the data stored in the target disk to the disk to be used according to a predetermined concurrency level, the method further includes:
[0036] Real-time monitoring of request response latency of management devices used to control disks;
[0037] When the detected request response latency is higher than the predetermined latency, the predetermined concurrency level is reduced, and the data stored in the target disk is migrated to the disk to be used according to the reduced predetermined concurrency level.
[0038] The downsizing process includes reducing the predetermined concurrency by a predetermined amount to obtain the reduced predetermined concurrency, or calculating the reduced predetermined concurrency based on the values used in the management device to characterize the processor's operating state.
[0039] Optionally, in one implementation, the number of target disks to be analyzed is multiple;
[0040] The step of triggering a data migration process for the target disk when the obtained health score meets a predetermined trigger condition includes:
[0041] When the health scores obtained for at least two target disks meet the predetermined triggering conditions, the data migration process for each of the at least two target disks is triggered sequentially according to the priority of the at least two target disks.
[0042] Among the at least two target disks, the target disk with higher priority has a lower health score compared to the target disk with lower priority.
[0043] Optionally, in one implementation, the method further includes:
[0044] After the data migration process of the target disk with the same priority is completed, a cooling verification process is performed on the data migrated from the target disk with the same priority.
[0045] Optionally, in one implementation, the predetermined triggering condition further includes that the health score obtained in multiple consecutive predetermined collection cycles is lower than the predetermined score threshold, and the variance of the health score of the target disk in the multiple predetermined collection cycles is less than a predetermined variance.
[0046] In a second aspect provided in this application, a disk management device is also provided, the device comprising:
[0047] The information acquisition module is used to acquire hardware status information of the target disk to be analyzed according to a predetermined acquisition cycle; wherein, the hardware status information includes feature values of multiple target features of the target disk;
[0048] The feature value recognition module is used to identify whether there are any feature values that are outliers in the currently acquired hardware status information in response to the acquired hardware status information.
[0049] The weight acquisition module is used to acquire the current weight corresponding to each target feature when it exists; wherein, the current weight of each target feature is the weight set for the target feature within the current predetermined period; the period length of the predetermined period is greater than the period length of the predetermined collection period;
[0050] The score calculation module is used to calculate a health score that characterizes the health status of the target disk based on the feature value of each target feature in the currently collected hardware status information and the obtained current weight; wherein, the health score is used to characterize the probability that the target disk has the risk of failure in the future time period;
[0051] The data migration module is used to trigger a data migration process for the target disk when the obtained health score meets a predetermined trigger condition; wherein, the predetermined trigger condition includes the health score being lower than a predetermined score threshold.
[0052] In a third aspect provided in the embodiments of this application, an electronic device is also provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; the memory is used to store computer programs; and the processor is used to implement any of the disk management methods provided in the first aspect when executing the programs stored in the memory.
[0053] In another aspect provided by the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored therein, and when the computer program is executed by a processor, it implements any of the disk management methods provided in the first aspect above.
[0054] In another aspect provided by the embodiments of this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the disk management methods provided in the first aspect above.
[0055] As can be seen from the above, in the disk management method provided in this application embodiment, the current weight of each target feature is a weight set for that target feature within the current predetermined period. That is, the current weight of the target feature is different depending on the current predetermined period. Thus, based on the feature value of each target feature in the currently collected hardware status information and the obtained current weight, a health score can be calculated to characterize the health status of the target disk within the current predetermined period. This health score is used to characterize the probability of the target disk having a risk of failure in the future time period. Therefore, when the obtained health score meets the predetermined triggering conditions, a data migration process for the target disk is triggered. As can be seen, this solution periodically collects the hardware status information of the disk, and when anomalies are detected in the collected hardware status information, it combines the weight of each target feature in the hardware status information within the current predetermined period with the feature value of each target feature collected to calculate the disk health score in a timely and accurate manner. Thus, when the calculated health score is lower than a predetermined score threshold, proactive data migration is performed on disks that may be faulty in advance. This reduces the possibility of insufficient data security caused by data migration after actual disk failure, and improves the data reliability of the data stored on the disk. Attached Figure Description
[0056] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.
[0057] Figure 1 A schematic flowchart illustrating a disk management method provided in an embodiment of this application;
[0058] Figure 2 A flowchart illustrating another disk management method provided in an embodiment of this application;
[0059] Figure 3 A flowchart illustrating a method for determining the weight of each target feature within a predetermined period, as provided in an embodiment of this application;
[0060] Figure 4 A schematic diagram of the architecture of a specific embodiment provided in this application;
[0061] Figure 5 A flowchart illustrating a specific embodiment of this application is provided.
[0062] Figure 6 This is a schematic diagram of the structure of a disk management device provided in an embodiment of this application;
[0063] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0064] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of this application.
[0065] To address the aforementioned technical problems, embodiments of this application provide a disk management method, apparatus, and electronic device. This method is applicable to various application scenarios involving disk management, such as managing disk online and offline operations. Furthermore, the method is applied to management devices capable of disk management, which can be various electronic devices such as laptops and desktop computers. Therefore, embodiments of this application do not limit the application scenarios or the executing entity of the method.
[0066] This application provides a disk management method that may include the following steps:
[0067] According to a predetermined collection cycle, the hardware status information of the target disk to be analyzed is collected; wherein, the hardware status information includes feature values of multiple target features of the target disk;
[0068] In response to the acquired hardware status information, identify whether there are any feature values that are outliers in the currently acquired hardware status information;
[0069] If present, obtain the current weight corresponding to each target feature; wherein, the current weight of each target feature is the weight set for that target feature within the current predetermined period; the period length of the predetermined period is greater than the period length of the predetermined collection period;
[0070] Based on the feature value of each target feature in the currently collected hardware status information and the obtained current weight, a health score is calculated to characterize the health status of the target disk; wherein, the health score is used to characterize the probability that the target disk has the risk of failure in the future time period;
[0071] When the obtained health score meets the predetermined triggering conditions, the data migration process for the target disk is triggered; wherein, the predetermined triggering conditions include the health score being lower than a predetermined score threshold.
[0072] As can be seen from the above, in the disk management method provided in this application embodiment, the current weight of each target feature is a weight set for that target feature within the current predetermined period. That is, the current weight of the target feature is different depending on the current predetermined period. Thus, based on the feature value of each target feature in the currently collected hardware status information and the obtained current weight, a health score can be calculated to characterize the health status of the target disk within the current predetermined period. This health score is used to characterize the probability of the target disk having a risk of failure in the future time period. Therefore, when the obtained health score meets the predetermined triggering conditions, a data migration process for the target disk is triggered. As can be seen, this solution periodically collects the hardware status information of the disk, and when anomalies are detected in the collected hardware status information, it combines the weight of each target feature in the hardware status information within the current predetermined period with the feature value of each target feature collected to calculate the disk health score in a timely and accurate manner. Thus, when the calculated health score is lower than a predetermined score threshold, proactive data migration is performed on disks that may be faulty in advance. This reduces the possibility of insufficient data security caused by data migration after actual disk failure, and improves the data reliability of the data stored on the disk.
[0073] Figure 1 This is a flowchart illustrating a disk management method provided in an embodiment of this application, as shown below. Figure 1 As shown, the method may include the following steps:
[0074] S101: Collect hardware status information of the target disk to be analyzed according to the predetermined collection cycle.
[0075] The hardware status information includes feature values of multiple target characteristics of the target disk.
[0076] The purpose of this application is to enable proactive data migration. Therefore, the management device used to control the disk can collect feature values of multiple target features of the target disk to be analyzed according to a predetermined collection cycle, as the hardware status information of the target disk.
[0077] For example, the management device collects feature values of multiple target characteristics of the target disk to be analyzed through the NVMe-MI (NVMe Management Interface, an out-of-band management interface in the NVMe protocol specifically designed for storage device management) protocol according to a predetermined collection cycle, which serve as the hardware status information of the target disk. The predetermined collection cycle can be one hour, two hours, one day, two days, etc., and this application does not specifically limit it.
[0078] It should be noted that different types of disks have different target characteristics. For example, disk types may include hard disk drives (HDDs), solid-state drives (SSDs), and solid-state hybrid drives (SSHDs). Specifically, the target characteristics set for hard disk drives may include some or all of the following: reallocation sector count, current unmapped sector count, head load / unload cycle count, and disk temperature. The target characteristics set for solid-state drives may include some or all of the following: reallocation sector count, wear leveling count, programming failure count, and disk temperature. The target characteristics set for solid-state drives may include some or all of the following: reallocation sector count, current unmapped sector count, head load / unload cycle count, wear leveling count, and disk temperature. The above examples are only partial examples from practical applications and do not limit the disk types involved in this application or the target characteristics set for each disk type.
[0079] Optionally, in addition to the above-mentioned features, the target characteristics of different types of disks may also include some or all of the SMART (Self-Monitoring Analysis and Reporting Technology) features such as NAND wear WAF (Write Amplification Factor), PE (Program / Erase) cycle jitter rate, cross-plane interference index, cell electrical stress accumulation value, charge leakage rate, and total physical write volume. In this regard, the embodiments of this application do not make specific limitations.
[0080] The aforementioned NAND wear rating WAF is determined based on the ratio of actual written data to host request volume, reflecting write efficiency. The aforementioned PE cycle jitter rate is determined based on the standard deviation of the program / erase cycle duration fluctuation. The aforementioned cross-plane interference index is determined based on the charge interference intensity between adjacent storage cells. Taking a solid-state drive as an example, the aforementioned cumulative cell electrical stress value is determined based on the historical integral of the voltage stress borne by the floating gate transistor. The aforementioned charge leakage rate is determined based on the abnormal flow velocity feature extraction layer of charge in the floating gate (a component of the floating gate transistor). The aforementioned total physical write volume is determined based on the actual written data volume within the lifetime. This application does not limit the method of obtaining the target feature; any method that can obtain the feature value of the target feature is applicable to this application.
[0081] It should be emphasized that in practical applications, the selection or setting of various target characteristics of the target disk can be set according to the actual situation, and no restrictions are imposed here.
[0082] Based on this, optionally, according to the disk type to which the target disk belongs, the feature values of multiple target features of the target disk are collected according to a predetermined collection period, as the hardware status information of the target disk.
[0083] S102: In response to the acquired hardware status information, identify whether there are any feature values that are outliers in the currently acquired hardware status information.
[0084] In this application, in response to the collected hardware status information, the feature value of any target feature in the currently collected hardware status information is identified, and compared with the historical feature value of the target feature within the current predetermined collection period, it is determined whether it is an outlier. If it is, it indicates that the target disk currently has a risk of failure, and in this case, subsequent steps are performed to determine the health status of the target disk. Otherwise, it indicates that the target disk currently does not have a risk of failure, and in this case, the current health status of the target disk can be considered healthy, and it is not necessary to continue performing subsequent steps to determine the health status of the target disk.
[0085] Optionally, historical feature values for each target feature are pre-acquired to determine the historical feature mean for that target feature. Thus, when there is a numerical difference between the currently acquired feature value of the target feature and the determined historical feature mean, it can be determined that the currently acquired feature value of the target feature is an outlier, that is, it is identified that there is an outlier feature value in the currently acquired hardware status information.
[0086] Alternatively, in one implementation, such as Figure 2As shown, in step S102 above, identifying whether there are any feature values that are outliers in the currently acquired hardware status information may include the following steps:
[0087] S1021: Obtain the time series features corresponding to each target feature;
[0088] Among them, the time series feature corresponding to each target feature is used to characterize the feature distribution of the target feature within the target time period;
[0089] S1022: For each target feature, based on the time series features corresponding to the target feature, analyze whether the feature value of the target feature in the currently acquired hardware status information is an outlier, and obtain the analysis result corresponding to the target feature;
[0090] S1023: Based on the analysis results corresponding to each target feature, determine whether there are any feature values that are outliers in the currently collected hardware status information.
[0091] In this implementation, in response to the collected hardware status information, for each target feature in the hardware status information, the feature distribution representing the target feature within the target time period is obtained as the time series feature of the target feature. Then, based on the target feature, it is analyzed whether the corresponding feature value of the target feature in the current predetermined collection period is an outlier relative to the historical feature value in the time series feature of the target feature, so as to obtain the analysis result corresponding to the target feature.
[0092] In this implementation, considering that the hardware state of the disk changes over time, the feature values of each target feature used to characterize the hardware state also have certain dynamic patterns observed at consecutive time points. Therefore, the dynamic patterns extracted from the feature values observed at consecutive time points based on the target feature can be used as the time series features corresponding to the target feature, so as to use the obtained time series features to characterize the feature distribution of the target feature within the target time period.
[0093] In this way, based on the time series features corresponding to each target feature, it is possible to analyze whether the feature value of the target feature in the currently acquired hardware status information is an outlier, thus obtaining the analysis result corresponding to the target feature. Based on the analysis results corresponding to each target feature, it can be determined whether there are any outlier feature values in the currently acquired hardware status information. That is, if any analysis result corresponding to each target feature indicates that the corresponding feature value is an outlier within the current predetermined acquisition period, it can be determined that there are outlier feature values in the currently acquired hardware status information.
[0094] For example, for each target feature, the feature values of the target feature collected at consecutive time points within the target time period are input into an LSTM (Long Short-Term Memory) model to extract time-series features of the target feature. Then, the extracted time-series features are input into a 1D-CNN (one-dimensional convolutional neural network) with a convolution kernel size of 3. The convolution kernels in the 1D-CNN capture local feature vector values at three consecutive time steps. By comparing the feature values (also called feature vector values) of the target feature collected within the current predetermined collection period with the baseline reference feature values (also called feature vector values) determined from the anomaly-free historical data of the target feature, it is determined whether the collected feature values of the target feature are outliers.
[0095] The feature value of the target feature captured within the current predetermined acquisition period can be understood as the value under the "local mode to be detected", and the reference feature value determined based on the target feature can be understood as the value under the "normal local mode". Thus, by comparing the values of the same target feature in the two modes, it is determined whether the acquired feature value and the feature value of the target feature belong to anomalies.
[0096] In this implementation, for each target feature, the time series features corresponding to that target feature are used to identify outliers. This allows for full utilization of the time dimension information of the data to capture the dynamic patterns and dependencies of the data over time, thereby improving the reliability of the identified outliers.
[0097] S103: If it exists, obtain the current weight corresponding to each target feature.
[0098] The current weight of each target feature is the weight set for that target feature within the current predetermined period; the period length of the predetermined period is longer than the period length of the predetermined collection period.
[0099] Considering that the hardware condition of a disk tends to decline over time—for example, the longer a disk is used, the higher its wear and tear, and the worse its hardware condition—the wear and tear of the disk becomes. However, the wear and tear of the disk does not change significantly within a certain time frame; that is, the change in the characteristic value used to characterize the wear and tear is not significant. Consequently, within this certain time frame, the impact of the wear and tear on the overall health of the disk is relatively fixed. Therefore, in this application, for each target feature, a weight is determined according to a predetermined period, and this weight is used as the current weight of the target feature within the current predetermined period. In other words, within a predetermined period, the impact of this target feature on the overall health of the disk is relatively fixed compared to the impact of other target features on the overall health of the disk; that is, within this predetermined period, the current weight corresponding to this target feature remains unchanged.
[0100] For example, within a predetermined period, the current weight of target feature 1 is 0.5, the current weight of target feature 2 is 0.2, and the current weight of target feature 3 is 0.3. This can be understood as the degree of influence of the above three target features on the health status of the disk within the predetermined period being: target feature 1 > target feature 3 > target feature 2. In other words, within the predetermined period, the degree of influence of target features 1, 2, and 3 on the health status of the disk is relatively fixed. The process of determining the current weight of each target feature can be determined using an Attention model.
[0101] In this way, if there are outlier feature values in the currently collected hardware status information, the current weight corresponding to each target feature can be directly obtained.
[0102] It should be noted that when the current predetermined period is the first predetermined period, the current weight of each target feature is its corresponding initial weight. When the current predetermined period is any other predetermined period besides the first predetermined period, the current weight of each target feature is the weight set for that target feature in the current predetermined period as determined by steps S301-S303 below.
[0103] Furthermore, the duration of the predetermined period used to determine the current weight of the target feature is longer than the duration of the predetermined collection period. For example, the duration of the predetermined period can be one week, one month, etc.
[0104] S104: Based on the feature value of each target feature in the currently acquired hardware status information and the obtained current weight, calculate the health score used to characterize the health status of the target disk; wherein, the health score is used to characterize the probability that the target disk will have a risk of failure in the future time period;
[0105] S105: When the obtained health score meets the predetermined triggering conditions, trigger the data migration process for the target disk.
[0106] The predetermined triggering conditions include a health score that is lower than a predetermined score threshold.
[0107] In this application, a health score is calculated based on the feature value of each target feature in the currently collected hardware status information and the obtained current weight. This health score quantifies the health status of the target disk, meaning it represents the probability that the target disk will have a failure risk in the future. Furthermore, considering that the purpose of this application is to predict the existence of a failure risk by detecting the disk's health status, a predetermined score threshold is set. This threshold represents the health score with the highest probability of a failure risk within a certain period. Specifically, when the obtained health score is lower than the predetermined score threshold, it indicates a higher probability that the target disk will have a failure risk within a certain period. In this case, a health score below the predetermined threshold can be set as a predetermined trigger condition. Therefore, when the obtained health score meets the predetermined trigger condition, a data migration process for the target disk is triggered to migrate the data on the target disk in advance.
[0108] Optionally, in one implementation, the predetermined triggering conditions further include that the health scores obtained in multiple consecutive predetermined collection cycles are all lower than a predetermined score threshold, and the variance of the health scores of the target disk in multiple predetermined collection cycles is less than a predetermined variance.
[0109] Considering the possibility of errors in the detected hardware status information, or the fact that the current weight of the target feature may not be suitable for the actual degradation of the target disk when entering the next predetermined cycle, the obtained health score may not effectively reflect the health status of the target disk. Therefore, in this implementation, in addition to the predetermined triggering condition that the health score is lower than a predetermined score threshold, it also includes that the health score obtained in multiple consecutive predetermined collection cycles is lower than the predetermined score threshold, and the variance of the target disk's health score in multiple predetermined collection cycles is less than a predetermined variance. This improves the standardization of the triggering conditions for the data migration process of the target disk, reduces the frequency of data migration for the target disk, and minimizes the waste of computing resources.
[0110] Optionally, in one implementation, step S104 above, which calculates a health score to characterize the health status of the target disk based on the feature value of each target feature in the currently acquired hardware status information and the obtained current weight, may include the following steps:
[0111] Step A1: For each target feature, obtain the historical mean and historical standard deviation of the target disk for that target feature, and calculate the target index value corresponding to the target feature based on the feature value, historical mean and historical standard deviation of the target feature in the currently collected hardware status information.
[0112] Wherein, the historical standard deviation of the target feature is the standard deviation corresponding to each historical feature value of the target feature, and the target index value corresponding to the target feature is used to characterize: the degree of impact of the target feature on disk damage when the target feature has the feature value in the currently collected hardware status information;
[0113] Step A2: Based on the target index value corresponding to each target feature and the current weight corresponding to that target feature, calculate the health score used to characterize the health status of the target disk.
[0114] In this implementation, for each target feature, the mean of all historical feature values of that target feature on the target disk is obtained as the historical mean of that target feature, representing the baseline average value of that target feature under normal conditions. The standard deviation of each historical feature value of that target feature is also obtained as the historical standard deviation, measuring the dispersion of the feature values and quantifying the deviation range of that target feature. Then, based on the feature values, historical mean, and historical standard deviation of that target feature in the currently collected hardware status information, a target index value is calculated to characterize the impact of that target feature on disk damage when the target feature has the feature values in the currently collected hardware status information.
[0115] It should be noted that the target indicator value can be understood as the standardized deviation term corresponding to the target feature when the target feature has the feature value in the currently collected hardware status information. For example, the feature value is normalized to the z-score (standard score), and a positive value indicates that the feature value is better than the mean, that is, when the target feature has the feature value in the currently collected hardware status information, the impact on disk damage is small, or even no damage to the disk; a negative value indicates that it is lower than the mean, that is, when the target feature has the feature value in the currently collected hardware status information, the impact on disk damage is large, and the target disk has a risk of deterioration (damage).
[0116] Then, based on the target index value corresponding to each feature and the current weight corresponding to that target feature, a health score is calculated to characterize the health status of the target disk.
[0117] Optionally, in one implementation, step A2 above, which calculates a health score to characterize the health status of the target disk based on the target index value corresponding to each target feature and the current weight corresponding to that target feature, may include the following steps:
[0118] Using the first formula, a health score is calculated based on the target indicator value corresponding to each target feature and the current weight corresponding to that target feature; wherein the first formula includes:
[0119] ;
[0120] Score represents the health score of the target disk, and i represents the i-th target feature of the target disk. F represents the current weight corresponding to the i-th target feature. i The target index value that represents the i-th target feature.
[0121] In this implementation, the target index value corresponding to each target feature, along with the current weight corresponding to that target feature, is substituted into the first formula mentioned above to calculate a health score used to characterize the health status of the target disk. Furthermore, the first formula as a whole, through the design of weighted standardized deviation (target index value) and subtraction from 100, ensures that high-bias features (e.g., outliers or abnormal patterns) have a greater negative impact on the health score, thus accurately reflecting the health status of the target disk.
[0122] Optionally, in one embodiment, the calculation method for the target index value corresponding to each target feature may include the following steps: calculating the target index value corresponding to each target feature according to a second formula; wherein, the second formula includes:
[0123] ;
[0124] F i The target index value that represents the i-th target feature. The feature value representing the i-th target feature in the currently acquired hardware status information. The historical mean representing the feature of the i-th target. The historical standard deviation characterizing the i-th target feature.
[0125] Among them, the above That is The standard deviation criterion (also known as the Raida criterion) is a statistical method based on the normal distribution used to identify and remove gross errors or outliers in data. Its core is to determine the reasonable range of data by calculating the standard deviation and regard data that exceeds this range as outliers.
[0126] As can be seen from the above, in the disk management method provided in this application embodiment, the current weight of each target feature is a weight set for that target feature within the current predetermined period. That is, the current weight of the target feature is different depending on the current predetermined period. Thus, based on the feature value of each target feature in the currently collected hardware status information and the obtained current weight, a health score can be calculated to characterize the health status of the target disk within the current predetermined period. This health score is used to characterize the probability of the target disk having a risk of failure in the future time period. Therefore, when the obtained health score meets the predetermined triggering conditions, a data migration process for the target disk is triggered. As can be seen, this solution periodically collects the hardware status information of the disk, and when anomalies are detected in the collected hardware status information, it combines the weight of each target feature in the hardware status information within the current predetermined period with the feature value of each target feature collected to calculate the disk health score in a timely and accurate manner. Thus, when the calculated health score is lower than a predetermined score threshold, proactive data migration is performed on disks that may be faulty in advance. This reduces the possibility of insufficient data security caused by data migration after actual disk failure, and improves the data reliability of the data stored on the disk.
[0127] Optionally, in one embodiment, each target feature corresponds to an initial weight; such as Figure 3 As shown, the method for determining the weight of each target feature within a predetermined period may include the following steps:
[0128] S301: When entering the predetermined cycle, obtain the feature value sequence of the target feature of the target disk;
[0129] The feature value sequence of the target feature includes the feature values of the target feature that have been collected;
[0130] S302: Based on the feature value sequence of the target feature, estimate the degree of impact of the target feature on the health status of the disk in the current predetermined period;
[0131] S303: Based on the estimated degree of impact, adjust the current weight of the target feature according to the predetermined adjustment method, and use the adjusted weight as the current weight set for the target feature within the predetermined period.
[0132] The predetermined adjustment method includes: if the estimated impact level in the predetermined period does not match the estimated impact level in the previous predetermined period, the weight is increased in response to the increase in the estimated impact level in the previous predetermined period; otherwise, the weight is decreased.
[0133] In this embodiment, each target feature has an initial weight, which can be set based on prior experience or based on the default weight ratio, both of which are reasonable.
[0134] When determining the current weight corresponding to each target feature, it can be determined whether a new predetermined period has been entered. If a new predetermined period has not been entered, the weight set for the target feature within the current predetermined period can be directly determined as the current weight of the target feature.
[0135] If a new predetermined period is to be entered, the weights set for the target feature within the new predetermined period need to be redefined and used as the current weights for that target feature. Specifically:
[0136] Upon entering the predetermined cycle, a feature value sequence of the target feature value of the target disk is acquired. This sequence includes the previously collected feature values of the target feature. Based on this feature value sequence, the correlation between the trend of feature value changes in the target feature during the previous predetermined cycle and the disk's health status at the end of the previous cycle is analyzed. Therefore, based on the analyzed correlation, the degree of influence of the target feature's feature value on the disk's health status is determined, thus predicting the degree of influence of the target feature on the disk's health status in the currently entered predetermined cycle. In other words, based on the degree of influence of the target feature on the disk's health status in the previous predetermined cycle, the degree of influence of the target feature on the disk's health status in the currently entered predetermined cycle is predicted. For example, if the disk's wear level was high in the previous predetermined cycle, it can be predicted that the wear level will have a greater impact on the disk's health status in the currently entered predetermined cycle.
[0137] Therefore, based on the estimated impact level, the current weight of the target feature is adjusted according to a predetermined adjustment method. The adjusted weight is then used as the current weight for that target feature within the currently entered predetermined period. For example, regarding the target feature of wear and tear, if the estimated impact of wear and tear on the disk's health status within the currently entered predetermined period does not match the estimated impact in the previous predetermined period, the current weight of wear and tear in the currently entered predetermined period is increased in response to an increase in the estimated impact relative to the previous predetermined period; conversely, the current weight of wear and tear in the currently entered predetermined period is decreased in response to a decrease in the estimated impact relative to the previous predetermined period. Of course, it is understandable that when the estimated impact of wear and tear on the disk's health status within the currently entered predetermined period matches the estimated impact in the previous predetermined period, there is no need to adjust the current weight of wear and tear.
[0138] The statement that the estimated impact level in the current predetermined period matches the estimated impact level in the previous predetermined period can be understood as the difference between the estimated impact levels in the two predetermined periods being within a predetermined allowable range. Conversely, if the difference between the two is not within the predetermined allowable range, it indicates that the estimated impact level in the current predetermined period does not match the estimated impact level in the previous predetermined period. The above matching can also be understood as the estimated impact level in the current predetermined period being exactly the same as the estimated impact level in the previous predetermined period, which is reasonable.
[0139] Optionally, in one implementation, the weight is increased by a predetermined step size in response to an increase in the estimated impact relative to the previous predetermined period, and the weight is decreased by a predetermined step size in response to a decrease in the estimated impact relative to the previous predetermined period.
[0140] Optionally, for each target feature, historical data of that target feature is acquired to establish the correlation between the target feature and the health status of the disk under different time-series modes. This allows for the determination of a first baseline correlation between the trend of the target feature's value change and the health status of the disk under normal conditions, and a second and third baseline correlation between the trend of the target feature's value change and the health status of the disk under abnormal conditions.
[0141] Thus, after analyzing the correlation between the trend of the target feature reflected by the feature value sequence in the previous predetermined period and the health status of the disk after the end of the previous predetermined period, the similarity between the analyzed correlation and the first, second, and third benchmark correlations is compared.
[0142] If the similarity with the first benchmark mentioned above is high, it can be considered that the impact of the target feature on the health status of the disk in the new predetermined period is similar to the impact on the health status of the disk in the previous predetermined period, and the weight of the target feature in the new predetermined period can continue to use the weight corresponding to the previous predetermined period.
[0143] If the similarity with the second benchmark mentioned above is high, it can be considered that the impact of the target feature on the health status of the disk in the new predetermined period is higher than that in the previous predetermined period. The weight of the target feature in the new predetermined period needs to be adjusted. At this time, the current weight of the target feature can be increased according to the predetermined step size, and the adjusted weight can be used as the current weight set for the target feature in the current predetermined period.
[0144] If the similarity with the third benchmark mentioned above is high, it can be considered that the impact of the target feature on the health status of the disk in the new predetermined period is lower than that in the previous predetermined period. The weight of the target feature in the new predetermined period needs to be adjusted. At this time, the current weight of the target feature can be reduced according to the predetermined step size, and the adjusted weight can be used as the current weight set for the target feature in the current predetermined period.
[0145] It should be noted that for multiple target features of a target disk, the degree of influence of each target feature on the disk's health status varies within different predetermined periods. Correspondingly, the current weight of that target feature also varies within a predetermined period. Therefore, to reflect the degree of influence of each target feature on the health status within a predetermined period, a predetermined value can be optionally used as the sum of the current weights of the multiple target features of the disk, for example, a value of 1 or 100. Thus, based on the current weight of each target feature within a predetermined period, the degree of influence of that target feature on the disk's health status within the current predetermined period can be clearly determined. For example, if the predetermined value is 1, then within a predetermined period, if the current weight of target feature 1 is 0.5, the current weight of target feature 2 is 0.2, and the current weight of target feature 3 is 0.3, then the degree of influence of the above three target features on the disk's health status within the current predetermined period can be determined as: Target Feature 1 > Target Feature 3 > Target Feature 2.
[0146] In this embodiment, the dynamic determination of the current weight is used to reflect the main influence of different target features on the health status of the disk over different time periods as usage time progresses, thereby improving the accuracy of the determined health score.
[0147] Optionally, in one embodiment, the data migration process for the target disk may include the following steps:
[0148] Step B: Select the disk to which the data stored in the target disk is to be migrated according to the predetermined disk selection criteria, obtain the disk to be used, and migrate the data stored in the target disk to the disk to be used according to the predetermined concurrency.
[0149] The predetermined disk selection criteria include: the most recent health score is higher than the target score, where the target score represents the health score with the highest probability of no disk failure risk within a specified time period; and the predetermined concurrency level represents the amount of data that can be migrated in parallel for each data migration process.
[0150] Considering the computational resource consumption of data migration and its impact on other business processes, this embodiment sets predetermined disk selection criteria. Specifically, the disks to be migrated from the target disk are those with a more recent health score than the target score. In other words, the disks to be utilized must meet the condition of having the highest probability of no disk failure risk within a specified time period. Thus, after selecting the disks to be migrated from the target disk according to the predetermined disk selection criteria, the data stored on the target disk can be migrated to the disks to be utilized according to the number of parallel migrations supported by each data migration process, as indicated by the predetermined concurrency level.
[0151] In this embodiment, by setting predetermined disk selection conditions, the selected disk to be used is the disk with the highest probability of not having a disk failure risk within a specified period of time. This reduces the possibility of performing the data migration process on the disk to be used again in the short term due to the health status of the disk after the data has been migrated to it. This reduces the execution frequency of the data migration process and further reduces the computing resources required for frequent data migration and the impact on the processing of other services.
[0152] Optionally, different types of disks support different types of data storage or processing. Therefore, when selecting the disk to be used corresponding to the target disk, it is necessary to select a disk with the same data processing type supported by the target disk as the disk to be used for the target disk.
[0153] Optionally, during data migration, to achieve a balance between reliability and efficiency, distributed storage systems can force multiple copies of data stored on a single disk to be physically isolated and stored in different independent fault domains to maximize resilience against systemic failure risks. In other words, during data migration, multiple copies of data stored on the target disk are forced to be stored on different storage devices (e.g., racks, data center areas, or server nodes) to reduce data risks caused by failures during the migration process.
[0154] Optionally, in one embodiment, after migrating the data stored in the target disk to the disk to be used according to a predetermined concurrency level in step B above, the disk management method provided in this application embodiment may further include the following steps:
[0155] Step C1: Monitor the request response latency of the management device used to control the disk in real time;
[0156] Step C2: When the detected request response latency is higher than the predetermined latency, the predetermined concurrency level is reduced, and the data stored in the target disk is migrated to the disk to be used according to the reduced predetermined concurrency level.
[0157] The reduction process includes either reducing the predetermined concurrency by a predetermined amount to obtain the reduced predetermined concurrency, or calculating the reduced predetermined concurrency based on a value in the management device used to characterize the processor's operating state.
[0158] In this embodiment, considering that the computing resources required for data migration may affect the processing of other services in the data storage system, for example, during the execution of the data migration task, the management device also needs to simultaneously receive and process access requests from the user terminal. That is, the user terminal sends a request to the management device to request access to the data stored in the disk controlled by the management device. If the management device prioritizes the data migration process, it will lead to uneven distribution of computing resources and increase the latency of request response to other services.
[0159] Therefore, a predetermined latency is set for the management device used to manage the disk. The request-response latency of the management device is monitored in real time. When the monitored request-response latency is higher than the predetermined latency, indicating that the execution of the data migration process has a significant impact on the processing of other services, the predetermined concurrency is reduced to lower the execution priority of the data migration process. Then, according to the reduced predetermined concurrency, the data stored in the target disk is migrated to the disk to be used.
[0160] The downsizing process may include: reducing the predetermined concurrency by a predetermined amount to obtain the reduced predetermined concurrency, thereby achieving rapid adjustment of the concurrency.
[0161] The aforementioned downsizing process may further include: obtaining a value from the management device used to characterize the processor's operating state, and then calculating a predetermined downsizing concurrency based on that value, so that the calculated predetermined downsizing concurrency is the maximum amount of data that can be migrated in parallel for each data migration process without affecting the processing of other services.
[0162] In this embodiment, by monitoring the request response latency of the management device in real time, the amount of data that can be migrated in parallel in each data migration process can be flexibly adjusted, so as to improve the execution efficiency of the data migration process while ensuring the processing of other services.
[0163] Optionally, in one embodiment, after the data migration process performed on the target disk is completed, the data stored in the target disk and the data stored in the disk to be used are compared to ensure the consistency of the data stored on the disk, thereby improving the security and reliability of the data stored on the disk.
[0164] For example, the Reed-Solomon matrix is used to verify the data stored in the target disk and the disk to be used.
[0165] Optionally, in one implementation, when there are multiple target disks to be analyzed, step S105 above may further include the following steps:
[0166] Step D: When the health scores obtained for at least two target disks meet the predetermined triggering conditions, the data migration process for each of the at least two target disks is triggered sequentially according to the priority of the at least two target disks;
[0167] Among the at least two target disks, the target disk with higher priority has a lower health score compared to the target disk with lower priority.
[0168] In this implementation, when there are multiple target disks to be analyzed, considering that the lower the health score, the higher the probability of the corresponding target disk having a risk of failure in the future, and the stronger the urgency of performing data migration on such target disks, the priority of executing the data migration process for each target disk can be determined according to their respective health scores when the health scores of at least two target disks meet the predetermined triggering conditions. Among the at least two target disks, the target disk with higher priority has a lower health score than the target disk with lower priority. Therefore, according to the priority of the at least two target disks, the data migration process for each of the at least two target disks is triggered sequentially.
[0169] There are several ways to set the priority of the target disks. For example, at least two target disks can be sorted (from low to high) according to their health scores, and priorities can be assigned based on their sorting position. The number of target disks belonging to the same priority can be a specified number, such as 1, 5, or 20. Another example is to pre-define a corresponding score range for each priority (the lower the score within the range, the higher the priority), and then configure the corresponding priority for each of the at least two target disks according to their respective health score ranges. This application does not specifically limit the method of priority setting, but regardless of the method, the health score of a target disk with higher priority will always be lower than that of a target disk with lower priority.
[0170] In this implementation, by setting priorities, the data migration process is performed first on target disks with low health scores. This reduces the possibility that target disks with more urgent data migration may be damaged while waiting for migration. It further reduces the possibility of insufficient data security when data migration is performed after actual disk damage, and improves the data reliability of the data stored on the disk.
[0171] Optionally, in one embodiment, the disk management method provided in this application may further include the following steps:
[0172] Step E: After the data migration process of the target disk with the same priority is completed, a cooling verification process is performed on the data migrated from the target disk with the same priority.
[0173] In this embodiment, the data migration process is performed on the target disk according to priority, and after the data migration process of the target disk of the same priority is completed, the data migrated on the target disk of the same priority is subjected to cooling verification processing.
[0174] For example, after each data migration process is completed for a target disk of the same priority, a full verification of the EC (Engineering Change) group is performed using the Reed-Solomon matrix verification method to ensure the security and reliability of the data involved in the migration. After verification, a forced cooling process is performed to forcibly cool down the original disk from which the data is being migrated, i.e., to forcibly control the original disk from performing any data processing or data storage operations. Then, after a certain period of time following the start of the forced cooling, the data migration process is executed again to ensure the integrity of the data involved in each migration.
[0175] To facilitate understanding, a specific example will be used below to describe in detail a disk management method provided in the embodiments of this application.
[0176] As the production scale of videos gradually expands, more and more video data is stored. For videos with low access frequency, such video data is usually placed in cold storage / low-frequency storage. However, this does not mean that the data will not be accessed. At certain times, some data may be accessed in a concentrated manner, such as when videos are re-produced. In this business scenario, since the purchased machines (disks used to store video data) are purchased in batches and in a concentrated manner, the lifespan of disks in the same batch is basically the same, and the probability of concentrated disk failure is high. This will lead to the systemic risk of disk retirement in large-scale distributed storage systems. For example, a sudden full migration will lead to an increase in disk load, increased latency, and decreased performance of the distributed storage system in the distributed storage cluster. In addition, the lack of lifespan prediction will lead to improper selection of the target disk for migration (i.e., the disk to be used in this application) or failure to provide early warning, thereby affecting the security and availability of stored data. In other words, there are exemplary problems in the prior art, such as: (1) services caused by migration; (2) reliability assurance of the target disk; (3) resolution of resource conflicts; and (4) optimization of data distribution.
[0177] To address the aforementioned issues, this specific embodiment innovatively proposes a five-dimensional management and control system, which includes: a health monitoring platform, a predictive decision engine, an intelligent scheduler, a migration executor, and a bandwidth allocation gateway.
[0178] To address the above problem (1), this specific embodiment uses dynamic QoS (Quality of Service) control to keep latency fluctuations within a small range; to address the above problem (2), this specific embodiment uses a hybrid model of LSTM (Long Short-Term Memory) and Attention to improve the accuracy of judging the health score of the disk; to address the above problem (3), this specific embodiment uses an adaptive concurrency control algorithm to control the response latency of other services to be no less than a predetermined latency; to address the above problem (4), the data stored in the same disk is stored in multiple disks located in different regions to achieve data migration and data storage across physical fault domains.
[0179] For example, Figure 4 This is a schematic diagram of the architecture of a specific embodiment provided in this application. The application scenario of this specific embodiment is a distributed storage cluster, which includes multiple disks.
[0180] A health monitoring platform is used to collect SMART features (also known as SMART monitoring data) from each disk in the entire distributed storage cluster. Then, the collected SMART features are delivered as monitoring data (i.e., hardware status information in this embodiment) to the prediction decision engine via a monitoring service. The prediction decision engine (also known as the intelligent prediction engine) collects the obtained monitoring data and uses it as input data for a Long Short-Term Memory (LSTM) network model. This allows the LTM network model to be trained and predict based on the received monitoring data to calculate the health score of each disk. The calculated health score is then input to the intelligent scheduler. Specifically, the health monitoring platform executes step S101 in this embodiment; the prediction decision engine executes steps S102-S105 in this embodiment.
[0181] It should be noted that the aforementioned intelligent prediction engine can be an engine containing a three-layer hybrid neural network structure, specifically:
[0182] The input layer processes the collected monitoring data. Then, the LSTM model unit is used to process the data to obtain the time series features corresponding to the target features represented by each monitoring data. The time series features corresponding to each target feature are then input into the 1D-CNN to extract the feature values belonging to the local anomaly pattern. If the time series features of the target feature can extract the feature values belonging to the local anomaly pattern, then each collected target feature is input into the decision layer.
[0183] The decision layer is used to pre-calculate the feature weight corresponding to each target feature in the current predetermined period using an attention mechanism (Attention model) (i.e., the current weight of each target feature in this application); after obtaining each target feature output by the input layer, it calculates the health score of each disk using a predetermined health scoring formula (i.e., the first formula and the second formula in this application).
[0184] Optionally, the current weight corresponding to each target feature can be dynamically calculated using the BP (Error Back Propagation Algorithm).
[0185] It should be noted that the monitoring data mentioned above (i.e., the hardware status information in this application embodiment) may include 26-dimensional SMART features. For example, the SMART features may include: underlying data read error rate, reallocation sector count, startup cycle count, bad block growth rate, write operation failure rate, erase failure count, temperature sensor reading, interface communication error rate, remaining reserved block ratio, average erase time, write throughput attenuation rate, command timeout event count, total physical write volume, NAND wear, PE cycle jitter rate, cross-plane interference index, etc. This specific embodiment does not list them all. Any feature used to characterize the health status of the disk can be used as monitoring data in this specific embodiment.
[0186] Then, the intelligent scheduler decides on data distribution and reorganization based on the calculated health score, and plans migration tasks (or migration routes). Specifically, based on the calculated health score, it determines whether there are disks among multiple disks that meet predetermined triggering conditions. If so, it selects the corresponding disk to be utilized for the disk that meets the predetermined triggering conditions, and dynamically plans the migration path for that disk. The planned migration path must meet the following requirements: the health score of the disk to be utilized must be higher than the target score; there must be multiple disks to be utilized located in different storage areas; and the amount of data migrated concurrently should be lower than the predetermined number of data blocks (i.e., data migration is performed according to a predetermined concurrency level). The intelligent scheduler is used to execute the step in this embodiment of the application of selecting the disk to which the data stored in the target disk is to be migrated according to the predetermined disk selection conditions, thus obtaining the disk to be utilized.
[0187] Subsequently, by invoking the migration executor, the data stored on the disk that meets the predetermined triggering conditions is migrated to the disk to be utilized. The migration executor is used to perform the step in this embodiment of the application of migrating the data stored on the target disk to the disk to be utilized according to a predetermined concurrency level.
[0188] It should be noted that during the data migration process, the bandwidth allocation gateway monitors the overall request-response latency of the entire distributed storage cluster in real time. When the request-response latency exceeds a predetermined latency, the concurrency of the data migration is adjusted to reduce the bandwidth consumed during the migration process, as well as the execution priority of the data migration. This achieves dynamic QoS control, keeping latency fluctuations within a small range. The aforementioned bandwidth allocation gateway is used to execute steps C1-C2 in the embodiments of this application.
[0189] In combination with the above Figure 4 The architecture shown in this application provides a flowchart of a specific embodiment, as illustrated below. Figure 5 As shown, it includes the following stages:
[0190] Phase 1, 501: Health Monitoring Phase. Specifically: SMART features are collected hourly via the NVMe-MI protocol. The collected SMART features are then aggregated, and the aggregated data is input into an LSTM for training to predict the disk's health score. If the obtained health score is below 30 for three consecutive times, and the variance of these three consecutive health scores is below 5, the data migration process for the disk is triggered. The above-mentioned health monitoring phase is steps S101-S105 in the embodiments of this application.
[0191] Phase Two 502: Task Planning Phase. Specifically: Before executing the data migration process for the disks, the CRUSH algorithm (Controlled Replication Under Scalable Hashing) is used to recalculate the data distribution of the data to be migrated stored on the disks, as well as the location of the disks to which the selected data to be migrated should be placed, determining the migration path between the two disks, and providing the execution basis for the subsequent Phase Three. The above-mentioned task planning phase is step B in the embodiments of this application.
[0192] Phase 3, 503: Migration Execution Phase. Specifically, during the data migration process, it is necessary to monitor in real time the changes in parameters such as processor, memory, and bandwidth of the management device of the distributed storage cluster to calculate the request response latency of other services (such as production services) within the entire distributed storage cluster besides the data migration process. Therefore, when the calculated request response latency exceeds a predetermined latency (30 milliseconds), the execution priority of the data migration service is automatically reduced by dynamically adjusting the concurrency of the data migration service. The above-mentioned migration execution phase is steps C1-C2 in the embodiments of this application.
[0193] The concurrency of the data migration service is calculated using the following formula:
[0194] ;
[0195] Here, Concurrency refers to the degree of concurrency, and CPU represents the current value of the processor.
[0196] Phase 4 (504): Cooling-off Verification Phase. Specifically: After each data migration process is completed, a full EC group verification is performed using the Reed-Solomon matrix verification method to ensure the security and reliability of the data involved in the migration. After verification, the original disk from which the data was migrated is subjected to forced cooling, i.e., the original disk is forcibly controlled to refrain from any data processing or data storage operations. Then, after a certain period of forced cooling, the data migration process is executed again to ensure the integrity of the data involved in the migration. The aforementioned cooling-off verification phase is step E in the embodiments of this application.
[0197] In this specific embodiment, dynamic concurrency control is used to control and reduce the delay impact on production operations during the migration process. Furthermore, by monitoring the health status of the disks, data in disks with poor health can be migrated in advance to reduce the probability of concentrated disk failures and meet the SLA (Service Level Agreement, a formal agreement between service providers and users that clarifies service standards, quality, and responsibilities) of the distributed storage cluster.
[0198] Furthermore, by monitoring the health status of disks and enabling proactive data migration, the impact of disk failures on stored data can be reduced, thereby improving data availability and security, reducing disk maintenance costs in distributed storage clusters, and increasing operational efficiency.
[0199] Based on the above method embodiments, this application provides a disk management device, such as... Figure 6 As shown, the device includes:
[0200] The information acquisition module 610 is used to acquire hardware status information of the target disk to be analyzed according to a predetermined acquisition cycle; wherein, the hardware status information includes feature values of multiple target features of the target disk.
[0201] The feature value recognition module 620 is used to identify, in response to the acquired hardware status information, whether there are any feature values that are out of the ordinary value in the currently acquired hardware status information.
[0202] The weight acquisition module 630 is used to acquire the current weight corresponding to each target feature when it exists; wherein, the current weight of each target feature is the weight set for the target feature within the current predetermined period; the period length of the predetermined period is greater than the period length of the predetermined collection period;
[0203] The score calculation module 640 is used to calculate a health score that characterizes the health status of the target disk based on the feature value of each target feature in the currently collected hardware status information and the obtained current weight; wherein, the health score is used to characterize the probability that the target disk has a risk of failure in the future time period;
[0204] The data migration module 650 is used to trigger a data migration process for the target disk when the obtained health score meets a predetermined trigger condition; wherein, the predetermined trigger condition includes the health score being lower than a predetermined score threshold.
[0205] As can be seen from the above, in the disk management method provided in this application embodiment, the current weight of each target feature is a weight set for that target feature within the current predetermined period. That is, the current weight of the target feature is different depending on the current predetermined period. Thus, based on the feature value of each target feature in the currently collected hardware status information and the obtained current weight, a health score can be calculated to characterize the health status of the target disk within the current predetermined period. This health score is used to characterize the probability of the target disk having a risk of failure in the future time period. Therefore, when the obtained health score meets the predetermined triggering conditions, a data migration process for the target disk is triggered. As can be seen, this solution periodically collects the hardware status information of the disk, and when anomalies are detected in the collected hardware status information, it combines the weight of each target feature in the hardware status information within the current predetermined period with the feature value of each target feature collected to calculate the disk health score in a timely and accurate manner. Thus, when the calculated health score is lower than a predetermined score threshold, proactive data migration is performed on disks that may be faulty in advance. This reduces the possibility of insufficient data security caused by data migration after actual disk failure, and improves the data reliability of the data stored on the disk.
[0206] Optionally, in one implementation, the feature value recognition module 620 is specifically used for:
[0207] Obtain the time series features corresponding to each target feature; wherein, the time series features corresponding to each target feature are used to characterize the feature distribution of the target feature within the target time period;
[0208] For each target feature, based on the time series features corresponding to that target feature, analyze whether the feature value of that target feature in the currently acquired hardware status information is an outlier, and obtain the analysis result corresponding to that target feature;
[0209] Based on the analysis results corresponding to each target feature, determine whether there are any feature values that are outliers in the currently collected hardware status information.
[0210] Optionally, in one implementation, the score calculation module 640 includes:
[0211] The indicator value calculation submodule is used to obtain the historical mean and historical standard deviation of the target feature for each target feature, and calculate the target indicator value corresponding to the target feature based on the feature value of the target feature in the currently collected hardware status information, the historical mean, and the historical standard deviation; wherein, the historical standard deviation of the target feature is the standard deviation corresponding to each historical feature value of the target feature, and the target indicator value corresponding to the target feature is used to characterize: the degree of impact of the target feature on disk damage when the target feature has the feature value in the currently collected hardware status information;
[0212] The score calculation submodule is used to calculate a health score that characterizes the health status of the target disk based on the target index value corresponding to each target feature and the current weight corresponding to that target feature.
[0213] Optionally, in one implementation, the fraction calculation submodule is specifically used for:
[0214] Using the first formula, a health score representing the health status of the target disk is calculated based on the target index value corresponding to each target feature and the current weight corresponding to that target feature; wherein, the first formula includes:
[0215] ;
[0216] Wherein, Score represents the health score of the target disk, and i represents the i-th target feature of the target disk. F represents the current weight corresponding to the i-th target feature. i The target index value that represents the i-th target feature.
[0217] Optionally, in one implementation, the calculation method for the target index value corresponding to each target feature includes:
[0218] Calculate the target index value corresponding to each target feature according to the second formula;
[0219] The second formula includes:
[0220] ;
[0221] F i The target index value that represents the i-th target feature. The feature value representing the i-th target feature in the currently acquired hardware status information. The historical mean representing the feature of the i-th target. The historical standard deviation characterizing the i-th target feature.
[0222] Optionally, in one implementation, each target feature corresponds to an initial weight; the method for determining the weight of each target feature within a predetermined period includes:
[0223] Each time a predetermined period is entered, the feature value sequence of the target feature of the target disk is acquired; wherein, the feature value sequence of the target feature includes the feature values of the target feature that have been collected;
[0224] Based on the feature value sequence of the target feature, the degree of impact of the target feature on the health status of the disk in the current predetermined period is estimated.
[0225] Based on the estimated degree of impact, the current weight of the target feature is adjusted according to a predetermined adjustment method, and the adjusted weight is used as the current weight set for the target feature within the current predetermined period. The predetermined adjustment method includes: if the estimated degree of impact within the current predetermined period does not match the estimated degree of impact in the previous predetermined period, the weight is increased in response to the increase in the estimated degree of impact relative to the previous predetermined period; otherwise, the weight is decreased.
[0226] Optionally, in one implementation, the data migration process for the target disk includes:
[0227] According to the predetermined disk selection criteria, select the disk to which the data stored in the target disk is to be migrated, obtain the disk to be used, and migrate the data stored in the target disk to the disk to be used according to the predetermined concurrency.
[0228] The predetermined disk selection criteria include: the most recent health score is higher than the target score, where the target score represents the health score with the highest probability of no disk failure risk within a specified time period; and the predetermined concurrency level represents the amount of data that can be migrated in parallel for each data migration process.
[0229] Optionally, in one implementation, the apparatus further includes:
[0230] The concurrency adjustment module is used to migrate the data stored in the target disk to the disk to be used according to a predetermined concurrency level, and then monitor the request response latency of the management device used to manage the disk in real time. When the monitored request response latency is lower than the predetermined latency, the predetermined concurrency level is reduced, and the data stored in the target disk is migrated to the disk to be used according to the reduced predetermined concurrency level.
[0231] The downsizing process includes reducing the predetermined concurrency by a predetermined amount to obtain the reduced predetermined concurrency, or calculating the reduced predetermined concurrency based on the values used in the management device to characterize the processor's operating state.
[0232] Optionally, in one implementation, the number of target disks to be analyzed is multiple;
[0233] The data migration module 650 is specifically used for:
[0234] When the health scores obtained for at least two target disks meet the predetermined triggering conditions, the data migration process for each of the at least two target disks is triggered sequentially according to the priority of the at least two target disks.
[0235] Among the at least two target disks, the target disk with higher priority has a lower health score compared to the target disk with lower priority.
[0236] Optionally, in one implementation, the apparatus further includes:
[0237] The cooling verification module is used to perform cooling verification on the data migrated from the target disk of the same priority after the data migration process of the target disk of the same priority has been completed.
[0238] Optionally, in one implementation, the predetermined triggering condition further includes that the health score obtained in multiple consecutive predetermined collection cycles is lower than the predetermined score threshold, and the variance of the health score of the target disk in the multiple predetermined collection cycles is less than a predetermined variance.
[0239] This application also provides an electronic device, such as... Figure 7 As shown, it includes a processor 701, a communication interface 702, a memory 703, and a communication bus 704, wherein the processor 701, the communication interface 702, and the memory 703 communicate with each other through the communication bus 704.
[0240] Memory 703 is used to store computer programs;
[0241] When the processor 701 executes the program stored in the memory 703, it implements any of the disk management methods described in the above embodiments.
[0242] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0243] The communication interface is used for communication between the aforementioned terminal and other devices.
[0244] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0245] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0246] In another embodiment provided in this application, a computer-readable storage medium is also provided, wherein a computer program is stored therein, and when the computer program is executed by a processor, it implements any of the disk management methods described in the above embodiments.
[0247] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the disk management methods described in the above embodiments.
[0248] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state disk (SSD)).
[0249] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0250] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device embodiments, electronic device embodiments, computer-readable storage medium embodiments, and computer program product embodiments are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0251] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.
Claims
1. A disk management method, characterized in that, include: According to a predetermined collection cycle, the hardware status information of the target disk to be analyzed is collected; wherein, the hardware status information includes feature values of multiple target features of the target disk; In response to the acquired hardware status information, identify whether there are any feature values that are outliers in the currently acquired hardware status information; If present, obtain the current weight corresponding to each target feature; wherein, the current weight of each target feature is the weight set for that target feature within the current predetermined period; the period length of the predetermined period is greater than the period length of the predetermined collection period; Based on the feature value of each target feature in the currently collected hardware status information and the obtained current weight, a health score is calculated to characterize the health status of the target disk; wherein, the health score is used to characterize the probability that the target disk has the risk of failure in the future time period; When the obtained health score meets the predetermined triggering conditions, the data migration process for the target disk is triggered; wherein, the predetermined triggering conditions include the health score being lower than a predetermined score threshold.
2. The method according to claim 1, characterized in that, The process of identifying whether there are any outlier features in the currently acquired hardware status information includes: Obtain the time series features corresponding to each target feature; wherein, the time series features corresponding to each target feature are used to characterize the feature distribution of the target feature within the target time period; For each target feature, based on the time series features corresponding to that target feature, analyze whether the feature value of that target feature in the currently acquired hardware status information is an outlier, and obtain the analysis result corresponding to that target feature; Based on the analysis results corresponding to each target feature, determine whether there are any feature values that are outliers in the currently collected hardware status information.
3. The method according to claim 1, characterized in that, The calculation of a health score, representing the health status of the target disk, based on the feature value of each target feature in the currently acquired hardware status information and the obtained current weight, includes: For each target feature, the historical mean and historical standard deviation of the target feature are obtained, and the target index value corresponding to the target feature is calculated based on the feature value of the target feature in the currently collected hardware status information, the historical mean, and the historical standard deviation; wherein, the historical standard deviation of the target feature is the standard deviation corresponding to each historical feature value of the target feature, and the target index value corresponding to the target feature is used to characterize: the degree of impact of the target feature on disk damage when the target feature has the feature value in the currently collected hardware status information; Based on the target index value corresponding to each target feature and the current weight corresponding to that target feature, a health score is calculated to characterize the health status of the target disk.
4. The method according to claim 3, characterized in that, The calculation of a health score, which characterizes the health status of the target disk, based on the target indicator value corresponding to each target feature and the current weight corresponding to that target feature includes: Using the first formula, a health score representing the health status of the target disk is calculated based on the target index value corresponding to each target feature and the current weight corresponding to that target feature; wherein, the first formula includes: ; Wherein, Score represents the health score of the target disk, and i represents the i-th target feature of the target disk. F represents the current weight corresponding to the i-th target feature. i The target index value that represents the i-th target feature.
5. The method according to claim 4, characterized in that, The calculation methods for the target index value corresponding to each target feature include: Calculate the target index value corresponding to each target feature according to the second formula; The second formula includes: ; F i The target index value that represents the i-th target feature. The feature value representing the i-th target feature in the currently acquired hardware status information. The historical mean representing the feature of the i-th target. The historical standard deviation characterizing the i-th target feature.
6. The method according to any one of claims 1-5, characterized in that, Each target feature has an initial weight; The method for determining the weight of each target feature within a predetermined period includes: When the predetermined period is entered, the feature value sequence of the target feature of the target disk is obtained; wherein, the feature value sequence of the target feature includes the feature values of the target feature that have been collected; Based on the feature value sequence of the target feature, the degree of impact of the target feature on the health status of the disk in the current predetermined period is estimated. Based on the estimated degree of impact, the current weight of the target feature is adjusted according to a predetermined adjustment method, and the adjusted weight is used as the current weight set for the target feature within the current predetermined period. The predetermined adjustment method includes: if the estimated degree of impact within the current predetermined period does not match the estimated degree of impact in the previous predetermined period, the weight is increased in response to the increase in the estimated degree of impact relative to the previous predetermined period; otherwise, the weight is decreased.
7. The method according to any one of claims 1-5, characterized in that, The data migration process for the target disk includes: According to the predetermined disk selection criteria, select the disk to which the data stored in the target disk is to be migrated, obtain the disk to be used, and migrate the data stored in the target disk to the disk to be used according to the predetermined concurrency. The predetermined disk selection criteria include: the most recent health score is higher than the target score, where the target score represents the health score with the highest probability of no disk failure risk within a specified time period; and the predetermined concurrency level represents the amount of data that can be migrated in parallel for each data migration process.
8. The method according to claim 7, characterized in that, After migrating the data stored in the target disk to the disk to be used according to a predetermined concurrency level, the method further includes: Real-time monitoring of request response latency of management devices used to control disks; When the detected request response latency is higher than the predetermined latency, the predetermined concurrency level is reduced, and the data stored in the target disk is migrated to the disk to be used according to the reduced predetermined concurrency level. The downsizing process includes reducing the predetermined concurrency by a predetermined amount to obtain the reduced predetermined concurrency, or calculating the reduced predetermined concurrency based on the values used in the management device to characterize the processor's operating state.
9. The method according to any one of claims 1-5, characterized in that, The number of target disks to be analyzed is multiple; The step of triggering a data migration process for the target disk when the obtained health score meets a predetermined trigger condition includes: When the health scores obtained for at least two target disks meet the predetermined triggering conditions, the data migration process for each of the at least two target disks is triggered sequentially according to the priority of the at least two target disks. Among the at least two target disks, the target disk with higher priority has a lower health score compared to the target disk with lower priority.
10. The method according to claim 9, characterized in that, The method further includes: After the data migration process of the target disk with the same priority is completed, a cooling verification process is performed on the data migrated from the target disk with the same priority.
11. The method according to claim 1, characterized in that, The predetermined triggering conditions also include that the health scores obtained in multiple consecutive predetermined collection cycles are all lower than the predetermined score threshold, and the variance of the health scores of the target disk in the multiple predetermined collection cycles is less than a predetermined variance.
12. A disk management device, characterized in that, include: The information acquisition module is used to acquire hardware status information of the target disk to be analyzed according to a predetermined acquisition cycle; wherein, the hardware status information includes feature values of multiple target features of the target disk; The feature value recognition module is used to identify, in response to the collected hardware status information, whether there are any feature values that are out of the ordinary value in the currently collected hardware status information. The weight acquisition module is used to acquire the current weight corresponding to each target feature when it exists; wherein, the current weight of each target feature is the weight set for the target feature within the current predetermined period; the period length of the predetermined period is greater than the period length of the predetermined collection period; The score calculation module is used to calculate a health score that characterizes the health status of the target disk based on the feature value of each target feature in the currently collected hardware status information and the obtained current weight; wherein, the health score is used to characterize the probability that the target disk has the risk of failure in the future time period; The data migration module is used to trigger a data migration process for the target disk when the obtained health score meets a predetermined trigger condition; wherein, the predetermined trigger condition includes the health score being lower than a predetermined score threshold.
13. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method described in any one of claims 1-11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1-11.