A method, device, storage medium, and program product for creating a disk array
By obtaining the probability of disk failure mounted on the disk array card and grouping selection, the problem of high risk of disk array data loss under the same disk conditions is solved, and a more reliable and secure disk array creation is achieved.
Patent Information
- Application Number
- CN202510495607.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-21
AI Technical Summary
Under the same disk conditions, how to reduce the risk of data loss in disk arrays, especially in RAID configurations, how to select disk combinations to minimize the probability of data loss.
By obtaining the probability of disk failure mounted on the disk array card, grouping disks based on configuration information, selecting the combination with the lowest probability of data loss, and creating a disk array of the corresponding level.
Improves the reliability and security of disk arrays and reduces the probability of data loss.
Smart Images

Figure CN120010793B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of disk array creation, and in particular, to a method, device, storage medium, and program product for creating a disk array. Background Art
[0002] RAID (Redundant Arrays of Independent Disks) is a data storage technology that can combine multiple physical disks into a logical disk to improve data read / write speed and fault tolerance. Although disk arrays have a certain degree of fault tolerance, they only allow a certain number of disks to fail. If the number of failed disks exceeds the specified limit, data loss will occur. For example, RAID5 can allow one disk to fail without data loss.
[0003] Suppose there are 6 disks of the same size mounted under a disk array card, numbered from 0 to 5. At this time, if you want to create 2 disk arrays of RAID5 level, each requiring 3 disks. Since there is no limit on which disks are selected to form a disk array, and the probability of each disk failing is different, and the number of disks that can fail in each disk array is limited. If the number of failed disks exceeds the allowable value, data loss will occur in that disk array. Therefore, the data loss probabilities of disk arrays composed of different combinations of disks are different.
[0004] Taking the probability of disk 0 and disk 1 failing as 0.08, and the probability of the remaining disks failing as 0.02 as an example, a RAID5-level disk array can allow at most one disk to fail without data loss. In this case, the data loss probability of combinations such as [0, 1, 2] and [3, 4, 5] will be higher than that of combinations such as [0, 2, 3] and [1, 4, 5]. Under the same disk conditions, selecting combinations such as [0, 1, 2] and [3, 4, 5] to create a disk array will bring a higher data loss risk.
[0005] It can be seen that how to make the created disk array have a lower data loss risk under the same disk conditions is a problem that needs to be solved by those skilled in the art. Summary of the Invention
[0006] The purpose of the embodiments of the present invention is to provide a method, device, storage medium, and program product for creating a disk array, which can make the created disk array have a lower data loss risk under the same disk conditions. The specific solutions are as follows:
[0007] In a first aspect, the present invention provides a method for creating a disk array, including:
[0008] Obtain the configuration information of the disk array to be created, and determine the failure probability of each disk mounted under the disk array card; the configuration information includes the total number of disk arrays to be created, the level of each disk array to be created, and the number of disks required for each;
[0009] Use the configuration information and based on the failure probability of each disk, group each disk to obtain a target disk combination; the target disk combination is the combination with the minimum data loss probability among several disk combinations obtained by grouping; among them, each disk combination includes sub-disk combinations corresponding to each disk array to be created respectively, and the data loss probability of each disk combination is a value determined based on the failure probability of the disks in each disk combination;
[0010] Use the level of each disk array to be created and based on each sub-disk combination in the target disk combination, create a disk array of the corresponding level.
[0011] Optionally, determining the failure probability of each disk mounted under the disk array card includes:
[0012] Obtain the current status information of each disk mounted under the disk array card;
[0013] Based on the current status information of each disk, determine the current stage where each disk is located;
[0014] Based on the current stage where each disk is located, determine the failure probability of each disk.
[0015] Optionally, based on the current status information of each disk, determining the current stage where each disk is located includes:
[0016] Based on the write times and erase-write utilization rate of each storage particle in any one disk, determine the current stage where any one disk is located;
[0017] Among them, any one disk is any one of each disk.
[0018] Optionally, based on the write times and erase-write utilization rate of each storage particle in any one disk, determining the current stage where any one disk is located includes:
[0019] Judge whether each storage particle in any one disk has been written once;
[0020] If any one storage particle in any one disk has not been written once, determine that the current stage where any one disk is located is the early stage;
[0021] If each storage particle in any one disk has been written once, judge whether the erase-write utilization rate of each storage particle in any one disk has reached the preset utilization rate;
[0022] If the write-erase usage rate of any storage particle in any disk does not reach the preset usage rate, determine that the current stage of any disk is the stable stage;
[0023] If the write-erase usage rates of all storage particles in any disk reach the preset usage rate, determine that the current stage of any disk is the late stage.
[0024] Optionally, based on the current stage of each disk, determine the failure probability of each disk, including:
[0025] If the current stage of any disk is the early stage, determine the failure probability of any disk based on the first relative coefficient and the preset failure probability;
[0026] If the current stage of any disk is the stable stage, determine the preset failure probability as the failure probability of any disk;
[0027] If the current stage of any disk is the late stage, determine the failure probability of any disk based on the second relative coefficient and the preset failure probability;
[0028] Among them, both the first relative coefficient and the second relative coefficient are greater than 1.
[0029] Optionally, determine whether all storage particles in any disk have been written once, including:
[0030] By determining whether the total written data volume of any disk is greater than the preset data volume, determine whether all storage particles in any disk have been written once;
[0031] Among them, the preset data volume is a value determined based on the number of all storage particles in any disk and the unit written data volume of each storage particle.
[0032] Optionally, determine the failure probability of any disk based on the first relative coefficient and the preset failure probability, including:
[0033] Based on the ratio between the total written data volume of any disk and the preset data volume, determine the first relative coefficient;
[0034] According to the first relative coefficient and the preset failure probability, determine the failure probability of any disk.
[0035] Optionally, determine whether all storage particles in any disk have been written once, including:
[0036] By determining whether the power-on time of any disk is greater than the preset time, determine whether all storage particles in any disk have been written once;
[0037] Among them, the preset time is a value determined based on the preset number of days and the number of hours per day.
[0038] Optionally, determining the failure probability of any disk based on the first relative coefficient and the preset failure probability includes:
[0039] Determining the first relative coefficient based on the ratio between the power-on time of any disk and the preset time;
[0040] Determining the failure probability of any disk according to the first relative coefficient and the preset failure probability.
[0041] Optionally, determining the failure probability of any disk based on the second relative coefficient and the preset failure probability includes:
[0042] Determining the second relative coefficient based on the erasure and write usage rate of each storage particle in any disk and the preset usage rate;
[0043] Determining the failure probability of any disk according to the second relative coefficient and the preset failure probability.
[0044] Optionally, determining the second relative coefficient based on the erasure and write usage rate of each storage particle in any disk and the preset usage rate includes:
[0045] Determining the average usage rate based on the erasure and write usage rate of each storage particle in any disk to obtain the erasure and write usage rate of any disk;
[0046] Determining the second relative coefficient according to the difference between the erasure and write usage rate of any disk and the preset usage rate.
[0047] Optionally, the process of determining the data loss probability of each disk combination includes:
[0048] Determining the data loss probability of any sub-disk combination based on the failure probability of the disks in any sub-disk combination; any sub-disk combination is any one of the sub-disk combinations included in each disk combination;
[0049] Determining the data loss probability of each disk combination according to the data loss probabilities corresponding to the sub-disk combinations included in each disk combination respectively.
[0050] Optionally, determining the data loss probability of any sub-disk combination based on the failure probability of the disks in any sub-disk combination includes:
[0051] Determining the target disk array to be created corresponding to any sub-disk combination;
[0052] Determining the allowable number of failed disks based on the level corresponding to the target disk array to be created; the allowable number of failed disks indicates that data loss occurs when the number of failed disks in the target disk array to be created is greater than the allowable number of failed disks;
[0053] Determine the data loss probability of any sub-disk combination by using the allowable number of failed disks and based on the failure probability of the disks in any sub-disk combination.
[0054] Optionally, determining the data loss probability of any sub-disk combination by using the allowable number of failed disks and based on the failure probability of the disks in any sub-disk combination includes:
[0055] Based on the number of disks required for the target disk array to be created and the allowable number of failed disks, determine each target number; the target number is a positive integer not greater than the number of disks required for the target disk array to be created and greater than the allowable number of failed disks;
[0056] Based on the failure probability of the disks in any sub-disk combination, determine the probability when the target number of disks in the target disk array to be created fails;
[0057] According to the probability when the target number of disks in the target disk array to be created fails, determine the data loss probability of any sub-disk combination.
[0058] Optionally, group each disk by using the configuration information and based on the failure probability of each disk to obtain the target disk combination, including:
[0059] Based on the total number of disk arrays to be created, create several disk creation tasks; the number of several disk creation tasks is one more than the total number of disk arrays to be created;
[0060] Use the first disk creation task in several disk creation tasks and based on the number of disks required for each disk array to be created, group each disk to obtain the current disk combination; the number of the first disk creation tasks is the same as the total number of disk arrays to be created;
[0061] Determine the data loss probability of the current disk combination through the second disk creation task in several disk creation tasks, determine the new current minimum data loss probability based on the data loss probability of the current disk combination and the current minimum data loss probability, and when the current grouping situation does not meet the preset end condition, jump to the step of using the first disk creation task in several disk creation tasks and grouping each disk based on the number of disks required for each disk array to be created until the current grouping situation meets the preset end condition, so as to determine the target disk combination based on the disk combination corresponding to the latest current minimum data loss probability;
[0062] Wherein, the current minimum data loss probability is a preset loss probability not less than 1 initially.
[0063] Optionally, create a first disk creation task among several disk creation tasks and group each disk based on the number of disks required for each disk array to be created, so as to obtain the current disk combination, including:
[0064] Through each first disk creation task, perform disk grouping operations in sequence based on the number of disks required for the corresponding disk array to be created, so as to obtain the current remaining disks by using the current first disk creation task, and take out a corresponding number of disks from the current remaining disks based on the number of disks required for the current disk array to be created, so as to obtain a sub-disk combination corresponding to the current disk array to be created and update the current remaining disks; the current disk array to be created corresponds to the current first disk creation task; the current remaining disks are initially all the disks mounted under the disk array card;
[0065] When the current first disk creation task is the last disk creation task among all the first disk creation tasks, determine the current disk combination based on the sub-disk combinations corresponding to each disk array to be created obtained by each first disk creation task.
[0066] Optionally, the preset end condition includes that the current grouping cumulative times reach the maximum grouping times or the current cumulative number of types of the disk combination reaches the maximum number of types; both the maximum grouping times and the maximum number of types are values determined based on the maximum number of types, and the maximum number of types is the maximum number of types of disk combinations that can be obtained after grouping each disk.
[0067] In a second aspect, the present invention provides an electronic device, including:
[0068] A memory for storing a computer program;
[0069] A processor for executing the computer program to implement the steps of the foregoing disk array creation method.
[0070] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the foregoing disk array creation method are implemented.
[0071] In a fourth aspect, the present invention provides a computer program product, including computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the foregoing disk array creation method are implemented.
[0072] In the present invention, configuration information of a disk array to be created is obtained, and the failure probability of each disk mounted under a disk array card is determined; the configuration information includes the total number of disk arrays to be created, the level of each disk array to be created, and the number of disks required for each; using the configuration information and based on the failure probability of each disk, each disk is grouped to obtain a target disk combination; the target disk combination is the combination with the minimum data loss probability among several disk combinations obtained by grouping; wherein, each disk combination includes a sub-disk combination corresponding to each disk array to be created respectively, and the data loss probability of each disk combination is a value determined based on the failure probability of the disks in each disk combination; using the level of each disk array to be created and based on each sub-disk combination in the target disk combination, a disk array of the corresponding level is created.
[0073] Advantageous effects: By using the configuration information of the disk array to be created and based on the failure probability of each disk mounted under the disk array card, each disk is grouped, and the combination with the minimum data loss probability is determined from several disk combinations obtained by grouping to obtain a target disk combination, and then a disk array is created based on the sub-disk combinations corresponding to each disk array to be created in the target disk combination. In this way, based on the failure probability of each disk mounted under the disk array card, the target disk combination with the minimum data loss probability is searched from the grouping of each disk, so that the disk array created based on the target disk combination has the minimum data loss probability, and the reliability and security of the disk array are improved. Description of the Drawings
[0074] To more clearly illustrate the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0075] Figure 1 It is a flowchart of a method for creating a disk array provided by an embodiment of the present invention;
[0076] Figure 2 It is a schematic diagram of a failure bathtub curve provided by an embodiment of the present invention;
[0077] Figure 3 It is a flowchart for determining the failure probability of a disk provided by an embodiment of the present invention;
[0078] Figure 4 It is a flowchart for creating a disk array provided by an embodiment of the present invention;
[0079] Figure 5 It is a structural diagram of an electronic device provided by an embodiment of the present invention. Specific embodiments
[0080] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0081] The terms "including" and "having" in the specification of the present invention and the above accompanying drawings, and any variations related to "including" and "having", are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may include steps or units not listed.
[0082] To enable those skilled in the art of the present technology to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0083] When constructing a disk array based on the disks mounted under a disk array card, there is no limit to which disks are selected to form the disk array. Moreover, the probability of each disk failing is different, and the number of disks that each disk array can allow to fail is limited. If the number of failed disks exceeds the allowable value, data loss will occur in the disk array. Therefore, the data loss probabilities of disk arrays composed of different combinations of disks are different, and the disk array with a higher data loss probability is more likely to bring more serious security problems. For this reason, the present invention provides a method for creating a disk array, which, based on the failure probabilities of the disks mounted under the disk array card, searches for the target disk combination with the minimum data loss probability from the groups of each disk, so that the disk array created based on the target disk combination has the minimum data loss probability, and plays a role in improving the reliability and security of the disk array.
[0084] See Figure 1 As shown, the embodiments of the present invention disclose a method for creating a disk array, including:
[0085] Step S11, obtain the configuration information of the disk array to be created, and determine the failure probabilities of the disks mounted under the disk array card; the configuration information includes the total number of disk arrays to be created, the levels of each disk array to be created, and the number of disks required for each.
[0086] The disk array creation method proposed in the embodiments of the present invention can be deployed on the host as an independent software. That is, the host obtains the configuration information of the disk array to be created input by the user, determines the failure probability of each disk mounted under the disk array card, and thus determines the target disk combination with the minimum data loss probability for creating the most reliable disk array.
[0087] In addition, the disk array creation method proposed in the embodiments of the present invention can be added as an optional function of the disk array card to the process of creating a disk array by the disk array card. That is, the disk array card obtains the configuration information of the disk array to be created sent by the host, determines the failure probability of each disk mounted under the disk array card, and thus determines the target disk combination with the minimum data loss probability for creating the most reliable disk array.
[0088] Among them, obtaining the configuration information of the disk array to be created may specifically include: obtaining the total number of disk arrays to be created, the level of each disk array to be created, and the number of disks required for each. In addition, the configuration information of the disk array to be created also includes the naming name of each disk array to be created.
[0089] It should be noted that the levels of the disk arrays to be created may include RAID0, RAID1, RAID5, RAID6, and RAID10, etc. RAID0 requires at least 2 disks to form, and the failure of any one disk will cause data loss, that is, the number of disks allowed to fail in RAID0 is 0. RAID1 also requires at least 2 disks to form, but the number of disks must be even, and at most one disk is allowed to fail without causing data loss, that is, the number of disks allowed to fail in RAID1 is 1. RAID5 requires at least 3 disks to form, and at most one disk is allowed to fail without causing data loss, that is, the number of disks allowed to fail in RAID5 is 1. RAID6 requires at least 4 disks to form, and the number of disks needs to be greater than or equal to 4, and at most two disks are allowed to fail without causing data loss, that is, the number of disks allowed to fail in RAID6 is 2. RAID10 also requires at least 4 disks to form, but the number of disks must be an even number greater than or equal to 4, and at most half of the disks are allowed to fail without causing data loss, that is, the number of disks allowed to fail in RAID10 is the total number of disks in RAID10 / 2.
[0090] Exemplarily, assuming that RAID10 consists of 8 disks, since RAID10 allows at most half of the disks to fail without causing data loss, the number of disks allowed to fail in RAID10 is equal to the total number of disks in RAID10 / 2 = 8 / 2 = 4, that is, the number of disks allowed to fail in RAID10 is 4.
[0091] Considering that the probability of a disk failure is related to the current stage the disk is in, and the failure probabilities of disks in different stages are different. Also, since disks are electronic products, electronic products generally have the characteristics of the failure bathtub curve as shown in Figure 2 , that is, the failure probabilities of disks in the early stage and the late stage are relatively high, while the failure probability of disks in the stable stage is relatively low.
[0092] Based on this, for determining the failure probabilities of each disk mounted under a disk array card, it can specifically include: obtaining the current status information of each disk mounted under the disk array card, and based on the current status information of each disk, determining the current stage each disk is in, and then based on the current stage each disk is in, determining the failure probability of each disk.
[0093] It should be noted that for the current status information of each disk, the running status of each disk can be monitored in real time by using SMART (Self-Monitoring Analysis and Reporting Technology) technology, so as to obtain the current status information of each disk.
[0094] Since a disk is mainly composed of a controller and NAND (NOT AND) particles, and NAND particles can also be referred to as storage particles, the current status information of a disk includes, but is not limited to, the write times and erase / write usage rates of each storage particle in the disk.
[0095] Specifically, in the process of determining the current stage each disk is in based on the current status information of each disk, the current stage of any disk is determined based on the write times and erase / write usage rates of each storage particle in any disk; where any disk is any one of each disk.
[0096] It can be understood that for any one of each disk mounted under a disk array card, the current stage of the disk can be determined based on the write times and erase / write usage rates of each storage particle in the disk. Among them, the current stage includes the early stage, the stable stage, and the late stage.
[0097] It should be noted that when the disk is in the early stage, the main reasons for disk failures are the quality of the disk controller and the storage particles themselves; when the disk is in the stable and late stages, the aging of the disk components is the main cause of disk failures. Among them, the reasons for the aging of disk components mainly include wear and tear aging caused by normal use of the disk under normal conditions, and also include accelerated aging caused by the use of the disk under abnormal conditions; abnormal conditions include, but are not limited to, abnormal environmental factors, such as too high or too low temperature, too high humidity, abnormal voltage, etc.
[0098] In the process of determining the current stage of any disk based on the write times and erase-write utilization rates of the storage particles in any disk, first determine whether each storage particle in any disk has been written once; if any storage particle in any disk has not been written once, determine that the current stage of any disk is the early stage; if each storage particle in any disk has been written once, then determine whether the erase-write utilization rates of each storage particle in any disk have all reached the preset utilization rate; if the erase-write utilization rate of any storage particle in any disk has not reached the preset utilization rate, determine that the current stage of any disk is the stable stage; if the erase-write utilization rates of each storage particle in any disk have all reached the preset utilization rate, determine that the current stage of any disk is the late stage.
[0099] It should be noted that when writing data to the disk, specifically, data is written to each storage particle in the disk in a cyclic write manner, that is, data is written sequentially starting from the first storage particle in the disk until the last storage particle is written, and after the last storage particle is written, it returns to the first storage particle to continue writing data.
[0100] Exemplarily, taking a disk with three storage particles as an example, first write data to the first storage particle, then write data to the second storage particle, then write data to the third storage particle, and then return to the first storage particle to write data, and so on by this rule, so as to realize data writing to the disk.
[0101] Based on this, for the early stage, it is possible to determine whether the current stage of any disk is the early stage by determining whether each storage particle in any disk has been written once, that is, by determining whether the write times of each storage particle in any disk are all greater than or equal to 1.
[0102] It should be noted that if some storage particles in any disk have not been written once, that is, the write times of some storage particles in any disk are less than 1, then determine that the current stage of any disk is the early stage.
[0103] In addition, if each storage particle in any disk has been written once, that is, the write times of each storage particle in any disk are greater than or equal to 1, it indicates that any disk has passed the early stage, and it is necessary to further determine whether any disk is in the stable stage or the late stage.
[0104] For the stable stage and the late stage, specifically, it is determined whether any disk is in the stable stage or the late stage by judging whether the erase-write utilization rate of each storage particle in any disk reaches the preset utilization rate.
[0105] It should be noted that the erase-write utilization rate of the storage particle is positively correlated with the aging degree of the disk to a certain extent, that is, the higher the erase-write utilization rate of the storage particle, the more serious the aging degree of the disk. Generally, when the erase-write utilization rate of the storage particle reaches more than 90%, the disk enters the late stage from the stable stage, and when the erase-write utilization rate of the storage particle reaches more than 95%, the disk is basically about to break down. Therefore, the preset utilization rate can be set with reference to 90%.
[0106] Specifically, if the erase-write utilization rate of some storage particles in any disk does not reach the preset utilization rate, it is determined that the current stage of any disk is the stable stage; if the erase-write utilization rate of each storage particle in any disk reaches the preset utilization rate, it is determined that the current stage of any disk is the late stage.
[0107] After determining the current stage of each disk, it is necessary to further determine the failure probability of each disk based on the current stage of each disk. Among them, the failure probability of the disk is the probability that the disk fails.
[0108] Specifically, for any disk among each disk, if the current stage of any disk is the early stage, the failure probability of any disk is determined based on the first relative coefficient and the preset failure probability; if the current stage of any disk is the stable stage, the preset failure probability is directly determined as the failure probability of any disk; if the current stage of any disk is the late stage, the failure probability of any disk is determined based on the second relative coefficient and the preset failure probability; where the first relative coefficient and the second relative coefficient are both greater than 1.
[0109] It should be noted that for the preset failure probability, the failure probabilities of several disks in different stages (early stage, stable stage, late stage) can be statistically analyzed to draw such as Figure 2The failure bathtub curve shown, and determine a preset failure probability according to the failure probability in the stable stage of the failure bathtub curve. Specifically, by uniformly sampling the failure bathtub curve in the stable stage to obtain a preset number of sampling points, and determining the preset failure probability based on the failure probabilities of the preset number of sampling points. For example, calculate the average value of the failure probabilities of the preset number of sampling points and determine the average value as the preset failure probability.
[0110] In addition, an approximately smooth curve can also be determined from the failure bathtub curve in the stable stage, and the preset failure probability is determined based on the failure probability corresponding to the approximately smooth curve. Specifically, by uniformly sampling the approximately smooth curve to obtain a preset number of sampling points, and determining the preset failure probability based on the failure probabilities of the preset number of sampling points.
[0111] In this way, by determining the preset failure probability in the above manner, the selection of the preset failure probability can be made scientific and reasonable to a certain extent, and it is more in line with the characteristics of the disk in the stable stage, ultimately playing a role in improving the accuracy of determining the failure probability of the disk.
[0112] It should also be noted that since the failure probability of the disk in the early stage and the late stage is generally greater than the failure probability of the disk in the stable stage, therefore, in the embodiments of the present invention, a first relative coefficient and a second relative coefficient greater than 1 are used to calculate the preset failure probability, so that the failure probability calculated for the disk in the early stage and the late stage is greater than the preset failure probability, and correspondingly greater than the failure probability calculated for the disk in the stable stage.
[0113] For determining whether each storage particle in any disk has been written once, in a specific implementation manner, it may specifically include: determining whether the total written data volume of any disk is greater than a preset data volume to determine whether each storage particle in any disk has been written once; wherein, the preset data volume is a value determined based on the number of each storage particle in any disk and the unit written data volume of each storage particle.
[0114] It can be understood that first, the total amount of data written to any disk is obtained through the Data Units Written field in the SMART information of any disk, and a preset data volume is determined based on the number of storage particles in any disk and the unit data write volume of each storage particle. Herein, the unit data write volume of each storage particle refers to the amount of data to be written at one time when writing data to each storage particle. Then, it is determined whether the total amount of data written to any disk is greater than the preset data volume. If the total amount of data written to any disk is greater than the preset data volume, it is determined that each storage particle in any disk has been written once. If the total amount of data written to any disk is less than or equal to the preset data volume, it is determined that some storage particles in any disk have not been written once.
[0115] Further, when it is determined that some storage particles in any disk have not been written once through the total amount of data written to any disk being less than or equal to the preset data volume, the current stage of any disk is determined to be the early stage. At this time, it is necessary to determine the failure probability of any disk based on the first relative coefficient and the preset failure probability.
[0116] According to one embodiment, if the current stage of any disk is the early stage, the first relative coefficient can be determined based on the ratio between the total amount of data written to any disk and the preset data volume, and the failure probability of any disk can be determined according to the first relative coefficient and the preset failure probability.
[0117] Specifically, the ratio between the total amount of data written to any disk and the preset data volume is determined, and the difference between 1 and the ratio is determined. Then, the product of the difference and 10 is determined as the first relative coefficient, and the failure probability of any disk is determined according to the product of the first relative coefficient and the preset failure probability. That is, the first relative coefficient = 10×(1 - the total amount of data written to any disk / the preset data volume), and the failure probability of any disk = the first relative coefficient × the preset failure probability.
[0118] For determining whether each storage particle in any disk has been written once, in another specific implementation, it may specifically include: determining whether the power-on time of any disk is greater than the preset time to determine whether each storage particle in any disk has been written once; wherein, the preset time is a value determined based on the preset number of days and the number of hours per day.
[0119] It can be understood that first, the power-on time of any disk is obtained through the power on hours field in the SMART information of any disk, and a preset time is determined based on a preset number of days and the number of hours per day. Here, both the power-on time and the preset time are in hours, and the number of hours per day is set to 24 hours per day. Then, it is judged whether the power-on time of any disk is greater than the preset time. If the power-on time of any disk is greater than the preset time, it is determined that each storage particle in any disk has been written once. If the power-on time of any disk is less than or equal to the preset time, it is determined that some storage particles in any disk have not been written once.
[0120] Among them, the preset number of days can be set according to the prior knowledge of the disk. For example, it is set to 90 days. At this time, the preset time is 90×24 = 2160 hours.
[0121] Furthermore, when it is determined that some storage particles in any disk have not been written once because the power-on time of any disk is less than or equal to the preset time, the current stage of any disk is determined to be the early stage. At this time, it is necessary to determine the failure probability of any disk based on the first relative coefficient and the preset failure probability.
[0122] According to one embodiment, if the current stage of any disk is the early stage, the first relative coefficient can be determined based on the ratio between the power-on time and the preset time of any disk, and the failure probability of any disk can be determined according to the first relative coefficient and the preset failure probability.
[0123] Specifically, the ratio between the power-on time and the preset time of any disk is determined, and the difference between 1 and the ratio is determined. Then, the product of the difference and 10 is determined as the first relative coefficient, and the failure probability of any disk is determined according to the product of the first relative coefficient and the preset failure probability. That is, the first relative coefficient = 10×(1 - the power-on time of any disk / the preset time), and the failure probability of any disk = the first relative coefficient × the preset failure probability.
[0124] For judging whether each storage particle in any disk has been written once, in another specific implementation manner, it may specifically include: obtaining the number of written storage particles in any disk, and judging whether each storage particle in any disk has been written once by determining whether the number of written storage particles in any disk is equal to the total number of all storage particles in any disk.
[0125] Among them, if the number of written storage particles in any disk is less than the total number of all storage particles in any disk, it is determined that some storage particles in any disk have not been written once; if the number of written storage particles in any disk is equal to the total number of all storage particles in any disk, it is determined that each storage particle in any disk has been written once.
[0126] Further, when the number of written storage particles in any disk is less than the total number of storage particles in any disk, and it is determined that some storage particles in any disk have not been written once, the current stage of any disk is determined to be the early stage. At this time, it is necessary to determine the failure probability of any disk based on the first relative coefficient and the preset failure probability.
[0127] According to one embodiment, if the current stage of any disk is the early stage, the first relative coefficient can be determined based on the ratio between the number of written storage particles in any disk and the total number of storage particles in any disk, and the failure probability of any disk can be determined according to the first relative coefficient and the preset failure probability.
[0128] Specifically, determine the ratio between the number of written storage particles in any disk and the total number of storage particles in any disk, and determine the difference between 1 and the ratio. Then, determine the product of the difference and 10 as the first relative coefficient, and determine the failure probability of any disk according to the product of the first relative coefficient and the preset failure probability. That is, the first relative coefficient = 10×(1 - the number of written storage particles in any disk / the total number of storage particles in any disk), and the failure probability of any disk = the first relative coefficient × the preset failure probability.
[0129] For any one of the disks, when the current stage of any disk is the late stage, the failure probability of any disk can be determined based on the product of a second relative coefficient greater than 1 and the preset failure probability.
[0130] Specifically, when the current stage of any disk is the late stage, the second relative coefficient is determined based on the erasure usage rate of each storage particle in any disk and the preset usage rate, and the failure probability of any disk is determined according to the second relative coefficient and the preset failure probability.
[0131] In the process of determining the second relative coefficient based on the erasure usage rate of each storage particle in any disk and the preset usage rate, first determine the average usage rate based on the erasure usage rate of each storage particle in any disk to obtain the erasure usage rate of any disk, and then determine the second relative coefficient according to the difference between the erasure usage rate of any disk and the preset usage rate.
[0132] According to one embodiment, when the current stage of any disk is the late stage, calculate the average of the erasure usage rates of the storage particles in any disk to obtain the average usage rate, and determine the average usage rate as the erasure usage rate of any disk. Then, determine the difference between the erasure usage rate of any disk and the preset usage rate, determine the ratio between the difference and 5, and based on the product of the ratio and 10, determine the second relative coefficient. Finally, based on the product of the second relative coefficient and the preset failure probability, determine the failure probability of any disk. That is, the second relative coefficient = 10×[(the erasure usage rate of any disk - the preset usage rate) / 5], and the failure probability of any disk = the second relative coefficient × the preset failure probability.
[0133] Step S12: Use the configuration information and based on the failure probabilities of the respective disks, group the respective disks to obtain a target disk combination; the target disk combination is the combination with the minimum data loss probability among several disk combinations obtained by grouping; wherein, each disk combination includes sub-disk combinations corresponding to the respective disks to be created in the disk array, and the data loss probability of each disk combination is a value determined based on the failure probabilities of the disks in each disk combination.
[0134] In the embodiment of the present invention, use the configuration information of the disk array to be created and based on the failure probabilities of the respective disks mounted under the disk array card, group the respective disks, and determine the combination with the minimum data loss probability from several disk combinations obtained by grouping to obtain the target disk combination.
[0135] Since the configuration information of the disk array to be created includes the total number of the disk arrays to be created, each disk combination obtained by grouping includes sub-disk combinations corresponding to the respective disks to be created in the disk array. That is, the number of sub-disk combinations included in each disk combination is equal to the total number of the disk arrays to be created, and different sub-disk combinations in each disk combination correspond to different disk arrays to be created.
[0136] Exemplarily, if the total number of the disk arrays to be created is 2, each disk combination obtained by grouping includes 2 sub-disk combinations, and the sub-disk combinations correspond one-to-one to the disk arrays to be created.
[0137] Since the configuration information of the disk array to be created also includes the number of disks required for each disk array to be created, for the sub-disk combinations corresponding to the respective disks to be created in each disk combination obtained by grouping, the number of disks included in the sub-disk combination is the same as the number of disks required for the corresponding disk array to be created.
[0138] Exemplarily, if the total number of disk arrays to be created is 2, and the number of disks required for the first disk array to be created is 1, and the number of disks required for the second disk array to be created is 3, then each disk combination obtained by grouping includes 2 sub-disk combinations. One sub-disk combination corresponds to the first disk array to be created and contains 2 disks; the other sub-disk combination corresponds to the second disk array to be created and contains 3 disks.
[0139] In addition, since the configuration information of the disk arrays to be created includes the levels of each disk array to be created, and different levels of disk arrays correspond to different allowable number of failed disks. Among them, the allowable number of failed disks indicates that data loss occurs when the number of failed disks is greater than the allowable number of failed disks in the disk array, that is, the data loss probability of the disk array is associated with the failure probability of the disks included in the disk array. Therefore, for each disk combination obtained by grouping, the data loss probability of each disk combination can be determined based on the failure probability of the disks included in each disk combination.
[0140] Exemplarily, if the allowable number of failed disks in the disk array is 1, it means that data loss occurs when the number of failed disks is greater than 1, that is, the disk array allows at most 1 disk to fail without causing data loss.
[0141] It should be noted that the process of determining the data loss probability of each disk combination may specifically include: determining the data loss probability of any one sub-disk combination based on the failure probability of the disks in any one sub-disk combination; where any one sub-disk combination is any one of the sub-disk combinations included in each disk combination; determining the data loss probability of each disk combination according to the data loss probabilities corresponding to the sub-disk combinations included in each disk combination.
[0142] It can be understood that for any one of the sub-disk combinations included in each disk combination, the data loss probability of any one sub-disk combination can be determined based on the failure probability of the disks included in any one sub-disk combination, and after determining the data loss probabilities corresponding to the sub-disk combinations included in each disk combination, the data loss probability of each disk combination can be determined based on the data loss probabilities corresponding to the sub-disk combinations included in each disk combination.
[0143] In the process of determining the data loss probability of any sub-disk combination based on the failure probability of the disks in any sub-disk combination, it may specifically include: determining the target disk array to be created corresponding to any sub-disk combination; determining the allowable number of failed disks based on the level corresponding to the target disk array to be created; where the allowable number of failed disks indicates that data loss occurs when the number of failed disks in the target disk array to be created is greater than the allowable number of failed disks; using the allowable number of failed disks and based on the failure probability of the disks in any sub-disk combination, determining the data loss probability of any sub-disk combination.
[0144] Specifically, for using the allowable number of failed disks and based on the failure probability of the disks in any sub-disk combination to determine the data loss probability of any sub-disk combination, it may specifically include: determining each target quantity based on the number of disks required for the target disk array to be created and the allowable number of failed disks; where the target quantity is a positive integer not greater than the number of disks required for the target disk array to be created and greater than the allowable number of failed disks; determining the probability that the target disk array to be created fails when the target quantity of disks fails based on the failure probability of the disks in any sub-disk combination, and determining the data loss probability of any sub-disk combination according to the probability that the target disk array to be created fails when the target quantity of disks fails.
[0145] Exemplarily, if the number of disks required for the target disk array to be created corresponding to any sub-disk combination is 3, and the allowable number of failed disks corresponding to the target disk array to be created is 1, then according to the definition that the target quantity is a positive integer not greater than the number of disks required for the target disk array to be created and greater than the allowable number of failed disks, it can be determined that the target quantities are 2 and 3. Then, based on the failure probability of the disks in any sub-disk combination, determine the probability that the target disk array to be created fails when 2 disks fail, and determine the probability that the target disk array to be created fails when 3 disks fail. After that, determine the data loss probability of any sub-disk combination according to the probability that the target disk array to be created fails when 2 disks fail and the probability that the target disk array to be created fails when 3 disks fail.
[0146] Since the number of disks required for the target disk array to be created corresponding to any sub-disk combination is 3, any sub-disk combination contains three disks, denoted as disk 0, disk 1, and disk 2. Assume that the failure probability of disk 0 is 0.2, the failure probability of disk 1 is 0.2, and the failure probability of disk 2 is 0.1. At this time, the probability of the target disk array to be created when 2 disks fail = the failure probability of disk 0 × the failure probability of disk 1 × the normal probability of disk 2 + the failure probability of disk 0 × the normal probability of disk 1 × the failure probability of disk 2 + the normal probability of disk 0 × the failure probability of disk 1 × the failure probability of disk 2 = 0.2×0.2×0.9 + 0.2×0.8×0.1 + 0.8×0.2×0.1 = 0.036 + 0.016 + 0.016 = 0.068. The probability of the target disk array to be created when 3 disks fail = the failure probability of disk 0 × the failure probability of disk 1 × the failure probability of disk 2 = 0.2×0.2×0.1 = 0.004. The data loss probability of any sub-disk combination = the probability of the target disk array to be created when 2 disks fail + the probability of the target disk array to be created when 3 disks fail = 0.068 + 0.004 = 0.072. It should be noted that the normal probability of a disk = 1 - the failure probability of the disk.
[0147] After determining the data loss probability corresponding to each sub-disk combination in each disk combination, taking each disk combination containing 2 sub-disk combinations as an example, where the data loss probability of the first sub-disk combination is 0.1 and the data loss probability of the second sub-disk combination is 0.2, calculate the data loss probability of each disk combination.
[0148] In a specific implementation manner, the data loss probability of each disk combination = 1 - (1 - the data loss probability of the first sub-disk combination) × (1 - the data loss probability of the second sub-disk combination) = 1 - 0.9×0.8 = 1 - 0.72 = 0.28.
[0149] In another specific implementation manner, the data loss probability of each disk combination = the data loss probability of the first sub-disk combination × (1 - the data loss probability of the second sub-disk combination) + (1 - the data loss probability of the first sub-disk combination) × the data loss probability of the second sub-disk combination + the data loss probability of the first sub-disk combination × the data loss probability of the second sub-disk combination = 0.1×0.8 + 0.9×0.2 + 0.1×0.2 = 0.08 + 0.18 + 0.02 = 0.28.
[0150] Based on the above calculation method for the data loss probability of the sub-disk combinations and the above calculation method for the data loss probability of each disk combination, the data loss probabilities of the several disk combinations obtained by grouping can be calculated, and thus the disk combination with the minimum data loss probability among the several disk combinations can be determined as the target disk combination.
[0151] It should be noted that if there is more than one combination with the minimum data loss probability among the several disk combinations, any one of the combinations with the minimum data loss probability among the several disk combinations can be determined as the target disk combination.
[0152] In the embodiments of the present invention, for the determination of the target disk combination, specifically, several disk creation tasks can be created first based on the total number of disk arrays to be created, and the target disk combination can be determined by using the several disk creation tasks. Among them, the number of the several disk creation tasks is one more than the total number of disk arrays to be created.
[0153] Specifically, create several disk creation tasks that are one more than the total number of disk arrays to be created to obtain several disk creation tasks; use the first disk creation task among the several disk creation tasks and based on the number of disks required for each disk array to be created, group each disk to obtain the current disk combination; where the number of the first disk creation tasks is the same as the total number of disk arrays to be created; determine the data loss probability of the current disk combination through the second disk creation task among the several disk creation tasks, determine the new current minimum data loss probability based on the data loss probability of the current disk combination and the current minimum data loss probability, and when the current grouping situation does not meet the preset end condition, jump back to the above step of using the first disk creation task among the several disk creation tasks and grouping each disk based on the number of disks required for each disk array to be created until the current grouping situation meets the preset end condition, so as to determine the target disk combination based on the disk combination corresponding to the latest current minimum data loss probability. It should be noted that the current minimum data loss probability is a preset loss probability not less than 1 at the beginning, and the number of the second disk creation tasks is one.
[0154] Among them, for determining the new current minimum data loss probability based on the data loss probability of the current disk combination and the current minimum data loss probability, if the data loss probability of the current disk combination is less than the current minimum data loss probability, then determine the data loss probability of the current disk combination as the new current minimum data loss probability, and if the data loss probability of the current disk combination is greater than or equal to the current minimum data loss probability, then keep the current minimum data loss probability unchanged to obtain the new current minimum data loss probability.
[0155] It can be understood that a task is created using the first disk, and based on the number of disks required for each disk array to be created, each disk is grouped to obtain the first disk combination. Then, the data loss probability of the first disk combination is determined through the second disk creation task. Based on the data loss probability of the first disk combination and the current minimum data loss probability (which is a preset loss probability not less than 1 at this time), a new current minimum data loss probability is determined. When the current grouping situation does not meet the preset end condition, the first disk creation task is used again, and based on the number of disks required for each disk array to be created, each disk is grouped to obtain the second disk combination, and the data loss probability of the second disk combination is determined again through the second disk creation task. Based on the data loss probability of the second disk combination and the current minimum data loss probability, a new current minimum data loss probability is determined, and so on, until the current grouping situation meets the preset end condition. Finally, the target disk combination is determined based on the disk combination corresponding to the latest current minimum data loss probability.
[0156] In the process of using the first disk creation task in several disk creation tasks and grouping each disk based on the number of disks required for each disk array to be created to obtain the current disk combination, the disk grouping operation is performed through each first disk creation task in turn based on the number of disks required for its corresponding disk array to be created, so as to obtain the current remaining disks using the current first disk creation task. Then, based on the number of disks required for the current disk array to be created, the corresponding number of disks is taken out from the current remaining disks to obtain the sub-disk combination corresponding to the current disk array to be created and update the current remaining disks. Among them, different first disk creation tasks correspond to different disk arrays to be created, and the current disk array to be created corresponds to the current first disk creation task. And the current remaining disks are initially the disks mounted under the disk array card. Further, when the current first disk creation task is the last disk creation task among all the first disk creation tasks, the current disk combination is determined based on the sub-disk combinations corresponding to each disk array to be created obtained from each first disk creation task.
[0157] It can be understood that the first disk creation task among the first disk creation tasks for each creates a disk grouping operation based on the number of disks required for the first disk array to be created corresponding to itself, so as to obtain the current remaining disks based on each disk mounted under the disk array card, and based on the number of disks required for the first disk array to be created, take out the corresponding number of disks from the current remaining disks to obtain a sub-disk combination corresponding to the first disk array to be created, and update the current remaining disks to the disks not taken out, and then send the current remaining disks and the sub-disk combination corresponding to the first disk array to be created to the second disk creation task among the first disk creation tasks. The second disk creation task creates a disk grouping operation based on the number of disks required for the second disk array to be created corresponding to itself, so as to take out the corresponding number of disks from the current remaining disks based on the number of disks required for the second disk array to be created to obtain a sub-disk combination corresponding to the second disk array to be created, and update the current remaining disks to the disks not taken out, and then send the current remaining disks, the sub-disk combination corresponding to the first disk array to be created, and the sub-disk combination corresponding to the second disk array to be created to the next disk creation task among the first disk creation tasks. And so on, until after the last disk creation task among the first disk creation tasks completes the disk grouping operation, determine the current disk combination based on the sub-disk combinations corresponding to each disk array to be created obtained by each first disk creation task.
[0158] It should be noted that for the preset end condition, it can include both that the current grouping cumulative times reach the maximum grouping times and that the current cumulative number of types of the disk combination reaches the maximum number of types; among them, both the maximum grouping times and the maximum number of types are values determined based on the maximum number of types, and the maximum number of types is the maximum number of types of disk combinations that can be obtained after grouping each disk.
[0159] Exemplarily, if the number of each disk mounted under the disk array card is 3, the total number of disk arrays to be created is 2, and the number of disks required for one of the disk arrays to be created is 1, and the number of disks required for the other disk array to be created is 2, then there are at most 3 types of disk combinations that can be obtained after grouping each disk mounted under the disk array card, such as [1] and [2, 3], [2] and [1, 3], [3] and [1, 2]. At this time, the maximum number of types is 3, and correspondingly, both the maximum grouping times and the maximum number of types are also 3.
[0160] Step S13: Create disk arrays of corresponding levels by using the levels of each disk array to be created and based on each sub-disk combination in the target disk combination.
[0161] In an embodiment of the present invention, after determining the target disk combination, disk arrays of corresponding levels are created by using the levels of each disk array to be created and based on the sub-disk combinations corresponding to each disk array to be created in the target disk combination.
[0162] Beneficial effects: By using the configuration information of the disk arrays to be created and based on the failure probabilities of the disks mounted under the disk array card, the present invention groups each disk, determines the combination with the lowest data loss probability from several disk combinations obtained by grouping, so as to obtain the target disk combination, and then creates the corresponding disk arrays based on the sub-disk combinations corresponding to each disk array to be created in the target disk combination. In this way, based on the failure probabilities of the disks mounted under the disk array card, the present invention searches for the target disk combination with the lowest data loss probability from the grouped disks, so that the disk array created based on the target disk combination has the lowest data loss probability, and it plays a role in improving the reliability and security of the disk array.
[0163] See Figure 3 and Figure 4 As shown, an embodiment of the present invention discloses a method for creating a disk array, including:
[0164] Obtain the SMART information of each disk mounted under the disk array card. For any disk among the disks mounted under the disk array card, obtain the total written data volume of any disk through the Data Units Written field in the SMART information of any disk, and obtain the power-on time of any disk through the power on hours field in the SMART information of any disk.
[0165] Judge whether the total written data volume of any disk is greater than a preset data volume, and judge whether the power-on time of any disk is greater than a preset time.
[0166] If the total written data volume of any disk is less than or equal to the preset data volume, or the power-on time of any disk is less than or equal to the preset time, then determine that the current stage of any disk is the early stage, and determine the first relative coefficient = 10×(1 - the total written data volume of any disk / the preset data volume), or the first relative coefficient = 10×(1 - the power-on time of any disk / the preset time), and then determine the failure probability of any disk according to the product between the first relative coefficient and the preset failure probability.
[0167] If the total written data volume of any disk is greater than the preset data volume, and the power-on time of any disk is greater than the preset time, then judge whether the erase-write utilization rates of the storage particles in any disk all reach the preset utilization rate.
[0168] If the erasure and write usage rate of any storage particle in any disk does not reach the preset usage rate, determine that the current stage of any disk is the stable stage, and directly determine the preset failure probability as the failure probability of any disk.
[0169] If the erasure and write usage rates of all storage particles in any disk reach the preset usage rate, determine that the current stage of any disk is the late stage, and determine the average usage rate based on the erasure and write usage rates of all storage particles in any disk, so as to determine the average usage rate as the erasure and write usage rate of any disk. Then determine the second relative coefficient = 10×[(the erasure and write usage rate of any disk - the preset usage rate) / 5], and determine the failure probability of any disk according to the product between the second relative coefficient and the preset failure probability.
[0170] After obtaining the configuration information of the disk array to be created (the total number of disk arrays to be created, the levels of each disk array to be created, the number of disks required for each disk array to be created), and determining the failure probabilities of each disk mounted under the disk array card, create one more disk creation task than the total number of disk arrays to be created based on the total number of disk arrays to be created, so as to obtain several disk creation tasks. Among them, several disk creation tasks include at least one first disk creation task and one second disk creation task. The number of at least one first disk creation task is the same as the total number of disk arrays to be created, and different first disk creation tasks correspond to different disk arrays to be created.
[0171] In the process of determining the current disk combination based on the first disk creation task, perform disk grouping operations through each first disk creation task based on the number of disks required for the corresponding disk array to be created in turn, so as to obtain the current remaining disks by using the current first disk creation task, and take out the corresponding number of disks from the current remaining disks based on the number of disks required for the current disk array to be created, so as to obtain the sub-disk combination corresponding to the current disk array to be created and update the current remaining disks; where the current disk array to be created is the disk array to be created corresponding to the current first disk creation task, and the current remaining disks are initially the disks mounted under the disk array card. Further, when the current first disk creation task is the last disk creation task among all the first disk creation tasks, determine the current disk combination based on the sub-disk combinations corresponding to each disk array to be created obtained by each first disk creation task.
[0172] Determine the data loss probability of the current disk combination through the second disk creation task. Based on the data loss probability of the current disk combination and the current minimum data loss probability, determine the new current minimum data loss probability. When the current grouping situation does not meet the preset end condition, jump back to the above process of determining the current disk combination based on the first disk creation task until the current grouping situation meets the preset end condition, so as to determine the target disk combination based on the disk combination corresponding to the latest current minimum data loss probability. It should be noted that the current minimum data loss probability is a preset loss probability not less than 1 at the beginning.
[0173] After determining the target disk combination, use the levels of each disk array to be created and create disk arrays of corresponding levels based on the sub-disk combinations corresponding to each disk array to be created in the target disk combination.
[0174] Beneficial effects: By using the configuration information of the disk arrays to be created and based on the failure probabilities of the disks mounted under the disk array card, the present invention groups each disk and determines the combination with the minimum data loss probability from several disk combinations obtained by grouping, so as to obtain the target disk combination. Then, based on the sub-disk combinations corresponding to each disk array to be created in the target disk combination, corresponding disk arrays are created. In this way, based on the failure probabilities of the disks mounted under the disk array card, the present invention searches for the target disk combination with the minimum data loss probability from the grouping of each disk, so that the disk array created based on the target disk combination has the minimum data loss probability, and plays a role in improving the reliability and security of the disk array.
[0175] Furthermore, the embodiment of the present application also discloses an electronic device. Figure 5 It is a structural diagram of an electronic device shown according to an exemplary embodiment. The content in the figure cannot be considered as any limitation on the scope of use of the present application. The electronic device may specifically include: at least one processor 11, at least one memory 12, a power supply 13, a communication interface 14, an input / output interface 15, and a communication bus 16. Among them, the memory 12 is used to store a computer program, and the computer program is loaded and executed by the processor 11 to implement the relevant steps in the disk array creation method disclosed in any of the foregoing embodiments. In addition, the electronic device in this embodiment may specifically be an electronic computer.
[0176] In this embodiment, the power supply 13 is used to provide operating voltages for each hardware device on the electronic device; the communication interface 14 can create a data transmission channel between the electronic device and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and no specific limitation is imposed thereon here; the input / output interface 15 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application requirements, and no specific limitation is imposed here.
[0177] In addition, the memory 12, as a carrier for resource storage, can be a read-only memory, a random access memory, a magnetic disk, an optical disk, etc. The resources stored thereon can include an operating system 121, a computer program 122, etc., and the storage method can be temporary storage or permanent storage.
[0178] Among them, the operating system 121 is used to manage and control each hardware device and the computer program 122 on the electronic device, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the disk array creation method executed by the electronic device disclosed in any of the foregoing embodiments, the computer program 122 can further include computer programs that can be used to complete other specific tasks.
[0179] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the disk array creation method disclosed above is implemented. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details are not described herein again.
[0180] Furthermore, this application also discloses a computer program product including a computer program / instructions; wherein, when the computer program / instructions are executed by a processor, the disk array creation method disclosed above is implemented. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details are not described herein again.
[0181] In this specification, the various embodiments are described in a progressive manner, and the key point of each embodiment is the difference from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0182] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0183] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0184] Finally, it should also be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0185] The technical solutions provided in this application have been introduced in detail above. Specific examples have been used in this article to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A method for creating a disk array, characterized in that, Including: Obtain the configuration information of the disk array to be created, and determine the failure probability of each disk mounted under the disk array card; the configuration information includes the total number of the disk arrays to be created, the level of each disk array to be created, and the number of disks required for each; Group the disks based on the configuration information and based on the failure probability of each disk to obtain a target disk combination; the target disk combination is the combination with the minimum data loss probability among several disk combinations obtained by grouping; wherein, each disk combination includes a sub-disk combination corresponding to each disk array to be created, and the data loss probability of each disk combination is a value determined based on the failure probability of the disks in each disk combination; Create disk arrays of corresponding levels by using the levels of each disk array to be created and based on each sub-disk combination in the target disk combination; Among them, the step of grouping the disks based on the configuration information and based on the failure probability of each disk to obtain a target disk combination includes: Group the disks based on the number of disks required for each disk array to be created to obtain a current disk combination; Determine the data loss probability of the current disk combination, determine a new current minimum data loss probability based on the data loss probability of the current disk combination and the current minimum data loss probability, and when the current grouping situation does not meet the preset end condition, jump to the step of grouping the disks based on the number of disks required for each disk array to be created until the current grouping situation meets the preset end condition, so as to determine the target disk combination based on the disk combination corresponding to the latest current minimum data loss probability.
2. The method for creating a disk array according to claim 1, wherein The step of determining the failure probability of each disk mounted under the disk array card includes: Obtain the current status information of each disk mounted under the disk array card; Determine the current stage where each disk is located based on the current status information of each disk; Determine the failure probability of each disk based on the current stage where each disk is located.
3. The method for creating a disk array according to claim 2, wherein The step of determining the current stage where each disk is located based on the current status information of each disk includes: Determine the current stage where any one disk is located based on the write times and the erase-write utilization rate of each storage particle in any one disk; Wherein, any one disk is any one of the disks.
4. The method for creating a disk array according to claim 3, wherein, The step of determining the current stage where any one disk is located based on the write times and the erase-write utilization rate of each storage particle in any one disk includes: Judge whether each storage particle in any one disk has been written once; If any one storage particle in any one disk has not been written once, determine that the current stage where any one disk is located is the early stage; If each storage particle in any one disk has been written once, judge whether the erase-write utilization rate of each storage particle in any one disk has reached the preset utilization rate; If the erase-write utilization rate of any one storage particle in any one disk has not reached the preset utilization rate, determine that the current stage where any one disk is located is the stable stage; If the erase-write usage rate of each storage particle in any of the disks reaches the preset usage rate, determine that the current stage of the any disk is the late stage.
5. The method for creating a disk array according to claim 4, wherein The determining the failure probability of each disk based on the current stage where each disk is located includes: If the current stage of any disk is the early stage, determine the failure probability of the any disk based on the first relative coefficient and the preset failure probability; If the current stage of any disk is the stable stage, determine the preset failure probability as the failure probability of the any disk; If the current stage of any disk is the late stage, determine the failure probability of the any disk based on the second relative coefficient and the preset failure probability; Wherein, both the first relative coefficient and the second relative coefficient are greater than 1.
6. The method for creating a disk array according to claim 5, wherein The determining whether each storage particle in any disk has been written once includes: By determining whether the total written data volume of the any disk is greater than the preset data volume to determine whether each storage particle in the any disk has been written once; Wherein, the preset data volume is a value determined based on the number of each storage particle in the any disk and the unit written data volume of each storage particle.
7. The method for creating a disk array according to claim 6, wherein The determining the failure probability of the any disk based on the first relative coefficient and the preset failure probability includes: Determine the first relative coefficient based on the ratio between the total written data volume of the any disk and the preset data volume; Determine the failure probability of the any disk according to the first relative coefficient and the preset failure probability.
8. The method for creating a disk array according to claim 5, wherein The determining whether each storage particle in any disk has been written once includes: By determining whether the power-on time of the any disk is greater than the preset time to determine whether each storage particle in the any disk has been written once; Wherein, the preset time is a value determined based on the preset number of days and the number of hours per day.
9. The method for creating a disk array according to claim 8, wherein The determining the failure probability of the any disk based on the first relative coefficient and the preset failure probability includes: Determine the first relative coefficient based on the ratio between the power-on time of the any disk and the preset time; Determine the failure probability of the any disk according to the first relative coefficient and the preset failure probability.
10. The method for creating a disk array according to claim 5, wherein The determining the failure probability of the any disk based on the second relative coefficient and the preset failure probability includes: Determine the second relative coefficient based on the erase-write usage rate of each storage particle in the any disk and the preset usage rate; Determine the failure probability of the any disk according to the second relative coefficient and the preset failure probability.
11. The method for creating a disk array according to claim 10, wherein The determining the second relative coefficient based on the erase-write usage rate of each storage particle in the any disk and the preset usage rate includes: Determine the average usage rate based on the erase-write usage rate of each storage particle in the any disk to obtain the erase-write usage rate of the any disk; Determine the second relative coefficient according to the difference between the erase-write usage rate of the any disk and the preset usage rate.
12. The method for creating a disk array according to claim 1, wherein The determination process of the data loss probability of each disk combination includes: Determine the data loss probability of any sub-disk combination based on the failure probability of the disks in any sub-disk combination; any sub-disk combination is any one of the sub-disk combinations included in each disk combination. Determine the data loss probability of each disk combination according to the data loss probabilities respectively corresponding to the sub-disk combinations in each disk combination.
13. The method for creating a disk array according to claim 12, wherein The determining the data loss probability of any sub-disk combination based on the failure probability of the disks in any sub-disk combination includes: Determine the target disk array to be created corresponding to any sub-disk combination. Based on the level corresponding to the target disk array to be created, determine the allowable number of failed disks; the allowable number of failed disks indicates that data loss occurs when the number of failed disks in the target disk array to be created is greater than the allowable number of failed disks. Use the allowable number of failed disks and based on the failure probability of the disks in any sub-disk combination, determine the data loss probability of any sub-disk combination.
14. The method for creating a disk array according to claim 13, wherein The using the allowable number of failed disks and based on the failure probability of the disks in any sub-disk combination to determine the data loss probability of any sub-disk combination includes: Based on the number of disks required for the target disk array to be created and the allowable number of failed disks, determine each target number; the target number is a positive integer that is not greater than the number of disks required for the target disk array to be created and is greater than the allowable number of failed disks. Based on the failure probability of the disks in any sub-disk combination, determine the probability when the target number of disks in the target disk array to be created fails. According to the probability when the target number of disks in the target disk array to be created fails, determine the data loss probability of any sub-disk combination.
15. The method for creating a disk array according to any one of claims 1 to 14, characterized in that, The using the configuration information and based on the failure probability of each disk to group each disk to obtain a target disk combination includes: Based on the total number of disk arrays to be created, create several disk creation tasks; the number of the several disk creation tasks is one more than the total number of disk arrays to be created. Use the first disk creation task among the several disk creation tasks and based on the number of disks required for each disk array to be created, group each disk to obtain a current disk combination; the number of the first disk creation tasks is the same as the total number of disk arrays to be created. Determine the data loss probability of the current disk combination through the second disk creation task among the several disk creation tasks, determine a new current minimum data loss probability based on the data loss probability of the current disk combination and the current minimum data loss probability, and when the current grouping situation does not meet the preset end condition, jump to the step of using the first disk creation task among the several disk creation tasks and based on the number of disks required for each disk array to be created to group each disk, until the current grouping situation meets the preset end condition, so as to determine the target disk combination based on the disk combination corresponding to the latest current minimum data loss probability. Wherein, the current minimum data loss probability is a preset loss probability not less than 1 at the initial time.
16. The method for creating a disk array according to claim 15, wherein The method of creating a first disk creation task among the several disk creation tasks and grouping the respective disks based on the number of disks required for each disk array to be created to obtain a current disk combination includes: Performing disk grouping operations on each of the first disk creation tasks in sequence based on the number of disks required for the respective disk arrays to be created, so as to obtain current remaining disks by using the current first disk creation task, and taking out a corresponding number of disks from the current remaining disks based on the number of disks required for the current disk array to be created, so as to obtain a sub-disk combination corresponding to the current disk array to be created and update the current remaining disks; the current disk array to be created corresponds to the current first disk creation task; the current remaining disks are the respective disks mounted under the disk array card at the initial time. When the current first disk creation task is the last disk creation task among all the first disk creation tasks, determining the current disk combination based on the sub-disk combinations corresponding to the respective disk arrays to be created obtained by each of the first disk creation tasks.
17. The method for creating a disk array according to claim 16, wherein The preset end condition includes that the current cumulative grouping times reach the maximum grouping times or the current cumulative number of types of the disk combination reaches the maximum number of types; both the maximum grouping times and the maximum number of types are values determined based on the maximum number of types, and the maximum number of types is the maximum number of types of disk combinations that can be obtained by grouping the respective disks.
18. An electronic device, characterized in that, Including: A memory for storing a computer program; A processor for executing the computer program to implement the steps of the disk array creation method according to any one of claims 1 to 17.
19. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, the steps of the disk array creation method according to any one of claims 1 to 17 are implemented.
20. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps of the disk array creation method according to any one of claims 1 to 17 are implemented.
Citation Information
Patent Citations
Solid state disk error correction method and device, equipment and medium
CN117908798A
Grouping of storage media based on parameters associated with the storage media
US20050044313A1
Shifting wearout of storage disks
US20170115903A1