Disk array creation method and device, storage medium and program product
By obtaining the probability of disk failure and grouping it during the disk array creation process, we can find the disk combination with the lowest probability of data loss, which solves the problem of high risk of data loss in disk arrays under the same disk conditions, and improves the reliability and security of disk arrays.
Patent Information
- Application Number
- CN202510495607.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-21
AI Technical Summary
When creating a disk array, how to reduce the risk of data loss under the same disk conditions, especially in a RAID5-level disk array, selecting which disks form a disk array will lead to different probability of data loss.
By obtaining the configuration information of the disk array to be created, the failure probability of each disk is determined, and the disks are grouped based on these probabilities, and the target disk combination with the lowest probability of data loss is found, thereby creating a disk array with the lowest probability of data loss.
A disk array created under the same disk conditions has a lower risk of data loss, improving the reliability and security of the disk array.
Smart Images

Figure CN120010793A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of disk array creation, and in particular to a disk array creation method, device, storage medium and program product. Background Art
[0002] RAID (Redundant Arrays of Independent Disks) is a data storage technology that combines multiple physical disks into one logical disk to improve data read and write speed and fault tolerance. Although disk arrays have a certain degree of fault tolerance, they only allow some disks to fail. If the number exceeds the specified number, data loss will occur. For example, RAID5 can allow one disk to fail without losing data.
[0003] Assume that there are 6 disks of the same size mounted on the disk array card, numbered 0 to 5. At this time, you want to create 2 RAID5-level disk arrays, each requiring 3 disks. There is no restriction on which disks to choose to form the disk array, and the probability of failure of each disk is different. The number of failed disks that can be allowed in each disk array is limited. If the number of failed disks exceeds the allowed value, data loss will occur in the disk array. Therefore, the probability of data loss of disk arrays composed of different combinations of disks is different.
[0004] For example, if the probability of failure of disk 0 and disk 1 is 0.08 and the probability of failure of the remaining disks is 0.02, a RAID5 disk array can allow at most one disk to fail without data loss. In this case, the probability of data loss for the combination [0,1,2] and [3,4,5] is higher than that for the combination [0,2,3] and [1,4,5]. Under the same disk conditions, choosing the combination [0,1,2] and [3,4,5] to create a disk array will bring a higher risk of data loss.
[0005] It can be seen that how to make the created disk array have a lower risk of data loss under the same disk conditions is a problem that those skilled in the art need to solve. Summary of the invention
[0006] The purpose of the embodiments of the present invention is to provide a disk array creation method, device, storage medium and program product, which can make the created disk array have a lower data loss risk under the same disk conditions. The specific scheme is as follows:
[0007] In a first aspect, the present invention provides a disk array creation method, comprising:
[0008] Obtain configuration information of the disk array to be created, and determine the failure probability of each disk mounted under the disk array card; the configuration information includes the total number of disk arrays to be created, the level of each disk array to be created, and the number of disks required for each;
[0009] Using the configuration information and based on the failure probability of each disk, each disk is grouped to obtain a target disk combination; the target disk combination is a combination with the smallest data loss probability among the plurality of disk combinations obtained by grouping; wherein each disk combination includes a sub-disk combination corresponding to each disk array to be created, and the data loss probability of each disk combination is a value determined based on the failure probability of the disks in each disk combination;
[0010] A disk array of a corresponding level is created by utilizing the level of each disk array to be created and based on each sub-disk combination in the target disk combination.
[0011] Optionally, determine the failure probability of each disk mounted on the disk array card, including:
[0012] Get the current status information of each disk mounted under the disk array card;
[0013] Based on the current status information of each disk, determine the current stage of each disk;
[0014] Based on the current stage of each disk, the failure probability of each disk is determined.
[0015] Optionally, based on the current status information of each disk, the current stage of each disk is determined, including:
[0016] Based on the number of times each storage particle in any disk is written and the erase / write usage rate, the current stage of any disk is determined;
[0017] Among them, any disk is any disk among the disks.
[0018] Optionally, based on the number of times each storage particle in any disk is written and the erase / write usage rate, determining the current stage of any disk includes:
[0019] Determine whether each storage particle in any disk has been written once;
[0020] If any storage particle in any disk has not been written once, it is determined that the current stage of any disk is an early stage;
[0021] If each storage particle in any disk has been written once, it is determined whether the erase and write utilization rate of each storage particle in any disk has reached a preset utilization rate;
[0022] If the erase / write usage rate of any storage particle in any disk does not reach the preset usage rate, it is determined that the current stage of any disk is a stable stage;
[0023] If the erase and write utilization rates of the storage particles in any disk reach the preset utilization rate, it is determined that the current stage of any disk is the late stage.
[0024] Optionally, based on the current stage of each disk, the failure probability of each disk is determined, including:
[0025] If the current stage of any disk is an early stage, determining the failure probability of any disk based on the first relative coefficient and a preset failure probability;
[0026] If the current stage of any disk is a stable stage, the preset failure probability is determined as the failure probability of any disk;
[0027] If the current stage of any disk is the late stage, determining the failure probability of any disk based on the second relative coefficient and the preset failure probability;
[0028] Among them, the first relative coefficient and the second relative coefficient are both greater than 1.
[0029] Optionally, determining whether each storage particle in any disk has been written once includes:
[0030] By determining whether the total written data volume of any disk is greater than the preset data volume, it is determined whether each storage particle in any disk has been written once;
[0031] The preset data volume is a value determined based on the number of storage particles in any disk and the unit write data volume of each storage particle.
[0032] Optionally, determining the failure probability of any disk based on the first relative coefficient and a preset failure probability includes:
[0033] Determining a first relative coefficient based on a ratio between a total written data amount of any disk and a preset data amount;
[0034] The failure probability of any disk is determined according to the first relative coefficient and the preset failure probability.
[0035] Optionally, determining whether each storage particle in any disk has been written once includes:
[0036] By determining whether the power-on time of any disk is greater than a preset time, it is determined whether each storage particle in any disk has been written once;
[0037] The preset time is a value determined based on the preset number of days and hours per day.
[0038] Optionally, determining the failure probability of any disk based on the first relative coefficient and a preset failure probability includes:
[0039] Determining a first relative coefficient based on a ratio between a power-on time and a preset time of any disk;
[0040] The failure probability of any disk is determined according to the first relative coefficient and the preset failure probability.
[0041] Optionally, determining the failure probability of any disk based on the second relative coefficient and a preset failure probability includes:
[0042] Determine a second relative coefficient based on the erase / write usage rate of each storage particle in any disk and a preset usage rate;
[0043] The failure probability of any disk is determined according to the second relative coefficient and the preset failure probability.
[0044] Optionally, determining the second relative coefficient based on the erase / write usage rate of each storage particle in any disk and a preset usage rate includes:
[0045] Determine an average usage rate based on the erase and write usage rates of each storage particle in any disk to obtain the erase and write usage rate of any disk;
[0046] The second relative coefficient is determined according to the difference between the erase / write usage rate of any disk and the preset usage rate.
[0047] Optionally, the process of determining the data loss probability for each disk combination includes:
[0048] Determine the data loss probability of any sub-disk combination based on the failure probability of the disk in any sub-disk combination; any sub-disk combination is any sub-disk combination among the sub-disk combinations included in each disk combination;
[0049] The data loss probability of each disk combination is determined according to the data loss probabilities corresponding to each sub-disk combination in each disk combination.
[0050] Optionally, determining the data loss probability of any sub-disk combination based on the failure probability of a disk in any sub-disk combination includes:
[0051] Determine a target disk array to be created corresponding to any sub-disk combination;
[0052] Based on the level corresponding to the target disk array to be created, the number of allowed failed disks is determined; the number of allowed failed disks indicates that data loss will occur when the number of failed disks of the target disk array to be created is greater than the number of allowed failed disks;
[0053] The data loss probability of any sub-disk combination is determined using the allowed number of failed disks and based on the failure probability of the disks in any sub-disk combination.
[0054] Optionally, determining the data loss probability of any sub-disk combination by using the allowed number of failed disks and based on the failure probability of disks in any sub-disk combination includes:
[0055] Determine each target quantity based on the number of disks required for the target disk array to be created and the number of disks that can fail; the target quantity is a positive integer that is not greater than the number of disks required for the target disk array to be created and greater than the number of disks that can fail;
[0056] Based on the failure probability of the disks in any sub-disk combination, determining the probability of failure of the target number of disks of the target disk array to be created;
[0057] The data loss probability of any sub-disk combination is determined according to the probability of failure of the target number of disks of the target disk array to be created.
[0058] Optionally, the disks are grouped based on the failure probability of each disk using the configuration information to obtain a target disk combination, including:
[0059] Based on the total number of disk arrays to be created, a number of disk creation tasks is created; the number of the number of disk creation tasks is one more than the total number of disk arrays to be created;
[0060] Using the first disk creation task among the plurality of disk creation tasks and based on the number of disks required by each disk array to be created, grouping the disks to obtain a current disk combination; the number of the first disk creation tasks is the same as the total number of the disk arrays to be created;
[0061] Determine the data loss probability of the current disk combination by using the second disk creation task among the plurality of disk creation tasks, determine a new current minimum data loss probability based on the data loss probability of the current disk combination and the current minimum data loss probability, and when the current grouping situation does not meet the preset end condition, jump to the step of using the first disk creation task among the plurality of disk creation tasks and grouping each disk based on the number of disks required for each disk array to be created, until the current grouping situation meets the preset end condition, so as to determine the target disk combination based on the disk combination corresponding to the latest current minimum data loss probability;
[0062] The current minimum data loss probability is initially a preset loss probability not less than 1.
[0063] Optionally, the first disk creation task among the plurality of disk creation tasks is used to group the disks based on the number of disks required by each disk array to be created, so as to obtain a current disk combination, including:
[0064] By means of each first disk creation task, disk grouping operations are performed in turn based on the number of disks required by the disk array to be created, so as to obtain the current remaining disks by using the current first disk creation task, and based on the number of disks required by the disk array to be created, a corresponding number of disks are taken out from the current remaining disks to obtain a sub-disk combination corresponding to the disk array to be created and update the current remaining disks; the disk array to be created currently corresponds to the current first disk creation task; the current remaining disks are initially the disks mounted under the disk array card;
[0065] When the current first disk creation task is the last disk creation task among the first disk creation tasks, the current disk combination is determined based on the sub-disk combinations corresponding to the disk arrays to be created respectively obtained by the first disk creation tasks.
[0066] Optionally, the preset end conditions include the current cumulative number of groupings reaching the maximum number of groupings or the current cumulative number of types of disk combinations reaching the maximum number of types; the maximum number of groupings and the maximum number of types are both values determined based on the maximum number of types, and the maximum number of types is the maximum number of types of disk combinations that can be obtained after grouping each disk.
[0067] In a second aspect, the present invention provides an electronic device, comprising:
[0068] Memory for storing computer programs;
[0069] The processor is used to execute the computer program to implement the steps of the above-mentioned disk array creation method.
[0070] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the aforementioned disk array creation method are implemented.
[0071] In a fourth aspect, the present invention provides a computer program product, comprising a computer program / instruction, which implements the steps of the aforementioned disk array creation method when executed by a processor.
[0072] In the present invention, configuration information of a disk array to be created is obtained, and the failure probability of each disk mounted under a disk array card is determined; the configuration information includes the total number of disk arrays to be created, the level of each disk array to be created, and the number of disks required for each; the configuration information is used and based on the failure probability of each disk, each disk is grouped to obtain a target disk combination; the target disk combination is a combination with the smallest data loss probability among several disk combinations obtained by grouping; wherein each disk combination includes sub-disk combinations corresponding to each disk array to be created, and the data loss probability of each disk combination is a value determined based on the failure probability of the disk in each disk combination; and the level of each disk array to be created is used and based on each sub-disk combination in the target disk combination to create a disk array of the corresponding level.
[0073] Beneficial effect: The present invention uses the configuration information of the disk array to be created and groups the disks based on the failure probability of each disk mounted under the disk array card, and determines the combination with the minimum data loss probability from the several disk combinations obtained by the grouping to obtain the target disk combination, thereby creating the corresponding disk array based on the sub-disk combination corresponding to each disk array to be created in the target disk combination. In this way, the present invention searches for the target disk combination with the minimum data loss probability from the grouping of each disk based on the failure probability of each disk mounted under the disk array card, thereby making the disk array created based on the target disk combination have the minimum data loss probability, and plays a role in improving the reliability and security of the disk array. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] In order to more clearly illustrate the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0075] Figure 1 A flow chart of a disk array creation method provided by an embodiment of the present invention;
[0076] Figure 2 A schematic diagram of a failure bathtub curve provided by an embodiment of the present invention;
[0077] Figure 3 A flow chart for determining the failure probability of a disk provided by an embodiment of the present invention;
[0078] Figure 4 A flowchart of creating a disk array provided by an embodiment of the present invention;
[0079] Figure 5 A structural diagram of an electronic device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0080] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0081] The terms "including" and "having" in the specification of the present invention and the above-mentioned drawings, as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but may include steps or units that are not listed.
[0082] In order to enable those skilled in the art to better understand the solution of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0083] When a disk array is constructed based on disks mounted on a disk array card, there is no restriction on which disks are selected to form the disk array, and the probability of failure of each disk is different, while the number of failed disks that can be allowed in each disk array is limited. If the number of failed disks exceeds the allowed value, data loss will occur in the disk array. Therefore, the probability of data loss of disk arrays composed of disks of different combinations is different, and the disk array with a higher probability of data loss is likely to cause more serious security problems. To this end, the present invention provides a disk array creation method, which searches for a target disk combination with the lowest probability of data loss from the grouping of each disk based on the failure probability of each disk mounted on the disk array card, so that the disk array created based on the target disk combination has the lowest probability of data loss, and plays a role in improving the reliability and security of the disk array.
[0084] See also Figure 1 As shown, an embodiment of the present invention discloses a method for creating a disk array, comprising:
[0085] Step S11, obtaining configuration information of the disk array to be created, and determining the failure probability of each disk mounted on the disk array card; the configuration information includes the total number of disk arrays to be created, the level of each disk array to be created, and the number of disks required for each.
[0086] The disk array creation method proposed in the embodiment of the present invention can be deployed on a host as an independent software, that is, the host obtains the configuration information of the disk array to be created input by the user, and determines the failure probability of each disk mounted on the disk array card, thereby determining the target disk combination with the minimum probability of data loss, so as to create the most reliable disk array.
[0087] In addition, the disk array creation method proposed in the embodiment of the present invention can be added as an optional function of the disk array card to the process of creating a disk array by the disk array card, that is, the disk array card obtains the configuration information of the disk array to be created sent by the host, and determines the failure probability of each disk mounted on the disk array card, thereby determining the target disk combination with the minimum probability of data loss, so as to create the most reliable disk array.
[0088] The acquisition of configuration information of the disk array to be created may specifically include: acquiring the total number of disk arrays to be created, the level of each disk array to be created, and the number of disks required for each disk array. In addition, the configuration information of the disk array to be created also includes the name of each disk array to be created.
[0089] It should be noted that the levels of the disk array to be created may include RAID0, RAID1, RAID5, RAID6, and RAID10. RAID0 requires at least 2 disks, and any disk failure will cause data loss, that is, the number of disks that RAID0 allows to fail is 0. RAID1 also requires at least 2 disks, but the number of disks must be an even number, and at most one of the disks can fail without causing data loss, that is, the number of disks that RAID1 allows to fail is 1. RAID5 requires at least 3 disks, and at most one of the disks can fail without causing data loss, that is, the number of disks that RAID5 allows to fail is 1. RAID6 requires at least 4 disks, and the number of disks must be greater than or equal to 4, and at most two of the disks can fail without causing data loss, that is, the number of disks that RAID6 allows to fail is 2. RAID10 also requires at least 4 disks, but the number of disks must be an even number greater than or equal to 4, and at most half of the disks can fail without causing data loss, that is, the number of disks that RAID10 allows to fail is the total number of disks of RAID10 / 2.
[0090] For example, assume that RAID10 consists of 8 disks. Since RAID10 allows at most half of the disks to fail without causing data loss, the number of disks that RAID10 allows to fail is equal to the total number of disks in RAID10 / 2=8 / 2=4, that is, the number of disks that RAID10 allows to fail is 4.
[0091] Considering that the probability of disk failure is related to the current stage of the disk, the failure probability of disks in different stages is different. In addition, since disks are electronic products, electronic products generally have Figure 2 The characteristics of the failure bathtub curve shown in FIG. 1 , that is, the failure probability of the disk is relatively high when the disk is in the early stage and the late stage, and the failure probability of the disk is relatively low when the disk is in the stable stage.
[0092] Based on this, determining the failure probability of each disk mounted on the disk array card may specifically include: obtaining current status information of each disk mounted on the disk array card, and based on the current status information of each disk, determining the current stage of each disk, and then based on the current stage of each disk, determining the failure probability of each disk.
[0093] It should be noted that the current status information of each disk can be obtained by using SMART (Self-Monitoring Analysis and Reporting Technology) technology to monitor the operating status of each disk in real time.
[0094] Since the disk is mainly composed of a controller and NAND (NOT AND) particles, NAND particles can also be recorded as storage particles. Therefore, the current status information of the disk includes but is not limited to the number of times each storage particle in the disk is written and the erase and write usage rate.
[0095] Specifically, in the process of determining the current stage of each disk based on the current status information of each disk, the current stage of any disk is determined based on the number of writes and erase usage rate of each storage particle in any disk; wherein any disk is any one of the disks.
[0096] It is understandable that for any disk among the disks mounted under the disk array card, the current stage of the disk can be determined based on the number of writes and the erase and write usage rate of each storage particle in the disk. The current stage includes an early stage, a stable stage, and a late stage.
[0097] It should be noted that when the disk is in the early stage, the quality of the disk controller and storage particles is the main cause of disk failure; when the disk is in the stable stage and the late stage, the aging of the disk components is the main cause of disk failure. Among them, the causes of disk component aging mainly include wear and aging caused by normal use of the disk under normal conditions, and also include accelerated aging caused by use of the disk under abnormal conditions; abnormal conditions include but are not limited to abnormal environmental factors, such as too high or too low temperature, excessive humidity, abnormal voltage, etc.
[0098] In the process of determining the current stage of any disk based on the number of times each storage particle in any disk has been written and the erase / write usage rate, it is first determined whether each storage particle in any disk has been written once; if any storage particle in any disk has not been written once, it is determined that the current stage of any disk is an early stage; if each storage particle in any disk has been written once, it is determined whether the erase / write usage rate of each storage particle in any disk has reached a preset usage rate; if the erase / write usage rate of any storage particle in any disk has not reached the preset usage rate, it is determined that the current stage of any disk is a stable stage; if the erase / write usage rate of each storage particle in any disk has reached the preset usage rate, it is determined that the current stage of any disk is a late stage.
[0099] It should be noted that when writing data to the disk, the data is written in a circular manner to each storage particle in the disk, that is, data is written in sequence starting from the first storage particle of the disk until the last storage particle is written, and after the last storage particle is written, data is returned to the first storage particle to continue writing data.
[0100] For example, taking a disk containing three storage particles as an example, data is first written to the first storage particle, then to the second storage particle, and then to the third storage particle, and then back to the first storage particle to write data, and so on, thereby achieving data writing to the disk.
[0101] Based on this, for the early stage, it is possible to determine whether the current stage of any disk is an early stage by judging whether each storage particle in any disk has been written once, that is, judging whether the number of times each storage particle in any disk has been written is greater than or equal to 1.
[0102] It should be noted that if some storage particles in any disk have not been written once, that is, the number of times some storage particles in any disk have been written is less than 1, it is determined that the current stage of any disk is an early stage.
[0103] In addition, if each storage particle in any disk has been written once, that is, the number of times each storage particle in any disk has been written is greater than or equal to 1, it indicates that any disk has passed the early stage, and it is necessary to further determine whether any disk is in the stable stage or the late stage.
[0104] As for the stable stage and the late stage, it is specifically determined whether the current stage of any disk is the stable stage or the late stage by judging whether the erase and write utilization rate of each storage particle in any disk reaches the preset utilization rate.
[0105] It should be noted that the erase and write utilization rate of storage particles is positively correlated with the aging degree of the disk to a certain extent, that is, the higher the erase and write utilization rate of storage particles, the more serious the aging degree of the disk. Generally speaking, when the erase and write utilization rate of storage particles reaches more than 90%, the disk enters the late stage from the stable stage, and when the erase and write utilization rate of storage particles reaches more than 95%, the disk is basically about to break down. Therefore, the preset utilization rate can be set with 90% as a reference.
[0106] Specifically, if the erase / write usage rate of some storage particles in any disk does not reach the preset usage rate, the current stage of any disk is determined to be the stable stage; if the erase / write usage rate of each storage particle in any disk reaches the preset usage rate, the current stage of any disk is determined to be the late stage.
[0107] After determining the current stage of each disk, it is necessary to further determine the failure probability of each disk based on the current stage of each disk. The failure probability of the disk is the probability of the disk failing.
[0108] Specifically, for any disk among the disks, if the current stage of any disk is an early stage, the failure probability of any disk is determined based on the first relative coefficient and the preset failure probability; if the current stage of any disk is a stable stage, the preset failure probability is directly determined as the failure probability of any disk; if the current stage of any disk is a late stage, the failure probability of any disk is determined based on the second relative coefficient and the preset failure probability; wherein, the first relative coefficient and the second relative coefficient are both greater than 1.
[0109] It should be noted that for the preset failure probability, the failure probability of several disks in different stages (early stage, stable stage, late stage) can be statistically analyzed to draw the following Figure 2The failure bathtub curve shown in FIG. 1 is used, and the preset failure probability is determined according to the failure probability of the failure bathtub curve in the stable stage. Specifically, the failure bathtub curve in the stable stage is uniformly sampled to obtain a preset number of sampling points, and the preset failure probability is determined based on the failure probability of the preset number of sampling points, for example, the average value of the failure probability of the preset number of sampling points is calculated, and the average value is determined as the preset failure probability.
[0110] In addition, a curve that is approximately smooth can be determined from the failure bathtub curve in the stable phase, and the preset failure probability can be determined based on the failure probability corresponding to the curve that is approximately smooth. Specifically, the curve that is approximately smooth is uniformly sampled to obtain a preset number of sampling points, and the preset failure probability can be determined based on the failure probability of the preset number of sampling points.
[0111] In this way, by determining the preset failure probability in the above manner, the selection of the preset failure probability can be made scientific and reasonable, and more in line with the characteristics of the disk in the stable stage, and ultimately play a role in improving the accuracy of determining the failure probability of the disk.
[0112] It should also be noted that since the failure probability of the disk in the early and late stages is generally greater than the failure probability of the disk in the stable stage, the embodiment of the present invention uses a first relative coefficient and a second relative coefficient greater than 1 to calculate the preset failure probability, so that the failure probability calculated in the early and late stages of the disk is greater than the preset failure probability, and correspondingly greater than the failure probability calculated when the disk is in the stable stage.
[0113] For determining whether each storage particle in any disk has been written once, in a specific implementation, it can specifically include: determining whether the total amount of written data in any disk is greater than a preset data amount, so as to determine whether each storage particle in any disk has been written once; wherein the preset data amount is a value determined based on the number of storage particles in any disk and the unit write data amount of each storage particle.
[0114] It can be understood that the total amount of data written to any disk is first obtained through the Data Units Written field in the SMART information of any disk, and the preset data amount is determined based on the number of storage particles in any disk and the unit amount of data written to each storage particle, wherein the unit amount of data written to each storage particle refers to the amount of data to be written to each storage particle at one time. Then, it is determined whether the total amount of data written to any disk is greater than the preset amount of data. If the total amount of data written to any disk is greater than the preset amount of data, it is determined that each storage particle in any disk has been written once. If the total amount of data written to any disk is less than or equal to the preset amount of data, it is determined that some storage particles in any disk have not been written once.
[0115] Furthermore, when the total amount of data written through any disk is less than or equal to the preset data amount and it is determined that some storage particles in any disk have not been written once, it is determined that the current stage of any disk is an early stage. At this time, it is necessary to determine the failure probability of any disk based on the first relative coefficient and the preset failure probability.
[0116] According to one of the embodiments, if the current stage of any disk is an early stage, a first relative coefficient can be determined based on the ratio between the total amount of written data of any disk and the preset amount of data, and the failure probability of any disk can be determined based on the first relative coefficient and the preset failure probability.
[0117] Specifically, the ratio between the total amount of data written to any disk and the preset amount of data is determined, and the difference between 1 and the ratio is determined, and then the product between the difference and 10 is determined as the first relative coefficient, and the failure probability of any disk is determined according to the product between the first relative coefficient and the preset failure probability. That is, the first relative coefficient = 10 × (1-the total amount of data written to any disk / the preset amount of data), and the failure probability of any disk = the first relative coefficient × the preset failure probability.
[0118] Regarding determining whether each storage particle in any disk has been written once, in another specific implementation, it may specifically include: determining whether the power-on time of any disk is greater than a preset time to determine whether each storage particle in any disk has been written once; wherein the preset time is a value determined based on a preset number of days and hours per day.
[0119] It can be understood that the power-on time of any disk is first obtained through the power on hours field in the SMART information of any disk, and the preset time is determined based on the preset number of days and the number of hours per day, wherein the power-on time and the preset time are both in hours, and the number of hours per day is set to 24 hours per day. Then, it is determined whether the power-on time of any disk is greater than the preset time. If the power-on time of any disk is greater than the preset time, it is determined that each storage particle in any disk has been written once. If the power-on time of any disk is less than or equal to the preset time, it is determined that some storage particles in any disk have not been written once.
[0120] The preset number of days may be set based on prior knowledge of the disk, for example, 90 days, in which case the preset time is 90×24=2160 hours.
[0121] Furthermore, when the power-on time of any disk is less than or equal to the preset time and it is determined that some storage particles in any disk have not been written once, it is determined that the current stage of any disk is an early stage. At this time, it is necessary to determine the failure probability of any disk based on the first relative coefficient and the preset failure probability.
[0122] According to one of the embodiments, if the current stage of any disk is an early stage, a first relative coefficient can be determined based on the ratio between the power-on time of any disk and the preset time, and the failure probability of any disk can be determined based on the first relative coefficient and the preset failure probability.
[0123] Specifically, the ratio between the power-on time of any disk and the preset time is determined, and the difference between 1 and the ratio is determined, and then the product between the difference and 10 is determined as the first relative coefficient, and the failure probability of any disk is determined according to the product between the first relative coefficient and the preset failure probability. That is, the first relative coefficient=10×(1-power-on time of any disk / preset time), and the failure probability of any disk=first relative coefficient×preset failure probability.
[0124] For determining whether each storage particle in any disk has been written once, in another specific implementation, it can specifically include: obtaining the number of written storage particles in any disk, and determining whether the number of written storage particles in any disk is equal to the total number of storage particles in any disk, so as to determine whether each storage particle in any disk has been written once.
[0125] Among them, if the number of written storage particles in any disk is less than the total number of storage particles in any disk, it is determined that some storage particles in any disk have not been written once; if the number of written storage particles in any disk is equal to the total number of storage particles in any disk, it is determined that all storage particles in any disk have been written once.
[0126] Furthermore, when the number of written storage particles in any disk is less than the total number of storage particles in any disk and it is determined that some storage particles in any disk have not been written once, it is determined that the current stage of any disk is an early stage. At this time, it is necessary to determine the failure probability of any disk based on the first relative coefficient and the preset failure probability.
[0127] According to one of the embodiments, if the current stage of any disk is an early stage, a first relative coefficient can be determined based on the ratio between the number of written storage particles in any disk and the total number of storage particles in any disk, and the failure probability of any disk can be determined based on the first relative coefficient and a preset failure probability.
[0128] Specifically, the ratio between the number of written storage particles in any disk and the total number of storage particles in any disk is determined, and the difference between 1 and the ratio is determined, and then the product between the difference and 10 is determined as the first relative coefficient, and the failure probability of any disk is determined according to the product between the first relative coefficient and the preset failure probability. That is, the first relative coefficient = 10 × (1-the number of written storage particles in any disk / the total number of storage particles in any disk), and the failure probability of any disk = the first relative coefficient × the preset failure probability.
[0129] For any disk among the disks, when the current stage of any disk is the late stage, the failure probability of any disk may be determined based on the product of a second relative coefficient greater than 1 and a preset failure probability.
[0130] Specifically, when the current stage of any disk is the late stage, the second relative coefficient is determined based on the erase and write usage rate of each storage particle in any disk and the preset usage rate, and the failure probability of any disk is determined based on the second relative coefficient and the preset failure probability.
[0131] In the process of determining the second relative coefficient based on the erase and write usage rate of each storage particle in any disk and the preset usage rate, the average usage rate is first determined based on the erase and write usage rate of each storage particle in any disk to obtain the erase and write usage rate of any disk, and then the second relative coefficient is determined based on the difference between the erase and write usage rate of any disk and the preset usage rate.
[0132] According to one of the embodiments, when the current stage of any disk is the late stage, the average value of the erase and write usage rate of each storage particle in any disk is calculated to obtain the average usage rate, and the average usage rate is determined as the erase and write usage rate of any disk, and then the difference between the erase and write usage rate of any disk and the preset usage rate is determined, and the ratio between the difference and 5 is determined, and the second relative coefficient is determined based on the product between the ratio and 10, and finally the failure probability of any disk is determined based on the product between the second relative coefficient and the preset failure probability. That is, the second relative coefficient = 10 × [(erase and write usage rate of any disk - preset usage rate) / 5], and the failure probability of any disk = the second relative coefficient × preset failure probability.
[0133] Step S12: using the configuration information and based on the failure probability of each disk, grouping each disk to obtain a target disk combination; the target disk combination is a combination with the smallest probability of data loss among the several disk combinations obtained by grouping; wherein each disk combination includes sub-disk combinations corresponding to each disk array to be created, and the data loss probability of each disk combination is a value determined based on the failure probability of the disks in each disk combination.
[0134] In an embodiment of the present invention, the configuration information of the disk array to be created is used and based on the failure probability of each disk mounted on the disk array card, the disks are grouped, and the combination with the lowest probability of data loss is determined from the several disk combinations obtained by the grouping to obtain the target disk combination.
[0135] Since the configuration information of the disk array to be created includes the total number of disk arrays to be created, each disk combination obtained by grouping includes sub-disk combinations corresponding to each disk array to be created, that is, the number of sub-disk combinations included in each disk combination is equal to the total number of disk arrays to be created, and different sub-disk combinations in each disk combination correspond to different disk arrays to be created.
[0136] Exemplarily, if the total number of disk arrays to be created is 2, each disk combination obtained by grouping includes 2 sub-disk combinations, and the sub-disk combinations correspond one-to-one to the disk arrays to be created.
[0137] Since the configuration information of the disk array to be created also includes the number of disks required for each disk array to be created, for the sub-disk combinations corresponding to each disk array to be created in each disk combination obtained by grouping, the number of disks included in the sub-disk combinations is the same as the number of disks required for the corresponding disk array to be created.
[0138] Exemplarily, if the total number of disk arrays to be created is 2, and the number of disks required for the first disk array to be created is 1, and the number of disks required for the second disk array to be created is 3, then each disk combination obtained by grouping includes 2 sub-disk combinations, one sub-disk combination corresponds to the first disk array to be created, and contains 2 disks; the other sub-disk combination corresponds to the second disk array to be created, and contains 3 disks.
[0139] In addition, since the configuration information of the disk array to be created includes the level of each disk array to be created, and disk arrays of different levels correspond to different tolerable numbers of failed disks, where the tolerable number of failed disks indicates that data loss occurs when the number of failed disks of the disk array is greater than the tolerable number of failed disks, that is, the data loss probability of the disk array is associated with the failure probability of the disks included in the disk array. Therefore, for each disk combination obtained by grouping, the data loss probability of each disk combination can be determined based on the failure probability of the disks included in each disk combination.
[0140] Exemplarily, if the allowed number of failed disks of the disk array is 1, it indicates that data loss occurs in the disk array when the number of failed disks is greater than 1, that is, the disk array allows at most one disk to fail without causing data loss.
[0141] It should be noted that the process of determining the data loss probability of each disk combination may specifically include: determining the data loss probability of any sub-disk combination based on the failure probability of the disk in any sub-disk combination; wherein any sub-disk combination is any sub-disk combination among the sub-disk combinations contained in each disk combination; and determining the data loss probability of each disk combination according to the data loss probabilities corresponding to each sub-disk combination in each disk combination.
[0142] It can be understood that, for any sub-disk combination among the sub-disk combinations contained in each disk combination, the data loss probability of any sub-disk combination can be determined based on the failure probability of the disks contained in any sub-disk combination, and after determining the data loss probability corresponding to each sub-disk combination in each disk combination, the data loss probability of each disk combination can be determined based on the data loss probability corresponding to each sub-disk combination in each disk combination.
[0143] In the process of determining the data loss probability of any sub-disk combination based on the failure probability of the disks in any sub-disk combination, the process can specifically include: determining the target disk array to be created corresponding to any sub-disk combination; determining the number of allowed failure disks based on the level corresponding to the target disk array to be created; wherein the number of allowed failure disks represents the data loss that occurs in the target disk array to be created when the number of failed disks is greater than the number of allowed failure disks; determining the data loss probability of any sub-disk combination using the number of allowed failure disks and based on the failure probability of the disks in any sub-disk combination.
[0144] Specifically, for determining the data loss probability of any sub-disk combination by using the allowed number of failed disks and based on the failure probability of the disks in any sub-disk combination, it can specifically include: determining each target number based on the number of disks required for the target disk array to be created and the allowed number of failed disks; wherein the target number is a positive integer not greater than the number of disks required for the target disk array to be created and greater than the allowed number of failed disks; based on the failure probability of the disks in any sub-disk combination, determining the probability of the target disk array to be created when the target number of disks fails, and determining the data loss probability of any sub-disk combination based on the probability of the target disk array to be created when the target number of disks fails.
[0145] Exemplarily, if the number of disks required for the target disk array to be created corresponding to any sub-disk combination is 3, and the number of disks to be allowed to fail corresponding to the target disk array to be created is 1, then according to the definition that the target number is a positive integer not greater than the number of disks required for the target disk array to be created and greater than the number of disks to be allowed to fail, the target numbers can be determined to be 2 and 3, and then based on the failure probability of the disks in any sub-disk combination, the probability of the target disk array to be created when 2 disks fail is determined, and the probability of the target disk array to be created when 3 disks fail is determined, and then the data loss probability of any sub-disk combination is determined based on the probability of the target disk array to be created when 2 disks fail and the probability when 3 disks fail.
[0146] Since the number of disks required for the target disk array to be created corresponding to any sub-disk combination is 3, any sub-disk combination contains three disks, denoted as disk 0, disk 1, and disk 2. Assuming that the failure probability of disk 0 is 0.2, the failure probability of disk 1 is 0.2, and the failure probability of disk 2 is 0.1, the probability of the target disk array to be created when two disks fail = the failure probability of disk 0 × the failure probability of disk 1 × the normal probability of disk 2 + the failure probability of disk 0 × the normal probability of disk 1 × the failure probability of disk 2 + the normal probability of disk 0 × the failure probability of disk 1 × the failure probability of disk 2 = 0.2 × 0.2 × 0.9 + 0.2 × 0.8 × 0.1 + 0 .8×0.2×0.1=0.036+0.016+0.016=0.068, the probability of the target disk array to be created having three disks failing = the probability of the failure of disk 0 × the probability of the failure of disk 1 × the probability of the failure of disk 2 = 0.2×0.2×0.1=0.004, the probability of data loss of any sub-disk combination = the probability of the target disk array to be created having two disks failing + the probability of the target disk array to be created having three disks failing = 0.068+0.004=0.072. It should be noted that the probability of a disk being normal = 1-the probability of a disk failing.
[0147] After determining the data loss probability corresponding to each sub-disk combination in each disk combination, the data loss probability of each disk combination is calculated by taking the example that each disk combination contains 2 sub-disk combinations, and the data loss probability of the first sub-disk combination is 0.1, and the data loss probability of the second sub-disk combination is 0.2.
[0148] In a specific implementation, the data loss probability of each disk combination=1-(1-data loss probability of the first sub-disk combination)×(1-data loss probability of the second sub-disk combination)=1-0.9×0.8=1-0.72=0.28.
[0149] In another specific implementation, the data loss probability of each disk combination = the data loss probability of the first sub-disk combination × (1-the data loss probability of the second sub-disk combination) + (1-the data loss probability of the first sub-disk combination) × the data loss probability of the second sub-disk combination + the data loss probability of the first sub-disk combination × the data loss probability of the second sub-disk combination = 0.1 × 0.8 + 0.9 × 0.2 + 0.1 × 0.2 = 0.08 + 0.18 + 0.02 = 0.28.
[0150] Based on the above-mentioned method of calculating the data loss probability of sub-disk combinations and the above-mentioned method of calculating the data loss probability of each disk combination, the data loss probability of each of the several disk combinations obtained by grouping can be calculated, so that the disk combination with the smallest data loss probability among the several disk combinations can be determined as the target disk combination.
[0151] It should be noted that, if there is more than one combination with the lowest probability of data loss among the several disk combinations, any one combination with the lowest probability of data loss among the several disk combinations may be determined as the target disk combination.
[0152] In the embodiment of the present invention, the target disk combination can be determined by first creating a plurality of disk creation tasks based on the total number of disk arrays to be created, and then using the plurality of disk creation tasks to determine the target disk combination, wherein the number of the plurality of disk creation tasks is one more than the total number of disk arrays to be created.
[0153] Specifically, create one more disk creation task than the total number of disk arrays to be created to obtain several disk creation tasks; use the first disk creation task in the several disk creation tasks and group each disk based on the number of disks required for each disk array to be created to obtain the current disk combination; wherein the number of the first disk creation tasks is the same as the total number of disk arrays to be created; determine the data loss probability of the current disk combination through the second disk creation task in the several disk creation tasks, determine a new current minimum data loss probability based on the data loss probability of the current disk combination and the current minimum data loss probability, and when the current grouping situation does not meet the preset end condition, jump back to the above step of grouping each disk using the first disk creation task in the several disk creation tasks and based on the number of disks required for each disk array to be created, until the current grouping situation meets the preset end condition, so as to determine the target disk combination based on the disk combination corresponding to the latest current minimum data loss probability. It should be noted that the current minimum data loss probability is a preset loss probability not less than 1 at the initial stage, and the number of the second disk creation tasks is one.
[0154] Among them, for the data loss probability based on the current disk combination and the current minimum data loss probability, a new current minimum data loss probability is determined. If the data loss probability of the current disk combination is less than the current minimum data loss probability, the data loss probability of the current disk combination is determined as the new current minimum data loss probability. If the data loss probability of the current disk combination is greater than or equal to the current minimum data loss probability, the current minimum data loss probability is kept unchanged to obtain a new current minimum data loss probability.
[0155] It can be understood that the first disk creation task is used and each disk is grouped based on the number of disks required for each disk array to be created to obtain a first disk combination, and then the data loss probability of the first disk combination is determined through the second disk creation task, and a new current minimum data loss probability is determined based on the data loss probability of the first disk combination and the current minimum data loss probability (at this time, a preset loss probability of not less than 1), and when the current grouping situation does not meet the preset end condition, the first disk creation task is reused and each disk is grouped based on the number of disks required for each disk array to be created to obtain a second disk combination, and the data loss probability of the second disk combination is determined again through the second disk creation task, and a new current minimum data loss probability is determined based on the data loss probability of the second disk combination and the current minimum data loss probability, and so on, until the current grouping situation meets the preset end condition, and finally the target disk combination is determined based on the disk combination corresponding to the latest current minimum data loss probability.
[0156] In the process of using the first disk creation task among several disk creation tasks and grouping each disk based on the number of disks required by each disk array to be created to obtain the current disk combination, each first disk creation task sequentially performs disk grouping operations based on the number of disks required by the corresponding disk array to be created, so as to obtain the current remaining disks using the current first disk creation task, and based on the number of disks required by the current disk array to be created, take out a corresponding number of disks from the current remaining disks to obtain the sub-disk combination corresponding to the current disk array to be created and update the current remaining disks; wherein different first disk creation tasks correspond to different disk arrays to be created, and the current disk array to be created corresponds to the current first disk creation task; and the current remaining disks are initially the disks mounted under the disk array card. Further, when the current first disk creation task is the last disk creation task among the first disk creation tasks, the current disk combination is determined based on the sub-disk combinations corresponding to the disk arrays to be created obtained by the first disk creation tasks.
[0157] It can be understood that the first disk creation task in each first disk creation task performs a disk grouping operation based on the number of disks required by the first disk array to be created corresponding to itself, so as to obtain the current remaining disks based on the disks mounted under the disk array card, and based on the number of disks required by the first disk array to be created, a corresponding number of disks are taken out from the current remaining disks to obtain a sub-disk combination corresponding to the first disk array to be created, and the current remaining disks are updated to disks that have not been taken out, and then the current remaining disks and the sub-disk combination corresponding to the first disk array to be created are sent to the second disk creation task in each first disk creation task. The second disk creation task performs a disk grouping operation based on the number of disks required by the second disk array to be created corresponding to itself, so as to take out a corresponding number of disks from the current remaining disks based on the number of disks required by the second disk array to be created, so as to obtain a sub-disk combination corresponding to the second disk array to be created, and the current remaining disks are updated to disks that have not been taken out, and then the current remaining disks and the sub-disk combination corresponding to the first disk array to be created and the sub-disk combination corresponding to the second disk array to be created are sent to the next disk creation task in each first disk creation task. The same process is repeated until the last disk creation task in each first disk creation task completes the disk grouping operation, and the current disk combination is determined based on the sub-disk combinations corresponding to each disk array to be created obtained by each first disk creation task.
[0158] It should be noted that the preset end conditions may include both the current cumulative number of groupings reaching the maximum number of groupings and the current cumulative number of types of disk combinations reaching the maximum number of types; wherein the maximum number of groupings and the maximum number of types are both values determined based on the maximum number of types, and the maximum number of types is the maximum number of types of disk combinations that can be obtained after grouping each disk.
[0159] For example, if the number of disks mounted on the disk array card is 3, the total number of disk arrays to be created is 2, and the number of disks required for one of the disk arrays to be created is 1, and the number of disks required for the other disk array to be created is 2, then after grouping the disks mounted on the disk array card, there are at most 3 disk combinations that can be obtained, such as [1] and [2, 3], [2] and [1, 3], and [3] and [1, 2]. In this case, the maximum number of types is 3. Accordingly, the maximum number of groupings and the maximum number of types are also 3.
[0160] Step S13: Create a disk array of a corresponding level by using the level of each disk array to be created and based on each sub-disk combination in the target disk combination.
[0161] In the embodiment of the present invention, after determining the target disk combination, a disk array of a corresponding level is created using the level of each disk array to be created and based on the sub-disk combinations in the target disk combination corresponding to each disk array to be created.
[0162] Beneficial effect: The present invention uses the configuration information of the disk array to be created and groups the disks based on the failure probability of each disk mounted under the disk array card, and determines the combination with the minimum data loss probability from the several disk combinations obtained by the grouping to obtain the target disk combination, thereby creating the corresponding disk array based on the sub-disk combination corresponding to each disk array to be created in the target disk combination. In this way, the present invention searches for the target disk combination with the minimum data loss probability from the grouping of each disk based on the failure probability of each disk mounted under the disk array card, thereby making the disk array created based on the target disk combination have the minimum data loss probability, and plays a role in improving the reliability and security of the disk array.
[0163] See also Figure 3 and Figure 4 As shown, an embodiment of the present invention discloses a method for creating a disk array, comprising:
[0164] Get the SMART information of each disk mounted on the disk array card. For any disk mounted on the disk array card, get the total amount of data written to any disk through the Data Units Written field in the SMART information of any disk, and get the power-on time of any disk through the power on hours field in the SMART information of any disk.
[0165] It is determined whether the total amount of data written to any disk is greater than a preset amount of data, and it is determined whether the power-on time of any disk is greater than a preset time.
[0166] If the total amount of data written to any disk is less than or equal to the preset data amount, or the power-on time of any disk is less than or equal to the preset time, then the current stage of any disk is determined to be the early stage, and the first relative coefficient = 10×(1-total amount of data written to any disk / preset data amount), or the first relative coefficient = 10×(1-power-on time of any disk / preset time) is determined, and then the failure probability of any disk is determined based on the product of the first relative coefficient and the preset failure probability.
[0167] If the total written data volume of any disk is greater than the preset data volume, and the power-on time of any disk is greater than the preset time, it is determined whether the erase and write utilization rate of each storage particle in any disk reaches the preset utilization rate.
[0168] If the erase / write usage rate of any storage particle in any disk does not reach the preset usage rate, the current stage of any disk is determined to be a stable stage, and the preset failure probability is directly determined as the failure probability of any disk.
[0169] If the erase and write usage rates of each storage particle in any disk reaches the preset usage rate, it is determined that the current stage of any disk is the late stage, and the average usage rate is determined based on the erase and write usage rates of each storage particle in any disk, so as to determine the average usage rate as the erase and write usage rate of any disk, and then determine the second relative coefficient = 10 × [(erase and write usage rate of any disk - preset usage rate) / 5], and determine the failure probability of any disk based on the product of the second relative coefficient and the preset failure probability.
[0170] After obtaining the configuration information of the disk array to be created (the total number of disk arrays to be created, the level of each disk array to be created, and the number of disks required for each disk array to be created), and determining the failure probability of each disk mounted on the disk array card, based on the total number of disk arrays to be created, create disk creation tasks that are one more than the total number of disk arrays to be created, so as to obtain a plurality of disk creation tasks. The plurality of disk creation tasks include at least one first disk creation task and one second disk creation task, the number of the at least one first disk creation task is the same as the total number of disk arrays to be created, and different first disk creation tasks correspond to different disk arrays to be created.
[0171] In the process of determining the current disk combination based on the first disk creation task, each first disk creation task sequentially performs disk grouping operations based on the number of disks required by the corresponding disk array to be created, so as to use the current first disk creation task to obtain the current remaining disks, and based on the number of disks required by the current disk array to be created, take out a corresponding number of disks from the current remaining disks to obtain the sub-disk combination corresponding to the current disk array to be created and update the current remaining disks; wherein the current disk array to be created is the disk array to be created corresponding to the current first disk creation task, and the current remaining disks are initially the various disks mounted under the disk array card. Further, when the current first disk creation task is the last disk creation task among the first disk creation tasks, the current disk combination is determined based on the sub-disk combinations corresponding to the disk arrays to be created obtained by each first disk creation task.
[0172] The data loss probability of the current disk combination is determined by the second disk creation task, and a new current minimum data loss probability is determined based on the data loss probability of the current disk combination and the current minimum data loss probability. When the current grouping situation does not meet the preset end condition, the process of determining the current disk combination based on the first disk creation task is jumped again until the current grouping situation meets the preset end condition, so as to determine the target disk combination based on the disk combination corresponding to the latest current minimum data loss probability. It should be noted that the current minimum data loss probability is initially a preset loss probability not less than 1.
[0173] After the target disk combination is determined, a disk array of a corresponding level is created by using the level of each disk array to be created and based on the sub-disk combinations in the target disk combination that respectively correspond to each disk array to be created.
[0174] Beneficial effect: The present invention uses the configuration information of the disk array to be created and groups the disks based on the failure probability of each disk mounted under the disk array card, and determines the combination with the minimum data loss probability from the several disk combinations obtained by the grouping to obtain the target disk combination, thereby creating the corresponding disk array based on the sub-disk combination corresponding to each disk array to be created in the target disk combination. In this way, the present invention searches for the target disk combination with the minimum data loss probability from the grouping of each disk based on the failure probability of each disk mounted under the disk array card, thereby making the disk array created based on the target disk combination have the minimum data loss probability, and plays a role in improving the reliability and security of the disk array.
[0175] Furthermore, the present application also discloses an electronic device. Figure 5 This is a structural diagram of an electronic device according to an exemplary embodiment, and the content in the diagram cannot be considered as any limitation on the scope of use of this application. The electronic device may specifically include: at least one processor 11, at least one memory 12, a power supply 13, a communication interface 14, an input / output interface 15, and a communication bus 16. Among them, the memory 12 is used to store a computer program, and the computer program is loaded and executed by the processor 11 to implement the relevant steps in the disk array creation method disclosed in any of the aforementioned embodiments. In addition, the electronic device in this embodiment may specifically be an electronic computer.
[0176] In this embodiment, the power supply 13 is used to provide working voltage for each hardware device on the electronic device; the communication interface 14 can create a data transmission channel between the electronic device and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 15 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0177] In addition, the memory 12, as a carrier for storing resources, can be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon can include an operating system 121, a computer program 122, etc., and the storage method can be temporary storage or permanent storage.
[0178] The operating system 121 is used to manage and control various hardware devices on the electronic device and the computer program 122, which can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the disk array creation method performed by the electronic device disclosed in any of the aforementioned embodiments, the computer program 122 can further include a computer program that can be used to complete other specific tasks.
[0179] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein the computer program, when executed by a processor, implements the disk array creation method disclosed above. The specific steps of the method can refer to the corresponding contents disclosed in the above embodiments, and will not be repeated here.
[0180] Furthermore, the present application also discloses a computer program product, including a computer program / instruction; wherein, when the computer program / instruction is executed by a processor, the disk array creation method disclosed above is implemented. For the specific steps of the method, reference may be made to the corresponding contents disclosed in the above embodiments, and no further description will be given here.
[0181] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0182] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0183] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0184] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0185] The technical solution provided by the present application is introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technical personnel in this field, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A method for creating a disk array, characterized in that: include: Obtaining configuration information of the disk array to be created, and determining the failure probability of each disk mounted under the disk array card; the configuration information includes the total number of the disk arrays to be created, the level of each of the disk arrays to be created, and the number of disks required for each; Using the configuration information and based on the failure probability of each disk, the disks are grouped to obtain a target disk combination; the target disk combination is a combination with the smallest data loss probability among the plurality of disk combinations obtained by grouping; wherein each disk combination includes sub-disk combinations corresponding to each disk array to be created, and the data loss probability of each disk combination is a value determined based on the failure probability of the disks in each disk combination; A disk array of a corresponding level is created by utilizing the level of each disk array to be created and based on each sub-disk combination in the target disk combination.
2. The disk array creation method according to claim 1, characterized in that: Determining the failure probability of each disk mounted on the disk array card includes: Get the current status information of each disk mounted under the disk array card; Based on the current status information of each disk, determining the current stage of each disk; Based on the current stage of each disk, the failure probability of each disk is determined.
3. The disk array creation method according to claim 2, characterized in that: The determining the current stage of each disk based on the current status information of each disk includes: Determine the current stage of any disk based on the number of times each storage particle in any disk is written and the erase / write usage rate; The any disk is any one of the disks.
4. The disk array creation method according to claim 3, characterized in that: The determining the current stage of any disk based on the number of times each storage particle in any disk is written and the erase / write usage rate includes: Determine whether each storage particle in any of the disks has been written once; If any storage particle in any of the disks has not been written once, it is determined that the current stage of any of the disks is an early stage; If each storage particle in any of the disks has been written once, determining whether the erase and write usage rate of each storage particle in any of the disks has reached a preset usage rate; If the erase / write usage rate of any storage particle in any disk does not reach the preset usage rate, determining that the current stage of any disk is a stable stage; If the write-erase usage rate of each storage particle in any one of the disks reaches a preset usage rate, it is determined that the current stage of any one of the disks is a late stage.
5. The disk array creation method according to claim 4, characterized in that: The determining the failure probability of each disk based on the current stage of each disk includes: If the current stage of any disk is an early stage, determining the failure probability of any disk based on the first relative coefficient and a preset failure probability; If the current stage of any disk is a stable stage, the preset failure probability is determined as the failure probability of any disk; If the current stage of any disk is a late stage, determining the failure probability of any disk based on the second relative coefficient and the preset failure probability; Wherein, the first relative coefficient and the second relative coefficient are both greater than 1.
6. The disk array creation method according to claim 5, characterized in that: The determining whether each storage particle in any one of the disks has been written once includes: By determining whether the total amount of written data of any disk is greater than a preset amount of data, it is determined whether each storage particle in any disk has been written once; The preset data volume is a value determined based on the number of storage particles in any one of the disks and the unit write data volume of each storage particle.
7. The disk array creation method according to claim 6, characterized in that: The determining the failure probability of any disk based on the first relative coefficient and the preset failure probability includes: Determining a first relative coefficient based on a ratio between a total amount of written data of any one of the disks and the preset amount of data; The failure probability of any one of the disks is determined according to the first relative coefficient and a preset failure probability.
8. The disk array creation method according to claim 5, characterized in that: The determining whether each storage particle in any one of the disks has been written once includes: By determining whether the power-on time of any disk is greater than a preset time, it is determined whether each storage particle in any disk has been written once; The preset time is a value determined based on the preset number of days and hours per day.
9. The disk array creation method according to claim 8, characterized in that: The determining the failure probability of any disk based on the first relative coefficient and the preset failure probability includes: determining a first relative coefficient based on a ratio between the power-on time of any one of the disks and the preset time; The failure probability of any one of the disks is determined according to the first relative coefficient and a preset failure probability.
10. The disk array creation method according to claim 5, characterized in that: The determining the failure probability of any disk based on the second relative coefficient and the preset failure probability comprises: Determining a second relative coefficient based on the erase / write usage rate of each storage particle in any one of the magnetic disks and the preset usage rate; The failure probability of any one of the disks is determined according to the second relative coefficient and the preset failure probability.
11. The disk array creation method according to claim 10, characterized in that: The determining of the second relative coefficient based on the erase / write usage rate of each storage particle in any one of the magnetic disks and the preset usage rate includes: Determine an average usage rate based on the erase and write usage rate of each storage particle in any one of the disks to obtain the erase and write usage rate of any one of the disks; A second relative coefficient is determined according to a difference between the erase / write usage rate of any one of the magnetic disks and the preset usage rate.
12. The disk array creation method according to claim 1, characterized in that: The process of determining the probability of data loss for each disk combination includes: Determine the data loss probability of any sub-disk combination based on the failure probability of the disk in the sub-disk combination; the sub-disk combination is any sub-disk combination among the sub-disk combinations included in each disk combination; The data loss probability of each disk combination is determined according to the data loss probabilities corresponding to each sub-disk combination in each disk combination.
13. The disk array creation method according to claim 12, characterized in that: The determining of the data loss probability of any sub-disk combination based on the failure probability of the disk in any sub-disk combination includes: Determine a target disk array to be created corresponding to any sub-disk combination; Based on the level corresponding to the target disk array to be created, determining the number of allowed failed disks; the number of allowed failed disks indicates that data loss occurs when the number of failed disks of the target disk array to be created is greater than the number of allowed failed disks; The data loss probability of any sub-disk combination is determined by using the allowed number of failed disks and based on the failure probability of the disks in any sub-disk combination.
14. The disk array creation method according to claim 13, characterized in that: The determining the data loss probability of any sub-disk combination by using the number of allowed failure disks and based on the failure probability of the disks in any sub-disk combination comprises: Determine each target quantity based on the number of disks required for the target disk array to be created and the number of disks to be allowed to fail; the target quantity is a positive integer that is not greater than the number of disks required for the target disk array to be created and greater than the number of disks to be allowed to fail; Based on the failure probability of the disks in any one of the sub-disk combinations, determining the probability of the target disk array to be created failing when the target number of disks fails; The data loss probability of any sub-disk combination is determined according to the probability of the target disk array to be created failing when the target number of disks fails.
15. The disk array creation method according to any one of claims 1 to 14, characterized in that: The utilizing the configuration information and based on the failure probability of each disk to group the disks to obtain a target disk combination includes: Based on the total number of disk arrays to be created, creating a number of disk creation tasks; the number of the number of disk creation tasks is one more than the total number of disk arrays to be created; Using the first disk creation task among the plurality of disk creation tasks and based on the number of disks required by each disk array to be created, grouping the disks to obtain a current disk combination; the number of the first disk creation tasks is the same as the total number of the disk arrays to be created; Determine the data loss probability of the current disk combination by the second disk creation task among the plurality of disk creation tasks, determine a new current minimum data loss probability based on the data loss probability of the current disk combination and the current minimum data loss probability, and when the current grouping situation does not meet the preset end condition, jump to the step of grouping the disks by using the first disk creation task among the plurality of disk creation tasks and based on the number of disks required for each disk array to be created, until the current grouping situation meets the preset end condition, so as to determine the target disk combination based on the disk combination corresponding to the latest current minimum data loss probability; The current minimum data loss probability is initially a preset loss probability not less than 1.
16. The disk array creation method according to claim 15, characterized in that: The utilizing the first disk creation task among the plurality of disk creation tasks and grouping the disks based on the number of disks required by the disk arrays to be created to obtain a current disk combination includes: By means of each of the first disk creation tasks, disk grouping operations are performed in sequence based on the number of disks required by the disk array to be created, respectively, so as to obtain the current remaining disks by using the current first disk creation task, and based on the number of disks required by the disk array to be created, a corresponding number of disks are taken out from the current remaining disks to obtain a sub-disk combination corresponding to the disk array to be created and update the current remaining disks; the current disk array to be created corresponds to the current first disk creation task; the current remaining disks are initially the disks mounted under the disk array card; When the current first disk creation task is the last disk creation task among the first disk creation tasks, the current disk combination is determined based on the sub-disk combinations corresponding to the disk arrays to be created respectively obtained by the first disk creation tasks.
17. The disk array creation method according to claim 16, characterized in that: The preset end conditions include that the current cumulative number of groupings reaches the maximum number of groupings or the current cumulative number of types of disk combinations reaches the maximum number of types; the maximum number of groupings and the maximum number of types are both values determined based on the maximum number of types, and the maximum number of types is the maximum number of types of disk combinations that can be obtained after grouping the disks.
18. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to execute the computer program to implement the steps of the disk array creation method according to any one of claims 1 to 17.
19. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the disk array creation method according to any one of claims 1 to 17 are implemented.
20. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the disk array creation method described in any one of claims 1 to 17 are implemented.
Citation Information
Patent Citations
Storage resource management method and device
CN103617006A
Mounting method of redundant arrays of inexpensive disks (RAID), Android equipment and storage medium
CN107728946A
Solid state disk error correction method and device, equipment and medium
CN117908798A
Equipment fault prediction model determination method and device, equipment, storage medium and product
CN119473809A
Grouping of storage media based on parameters associated with the storage media
US20050044313A1