Data storage method and device, storage medium and electronic equipment
By setting the number of data chunks of data packets according to the strip width and disk number in the RAID6 disk array, and determining the storage disks for each data chunk, the problem of low disk storage performance caused by uneven distribution of data chunks is solved, and the uniform distribution of data volume and storage performance are improved.
Patent Information
- Application Number
- CN202412000517.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-12-31
AI Technical Summary
Due to the uneven distribution of data chunking in the RAID6 disk array, the disk storage performance is low, and the existing technology has failed to effectively solve this problem.
By setting the number of data chunks of a data packet based on the product of the stripe width and the number of disks, and determining the disk corresponding to each data chunk, the pre-distribution result of the data chunks of the data packet is generated, and the data chunks are finally stored evenly on multiple disks.
The data volume of data packets is evenly distributed on multiple disks, which improves disk storage performance, simplifies spatial distribution algorithms, and reduces process complexity and time-consuming conversion of storage address algorithms.
Smart Images

Figure CN120045127A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of storage technologies, and in particular, to a data storage method, apparatus, storage medium, and electronic device. Background Art
[0002] RAID6 (Redundant Array of Independent Disks with Double Parity) is usually an array composed of 4 - 16 disks. A stripe contains check block p, check block q, and several data blocks. Check block q is usually larger than other blocks. Therefore, in a stripe, the block data is not evenly distributed on multiple disks. Thus, the concept of a packet is introduced. Each packet consists of the number of disks * the stripe width blocks. The block data in each packet is evenly distributed on multiple disks, and the spatial distribution of each packet is the same. Since block q is larger than other blocks, it is also required that the spatial distribution of each packet is aligned, so the spatial distribution algorithm is complex. This further causes the complexity of the process and the time-consuming conversion of the spatial distribution address algorithm. It greatly affects the disk storage performance.
[0003] Therefore, in the related art, there is a problem of how to improve the disk storage performance.
[0004] In view of the problem in the related art of how to improve the disk storage performance, no effective solution has been proposed yet. Summary of the Invention
[0005] The embodiments of the present application provide a data storage method, apparatus, storage medium, and electronic device to at least solve the problem of how to improve the disk storage performance in the related art.
[0006] According to an embodiment of the present application, a data storage method is provided, which is applied to an independent disk array with double parity check. The independent disk array includes multiple disks, and the multiple disks are used to store data chunks of a stripe. The stripe width of the stripe is set based on the number of disks in the independent disk array. The data chunks of the stripe are respectively stored on different disks. The data chunks of the stripe include a first parity block, a second parity block, and multiple data blocks. The data volume of the first parity block is greater than that of the second parity block, and the data volume of the second parity block is the same as that of each data block. The method includes: setting the number of data chunks of a data packet according to the product of the stripe width and the number of disks, and the data chunks of the data packet come from different disks; determining the disk corresponding to each data chunk among the data chunks in multiple stripes, and generating a preliminary distribution result of the data chunks of the data packet according to each data chunk and the disk corresponding to each data chunk, wherein the total number of data chunks of the multiple stripes is equal to the number of data chunks of the data packet; storing each data chunk into the disk corresponding to each data chunk according to the preliminary distribution result, so that the data volume of the data packet is evenly distributed among the multiple disks.
[0007] In an exemplary embodiment, determining the disk corresponding to each data chunk among the data chunks in multiple stripes, and generating a preliminary distribution result of the data chunks of the data packet according to each data chunk and the disk corresponding to each data chunk includes: obtaining the stripe number of a target stripe among the multiple stripes, wherein the initial value of the stripe number is 0 and the increment is 1; determining the correspondence between the data chunks of the target stripe and the disks according to the stripe number of the target stripe and the number of disks; determining the preliminary distribution result according to the correspondence between the data chunks of the multiple stripes and the disks.
[0008] In an exemplary embodiment, determining the correspondence between the data chunks of the target stripe and the disks according to the stripe number of the target stripe and the number of disks includes: obtaining the disk number of the target disk among the multiple disks, where the initial value of the disk number is 0 and the increment is 1; performing a modulo operation on the number of disks according to the stripe number of the target stripe to obtain a first target value; calculating the difference between the disk number and the first target value to obtain a second target value; in the case where it is determined that the second target value is a non - negative number, determining the second target value as a third target value; or, in the case where it is determined that the second target value is a negative number, determining the sum value of the second target value and the number of disks as the third target value; in the case where it is determined that the third target value is 0, determining the target disk as the disk corresponding to the first parity block in the target stripe; or, in the case where it is determined that the third target value is 1, determining the target disk as the disk corresponding to the second parity block in the target stripe; or, in the case where it is determined that the third target value is greater than 1, determining the target disk as the disk corresponding to the data block in the target stripe.
[0009] In an exemplary embodiment, determining the correspondence between the data chunks of the target stripe and the disks according to the stripe number of the target stripe and the number of disks further includes: determining a fourth target value according to the sum value of a first preset value and a target product, where the target product represents the product of a second preset value and the stripe number; in the case where it is determined that a fifth target value is a non - negative number, determining the disk with the disk number being the fifth target value among the multiple disks as the disk corresponding to the first parity block in the target stripe, where the fifth target value represents the difference between the number of disks and the fourth target value; and, determining the disk with the disk number being a sixth target value among the multiple disks as the disk corresponding to the second parity block in the target stripe, where the sixth target value represents the difference between the fifth target value and the first preset value; and, determining the disks among the multiple disks whose disk numbers are not the fifth target value and the sixth target value as the disks corresponding to the data blocks in the target stripe.
[0010] In an exemplary embodiment, the method further includes: in the case where it is determined that the fifth target value is a negative number, determining the parity of the number of disks; determining the correspondence between the data chunks of the target stripe and the disks according to the fifth target value and the parity of the number of disks.
[0011] In an exemplary embodiment, determining the correspondence between the data chunks of the target stripe and the disks according to the fifth target value and the parity of the number of disks includes: when determining that the number of disks is odd, determining the disk with the seventh target value among the multiple disks as the disk corresponding to the first parity block in the target stripe, where the seventh target value represents the sum value of the number of disks and the fifth target value; and determining the disk with the eighth target value among the multiple disks as the disk corresponding to the second parity block in the target stripe, where the eighth target value represents the difference value between the seventh target value and the first preset value; and determining the disks among the multiple disks whose disk numbers are not the seventh target value and the eighth target value as the disks corresponding to the data blocks in the target stripe.
[0012] In an exemplary embodiment, determining the correspondence between the data chunks of the target stripe and the disks according to the fifth target value and the parity of the number of disks further includes: when determining that the number of disks is even, determining the disk with the ninth target value among the multiple disks as the disk corresponding to the second parity block in the target stripe, where the ninth target value represents the sum value of the number of disks and the fifth target value; and determining the disk with the tenth target value among the multiple disks as the disk corresponding to the first parity block in the target stripe, where the tenth target value represents the difference value between the ninth target value and the first preset value; and determining the disks among the multiple disks whose disk numbers are not the ninth target value and the tenth target value as the disks corresponding to the data blocks in the target stripe.
[0013] According to another embodiment of the present application, a data storage device is provided, which is applied to an independent disk array with double parity check. The independent disk array includes a plurality of disks, and the plurality of disks are used to store data chunks of a stripe. The stripe width of the stripe is set based on the number of disks in the independent disk array. The data chunks of the stripe are respectively stored on different disks. The data chunks of the stripe include a first check block, a second check block, and a plurality of data blocks. The data volume of the first check block is greater than that of the second check block, and the data volume of the second check block is the same as that of each data block. The device includes: a setting module, configured to set the number of data chunks of a data packet according to the product of the stripe width and the number of disks, and the data chunks of the data packet come from different disks; a generating module, configured to determine the disk corresponding to each data chunk among the data chunks in a plurality of stripes, and generate a pre-distribution result of the data chunks of the data packet according to each data chunk and the disk corresponding to each data chunk, wherein the total number of data chunks of the plurality of stripes is equal to the number of data chunks of the data packet; a storage module, configured to store each data chunk into the disk corresponding to each data chunk according to the pre-distribution result, so that the data volume of the data packet is evenly distributed on the plurality of disks.
[0014] According to still another embodiment of the present application, a computer-readable storage medium is further provided. A computer program is stored in the computer-readable storage medium. The computer program is configured to execute the steps in any one of the above method embodiments when running.
[0015] According to still another embodiment of the present application, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0016] According to still another embodiment of the present application, a computer program product is further provided, including a computer program. The computer program implements the steps in the above method embodiments when executed by a processor.
[0017] Through the present application, the number of data chunks of a data packet can be set according to the product of the stripe width and the number of disks, then the disk corresponding to each data chunk in the data packet is determined to generate a pre-distribution result of the data chunks of the data packet, and finally each data chunk is stored into the corresponding disk according to the pre-distribution result, so that the data volume of the data packet is evenly distributed on the plurality of disks. Furthermore, the problem of how to improve the disk storage performance in the related art can be solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1It is a hardware structure block diagram of a server device for a data storage method according to an embodiment of the present application;
[0019] Figure 2 It is a flowchart of a data storage method according to an embodiment of the present application;
[0020] Figure 3 It is a schematic flowchart of a data storage method according to an embodiment of the present application;
[0021] Figure 4 It is a structure block diagram of a data storage device according to an embodiment of the present application. Detailed implementation manners
[0022] In the following, embodiments of the present application will be described in detail with reference to the drawings and in conjunction with the embodiments.
[0023] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence.
[0024] The method embodiments provided in the embodiments of the present application can be executed on a server device or a similar computing device. Taking running on a server device as an example, Figure 1 It is a hardware structure block diagram of a server device for a data storage method according to an embodiment of the present application. As Figure 1 shown, the server device may include one or more ( Figure 1 only one is shown in Figure 1 processors 102 (the processors 102 may include, but are not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above-mentioned server device may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown in Figure 1 is only schematic and does not limit the structure of the above-mentioned server device. For example, the server device may further include more or fewer components than
[0025] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the data storage method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above-mentioned data storage method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the server device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0026] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the server device. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0027] In this embodiment, a data storage method is provided, which is applied to a redundant array of independent disks with double parity check. The redundant array of independent disks includes a plurality of disks, and the plurality of disks are used to store data chunks of a stripe. The stripe width of the stripe is set based on the number of disks in the redundant array of independent disks. The data chunks of the stripe are respectively stored on different disks. The data chunks of the stripe include a first check block, a second check block, and a plurality of data blocks. The data volume of the first check block is greater than the data volume of the second check block, and the data volume of the second check block is the same as the data volume of each data block. Figure 2 It is a flowchart of the data storage method according to the embodiments of the present application, as Figure 2 shown, and this process includes the following steps:
[0028] Step S202, set the number of data chunks of a data packet according to the product of the stripe width and the number of disks. The data chunks of the data packet come from different disks;
[0029] Step S204: Determine the disk corresponding to each data block in the data blocks of multiple stripes, and generate a preliminary distribution result of the data blocks of the data packet according to each data block and the disk corresponding to each data block, where the total number of data blocks of the multiple stripes is equal to the number of data blocks of the data packet;
[0030] Step S206: Store each data block into the disk corresponding to each data block according to the preliminary distribution result, so that the data volume of the data packet is evenly distributed among the multiple disks.
[0031] In an alternative embodiment, assume that there are 6 disks in an independent disk array of RAID6, then the stripe width is also 6, that is, there are 6 data blocks in one stripe, specifically including parity block q (equivalent to the first parity block), parity block p (equivalent to the second parity block) and 4 data blocks, and the 6 data blocks are respectively distributed on 6 disks. There are a total of 36 data blocks in one data packet. Since parity block q is larger than other data blocks, the data volume of one stripe is not evenly distributed on 6 disks. However, in one data packet, by reasonably distributing the data blocks, it is possible to achieve the same amount of data stored on each disk.
[0032] Through the above steps, the number of data blocks of the data packet can be set according to the product of the stripe width and the number of disks, then determine the disk corresponding to each data block in the data packet to generate a preliminary distribution result of the data blocks of the data packet, and finally store each data block into the corresponding disk according to the preliminary distribution result, so that the data volume of the data packet is evenly distributed among the multiple disks. Furthermore, it can solve the problem of how to improve the disk storage performance in the related art.
[0033] In an exemplary embodiment, determining the disk corresponding to each data block in the data blocks of multiple stripes and generating a preliminary distribution result of the data blocks of the data packet according to each data block and the disk corresponding to each data block includes: obtaining the stripe number of the target stripe in the multiple stripes, where the initial value of the stripe number is 0 and the increment is 1; determining the correspondence between the data blocks of the target stripe and the disks according to the stripe number of the target stripe and the number of disks; determining the preliminary distribution result according to the correspondence between the data blocks of the multiple stripes and the disks.
[0034] Optionally, in the above embodiment, assume that there are 6 disks in an independent disk array of RAID6, then the stripe width is also 6, there are 36 data packets in one data packet, that is, corresponding to 6 stripes, and the disk numbers are disk 0, disk 1..... disk 5, and the stripe numbers are stripe 0, stripe 1..... stripe 5.
[0035] In an exemplary embodiment, determining the correspondence between the data chunks of the target stripe and the disks according to the stripe number of the target stripe and the number of disks includes: obtaining the disk number of the target disk among the multiple disks, where the initial value of the disk number is 0 and the increment is 1; performing a modulo operation on the number of disks according to the stripe number of the target stripe to obtain a first target value; calculating the difference between the disk number and the first target value to obtain a second target value; in the case where it is determined that the second target value is non - negative, determining the second target value as a third target value; or, in the case where it is determined that the second target value is negative, determining the sum value of the second target value and the number of disks as the third target value; in the case where it is determined that the third target value is 0, determining the target disk as the disk corresponding to the first parity block in the target stripe; or, in the case where it is determined that the third target value is 1, determining the target disk as the disk corresponding to the second parity block in the target stripe; or, in the case where it is determined that the third target value is greater than 1, determining the target disk as the disk corresponding to the data block in the target stripe.
[0036] Optionally, in the above - mentioned embodiment, in the case where it is determined that the third target value is greater than 1, determining the target disk as the disk corresponding to the data block in the target stripe includes: in the case where the difference between the third target value and 1 is N, determining the target disk as the disk corresponding to the Nth data block in the target stripe, where N is a positive integer.
[0037] In an alternative embodiment, as Figure 3 shown, it specifically includes the following steps:
[0038] Step S301: Input the stripe number strideNumber, the disk number ComponentIndex, and the total number of disks ComponentCount;
[0039] Step S302: Calculate Rotation (equivalent to the first target value) = strideNumber % ComponentCount; Step S303: Calculate StripNum (equivalent to the second target value) = ComponentIndex - Rotation; Step S304: Judge the positive or negative of StripNum, if it is negative, execute Step S305, otherwise execute Step S306; Step S305: StripNum (equivalent to the third target value) = StripNum + Componentcount;
[0040] Step S306: Return the result StripNum.
[0041] Optionally, assume that there are 6 disks in an independent disk array of RAID6, the strip width is also 6, a data packet has 36 data chunks, that is, it includes 6 strips. Each strip includes parity block q, parity block p, data block 2, data block 3, data block 4, and data block 5. Different return results correspond to different data chunks. The return result of 0 corresponds to parity block p, the return result of 1 corresponds to parity block q, the return result of 2 corresponds to data block 2, the return result of 3 corresponds to data block 3, the return result of 4 corresponds to data block 4, and the return result of 5 corresponds to data block 5. If the input strip number is 3 and the disk number is 2, and the calculated return result is 5, it means that the data stored in disk 2 for strip 3 is data block 5. After calculating all the data chunks of the 6 strips in the data packet, the obtained distribution result is shown in Table 1:
[0042] Table 1
[0043] Disk 0 Disk 1 Disk 2 Disk 3 Disk 4 Disk 5 Strip 0 p q 2 3 4 5 Strip 1 5 p q 2 3 4 Strip 2 4 5 p q 2 3 Strip 3 3 4 5 p q 2 Strip 4 2 3 4 5 p q Strip 5 q 2 3 4 5 p
[0044] Optionally, assume that there are 5 disks in an independent disk array of RAID6, the strip width is also 5, a data packet has 25 data chunks, that is, it includes 5 strips. Each strip includes parity block q, parity block p, data block 2, data block 3, and data block 4. Different return results correspond to different data chunks. The return result of 0 corresponds to parity block p, the return result of 1 corresponds to parity block q, the return result of 2 corresponds to data block 2, the return result of 3 corresponds to data block 3, and the return result of 4 corresponds to data block 4. If the input strip number is 1 and the disk number is 2, and the calculated return result is 0, it means that the data stored in disk 2 for strip 1 is parity block q. After calculating all the data chunks of the 5 strips in the data packet, the obtained distribution result is shown in Table 2:
[0045] Table 2
[0046] Disk 0 Disk 1 Disk 2 Disk 3 Disk 4 Strip 0 p q 2 3 4 Strip 1 4 p q 2 3 Strip 2 3 4 p q 2 Strip 3 2 3 4 p q Strip 4 q 2 3 4 p
[0047] Through the above embodiments, the distribution result of the data chunks in the data packet can be efficiently calculated, and then the data chunks can be stored according to the distribution result, ensuring that only one parity block q is distributed on each disk in a packet, achieving the effect of uniform storage, avoiding the bottleneck of the entire storage array caused by insufficient storage space on one of the disks, and maximizing the storage utilization efficiency.
[0048] In an exemplary embodiment, determining the correspondence between the data chunks of the target stripe and the disks according to the stripe number of the target stripe and the number of disks further includes: determining a fourth target value according to the sum value of a first preset value and a target product, where the target product represents the product of a second preset value and the stripe number; in the case where a fifth target value is determined to be a non-negative number, determining the disk with the disk number being the fifth target value among the multiple disks as the disk corresponding to the first parity block in the target stripe, where the fifth target value represents the difference between the number of disks and the fourth target value; and, determining the disk with the disk number being a sixth target value among the multiple disks as the disk corresponding to the second parity block in the target stripe, where the sixth target value represents the difference between the fifth target value and the first preset value; and, determining the disks among the multiple disks whose disk numbers are not the fifth target value and the sixth target value as the disks corresponding to the data blocks in the target stripe.
[0049] In an alternative embodiment, another method for calculating the correspondence between the data chunks of the target stripe and the disks is also provided. In the above embodiment, for example, the first preset value is 1, the second preset value is 2. Assume there are 6 disks in an independent disk array of RAID6 and the stripe width is also 6. If the input stripe number is 1, the calculated fifth target value = the number of disks 6 * (the first preset value 1 + the stripe number 1 * the second preset value 2) = 3. Then, the disk 3 is determined as the disk corresponding to the first parity block q in stripe 1. Then, calculate the sixth target value = the fifth target value 3 * the first preset value 1 = 2. The disk 2 is determined as the disk corresponding to the second parity block p in stripe 1. The other disks are determined as the disks corresponding to the data blocks. Through the above steps, the data chunk distribution result of stripe 0 * 2 can be calculated, and the obtained distribution result is shown in Table 3:
[0050] Table 3
[0051] Disk 0 Disk 1 Disk 2 Disk 3 Disk 4 Disk 5 Strip 0 2 3 4 5 p q Strip 1 4 5 p q 2 3 Strip 2 p q 2 3 4 5 Strip 3 / / / / / / Strip 4 / / / / / / Strip 5 / / / / / /
[0052] Optionally, in the above embodiment, for example, the first preset value is 1, the second preset value is 2. Assume there are 5 disks in an independent disk array of RAID6 and the stripe width is also 5. If the input stripe number is 1, the calculated fifth target value = the number of disks 65 * (the first preset value 1 + the stripe number 1 * the second preset value 2) = 32. Then, the disk 32 is determined as the disk corresponding to the first parity block q in stripe 1. Then, calculate the sixth target value = the fifth target value 32 - the first preset value 1 = 21. The disk 31 is determined as the disk corresponding to the second parity block p in stripe 1. The other disks are determined as the disks corresponding to the data blocks.
[0053] It should be noted that there is a special case in the above calculation process, that is, the total number of stripes is odd, and the input stripe number is the middle number of the total number of stripes. For example, there are stripes 0-4, the input stripe number is 2, and the calculated fifth target value is 0. Then, disk 0 is determined as the disk corresponding to the first parity block q in stripe 2. At this time, the calculated sixth target value is -1, and it is necessary to calculate the sum of the disk number and the sixth target value, that is, 5+(-1)=4. Disk 4 is determined as the disk corresponding to the second parity block p in stripe 2. Through the above steps, the data block distribution result of stripes 0-2 can be calculated, and the obtained distribution result is shown in Table 4:
[0054] Table 4
[0055] Disk 0 Disk 1 Disk 2 Disk 3 Disk 4 Strip 0 2 3 4 p q Strip 1 4 p q 2 3 Strip 2 q 2 3 4 p Strip 3 / / / / / Strip 4 / / / / /
[0056] In an exemplary embodiment, the method further includes: when it is determined that the fifth target value is negative, determining the parity of the disk number; and determining the correspondence between the data blocks of the target stripe and the disks according to the fifth target value and the parity of the disk number.
[0057] In an exemplary embodiment, determining the correspondence between the data blocks of the target stripe and the disks according to the fifth target value and the parity of the disk number includes: when it is determined that the disk number is odd, determining the disk with the disk number of the seventh target value among the multiple disks as the disk corresponding to the first parity block in the target stripe, where the seventh target value represents the sum of the disk number and the fifth target value; and determining the disk with the disk number of the eighth target value among the multiple disks as the disk corresponding to the second parity block in the target stripe, where the eighth target value represents the difference between the seventh target value and the first preset value; and determining the disks with disk numbers other than the seventh target value and the eighth target value among the multiple disks as the disks corresponding to the data blocks in the target stripe.
[0058] Optionally, in the above embodiments, for example, the first preset value is 1, the second preset value is 2. Assume that there are 5 disks in an independent disk array of RAID6 and the stripe width is also 5. If the input stripe number is 3, the calculated fifth target value = the number of disks 5 - (the first preset value 1 + the stripe number 3 * the second preset value 2) = -2, and the calculated seventh target value = the number of disks 5 + the fifth target value (*2) = 3. Then, the disk 3 is determined as the disk corresponding to the first parity block q in stripe 1. Then, the calculated eighth target value = the seventh target value 3 * the first preset value 1 = 2, and the disk 2 is determined as the disk corresponding to the second parity block p in stripe 1. The other disks are determined as the disks corresponding to the data blocks. Through the above steps, the data block distribution result of stripe 3 - 4 can be calculated. Combining with the calculation result in Table 4 in the above embodiments, the obtained distribution result is shown in Table 5:
[0059] Table 5
[0060] Disk 0 Disk 1 Disk 2 Disk 3 Disk 4 Strip 0 2 3 4 p q Strip 1 4 p q 2 3 Strip 2 q 2 3 4 p Strip 3 3 4 p q 2 Strip 4 p q 2 3 4
[0061] In an exemplary embodiment, determining the correspondence between the data blocks of the target stripe and the disks according to the parity of the fifth target value and the number of disks further includes: when it is determined that the number of disks is even, determining the disk with the ninth target value among the multiple disks as the disk corresponding to the second parity block in the target stripe, where the ninth target value represents the sum value of the number of disks and the fifth target value; and, determining the disk with the tenth target value among the multiple disks as the disk corresponding to the first parity block in the target stripe, where the tenth target value represents the difference value between the ninth target value and the first preset value; and, determining the disks among the multiple disks whose disk numbers are not the ninth target value and the tenth target value as the disks corresponding to the data blocks in the target stripe.
[0062] Optionally, in the above embodiments, for example, the first preset value is 1, the second preset value is 2. Assume that there are 6 disks in an independent disk array of RAID6 and the stripe width is also 6. If the input stripe number is 5, the calculated fifth target value = the number of disks 6 - (the first preset value 1 + the stripe number 5 * the second preset value 2) = *5, and the calculated ninth target value = the number of disks 6 + the fifth target value (*5) = 1. Then, the disk 2 is determined as the disk corresponding to the second parity block p in stripe 5. Then, the calculated tenth target value = the ninth target value 1 * the first preset value 1 = 0, and the disk 0 is determined as the disk corresponding to the first parity block q in stripe 5. The other disks are determined as the disks corresponding to the data blocks. Through the above steps, the data block distribution result of stripe 3 - 5 can be calculated. Combining with the calculation result in Table 3 in the above embodiments, the obtained distribution result is shown in Table 6:
[0063] Table 6
[0064] Disk 0 Disk 1 Disk 2 Disk 3 Disk 4 Disk 5 Strip 0 2 3 4 5 p q Strip 1 4 5 p q 2 3 Strip 2 p q 2 3 4 5 Strip 3 2 3 4 5 q p Strip 4 4 5 q p 2 3 Strip 5 q p 2 3 4 5
[0065] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application.
[0066] In this embodiment, a data storage device is also provided. The system is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can implement a predetermined function. Although the systems described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0067] Figure 4 is a structural block diagram of a data storage device according to an embodiment of the present application. As Figure 4 shown, the device includes:
[0068] A setting module 42, configured to set the number of data blocks of a data packet according to the product of the strip width and the number of disks, and the data blocks of the data packet come from different disks;
[0069] A generating module 44, configured to determine the disk corresponding to each data block in the data blocks of multiple strips, and generate a pre-distribution result of the data blocks of the data packet according to each data block and the disk corresponding to each data block, wherein the total number of data blocks of the multiple strips is equal to the number of data blocks of the data packet;
[0070] A storage module 46, configured to store each data block into the disk corresponding to each data block according to the pre-distribution result, so that the data volume of the data packet is evenly distributed on the multiple disks.
[0071] Through the above device, the number of data chunks of a data packet can be set according to the product of the strip width and the number of disks, and then the disks corresponding to each data chunk in the data packet can be determined to generate a preliminary distribution result of the data chunks of the data packet. Finally, each data chunk is stored in the corresponding disk according to the preliminary distribution result, so that the data volume of the data packet is evenly distributed among multiple disks. Furthermore, the problem of how to improve the disk storage performance in the related art can be solved.
[0072] In an exemplary embodiment, the generating module 44 is further configured to obtain the strip number of the target strip among the multiple strips, where the initial value of the strip number is 0 and the increment is 1; determine the correspondence between the data chunks of the target strip and the disks according to the strip number of the target strip and the number of disks; and determine the preliminary distribution result according to the correspondence between the data chunks of the multiple strips and the disks.
[0073] In an exemplary embodiment, the generating module 44 is further configured to obtain the disk number of the target disk among the multiple disks, where the initial value of the disk number is 0 and the increment is 1; perform a modulo operation on the number of disks according to the strip number of the target strip to obtain a first target value; calculate the difference between the disk number and the first target value to obtain a second target value; in the case where it is determined that the second target value is non - negative, determine the second target value as a third target value; or, in the case where it is determined that the second target value is negative, determine the sum value of the second target value and the number of disks as the third target value; in the case where it is determined that the third target value is 0, determine the target disk as the disk corresponding to the first check block in the target strip; or, in the case where it is determined that the third target value is 1, determine the target disk as the disk corresponding to the second check block in the target strip; or, in the case where it is determined that the third target value is greater than 1, determine the target disk as the disk corresponding to the data block in the target strip.
[0074] In an exemplary embodiment, the generating module 44 is further configured to determine a fourth target value according to the sum of a first preset value and a target product, where the target product represents the product of a second preset value and the stripe number; in the case where it is determined that the fifth target value is non - negative, determine the disk with the disk number being the fifth target value among the multiple disks as the disk corresponding to the first parity block in the target stripe, where the fifth target value represents the difference between the number of disks and the fourth target value; and, determine the disk with the disk number being the sixth target value among the multiple disks as the disk corresponding to the second parity block in the target stripe, where the sixth target value represents the difference between the fifth target value and the first preset value; and, determine the disks among the multiple disks whose disk numbers are not the fifth target value and the sixth target value as the disks corresponding to the data blocks in the target stripe.
[0075] In an exemplary embodiment, the generating module 44 is further configured to, in the case where it is determined that the fifth target value is negative, determine the parity of the number of disks; and determine the correspondence between the data blocks of the target stripe and the disks according to the fifth target value and the parity of the number of disks.
[0076] In an exemplary embodiment, the generating module 44 is further configured to, in the case where it is determined that the number of disks is odd, determine the disk with the disk number being the seventh target value among the multiple disks as the disk corresponding to the first parity block in the target stripe, where the seventh target value represents the sum of the number of disks and the fifth target value; and, determine the disk with the disk number being the eighth target value among the multiple disks as the disk corresponding to the second parity block in the target stripe, where the eighth target value represents the difference between the seventh target value and the first preset value; and, determine the disks among the multiple disks whose disk numbers are not the seventh target value and the eighth target value as the disks corresponding to the data blocks in the target stripe.
[0077] In an exemplary embodiment, the generating module 44 is further configured to, in the case where it is determined that the number of disks is even, determine the disk with the disk number being the ninth target value among the multiple disks as the disk corresponding to the second parity block in the target stripe, where the ninth target value represents the sum of the number of disks and the fifth target value; and, determine the disk with the disk number being the tenth target value among the multiple disks as the disk corresponding to the first parity block in the target stripe, where the tenth target value represents the difference between the ninth target value and the first preset value; and, determine the disks among the multiple disks whose disk numbers are not the ninth target value and the tenth target value as the disks corresponding to the data blocks in the target stripe.
[0078] It should be noted that the above-mentioned modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to this: the above-mentioned modules are all located in the same processor; or, the above-mentioned modules are respectively located in different processors in any combination form.
[0079] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. Wherein, the computer program is set to execute the steps in any one of the above method embodiments when running.
[0080] Optionally, in this embodiment, the above storage medium can be set to store program codes for executing the following steps:
[0081] S1, set the number of data chunks of the data packet according to the product of the strip width and the number of disks, and the data chunks of the data packet come from different disks;
[0082] S2, determine the disk corresponding to each data chunk in the data chunks of multiple strips, and generate a pre-distribution result of the data chunks of the data packet according to each data chunk and the disk corresponding to each data chunk, wherein the total number of data chunks of the multiple strips is equal to the number of data chunks of the data packet;
[0083] S3, store each data chunk into the disk corresponding to each data chunk according to the pre-distribution result, so that the data volume of the data packet is evenly distributed on the multiple disks.
[0084] In an exemplary embodiment, the above computer-readable storage medium may include but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks or optical discs that can store computer programs.
[0085] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is set to run the computer program to execute the steps in any one of the above method embodiments.
[0086] Optionally, in this embodiment, the above processor can be set to execute the following steps through a computer program:
[0087] S1, set the number of data chunks of the data packet according to the product of the strip width and the number of disks, and the data chunks of the data packet come from different disks;
[0088] S2. Determine the disk corresponding to each data block in the data blocks of multiple stripes, and generate a preliminary distribution result of the data blocks of the data packet according to each data block and the disk corresponding to each data block, where the total number of data blocks of the multiple stripes is equal to the number of data blocks of the data packet;
[0089] S3. Store each data block into the disk corresponding to each data block according to the preliminary distribution result, so that the data volume of the data packet is evenly distributed among the multiple disks.
[0090] An embodiment of the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any one of the above method embodiments are implemented.
[0091] Optionally, in this embodiment, the above computer program can be set to execute the following steps through the computer program:
[0092] S1. Set the number of data blocks of the data packet according to the product of the stripe width and the number of disks. The data blocks of the data packet come from different disks;
[0093] S2. Determine the disk corresponding to each data block in the data blocks of multiple stripes, and generate a preliminary distribution result of the data blocks of the data packet according to each data block and the disk corresponding to each data block, where the total number of data blocks of the multiple stripes is equal to the number of data blocks of the data packet;
[0094] S3. Store each data block into the disk corresponding to each data block according to the preliminary distribution result, so that the data volume of the data packet is evenly distributed among the multiple disks.
[0095] Specific examples in this embodiment can refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.
[0096] Obviously, those skilled in the art should understand that the above modules or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps of them can be made into a single integrated circuit module to implement. In this way, the present application is not limited to any specific combination of hardware and software.
[0097] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the principle of the present application shall be included within the protection scope of the present application.
Claims
1. A data storage method, characterized in that: An independent disk array for dual parity check, the independent disk array comprising a plurality of disks, the plurality of disks being used to store stripe data blocks, the stripe width of the stripe being set based on the number of disks of the independent disk array, the stripe data blocks being respectively stored on different disks, the stripe data blocks comprising a first check block, a second check block and a plurality of data blocks, the data volume of the first check block being greater than the data volume of the second check block, the data volume of the second check block being the same as the data volume of each data block, the method comprising: Setting the number of data blocks of the data packet according to the product of the stripe width and the number of disks, the data blocks of the data packet coming from different disks; Determine the disk corresponding to each data block in the data blocks of the multiple stripes, and generate a pre-distribution result of the data blocks of the data packet according to each data block and the disk corresponding to each data block, wherein the total number of data blocks of the multiple stripes is equal to the number of data blocks of the data packet; According to the pre-distribution result, each data block is stored in the disk corresponding to each data block, so that the data volume of the data packet is evenly distributed on the multiple disks.
2. The method according to claim 1, characterized in that Determining a disk corresponding to each of the data blocks in the plurality of stripes, and generating a pre-distribution result of the data blocks of the data packet according to each of the data blocks and the disk corresponding to each of the data blocks, including: Obtaining a stripe number of a target stripe among the plurality of stripes, wherein an initial value of the stripe number is 0 and an increment is 1; Determine the correspondence between the data blocks of the target stripe and the disks according to the stripe number of the target stripe and the number of disks; The pre-distribution result is determined according to the correspondence between the data blocks of the multiple stripes and the disks.
3. The method according to claim 2, characterized in that Determining the correspondence between the data blocks of the target stripe and the disks according to the stripe number of the target stripe and the number of disks includes: Obtaining a disk number of a target disk among the multiple disks, wherein the initial value of the disk number is 0 and the increment is 1; Performing a modulo operation on the number of disks according to the stripe number of the target stripe to obtain a first target value; Calculate the difference between the disk number and the first target value to obtain a second target value; When it is determined that the second target value is a non-negative number, determining the second target value as a third target value; Or, when it is determined that the second target value is a negative number, the sum of the second target value and the number of disks is determined as the third target value; When it is determined that the third target value is 0, the target disk is determined to be a disk corresponding to the first check block in the target stripe; Or, when it is determined that the third target value is 1, the target disk is determined to be the disk corresponding to the second check block in the target stripe; Or, when it is determined that the third target value is greater than 1, the target disk is determined to be a disk corresponding to the data block in the target stripe.
4. The method according to claim 2, characterized in that: Determining the correspondence between the data blocks of the target stripe and the disks according to the stripe number of the target stripe and the number of disks, further comprising: Determining a fourth target value according to a sum of the first preset value and a target product, wherein the target product represents a product of the second preset value and the stripe number; In the case where it is determined that the fifth target value is a non-negative number, a disk whose disk number is the fifth target value among the multiple disks is determined as a disk corresponding to the first check block in the target stripe, wherein the fifth target value represents a difference between the number of disks and the fourth target value; and, determining a disk whose disk number is a sixth target value among the multiple disks as a disk corresponding to a second check block in the target stripe, wherein the sixth target value represents a difference between the fifth target value and the first preset value; and determining the disks whose disk numbers are not the fifth target value and the sixth target value among the multiple disks as the disks corresponding to the data blocks in the target stripe.
5. The method according to claim 4, characterized in that The method further comprises: In the case where it is determined that the fifth target value is a negative number, determining the parity of the number of disks; The correspondence between the data blocks of the target stripe and the disks is determined according to the fifth target value and the parity of the number of disks.
6. The method according to claim 5, characterized in that Determining the correspondence between the data blocks of the target stripe and the disks according to the fifth target value and the parity of the number of disks includes: In the case where it is determined that the number of disks is an odd number, a disk whose disk number is a seventh target value among the multiple disks is determined as a disk corresponding to the first check block in the target stripe, wherein the seventh target value represents the sum of the number of disks and the fifth target value; and, determining a disk whose disk number is an eighth target value among the multiple disks as a disk corresponding to a second check block in the target stripe, wherein the eighth target value represents a difference between the seventh target value and the first preset value; and determining the disks whose disk numbers are not the seventh target value and the eighth target value among the multiple disks as the disks corresponding to the data blocks in the target stripe.
7. The method according to claim 5, characterized in that Determining the correspondence between the data blocks of the target stripe and the disks according to the fifth target value and the parity of the number of disks, further comprising: In the case where it is determined that the number of disks is an even number, a disk whose disk number is a ninth target value among the multiple disks is determined as a disk corresponding to a second check block in the target stripe, wherein the ninth target value represents a sum of the number of disks and the fifth target value; and, determining a disk whose disk number is a tenth target value among the multiple disks as a disk corresponding to a first check block in the target stripe, wherein the tenth target value represents a difference between the ninth target value and the first preset value; and determining the disks whose disk numbers are not the ninth target value and the tenth target value among the multiple disks as the disks corresponding to the data blocks in the target stripe.
8. A data storage device, characterized in that: An independent disk array for dual parity check, the independent disk array comprising a plurality of disks, the plurality of disks being used to store stripe data blocks, the stripe width of the stripe being set based on the number of disks of the independent disk array, the stripe data blocks being respectively stored on different disks, the stripe data blocks comprising a first check block, a second check block and a plurality of data blocks, the data volume of the first check block being greater than the data volume of the second check block, the data volume of the second check block being the same as the data volume of each data block, the device comprising: A setting module, used for setting the number of data blocks of a data packet according to the product of the stripe width and the number of disks, wherein the data blocks of the data packet come from different disks; A generation module, used to determine the disk corresponding to each data block in the data blocks of the multiple stripes, and generate a pre-distribution result of the data blocks of the data packet according to each data block and the disk corresponding to each data block, wherein the total number of data blocks of the multiple stripes is equal to the number of data blocks of the data packet; The storage module is used to store each data block in the disk corresponding to each data block according to the pre-distribution result, so that the data volume of the data packet is evenly distributed on the multiple disks.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method described in any one of claims 1 to 7 when executed by a processor.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Independent disk redundancy array construction method and device
CN101504623A
Disk array online capacity expansion method and device and computer readable storage medium
CN112130768A
Write data processing method and device for disk array, equipment and medium
CN116719484A
Data processing method and device of storage equipment, storage medium and electronic equipment
CN117193672A
Data storage method, computer program product, equipment and computer medium
CN118656039A