Data processing method and device, equipment, medium and program product
By setting up the data disk and the verification disk in the disk array, and using geometric verification rules to generate mapping relationships, the problem of inefficient data recovery when multiple disks fail at the same time is solved, and efficient and accurate data recovery is achieved.
Patent Information
- Application Number
- CN202510866424.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-26
AI Technical Summary
In the prior art, when multiple disks fail at the same time, it is difficult to directly recover data by dual distributed parity, resulting in low recovery efficiency and high computing resource consumption.
By setting up multiple data disks and n check disks in the disk array, the geometric verification rules are used to generate verification blocks and data blocks to establish mapping relationships, and in the event of a disk failure, the geometric verification rules are called forward and/or reversely according to the type and number of faults.
It realizes efficient data recovery in high fault tolerance scenarios for multiple failed disks, improves the efficiency and accuracy of data recovery, and ensures the availability of storage systems.
Smart Images

Figure CN120371595A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a data processing method, apparatus, device, medium, and program product. Background Art
[0002] With the continuous increase in storage requirements, the risk of disk damage and failure also rises. To ensure high reliability and availability of data, storage systems usually adopt dual distributed parity check technology to achieve data redundancy.
[0003] In the process of implementing the inventive concept, it is found that at least the following problems exist in the related art. When multiple disks fail simultaneously, it is difficult to directly recover data only relying on dual distributed parity check, and often requires relying on complex recursive calculations. This method not only consumes a large amount of computing resources, but also results in low recovery efficiency. Summary of the Invention
[0004] In view of the above problems, the present invention provides a data processing method, apparatus, device, medium, and program product.
[0005] According to a first aspect of the present invention, there is provided a data processing method, including: in response to a data writing operation, performing a check on the written data according to the geometric check rules corresponding to each of the above check disks, so as to store the check blocks obtained by the check to the corresponding check disks, and there is a mapping relationship between the check blocks and the written data stored in the above multiple data disks; if it is detected that the above disk fails, determining a recovery method according to the type of the failed disk and the number of disks that have failed; based on the above mapping relationship, forwardly and / or reversely calling the above geometric check rules according to the above recovery method to recover the data of the above failed disk.
[0006] A second aspect of the present invention provides a data processing apparatus, including: a data check module, configured to, in response to a data writing operation, perform a check on the written data according to the geometric check rules corresponding to each of the above check disks, so as to store the check blocks obtained by the check to the corresponding check disks, and there is a mapping relationship between the check blocks and the written data stored in the above multiple data disks; a method determination module, configured to, if it is detected that the above disk fails, determine a recovery method according to the type of the failed disk and the number of disks that have failed; a data recovery module, configured to, based on the above mapping relationship, forwardly and / or reversely call the above geometric check rules according to the above recovery method to recover the data of the above failed disk.
[0007] A third aspect of the present invention provides an electronic device, including: one or more processors; a memory, configured to store one or more computer programs, wherein the above one or more processors execute the above one or more computer programs to implement the steps of the above method.
[0008] The fourth aspect of the present invention further provides a computer-readable storage medium, on which a computer program or instruction is stored. When the computer program or instruction is executed by a processor, the steps of the above method are implemented.
[0009] The fifth aspect of the present invention further provides a computer program product, including a computer program or instruction. When the computer program or instruction is executed by a processor, the steps of the above method are implemented.
[0010] According to the embodiments of the present invention, by setting multiple data disks and n parity disks in a disk array, and using geometric parity rules to generate parity blocks for the written data, a mapping relationship is established between the parity blocks and the data blocks, which improves the limit of the fault tolerance ability of the storage system. Then, when a disk fails, according to the type and quantity of the faults, the geometric parity rules are called forward and / or backward based on the mapping relationship to recover the data, realizing the data recovery of multiple faulty disks. This method not only supports high-fault-tolerance scenarios with multiple disks failing simultaneously, but also effectively improves the efficiency and accuracy of data recovery through the combined invocation of the mapping relationship and the geometric parity rules. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Through the following description of the embodiments of the present invention with reference to the drawings, the above content and other objects, features, and advantages of the present invention will become clearer.
[0012] Figure 1 The application scenario diagram of the data processing method, device, equipment, medium, and program product according to the embodiments of the present invention is shown.
[0013] Figure 2 The flowchart of the data processing method according to the embodiments of the present invention is shown.
[0014] Figure 3 The schematic diagram of performing parity check according to the geometric parity rules in the embodiments of the present invention is shown.
[0015] Figure 4 The flowchart of performing data parity check in the embodiments of the present invention is shown.
[0016] Figure 5 The flowchart of three-disk fault recovery in the embodiments of the present invention is shown.
[0017] Figure 6 The structural block diagram of the data processing device according to the embodiments of the present invention is shown.
[0018] Figure 7 The block diagram of the electronic device suitable for implementing the data processing method according to the embodiments of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, numerous specific details are set forth to provide a comprehensive understanding of the embodiments of the present invention. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.
[0020] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0021] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted to have a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0022] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0023] In the technical solution of the present invention, the data involved (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. Moreover, for the processing of relevant data such as collection, storage, use, processing, transmission, provision, disclosure, and application, etc., all comply with relevant laws, regulations, and standards, necessary confidentiality measures are taken, it does not violate public order and good customs, and a corresponding operation entry is provided for the user to choose to authorize or reject.
[0024] An embodiment of the present invention provides a data processing method. In response to a data writing operation, according to the geometric verification rules corresponding to each verification disk, verify the written data to store the verified verification blocks into the corresponding verification disks. There is a mapping relationship between the verification blocks and the written data stored in multiple data disks. If it is detected that a disk fails, then determine the recovery method according to the type of the failed disk and the number of failed disks. Based on the mapping relationship, forward and / or reverse call the geometric verification rules according to the recovery method to recover the data of the failed disk.
[0025] Figure 1 The application scenario diagram of the data processing method, apparatus, device, medium and program product according to the embodiments of the present invention is shown.
[0026] As Figure 1 shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium to provide a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0027] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for examples).
[0028] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0029] The server 105 may be a server that provides various services, such as a background management server that supports the websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (only for examples). The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0030] It should be noted that the data processing method provided by the embodiments of the present invention can generally be executed by the server 105. Correspondingly, the data processing apparatus provided by the embodiments of the present invention can generally be set in the server 105. The data processing method provided by the embodiments of the present invention can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the data processing apparatus provided by the embodiments of the present invention can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0031] It should be understood that Figure 1 the numbers of the first terminal device, the second terminal device, the third terminal device, the network, and the server in are merely illustrative. According to the implementation requirements, there can be any number of the first terminal device, the second terminal device, the third terminal device, the network, and the server.
[0032] Based on the scenario described below Figure 1 through Figures 2 to 5 the data processing method of the disclosed embodiments will be described in detail.
[0033] Figure 2 FIG. shows a flowchart of the data processing method according to an embodiment of the present invention.
[0034] As Figure 2 shown, this embodiment includes operations S210 to S230.
[0035] In operation S210, in response to a data writing operation, according to the geometric check rules corresponding to each check disk, the written data is checked, so as to store the checked check blocks into the corresponding check disks, and there is a mapping relationship between the check blocks and the written data stored in multiple data disks.
[0036] In operation S220, if it is detected that a disk fails, then according to the type of the failed disk and the number of failed disks, a recovery method is determined.
[0037] In operation S230, based on the mapping relationship, the geometric check rules are called forward and / or backward according to the recovery method to recover the data of the failed disk.
[0038] According to an embodiment of the present invention, when the disk array receives a data writing operation, the written data is checked according to the geometric check rules corresponding to each check disk. The disk array is configured with n check disks in addition to multiple data disks, where n is an integer greater than 2. For example, 3 check disks (P, Q, R) can be configured, and each check disk corresponds to a different geometric check rule.
[0039] The written data will first be calculated according to the geometric check rules corresponding to each check disk to generate check blocks having a mapping relationship with the written data. These check blocks will be respectively stored into the corresponding check disks, while the original written data is stored into multiple data disks, thereby ensuring the integrity and reliability of the data through the check mechanism.
[0040] During the operation of the disk array, continuously monitor the status of each disk. When a disk failure is detected, immediately determine the type of the failed disk (whether it is a data disk or a parity disk), and at the same time count the number of failed disks. For example, if the data disk D1 is detected to have failed, record the failure type as a data disk and the failure number as 1.
[0041] Based on the type and number of failed disks, determine the corresponding recovery method. Different failure situations may correspond to different recovery strategies. For example, if only one data disk fails, data recovery can be directly performed using the P disk or the Q disk or the R disk. If multiple data disks fail simultaneously, it may be necessary to combine the parity information of the Q disk and the R disk for complex recovery to ensure the highest possible recovery efficiency while guaranteeing the accuracy of data recovery.
[0042] After determining the recovery method, based on the mapping relationship between the parity blocks and the written data, forward and / or reverse call the geometric parity rules according to the requirements of the recovery method. In this way, using the data blocks and parity blocks stored on the non-failed disks, recalculate and generate the data on the failed disks, thereby realizing the recovery of the data on the failed disks and enabling the disk array to resume normal operation as soon as possible to ensure the availability of data.
[0043] According to the embodiments of the present invention, by setting multiple data disks and n parity disks in the disk array and using the geometric parity rules to generate parity blocks for the written data, a mapping relationship is established between the parity blocks and the data blocks, which improves the limitation of the fault tolerance of the storage system. Then, when a disk failure occurs, based on the failure type and number, forward and / or reverse call the geometric parity rules based on the mapping relationship to recover the data, realizing the data recovery of multiple failed disks. This method not only supports the high fault tolerance scenario of multiple disks failing simultaneously, but also effectively improves the efficiency and accuracy of data recovery through the combined call of the mapping relationship and the geometric parity rules.
[0044] According to the embodiments of the present invention, perform parity check on the written data according to the geometric parity rules corresponding to each parity disk, including: determining the data blocks in each data stripe of the written data that are respectively stored in multiple data disks, where the data stripe is obtained by dividing the written data according to a preset stripe size; based on the distribution positions of the data blocks in the data stripe in multiple data disks, perform parallel parity check on the data stripe according to the geometric parity rules corresponding to each parity disk, and the geometric parity rules include horizontal parity rules, diagonal parity rules, and anti-diagonal parity rules.
[0045] The written data is segmented according to a preset strip size to obtain a number of data strips, and then the storage positions of each data block in each data strip among multiple data disks are determined. For example, assume the written data is a large file and the preset strip size is 64 KB. The file is segmented into multiple 64 - KB data strips. This process needs to combine the striping strategy of the storage system, split the data strips into multiple data blocks according to specific rules, and map them to different data disk address spaces, such as distributing them in a row - major or column - major manner to ensure the uniform distribution of data blocks in the disk array. For instance, in an array with 3 data disks, the data blocks can be stored on D1, D2, and D3 in sequence in a row - major manner.
[0046] After completing the distribution and positioning of the data blocks, based on the specific distribution positions of each data block among multiple data disks, parallel verification of the data strips is performed according to the geometric verification rules corresponding to each parity disk. Among them, the horizontal verification rule can perform verification calculations on the data blocks in the same row or the same horizontal dimension. For example, a horizontal parity block is generated through exclusive - OR operation. The diagonal verification rule focuses on the diagonal distribution path of the data blocks in the disk array and combines and verifies the data blocks in the diagonal direction with a specific slope. The anti - diagonal verification rule corresponds to the diagonal direction with the opposite slope and performs verification operations on the data blocks on this path.
[0047] Figure 3 It shows a schematic diagram of verification according to the geometric verification rules in an embodiment of the present invention.
[0048] As Figure 3 shown, during the process of verifying the written data, the horizontal verification rule 301, the diagonal verification rule 302, and the anti - diagonal verification rule 303 are respectively used to verify the data blocks 305 stored in multiple data disks 304. Specifically, the horizontal verification rule 301 verifies the data blocks from left to right, the diagonal verification rule 302 verifies the data blocks along a specific diagonal direction, and the anti - diagonal verification rule 303 verifies the data blocks along the direction opposite to the diagonal direction. Through the coordinated action of these three verification rules, the accuracy and integrity of the data are ensured.
[0049] In specific implementation, multi - threads or hardware parallel processing units can be used to simultaneously start horizontal, diagonal, and anti - diagonal verification calculation tasks for each data strip. Each verification task reads the corresponding data blocks from the data disks according to the corresponding geometric rules, performs operations according to the rules to generate parity blocks, and writes the parity blocks to the corresponding parity disks. For example, a certain parity disk is specifically used to store the parity blocks generated by the horizontal verification rule, and another parity disk stores the parity blocks of the diagonal verification rule. In this way, parallel verification processing based on geometric rules is realized, ensuring the integrity and reliability verification of the written data is efficiently completed during data storage.
[0050] According to an embodiment of the present invention, based on the distribution positions of data blocks in a data stripe among multiple data disks, according to geometric verification rules corresponding to each verification disk, the data stripe is verified in parallel, including: for the horizontal verification rule, performing an exclusive OR operation on multiple data blocks distributed in multiple data disks in the same data stripe; for the diagonal verification rule, based on the distribution positions, according to a preset diagonal index rule, selecting data blocks across data stripes to perform an exclusive OR operation; for the anti-diagonal verification rule, based on the distribution positions, according to a preset anti-diagonal index rule, selecting data blocks across data stripes to perform an exclusive OR operation; wherein, when selecting data blocks, the preset diagonal index rule and the preset anti-diagonal index rule perform virtual zero-padding processing for boundary positions.
[0051] In the parallel implementation of geometric verification rules based on the distribution positions of data blocks, a dedicated processing thread is started for the horizontal verification rule. For example, assume that the current data stripe contains data blocks D 11 、D 21 、D 31 , which are stored on data disks D1, D2, and D3 respectively. This thread will obtain the data block address mapping table of the current data stripe in multiple data disks from the storage controller, and then sequentially read the contents of all data blocks within the same data stripe. By repeatedly performing the exclusive OR operation, the binary values of each data block are bitwise exclusive ORed, and finally a horizontal verification block is generated. P1 = D 11 ⊕D 21 ⊕D 31 This process is performed in real time while the data is being written, ensuring that the verification block is completed synchronously with the data block writing operation. The specific calculation method is shown in formula (1).
[0052]
[0053] Among them, represents the data block of the j-th data stripe on the i-th disk (the j-th row and i-th column), and there are a total of (three verification disks) disks in the disk array, and there are a total of data stripes, and both represent the exclusive OR operation between data blocks, represents the verification block obtained by verifying the P verification disk.
[0054] For the diagonal verification rule, an independent calculation task is created based on the preset diagonal index rule. For example, assume that the position information of the current data stripe indicates that data blocks D 11 、D 22 、D 33This task calculates all the data block addresses on the diagonal path based on the position information of the current data stripe and in combination with the topology of the data disk array. During the calculation process, if a boundary position is encountered, the virtual zero-padding processing mechanism is automatically triggered, and virtual data blocks filled with all zeros are generated to replace the actually non-existent data blocks to participate in the operation. The calculation task traverses all the data blocks (including virtual zero-padding blocks) on the diagonal path and performs an exclusive OR operation to generate the diagonal parity block Q1 = D 11 ⊕D 22 ⊕D 33 , and writes the result to the corresponding parity disk area.
[0055] The implementation of the anti-diagonal parity rule is similar to that of the diagonal parity rule, but with the opposite indexing direction. A dedicated anti-diagonal calculation thread is started, and the data block positions across data stripes are determined according to the preset anti-diagonal indexing rule. For example, assume that it is necessary to calculate the data blocks D 13 、D 22 、D 31 on the anti-diagonal path. Similarly, virtual zero-padding processing is performed when encountering boundary conditions to ensure the continuity of the anti-diagonal path. The thread reads the actual data blocks from the data disks and performs an exclusive OR operation with the virtual zero-padding blocks to finally generate the anti-diagonal parity block R1 = D 13 ⊕D 22 ⊕D 31 . These three parity tasks are implemented in parallel through a thread pool, each independently calculating the parity block, so as to efficiently complete multi-dimensional parity protection without affecting the data writing performance and ensure the integrity and reliability of the data.
[0056] According to the embodiments of the present invention, when selecting data blocks, the preset diagonal indexing rule and the preset anti-diagonal indexing rule perform virtual zero-padding processing for boundary positions, including: if the data block position calculated according to the preset diagonal indexing rule or the preset anti-diagonal indexing rule exceeds the range of data stripes in multiple data disks, the coordinate value of the data block position is mapped to the range through a modulo operation to obtain the mapped position; the data block corresponding to the mapped position is determined as a zero value.
[0057] When implementing the boundary processing in the geometric parity rule, a range check is performed on the data block positions calculated by the diagonal indexing rule and the anti-diagonal indexing rule. For example, assume that the data stripes are distributed on 3 data disks, and each data disk has 3 data blocks, forming a 3×3 data matrix. When the calculated data block coordinate value exceeds the range of data stripes in multiple data disks, the modulo operation mechanism is automatically triggered. This mechanism performs a modulo calculation on the out-of-range coordinate value and the dimension parameter of the data disk array, thereby mapping the coordinate value back to the range to generate a legal mapped position.
[0058] For example, assume that the calculated data block position according to the diagonal indexing rule is (3, 4), which is outside the range of the 3×3 matrix. Through modulo operation, the coordinate values (3, 4) are mapped back to the range of the 3×3 matrix to generate a legal mapped position. For the data block corresponding to the mapped position, it is regarded as a virtual zero-filled block and assigned a zero value. This processing process is implemented through memory mapping. During data verification calculation, a virtual data block table is maintained. When it is detected that the data block at the mapped position actually does not exist, zero-value data is automatically obtained from this table to participate in the exclusive OR operation.
[0059] Table 1
[0060]
[0061] As shown in Table 1, when calculating the check block Q1 = D 11 ⊕D 22 ⊕D 33 if D 33 is out of bounds, it is mapped to a virtual zero-filled block through modulo operation, and the actual calculation becomes Q1 = D 11 ⊕D 22 ⊕0.
[0062] Specifically, the exclusive OR relationship between the data blocks and the check blocks in Table 1 above is: D 11 ⊕D 21 ⊕c1 = P1; D 12 ⊕D 22 ⊕D 32 = P2; D 13 ⊕D 23 ⊕D 33 = P3; D 11 ⊕D 22 ⊕D 33 =Q1; D 12 ⊕D 23 = Q2; D 13 ⊕0⊕D 31 =Q3; D 21 ⊕D 32 =Q4; D 11 ⊕D 33 ⊕0=R1; D 23 ⊕D 32 ⊕0=R2; D 13 ⊕D 22 ⊕D 31 =R3; D 21 ⊕D 12 ⊕0=R4.
[0063] This processing method not only ensures the continuity of the geometric verification rules but also avoids calculation errors caused by boundary overflows, ensuring the correct generation of verification blocks under various data distribution scenarios and effectively enhancing the data redundancy protection ability of the system.
[0064] According to an embodiment of the present invention, in the coordinate values of the data block positions, the abscissa represents the index of the disk; for the preset diagonal indexing rule, the ordinate in the coordinate values is obtained by subtracting the index of the disk from the index of the data stripe; for the preset anti-diagonal indexing rule, the ordinate in the coordinate values is obtained by adding the index of the data stripe and the index of the disk.
[0065] When implementing the coordinate calculation of the geometric verification rules, a two-dimensional coordinate with the index of the disk as the abscissa and a specific operation result as the ordinate is assigned to each data block. For the diagonal verification rule, the ordinate is calculated by subtracting the index i of the disk from the index j of the current data stripe, forming a diagonal path with a specific slope. For example, if the current processing is the 3rd data stripe and the data block corresponding to the index of the disk is 2, its ordinate is 3 - 2 = 1, thereby determining the position of the data block in the diagonal path.
[0066] For the anti-diagonal verification rule, the ordinate is calculated by adding the index j of the data stripe and the index i of the disk, forming a verification path with a reverse slope. For example, when processing the 3rd data stripe, the ordinate corresponding to the data block with the index of the disk being 2 is 3 + 2 = 5, and an anti-diagonal data block sequence is constructed accordingly. These two coordinate calculation methods efficiently map data blocks to verification paths in different directions through simple addition and subtraction operations, providing a clear data positioning mechanism for parallel execution of geometric verification.
[0067] According to an embodiment of the present invention, the coordinate values of the data block positions are mapped to a range through modulo operation to obtain the mapped position, including: comparing the ordinate in the coordinate values with the range of the data stripe to obtain a comparison result; if the comparison result indicates that the ordinate exceeds the range, a modulo operation is performed on the ordinate to determine the mapped position based on the calculated ordinate.
[0068] When implementing the boundary processing of the geometric verification rules, a range check is performed on the calculated data block coordinates. Specifically, when the ordinate of the data block is determined by the diagonal indexing rule or the anti-diagonal indexing rule, the ordinate value is compared with the valid range of the data stripe. This comparison process is completed by determining whether the ordinate is less than the range lower limit or greater than the range upper limit.
[0069] If the comparison result shows that the ordinate exceeds the valid range, the modulo operation mechanism is immediately activated. For the ordinate value that exceeds the range, perform a modulo operation with the total number of data stripes. For example, if the total number of data stripes is N and the calculated ordinate is Y, when Y is greater than or equal to N, the modulo operation result is Y%N, ensuring that the obtained result is positive and within the valid range.
[0070] The new ordinate value obtained through the modulo operation will replace the original ordinate that exceeds the range, and together with the original abscissa, form the mapping position. Access the corresponding data block according to this mapping position. If there is no physical data block at this position actually, it is regarded as a zero value according to the preset rule and participates in the verification calculation. This processing method ensures that the geometric verification rule can be continuously effective at the data stripe boundary, realizes the logic of virtual zero filling through mathematical mapping, not only ensures the correctness of the verification algorithm, but also avoids complex boundary condition judgments, improving the processing efficiency and reliability of the system.
[0071] After the modulo operation, the diagonal verification rule and the anti-diagonal verification rule verify the written data to generate the corresponding verification blocks. The specific calculation method of the diagonal verification rule is shown in formula (2), and the specific calculation method of the anti-diagonal verification rule is shown in formula (3).
[0072]
[0073] Among them, represents the verification block obtained by the Q verification disk verification, represents the verification block obtained by the R verification disk verification, represents the modulo operation.
[0074] Figure 4 shows the flowchart of data verification according to an embodiment of the present invention.
[0075] As Figure 4 shown, the encoding algorithm for verifying the written data using the verification disks P, Q, and R includes operations S410 to S440.
[0076] In operation S410, determine whether the variable j satisfies "0≤j≤n - 1". If not, the process ends. If satisfied, perform operation S420.
[0077] In operation S420, calculate the horizontal P verification of the j-th row and store the result in P[j]. This step involves performing a horizontal verification operation on the data of the j-th row to ensure the accuracy of the data in this dimension and save the verification result in the corresponding array element.
[0078] In operation S430, the diagonal Q check for the j-th one is calculated and the result is stored in Q[j]. This operation performs a check in the diagonal direction for a specific j-th data block, calculates the check value through a specific algorithm, and stores it in the corresponding position of the Q array.
[0079] In operation S440, the anti-diagonal R check for the j-th one is calculated and the result is stored in R[j]. This step performs a check in the anti-diagonal direction for the j-th data block, and stores the calculated check result in the corresponding position of the R array.
[0080] After completing the above three check calculation operations, the entire check process ends. This process clearly demonstrates that under specific conditions, data is checked in multiple directions and the check results are stored, ensuring the integrity of data checking.
[0081] According to an embodiment of the present invention, if a disk failure is detected, then according to the type of the failed disk and the number of failed disks, a recovery method is determined, including: if the type is a parity disk, it is determined that the recovery method is to positively call the corresponding geometric check rule based on the mapping relationship between data blocks and parity blocks to re-check the data blocks stored in multiple data disks.
[0082] When a disk failure is detected, first identify the type and number of the failed disk to determine the corresponding recovery method. When the failed disk is determined to be a parity disk, activate the forward recovery mechanism based on the geometric check rule. This mechanism relies on the mapping relationship between data blocks and parity blocks and restores the data of the failed parity disk by re-checking.
[0083] According to the type of geometric check rule corresponding to the parity disk (horizontal, diagonal or anti-diagonal), locate the relevant data blocks from multiple data disks. For the horizontal check rule, read the contents of all data blocks in the same data stripe and recalculate the horizontal parity block according to the exclusive OR operation rule. For the diagonal and anti-diagonal check rules, according to the preset index rule, select the data blocks participating in the operation from the data block set across data stripes, and combine the boundary processing mechanism (such as modulo operation and virtual zero padding) to ensure the accuracy of data block selection.
[0084] After obtaining all relevant data blocks, perform the corresponding geometric check operations in parallel to generate new parity blocks. These re-calculated parity blocks will be written into the spare parity disk or the new device after replacing the failed parity disk, thereby restoring the integrity of the parity data. The entire recovery process makes full use of the mathematical characteristics of the geometric check rule, and by positive calculation rather than the traditional data reconstruction method, significantly improves the recovery efficiency, reduces the I / O (input / output) overhead during the data recovery process, and ensures that the system can quickly return to the normal operating state.
[0085] According to an embodiment of the present invention, the number of disks that are allowed to fail simultaneously in a disk array is r, where r is an integer greater than or equal to 1 and less than or equal to n; if a disk failure is detected, the recovery method is determined based on the type of the failed disk and the number of failed disks, and further includes: if the type is a data disk, determining that the recovery method is to reversely call the geometric verification rules of each parity disk based on the number of disks to perform a back substitution process on the parity blocks stored in the parity disks.
[0086] When the disk array detects a failure of a data disk, since the number of disks that are allowed to fail simultaneously in the disk array is r (1 ≤ r ≤ n), a reverse recovery process is started based on the number of failed disks and the failure situation. At this time, the parity blocks stored in the parity disks can be used as key recovery data according to the architecture of the disk array and the corresponding geometric verification rules of each parity disk.
[0087] Perform reverse parsing on the geometric verification rules of each parity disk, and use the mathematical correlation relationship between the parity blocks and the data blocks to construct an equation set. For the horizontal verification rule, combine the parity blocks in the same horizontal verification group with the data blocks of other normal data disks, and try to deduce the original data of the failed data disk through the inverse operation of the exclusive OR operation. Assume that the parity block P1 is combined with the data blocks D 11 、D 21 to try to deduce the original data D 31 of the failed data disk through the inverse operation of the exclusive OR operation. The specific calculation is D 31 = P1 ⊕ D 11 ⊕ D 21 .
[0088] For the diagonal and anti - diagonal verification rules, locate the relevant parity blocks and normal data blocks according to the preset index rules, and also use the reverse logic of the verification rules to disassemble and back - substitute the information contained in the parity blocks. For example, in the diagonal verification rule, the parity block Q1 is combined with the normal data blocks D 11 、D 22 to deduce the failed data block D 33 through the reverse operation, and the calculation is D 33 = Q1 ⊕ D 11 ⊕ D 22 .
[0089] During the back - substitution process, if the number of failed disks is less than the maximum allowable number of failed disks r, use the multi - dimensional parity blocks provided by multiple parity disks to accurately restore the data block content on the failed data disk through a simultaneous reverse calculation method. For example, if the number of failed disks is 2 and r = 3, the parity blocks of the P, Q, and R parity disks can be used to accurately restore the data blocks on the failed data disk through a simultaneous equation set.
[0090] In boundary cases, if the number of faulty disks approaches r, the parity check rules are adaptively adjusted in combination with the extended logic of virtual zero-padding. For example, if three data disks fail simultaneously, approaching the fault tolerance limit of the array, through virtual zero-padding processing, it is ensured that even in such a case, the lost data of the faulty data disks can be recovered from the parity blocks of the parity disks by reversely invoking the geometric parity check rules, guaranteeing the integrity and availability of the data.
[0091] According to an embodiment of the present invention, if a disk failure is detected, the recovery method is determined based on the type of the faulty disk and the number of faulty disks, and further includes: if the types include data disks and parity disks, the geometric parity check rules are combinedly invoked forward and backward based on the number of disks and the mapping relationship between data blocks and parity blocks for data recovery.
[0092] When both data disks and parity disks fail simultaneously in the disk array, a hybrid recovery mechanism is started, and data recovery is performed by combining the forward and backward invocations of the geometric parity check rules. For example, assume there are three data disks (D1, D2, D3) and three parity disks (P, Q, R) in the disk array, and disks D2 and Q fail. In this case, first, it is evaluated whether the number of faulty disks is within the fault tolerance range allowed by the array (i.e., not exceeding r), and the mapping relationship between data blocks and parity blocks is analyzed to determine which data blocks and parity blocks are affected.
[0093] The affected data blocks and parity blocks are divided into different recovery groups. For each recovery group, the corresponding recovery strategy is selected according to its mapping relationship. For data blocks with sufficient parity blocks still available, the geometric parity check rules are invoked forward, and the data of the faulty data disks are recalculated using the parity blocks in the normal parity disks. For example, if the P disk corresponding to the horizontal parity check rule is normal, and the data block D 21 in the D2 disk fails, then the parity block P1 = D 11 ⊕D 21 ⊕D 31 can be used to regenerate the value of the faulty data block D 21 through exclusive OR operation, that is, D 21 = P1⊕D 11 ⊕D31.
[0094] For data blocks that cannot be directly recovered forward due to the failure of the parity disk, the geometric parity check rules are invoked backward, and back substitution calculation is performed in combination with the remaining parity blocks and normal data blocks. For example, when the diagonal parity disk Q fails, the relevant parity block R1 is obtained from the anti-diagonal parity check rule, and the original content of the faulty data block is gradually deduced by reversely analyzing the information contained in these parity blocks. Assume R1 = D 13 ⊕D 22 ⊕D 31 and D22 If a failure occurs, then through D 22 =R1⊕D 13 ⊕D 31 D can be deduced 22 value.
[0095] During the combined call process, a recovery matrix containing forward and reverse equations is constructed. Using the mapping relationship between data blocks and parity blocks, all the data on the failed disks is solved through matrix operations. If a boundary situation is encountered, the virtual zero-padding mechanism is automatically applied to ensure that the geometric parity rules can work properly in any case. In this way, the remaining parity information and data information can be efficiently utilized, while ensuring the recovery accuracy, minimizing the time and resource consumption of data recovery.
[0096] According to an embodiment of the present invention, the data processing method further includes: if it is detected that the failed disk needs to be reconstructed, the recovered data is written into the disk obtained by reconstruction.
[0097] When it is detected that the failed disk needs to be reconstructed, the reconstruction process is first started, and a new physical disk or an active spare disk is allocated in the storage array as the reconstruction target. For example, assuming that the data disk D2 in the disk array fails, a new spare disk D2' will be automatically allocated as the reconstruction target.
[0098] Subsequently, based on the previously recovered data content, the recovered data is written into the reconstructed disk in an orderly manner according to a specific writing strategy. For example, if the recovered data blocks are D 21 、D 22 、D 23 , these data blocks will be written to the corresponding positions of the reconstructed disk D2' according to the original distribution strategy.
[0099] During the writing process, according to the distribution rules of data stripes, the recovered data blocks are mapped to the corresponding positions of the reconstructed disk. For data stored in a striped manner, ensure that the data blocks in each data stripe are written to the correct sectors of the reconstructed disk according to the original distribution strategy. At the same time, maintain the mapping relationship between data blocks and parity blocks to ensure the consistency of the newly written data with the data on other normal disks.
[0100] To ensure data integrity, an integrity check is performed on the reconstructed disk after writing. This includes recalculating the parity blocks of the written data and comparing them with the parity blocks generated during the recovery process. If the comparison results are consistent, it is confirmed that the reconstruction is successful. If there are differences, the recovery and writing operations are performed again until the check passes. The entire reconstruction process will be recorded in the system log for administrators to track and audit. In this way, the reconstruction of the failed disk can be efficiently completed without affecting business continuity, restoring the redundancy ability and data reliability of the storage array.
[0101] According to an embodiment of the present invention, based on the number of disks, the geometric verification rules of each verification disk are called reversely to perform a back substitution process on the verification blocks stored in the verification disk, including: if the number of disks r is 1, a target disk is selected from n verification disks, and according to the geometric verification rule corresponding to the target disk, a reverse exclusive OR operation is performed on the verification blocks in the target disk to obtain the data of the failed disk.
[0102] When only one data disk fails in the disk array (i.e., r = 1), a target disk is intelligently selected from n verification disks for data recovery. This selection process is based on the data mapping relationship between the verification disk and the failed data disk, and the verification disk that has a direct mapping relationship with the failed data disk and has the optimal data distribution is preferentially selected as the target disk. For example, assume that data disk D2 fails, and among verification disks P, Q, and R, disk P has a direct mapping relationship with disk D2 and has the optimal data distribution, then disk P will be preferentially selected as the target disk.
[0103] After the target disk is selected, according to the geometric verification rule corresponding to the target disk (horizontal, diagonal, or anti-diagonal), the reverse exclusive OR operation process is started. For the horizontal verification rule, the data blocks of all other normal data disks in the same horizontal verification group are collected, and by continuously exclusive ORing the contents of these data blocks, the original data of the failed data disk is finally obtained. For the diagonal verification rule or the anti-diagonal verification rule, according to the preset index rule, the data blocks participating in the operation are screened out from the data block set across data stripes, and the failed data is also restored through the reverse exclusive OR operation.
[0104] During the entire back substitution process, boundary conditions are automatically processed. For example, when a virtual zero-padding position is encountered, it is regarded as a zero value to participate in the operation to ensure the accuracy of the reverse calculation. In this way, only by calling the geometric verification rules of a single verification disk, the data of the failed data disk can be efficiently restored, significantly reducing the I / O overhead and computational complexity during the data recovery process, and improving the recovery efficiency and availability of the system.
[0105] According to an embodiment of the present invention, based on the number of disks, the geometric verification rules of each verification disk are called reversely to perform a back substitution process on the verification blocks stored in the verification disk, and further include: if the number of disks r is greater than 1 and less than n, n - 1 target disks are selected from n verification disks; according to the geometric verification rules corresponding to the n - 1 target disks respectively, a system of equations is constructed; through reverse exclusive OR operation, the data of the failed disk is deduced from the system of equations.
[0106] When the number of failed disks in the disk array is greater than 1 and less than n, n-1 target disks are quickly selected from the n check disks. The selection process comprehensively considers the data relevance of each check disk and the failed data disk and the balance of data distribution to ensure that the selected target disk can provide the most effective information for data recovery. For example, if there are 3 check disks (P, Q, R) in the disk array and 2 data disks fail, 2 target disks (such as P and Q) are selected.
[0107] After selecting n-1 target disks, a data recovery equation group is constructed based on the geometric verification rules (horizontal, oblique or reverse oblique) corresponding to each target disk. For the horizontal verification rules, the relationship between the data blocks of the normal data disks and the verification blocks of the target disks in the same horizontal verification group is converted into an equation. For the oblique and reverse oblique verification rules, the logical relationship between the data blocks and the verification blocks across the data stripes is abstracted into equation expressions based on the preset index rules. These equations are associated through the XOR operation between the data blocks and the verification blocks to form a complete system of equations.
[0108] When deriving the data of the faulty disk, the reverse XOR operation characteristics are used to logically decompose the check block value in the equation group and the data block value of the normal data disk. By using the simultaneous equation group and the reversibility of the XOR operation, the redundant terms in the equation are gradually eliminated, and the data block content on the faulty disk is finally solved. During the calculation process, if the boundary position data is involved, the virtual zero padding mechanism is automatically enabled to ensure the integrity of the equation group and the accuracy of the calculation, so as to efficiently and accurately restore the data of the faulty disk and ensure that the disk array quickly restores the data integrity and availability.
[0109] According to an embodiment of the present invention, based on the number of disks, the geometric check rules of each check disk are reversely called to perform back-substitution processing on the check blocks stored in the check disk, and also includes: if the number of disks r is n, a group of equations is constructed according to the geometric check rules corresponding to the n target disks respectively; and the data of the faulty disk is derived from the group of equations through a reverse XOR operation.
[0110] When all n disks in the disk array fail at the same time (i.e., r=n), the full recovery mechanism is started, and a complete set of equations is constructed based on the geometric verification rules of all check disks for data recovery. At this time, all n check disks are regarded as target disks, and the geometric verification rules (horizontal, oblique, or reverse oblique) corresponding to each check disk are used, combined with the mapping relationship between data blocks and check blocks, to construct an equation set containing n independent equations.
[0111] Each equation corresponds to the parity check rule of a parity disk. The left side of the equation is the value of the check block, and the right side is the exclusive OR operation expression of the data blocks participating in this check rule. According to the preset index rule, the data blocks across data stripes are correctly incorporated into the corresponding equations, and virtual zero-padding is performed on the data blocks at the boundary positions to ensure the integrity of the equation set. By the reversibility of the reverse exclusive OR operation, the equation set is converted into a solvable mathematical model, and algorithms such as matrix operations or Gaussian elimination are used to gradually eliminate redundant information and solve the content of the data blocks on all faulty disks.
[0112] During the solution process, the mapping relationship between the parity check rules and the relevance of the data blocks are automatically processed to ensure the independence and solvability of each equation. Since the geometric parity check rules of all parity disks are utilized, the equation set has sufficient constraint conditions to uniquely determine the data of all faulty disks. Finally, the solution results are written into the reconstructed disks, and integrity verification is performed to ensure the accuracy of the restored data. This method makes full use of the redundant characteristics of the geometric parity check rules and can effectively restore data even in extreme cases, ensuring the high availability and data security of the storage system.
[0113] Figure 5 The flowchart of three-disk failure recovery according to an embodiment of the present invention is shown.
[0114] As Figure 5 shown, when 3 disks in the disk array fail, operations S501 to S512 are executed.
[0115] In operation S501, it is judged whether the faulty disks are 3 data disks. If so, operation S512 is executed. If not, operation S502 is executed.
[0116] In operation S502, it is judged whether the faulty disks are 3 parity disks. If so, operation S507 is executed. If not, operation S503 is executed.
[0117] In operation S503, it is judged whether the faulty disks are 2 data disks and 1 parity disk. If so, operation S506 is executed. If not, operation S504 is executed.
[0118] In operation S504, it is judged whether the faulty disks are 1 data disk and 2 parity disks. If so, operation S505 is executed. If not, the repair of the faulty disks is ended.
[0119] In operation S505, it is judged whether the faulty parity disks are the diagonal parity disk and the anti-diagonal parity disk. If not, operation S508 is executed. If so, operation S509 is executed.
[0120] In operation S506, determine whether the failed disk is a horizontal parity disk. If not, perform operation S510. If not, perform operation S510. If so, perform operation S511.
[0121] In operation S507, re-perform verification based on the data disks.
[0122] In operation S508, first recover the failed data disk and the horizontal parity disk according to the normally operating parity disks, and finally calculate the diagonal parity disk (or anti-diagonal parity disk).
[0123] In operation S509, recover the failed data disk according to the horizontal parity disk, and then recover the diagonal parity disk and the anti-diagonal parity disk according to all the recovered data disks.
[0124] In operation S510, recover the failed data disk according to the non-failed horizontal parity disk and the diagonal parity disk (or anti-diagonal parity disk), and then recover the failed diagonal parity disk (or anti-diagonal parity disk) according to the data disks and the horizontal parity disk.
[0125] In operation S511, recover the failed data disk according to the non-failed diagonal parity disk and the anti-diagonal parity disk, and then recover the failed horizontal parity disk according to the data disks.
[0126] In operation S512, first handle one data disk failure, and the remaining two perform data recovery through the recovered parity disks.
[0127] The above decoding algorithm is for the case where 3 disks in the disk array fail. It determines the type of the failed disks through a series of judgments, and then performs corresponding operations, such as recovering data disks using different parity disks, re-verifying, etc., to achieve data repair of the failed disks.
[0128] According to an embodiment of the present invention, the data processing method further includes: if the number of disks r is n, call the decoding algorithm corresponding to the recovery method, and perform the following operations: calculate the relative spacing parameters between the failed disks and the remaining disks according to the distribution positions of the failed disks in the disk array, so as to construct a composite equation set including horizontal parity rules, diagonal parity rules, and anti-diagonal parity rules; based on the calculated minimum iteration step size parameter, iteratively perform multiple rounds of XOR simplification on the composite equation set to gradually eliminate intermediate variables; perform back substitution calculation on the target equation set after eliminating intermediate variables to determine the missing data blocks in the failed disks.
[0129] In the extreme case where all disks (r = n) in the disk array fail, based on the mathematical properties of the geometric parity rules, a composite equation set is constructed for data recovery. First, according to the distribution positions of the failed disks in the array, the relative spacing parameters between them and the remaining disks are calculated. These parameters include the horizontal spacing, the oblique slope, and the anti - oblique slope, etc., which are used to determine the geometric relationships between data blocks. Based on these parameters, the horizontal parity rule, the oblique parity rule, and the anti - oblique parity rule are integrated to construct a composite equation set containing multiple equations. Each equation corresponds to a parity rule, and the relationship between the data block and the parity block is expressed through the exclusive - OR operation.
[0130] To simplify this complex equation set, the minimum iteration step - length parameter is calculated. This parameter is obtained by analyzing the relative position relationships between the failed disks and the periodic characteristics of the parity rules. Based on this minimum iteration step - length parameter, the composite equation set is iteratively simplified through multiple rounds of exclusive - OR operations. In each round of simplification, a specific combination of equations is selected for exclusive - OR operations to gradually eliminate intermediate variables and reduce the complexity of the equation set. This iterative simplification process makes full use of the reversibility and commutativity of the exclusive - OR operation to ensure that the solution of the equation set remains unchanged at each step of simplification.
[0131] After multiple rounds of simplification, the composite equation set is transformed into a target equation set that only contains the data blocks of the failed disks. At this time, back - substitution calculation is used to solve these unknown data blocks. The back - substitution process starts from the known boundary conditions or the virtual zero - padding positions and uses the simplified equations to gradually deduce the data block values on each failed disk. During the calculation process, modulo operations and boundary cases are automatically processed to ensure that all calculations are carried out within the valid range. In this way, the system can still recover the complete data in the extreme case where all disks fail, ensuring the high availability and data security of the storage system. The specific calculation process is as follows.
[0132] Suppose the indexes of the failed disks are x, y, z, where x < y < z. After x, y, z fail, the data blocks of the remaining disks are known. Among x, y, z, the data blocks are unknown and to be recovered. Among them, i and j represent the data block of the j - th stripe of the i - th disk, and there are disks in total, and stripes in total. Corresponding to the virtual data stripe, assume the values are all 0, and let the horizontal P parity disk of the virtual row .
[0133] Step 1: Calculate : .
[0134] Step 2: List the equation set and calculate , where .
[0135] Sub-step 2.1: For the th row, calculate using the horizontal P checksum of the th row: .
[0136] Sub-step 2.2: Calculate using the horizontal P checksum of the th row: .
[0137] Sub-step 2.3: Calculate using the backslash R checksum after : Determine the backslash R checksum of : . Calculate using checksum: .
[0138] Sub-step 2.4: Calculate using the slash Q checksum after : Determine the slash Q checksum involving : . Calculate using the Q checksum: .
[0139] Sub-step 2.5: Add the calculated in the above sub-steps 2.1 to 2.4: . Store .
[0140] Repeat the above sub-steps 2.1 to 2.5 until all are calculated, .
[0141] It should be noted that the additions in step 2 are all performed on the data block and the checksum block, so it is modulo 2 addition , and its essence is the exclusive OR operation of the data. All involved in the calculation, .
[0142] Step 3: Calculate . Respectively satisfy: Take .
[0143] If or , then swap the values of g and h, that is: .
[0144] If or , then ; otherwise ;
[0145] Step 4: For each row, perform pairwise summation simplification and calculate . For the th row, calculate and store .
[0146] Steps 3 and 4 are to ensure the consistency of subsequent processing. In the simulation, it is found that special cases such as m = m1 = m4 or m = m2 = m3 may occur, that is, among m1, m2, m3, and m4, two values are equal and both are equal to the minimum value m. Such special cases need to be carefully handled in the processing of Steps 3 and 4.
[0147] Step 5: Back-substitute to solve, restore the failed disk y, and perform iterative operations in pairwise summation.
[0148] Input: The calculated array , p, flag. Output: Data recovery of the failed disk y.
[0149] Specifically, for j = 0 to (p - 1), perform the following operations: r = (p - 1)+(j × 2h), k = r + 2h, calculate A[y, k]=B[r + flag × 2h]-A[y, r], and the loop ends.
[0150] In the above algorithm: Because in modulo 2 arithmetic, addition and subtraction are equivalent: .
[0151] So far, the data in the failed data disk has all been solved.
[0152] According to the embodiments of the present invention, based on the number of disks and the mapping relationship between data blocks and parity blocks, the geometric parity rules are called in a forward and reverse combination to perform data recovery, including: according to the parity blocks in the normally operating parity disk and the number of disks, the geometric parity rules are reversely called to restore the data of the data disk type in the failed disk; if it is determined that the data recovery of the data disk type in the failed disk is completed, the data blocks stored in multiple data disks are re-verified to restore the data of the parity disk type in the failed disk.
[0153] When both data disks and parity disks fail simultaneously in a disk array, based on the mapping relationship between the number of disks and data blocks - parity blocks, a combined call mechanism for forward and reverse parity check rules is initiated. First, extract the parity block data from the normally operating parity disks, and according to the current number of failed disks, reverse - resolve the mathematical relationship between each parity check rule and the data blocks. For example, for a failed data disk, based on the reverse logic of geometric parity check rules such as diagonal and anti - diagonal, taking the parity block as a known quantity, construct an equation system in combination with the data blocks of the remaining normal data disks, and deduce the original data of the failed data disk through the inverse process of exclusive - OR operation. During this process, calculate the relative spacing parameter according to the distribution position of the failed disks to ensure that the equation system can accurately cover all data blocks to be recovered, and perform virtual zero - padding processing on the boundary positions to ensure the integrity of the calculation.
[0154] After confirming that all the data in the failed data disks has been fully recovered, initiate the forward parity check process to recover the failed parity disks. At this time, the recovered data disks and other normal data disks form a complete data set, and perform exclusive - OR operations on all data blocks again according to the geometric parity check rules of horizontal, diagonal, and anti - diagonal. For example, for the horizontal parity check rule, perform parallel exclusive - OR calculations on the data blocks within the same data stripe to generate a new horizontal parity value. For the diagonal and anti - diagonal parity check rules, select data blocks across stripes according to the preset index rules for parity check operations. The newly generated parity values will be written to the corresponding positions of the failed parity disks according to the mapping relationship between data blocks and parity blocks, completing the recovery of parity check data. During the entire combined call process, improve the recovery efficiency through multi - thread parallel processing, and perform parity value comparison in each step to ensure the accuracy and integrity of data recovery.
[0155] According to an embodiment of the present invention, the data processing method further includes: creating a structure internal memory space for storing the topological relationship of the disk array and the status information of the failed disks, and a full - volume memory page for writing the recovered data; if the data recovery is completed, write the recovered data in the full - volume memory page to the failed disks.
[0156] First, allocate a dedicated structure space in the memory to store the topological relationship of the disk array and the status information of the failed disks. This structure includes the physical layout parameters of the disk array (such as the number of disks, stripe size, parity check rule type, etc.), the logical mapping relationship of each disk (such as the storage location index of data blocks and parity blocks), and the real - time status identifier of the failed disks (such as failure type, failure time, recovery progress, etc.). At the same time, apply for a continuous full - volume memory page, the size of which matches the total capacity of the failed disks, to temporarily store the complete data generated during the recovery process.
[0157] After the data recovery process is started, according to the forward or reverse calculation results of the geometric verification rules, the gradually recovered data blocks are written to the corresponding positions of the full memory pages according to the topology rules of the disk array. For example, the recovered data blocks will be accurately placed at the specified offset addresses of the memory pages according to the striping strategy and the coordinate mapping relationship in the geometric verification rules, ensuring that the physical layout of the data is consistent with the original array. During the recovery process, the fault status information in the structure is updated in real time to record the range of the recovered data blocks and the remaining tasks to be processed.
[0158] After the data of all faulty disks has been calculated through the geometric verification rules and integrated in the memory pages, a data writing operation is triggered. At this time, the storage controller will write the recovered data in the full memory pages to the newly replaced faulty disks or the target disks for reconstruction in sequence according to the disk topology information recorded in the structure. During the writing process, cache prefetching and batch writing strategies are adopted to improve the data writing efficiency, and the integrity verification of the disks is performed after the writing is completed. By comparing the check values of the memory page data and the actual stored data on the disks, it is ensured that the recovered data is accurately written to the faulty disks, and finally the normal operating state of the disk array is restored.
[0159] Based on the above data processing method, the present invention also provides a data processing device. The following will be combined with Figure 6 to describe this device in detail.
[0160] Figure 6 shows a structural block diagram of a data processing device according to an embodiment of the present invention.
[0161] As Figure 6 shown, the data processing device 600 of this embodiment includes a data verification module 610, a method determination module 620, and a data recovery module 630.
[0162] The data verification module 610 is configured to, in response to a data writing operation, verify the written data according to the geometric verification rules corresponding to each verification disk, so as to store the verified check blocks to the corresponding verification disks, and there is a mapping relationship between the check blocks and the written data stored in multiple data disks. In one embodiment, the data verification module 610 may be used to perform the operation S210 described above, which will not be elaborated here.
[0163] The method determination module 620 is configured to, if a disk failure is detected, determine a recovery method according to the type of the faulty disk and the number of faulty disks. In one embodiment, the method determination module 620 may be used to perform the operation S220 described above, which will not be elaborated here.
[0164] A data recovery module 630 is configured to call geometric verification rules forwardly and / or reversely according to a recovery method based on a mapping relationship to recover data of a faulty disk. In an embodiment, the data recovery module 630 may be configured to perform the operation S230 described above, which will not be elaborated herein.
[0165] According to an embodiment of the present invention, any plurality of modules among the data verification module 610, the method determination module 620, and the data recovery module 630 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the data verification module 610, the method determination module 620, and the data recovery module 630 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable means such as hardware or firmware by integrating or packaging circuits, or may be implemented in any one of the three implementation manners of software, hardware, and firmware or in a suitable combination of any several of them. Alternatively, at least one of the data verification module 610, the method determination module 620, and the data recovery module 630 may be at least partially implemented as a computer program module, which may perform corresponding functions when the computer program module is run.
[0166] It should be noted that the data processing device part in the embodiment of the present invention corresponds to the data processing method part in the embodiment of the present invention. For the description of the data processing device part, please refer to the data processing method part specifically, which will not be elaborated herein.
[0167] Figure 7 A block diagram of an electronic device suitable for implementing the data processing method according to an embodiment of the present invention is shown.
[0168] As Figure 7 shown, an electronic device 700 according to an embodiment of the present invention includes a processor 701, which may perform various appropriate actions and processes according to a program stored in a read only memory (ROM) 702 or a program loaded from a storage section 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 701 may also include on-board memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0169] In the RAM 703, various programs and data required for the operation of the electronic device 700 are stored. The processor 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. The processor 701 performs various operations of the method flow according to the embodiments of the present invention by executing programs in the ROM 702 and / or the RAM 703. It should be noted that the programs can also be stored in one or more memories other than the ROM 702 and the RAM 703. The processor 701 can also perform various operations of the method flow according to the embodiments of the present invention by executing programs stored in one or more memories.
[0170] According to an embodiment of the present invention, the electronic device 700 may further include an input / output (I / O) interface 705, and the input / output (I / O) interface 705 is also connected to the bus 704. The electronic device 700 may further include one or more of the following components connected to the input / output (I / O) interface 705: an input portion 706 including a keyboard, a mouse, etc.; an output portion 707 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 708 including a hard disk, etc.; and a communication portion 709 including a network interface card such as a LAN card, a modem, etc. The communication portion 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 710 as needed so that a computer program read from it can be installed into the storage portion 708 as needed.
[0171] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.
[0172] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the above-described ROM 702 and / or RAM 703 and / or one or more memories other than ROM 702 and RAM 703.
[0173] An embodiment of the present invention also includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the data processing method provided by the embodiment of the present invention.
[0174] When the computer program is executed by the processor 701, it executes the above functions defined in the system / apparatus of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0175] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 709, and / or be installed from the removable medium 711. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0176] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 709, and / or be installed from the removable medium 711. When the computer program is executed by the processor 701, it executes the above functions defined in the system of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0177] In accordance with embodiments of the present invention, program code for executing the computer programs provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0178] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0179] Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0180] The above describes the embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments are described separately above, this does not mean that the measures in the respective embodiments cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.
Claims
1. A data processing method, characterized in that, Applied to a disk array, the disk array consists of multiple disks, the multiple disks include multiple data disks and n parity disks, where n is an integer greater than 2, and the method includes: In response to a data write operation, according to the geometric parity rules corresponding to each of the parity disks, perform parity check on the written data to store the parity blocks obtained by the parity check to the corresponding parity disks, and there is a mapping relationship between the parity blocks and the written data stored in the multiple data disks; If it is detected that a disk fails, determine the recovery method according to the type of the failed disk and the number of failed disks; Based on the mapping relationship, forward and / or backward invoke the geometric parity rules according to the recovery method to recover the data of the failed disk.
2. The method according to claim 1, wherein The performing parity check on the written data according to the geometric parity rules corresponding to each of the parity disks includes: Determine the data blocks in which each data stripe in the written data is stored in the multiple data disks, and the data stripe is obtained by splitting the written data according to a preset stripe size; Based on the distribution positions of the data blocks in each data stripe in the multiple data disks, perform parallel parity check on the data stripe according to the geometric parity rules corresponding to each of the parity disks, and the geometric parity rules include a horizontal parity rule, a diagonal parity rule, and an anti-diagonal parity rule.
3. The method according to claim 2, wherein The performing parallel parity check on the data stripe based on the distribution positions of the data blocks in each data stripe in the multiple data disks according to the geometric parity rules corresponding to each of the parity disks includes: For the horizontal parity rule, perform exclusive OR operation on multiple data blocks distributed in the multiple data disks in the same data stripe; For the diagonal parity rule, based on the distribution position, select data blocks across data stripes to perform exclusive OR operation according to a preset diagonal index rule; For the anti-diagonal parity rule, based on the distribution position, select data blocks across data stripes to perform exclusive OR operation according to a preset anti-diagonal index rule; Wherein, when selecting data blocks, the preset diagonal index rule and the preset anti-diagonal index rule perform virtual zero-padding processing for boundary positions.
4. The method according to claim 3, wherein The performing virtual zero-padding processing for boundary positions by the preset diagonal index rule and the preset anti-diagonal index rule when selecting data blocks includes: If the data block position calculated according to the preset diagonal index rule or the preset anti-diagonal index rule exceeds the range of the data stripe in the multiple data disks, map the coordinate value of the data block position to the range through modulo operation to obtain a mapped position; Determine the data block corresponding to the mapped position as a zero value.
5. The method according to claim 4, characterized in that, The abscissa in the coordinate value of the data block position represents the index of the disk; For the preset diagonal index rule, the ordinate in the coordinate value is obtained by subtracting the index of the disk from the index of the data stripe; For the preset anti-diagonal index rule, the ordinate in the coordinate value is obtained by adding the index of the disk to the index of the data stripe.
6. The method according to claim 5, wherein Mapping the coordinate value of the data block position to the range through modulo operation to obtain a mapping position, including: Comparing the ordinate in the coordinate value with the range of the data strip to obtain a comparison result; If the comparison result indicates that the ordinate exceeds the range, perform a modulo operation on the ordinate to determine the mapping position according to the calculated ordinate.
7. The method according to claim 2, characterized in that If it is detected that the disk fails, determining a recovery method according to the type of the failed disk and the number of failed disks, including: If the type is a parity disk, determine that the recovery method is to forwardly call the corresponding geometric parity check rule based on the mapping relationship between the data block and the parity block to re-check the data blocks stored in the multiple data disks.
8. The method according to claim 2, characterized in that, The number of disks that the disk array allows to fail simultaneously is r, where r is an integer greater than or equal to 1 and less than or equal to n; if it is detected that the disk fails, determining a recovery method according to the type of the failed disk and the number of failed disks, further including: If the type is a data disk, determine that the recovery method is to reversely call the geometric parity check rules of each parity disk based on the number of disks to perform a back substitution process on the parity blocks stored in the parity disks.
9. The method according to claim 2, characterized in that, If it is detected that the disk fails, determining a recovery method according to the type of the failed disk and the number of failed disks, further including: If the type includes the data disk and the parity disk, based on the number of disks and the mapping relationship between the data block and the parity block, perform a combined forward and reverse call of the geometric parity check rule to perform data recovery.
10. The method according to claim 1, wherein The method further includes: If it is detected that the failed disk needs to be reconstructed, write the recovered data into the reconstructed disk.
11. The method according to claim 8, wherein The reversely calling the geometric parity check rules of each parity disk based on the number of disks to perform a back substitution process on the parity blocks stored in the parity disks, including: If the number of disks r is 1, select a target disk from the n parity disks, and perform a reverse exclusive OR operation on the parity blocks in the target disk according to the geometric parity check rule corresponding to the target disk to obtain the data of the failed disk.
12. The method according to claim 8, wherein The reversely calling the geometric parity check rules of each parity disk based on the number of disks to perform a back substitution process on the parity blocks stored in the parity disks, further including: If the number of disks r is greater than 1 and less than n, select n - 1 target disks from the n parity disks; Construct a system of equations according to the geometric parity check rules corresponding to the n - 1 target disks respectively; Derive the data of the failed disk from the system of equations through a reverse exclusive OR operation.
13. The method according to claim 8, wherein The reversely calling the geometric parity check rules of each parity disk based on the number of disks to perform a back substitution process on the parity blocks stored in the parity disks, further including: If the number of disks r is n, construct a system of equations according to the geometric parity check rules corresponding to the n parity disks respectively; Derive the data of the failed disk from the system of equations through a reverse exclusive OR operation.
14. The method according to claim 13, wherein The method further includes: If the number r of the disks is n, call the decoding algorithm corresponding to the recovery method and perform the following operations: According to the distribution positions of the faulty disks in the disk array, calculate the relative spacing parameters between the faulty disks and the remaining disks, so as to construct a composite equation set including the horizontal parity check rule, the diagonal parity check rule, and the anti-diagonal parity check rule; Based on the calculated minimum iteration step size parameter, iteratively perform multiple rounds of exclusive OR simplification on the composite equation set to gradually eliminate intermediate variables; By performing back substitution calculation on the target equation set after eliminating intermediate variables, determine the missing data blocks in the faulty disks.
15. The method according to claim 9, wherein Based on the number of disks and the mapping relationship between the data blocks and the parity check blocks, the combined forward and reverse calls of the geometric parity check rules for data recovery include: According to the parity check blocks in the normally operating parity check disks and the number of disks, restore the data of the data disks among the faulty disks by reversely calling the corresponding geometric parity check rules; If it is determined that the data recovery of the data disks among the faulty disks is completed, re-check the data blocks stored in the multiple data disks to restore the data of the parity check disks among the faulty disks.
16. The method according to claim 1, characterized in that The method further includes: Create a structure memory space for storing the topological relationship of the disk array and the status information of the faulty disks, and a full memory page for writing recovery data; If the data recovery is completed, write the recovery data in the full memory page to the faulty disks.
17. A data processing device, characterized in that, Applied to a disk array, the disk array is composed of multiple disks, the multiple disks include multiple data disks and n parity check disks, n is an integer greater than 2, and the device includes: A data verification module, configured to, in response to a data writing operation, verify the written data according to the geometric parity check rules corresponding to the respective parity check disks, so as to store the verified parity check blocks to the corresponding parity check disks, and there is a mapping relationship between the parity check blocks and the written data stored in the multiple data disks; A method determination module, configured to, if a disk failure is detected, determine a recovery method according to the type of the faulty disk and the number of faulty disks; A data recovery module, configured to, based on the mapping relationship, forwardly and / or reversely call the geometric parity check rules according to the recovery method to recover the data of the faulty disks.
18. An electronic device, including: One or more processors; A memory for storing one or more computer programs, Characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 16.
19. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, The computer program or instruction, when executed by the processor, implements the steps of the method according to any one of claims 1 to 16.
20. A computer program product, characterized in that, Including a computer program, the computer program, when executed by the processor, implements the method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Method and device for data backup and recovery in redundant array of inexpensive disks
CN101770409A
Method for constructing disk array by horizontal grouping parallel concentrated verification
CN101976175A
Device, program, recording medium, and method for extending service life of memory
US20170003890A1
Cited By
RAID data processing method and device, chip, electronic equipment, storage medium and computer program product
CN120704957A
Abnormal data processing method and device, storage medium and electronic equipment
CN120723520A