Data processing method, device, equipment, medium and program product

By setting up multiple data disks and check disks in the disk array, using geometric check rules to generate check blocks and recover data, the problem of low data recovery efficiency when multiple disks fail is solved, and efficient and accurate data recovery is achieved.

CN120371595BActive Publication Date: 2025-09-16INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510866424.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-16
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

In the prior art, when multiple disks fail simultaneously, it is difficult to directly recover data by relying solely on dual distributed parity checking, which consumes large computing resources and has low recovery efficiency.

Method used

By setting up multiple data disks and n check disks in the disk array, geometric check rules are used to generate check blocks for written data, a mapping relationship is established, and when a disk fails, the geometric check rules are called forward and/or reversely based on the mapping relationship to recover data according to the type and number of failures.

Benefits of technology

It achieves efficient data recovery in high-fault-tolerance scenarios with multiple failed disks, improves the efficiency and accuracy of data recovery, and ensures the availability of the storage system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371595B_ABST
    Figure CN120371595B_ABST
Patent Text Reader

Abstract

The present invention provides a data processing method applicable to the field of computer technology. The method comprises: in response to a data write operation, verifying the written data according to geometric verification rules corresponding to each verification disk, storing the verified verification blocks on the corresponding verification disk, and mapping the verification blocks to the written data stored on multiple data disks; if a disk failure is detected, determining a recovery method based on the type of failed disk and the number of failed disks; and based on the mapping relationship, forward and / or reversely invoking the geometric verification rules according to the recovery method to recover data on the failed disk. The present invention also provides a data processing device, apparatus, medium, and program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a data processing method, apparatus, device, medium and program product. Background Art

[0002] As storage demands continue to increase, the risk of disk damage and failure also increases. To ensure high data reliability and availability, storage systems typically use dual distributed parity technology to achieve data redundancy.

[0003] During the implementation of the present invention, we discovered that related technologies have at least the following problems: When multiple disks fail simultaneously, it's difficult to directly recover data using only dual distributed parity checking, often requiring complex recursive calculations. This approach not only consumes large amounts of computing resources but also results in low recovery efficiency. Summary of the Invention

[0004] In view of the above problems, the present invention provides a data processing method, apparatus, device, medium and program product.

[0005] According to the first aspect of the present invention, a data processing method is provided, comprising: in response to a data write operation, verifying the written data according to the geometric verification rules corresponding to each of the above-mentioned verification disks, so as to store the verification blocks obtained by verification in the corresponding verification disks, and there is a mapping relationship between the above-mentioned verification blocks and the written data stored in the above-mentioned multiple data disks; if it is detected that the above-mentioned disk fails, determining a recovery method according to the type of the failed disk and the number of failed disks; based on the above-mentioned mapping relationship, calling the above-mentioned geometric verification rules forwardly and / or reversely according to the above-mentioned recovery method to recover the data of the above-mentioned failed disk.

[0006] The second aspect of the present invention provides a data processing device, including: a data verification module, which is used to verify the written data in response to the data writing operation according to the geometric verification rules corresponding to each of the above-mentioned verification disks, so as to store the verification blocks obtained by verification to the corresponding verification disks, and there is a mapping relationship between the above-mentioned verification blocks and the written data stored to the above-mentioned multiple data disks; a method determination module, which is used to determine the recovery method according to the type of the failed disk and the number of failed disks if a failure of the above-mentioned disk is detected; a data recovery module, which is used to call the above-mentioned geometric verification rules forwardly and / or reversely according to the above-mentioned recovery method based on the above-mentioned mapping relationship to recover the data of the above-mentioned failed disk.

[0007] A third aspect of the present invention provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0008] The fourth aspect of the present invention further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.

[0009] The fifth aspect of the present invention further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor.

[0010] According to an embodiment of the present invention, by configuring multiple data disks and n check disks in a disk array and utilizing geometric check rules to generate check blocks for written data, a mapping relationship is established between the check blocks and the data blocks, thereby improving the storage system's fault tolerance. Subsequently, in the event of a disk failure, data is recovered by applying the geometric check rules forward and / or backward based on the mapping relationship, depending on the type and number of failures. This method not only supports high fault tolerance scenarios where multiple disks fail simultaneously, but also effectively improves the efficiency and accuracy of data recovery by combining the mapping relationship with the geometric check rules. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above contents and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings.

[0012] Figure 1 An application scenario diagram of a data processing method, apparatus, device, medium, and program product according to an embodiment of the present invention is shown.

[0013] Figure 2 A flow chart of a data processing method according to an embodiment of the present invention is shown.

[0014] Figure 3 A schematic diagram of performing verification according to geometric verification rules in an embodiment of the present invention is shown.

[0015] Figure 4 A flow chart of data verification according to an embodiment of the present invention is shown.

[0016] Figure 5 A flowchart of three-disk failure recovery according to an embodiment of the present invention is shown.

[0017] Figure 6 A structural block diagram of a data processing device according to an embodiment of the present invention is shown.

[0018] Figure 7 A block diagram of an electronic device suitable for implementing a data processing method according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0019] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessarily confusing the concept of the present invention.

[0020] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise", "include", etc. used herein indicate the presence of the features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.

[0021] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0022] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0023] In the technical solution of the present invention, the data involved (including but not limited to data used for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0024] An embodiment of the present invention provides a data processing method, which responds to a data write operation, verifies the written data according to the geometric verification rules corresponding to each verification disk, and stores the verification blocks obtained in the verification on the corresponding verification disk. There is a mapping relationship between the verification blocks and the written data stored on multiple data disks; if a disk failure is detected, a recovery method is determined according to the type of the failed disk and the number of failed disks; based on the mapping relationship, the geometric verification rules are called forward and / or reversely according to the recovery method to recover the data of the failed disk.

[0025] Figure 1 An application scenario diagram of a data processing method, apparatus, device, medium, and program product according to an embodiment of the present invention is shown.

[0026] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or optical fiber cables.

[0027] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a server 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).

[0028] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.

[0029] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.

[0030] It should be noted that the data processing method provided in the embodiment of the present invention can generally be executed by the server 105. Accordingly, the data processing device provided in the embodiment of the present invention can generally be set in the server 105. The data processing method provided in the embodiment of the present invention can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the data processing device provided in the embodiment of the present invention can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.

[0031] It should be understood that Figure 1 The number of the first terminal device, the second terminal device, the third terminal device, the network and the server is only . According to the implementation requirements, there can be any number of the first terminal device, the second terminal device, the third terminal device, the network and the server.

[0032] The following will be based on Figure 1 The scene described by Figures 2 to 5 The data processing method of the disclosed embodiment is described in detail.

[0033] Figure 2 A flow chart of a data processing method according to an embodiment of the present invention is shown.

[0034] like Figure 2 As shown, this embodiment includes operations S210 to S230.

[0035] In operation S210, in response to the data writing operation, the written data is verified according to the geometric verification rules corresponding to each verification disk, so that the verification block obtained is stored in the corresponding verification disk, and there is a mapping relationship between the verification block and the written data stored in multiple data disks.

[0036] In operation S220 , if a disk failure is detected, a recovery method is determined based on the type of failed disks and the number of failed disks.

[0037] In operation S230 , based on the mapping relationship, the geometry check rule is called forward and / or reversely according to the recovery method to recover the data of the failed disk.

[0038] According to an embodiment of the present invention, when a disk array receives a data write operation, it verifies the written data according to the geometric verification rules corresponding to each verification disk. In addition to the multiple data disks, the disk array is configured with n verification disks, where n is an integer greater than 2. For example, three verification disks (P, Q, and R) can be configured, each corresponding to a different geometric verification rule.

[0039] The written data will first be calculated according to the geometric verification rules corresponding to each verification disk to generate verification blocks that have a mapping relationship with the written data. These verification blocks will be stored in the corresponding verification disks respectively, and the original written data will be stored in multiple data disks, thereby ensuring the integrity and reliability of the data through the verification mechanism.

[0040] During disk array operation, the system continuously monitors the status of each disk. When a disk failure is detected, the system immediately determines the type of failed disk (data disk or parity disk) and counts the number of failed disks. For example, if data disk D1 is detected as faulty, the fault type is recorded as data disk and the number of failed disks is recorded as 1.

[0041] Determine the appropriate recovery method based on the type and number of failed disks. Different failure scenarios may require different recovery strategies. For example, if only one data disk fails, data recovery can be performed directly using the P drive, Q drive, or R drive. If multiple data disks fail simultaneously, a complex recovery may require combining the parity information of the Q and R drives to ensure maximum recovery efficiency while maintaining data recovery accuracy.

[0042] After determining the recovery method, the system uses the mapping between the check blocks and the written data to apply the geometric check rules in a forward and / or reverse manner, as required by the recovery method. This method uses the data blocks and check blocks stored on the surviving disks to recalculate and generate the data on the failed disk, thus restoring the data on the failed disk and enabling the disk array to return to normal operation as quickly as possible, ensuring data availability.

[0043] According to an embodiment of the present invention, by configuring multiple data disks and n check disks in a disk array and utilizing geometric check rules to generate check blocks for written data, a mapping relationship is established between the check blocks and the data blocks, thereby improving the storage system's fault tolerance. Subsequently, in the event of a disk failure, data is recovered by applying the geometric check rules forward and / or backward based on the mapping relationship, depending on the type and number of failures. This method not only supports high fault tolerance scenarios where multiple disks fail simultaneously, but also effectively improves the efficiency and accuracy of data recovery by combining the mapping relationship with the geometric check rules.

[0044] According to an embodiment of the present invention, the written data is verified according to the geometric verification rules corresponding to each verification disk, including: determining that each data stripe in the written data is respectively stored in the data blocks in multiple data disks, and the data stripes are obtained by dividing the written data according to a preset stripe size; based on the distribution position of each data block in the data stripe in the multiple data disks, the data stripes are verified in parallel according to the geometric verification rules corresponding to each verification disk, and the geometric verification rules include horizontal verification rules, oblique verification rules and reverse oblique verification rules.

[0045] The written data is split according to the preset stripe size to obtain several data stripes. The storage location of each data block in each data stripe on multiple data disks is then determined. For example, assuming the written data is a large file with a preset stripe size of 64KB, the file is split into multiple 64KB data stripes. This process requires combining the striping strategy of the storage system to split the data stripe into multiple data blocks according to specific rules and map them to different data disk address spaces, such as distributing them in a row-first or column-first manner to ensure that the data blocks are evenly distributed in the disk array. For example, in an array containing three data disks, data blocks can be stored in D1, D2, and D3 in a row-first manner.

[0046] After the data blocks are distributed and located, the data stripes are checked in parallel based on the specific distribution locations of each data block across multiple data disks and the geometric check rules corresponding to each check disk. Horizontal check rules can perform check calculations on data blocks in the same row or horizontal dimension, for example, by generating horizontal check blocks through an exclusive-or operation. Diagonal check rules focus on the paths along which data blocks are distributed diagonally within the disk array, performing combined check calculations on data blocks along diagonal lines with a specific slope. Reverse diagonal check rules correspond to diagonal lines with opposite slopes, performing check calculations on data blocks along that path.

[0047] Figure 3 A schematic diagram of performing verification according to geometric verification rules in an embodiment of the present invention is shown.

[0048] like Figure 3 As shown, during the verification of written data, data blocks 305 stored in multiple data disks 304 are verified using horizontal verification rules 301, diagonal verification rules 302, and reverse diagonal verification rules 303. Specifically, horizontal verification rules 301 verify data blocks from left to right, diagonal verification rules 302 verify data blocks along a specific diagonal direction, and reverse diagonal verification rules 303 verify data blocks in a direction opposite to the diagonal direction. The synergistic effect of these three verification rules ensures data accuracy and integrity.

[0049] In specific implementations, multithreading or hardware parallel processing units can be used to simultaneously initiate horizontal, diagonal, and reverse diagonal checksum calculation tasks for each data stripe. Each checksum task reads the corresponding data block from the data disk based on the corresponding geometric rule, performs calculations based on the rule to generate a checksum, and writes the checksum to the corresponding checksum disk. For example, one checksum disk can store checksums generated by the horizontal checksum rule, while another can store checksums generated by the diagonal checksum rule. This method enables parallel checksum processing based on geometric rules, ensuring that the integrity and reliability of written data are efficiently verified during data storage.

[0050] According to an embodiment of the present invention, based on the distribution position of each data block in the data stripe in multiple data disks, the data stripes are verified in parallel according to the geometric verification rules corresponding to each verification disk, including: for the horizontal verification rule, performing an XOR operation on multiple data blocks distributed in multiple data disks in the same data stripe; for the oblique verification rule, based on the distribution position, according to the preset oblique index rule, selecting the data blocks across the data stripes to perform an XOR operation; for the reverse oblique verification rule, based on the distribution position, according to the preset reverse oblique index rule, selecting the data blocks across the data stripes to perform an XOR operation; wherein, when selecting the data blocks, the preset oblique index rule and the preset reverse oblique index rule adopt virtual zero padding for the boundary positions.

[0051] In the parallel implementation of geometric validation rules based on the distribution of data blocks, a dedicated processing thread is started for the horizontal validation rules. For example, suppose the current data stripe contains data block D 11 、D 21 、D 31 , stored on data disks D1, D2, and D3 respectively. This thread obtains the data block address mapping table of the current data stripe in multiple data disks from the storage controller, and then reads the contents of all data blocks in the same data stripe in sequence. By looping and performing XOR operations, the binary value of each data block is XORed bit by bit, and finally a horizontal parity block is generated. P1=D 11 ⊕D 21 ⊕D 31 This process is performed in real time while the data is being written, ensuring that the parity block and data block writing operations are completed synchronously. The specific calculation method is shown in formula (1).

[0052]

[0053] in, Represents the data block of the jth data stripe of the i-th disk (i-th column, j-th row). There are a total of (three parity disks) disks, with a total of indivual, and Both represent the XOR operation between data blocks, Indicates the parity block obtained by parity disk P.

[0054] For the oblique verification rule, an independent calculation task is created based on the preset oblique index rule. For example, if the location information of the current data stripe indicates that the data block D on the oblique path needs to be calculated 11 、D 22 、D 33This task will calculate the addresses of all data blocks on the diagonal path based on the location information of the current data stripe and the topological structure of the data disk array. During the calculation process, if a boundary position is encountered, the virtual zero-padding processing mechanism will be automatically triggered to generate a virtual data block of all zeros to replace the actual data block that does not exist to participate in the calculation. The calculation task will traverse all data blocks on the diagonal path (including virtual zero-padding blocks) and perform an exclusive OR operation to generate a diagonal check block Q1=D 11 ⊕D 22 ⊕D 33 and write the results to the corresponding check disk area.

[0055] The implementation of the reverse oblique check rule is similar to the oblique check rule, but uses the opposite index direction. Start a dedicated reverse oblique calculation thread and determine the location of data blocks across data stripes according to the preset reverse oblique index rule. For example, suppose you need to calculate the data block D on the reverse oblique path. 13 、D 22 、D 31 Similarly, virtual zero padding is performed when encountering boundary conditions to ensure the continuity of the reverse oblique path. The thread will read the actual data block from the data disk and perform an XOR operation with the virtual zero padding block to finally generate the reverse oblique check block R1=D 13 ⊕D 22 ⊕D 31 These three verification tasks are executed in parallel through the thread pool, and each independently calculates the verification block, thereby efficiently completing multi-dimensional verification protection without affecting data writing performance, ensuring data integrity and reliability.

[0056] According to an embodiment of the present invention, when selecting a data block, the preset oblique indexing rule and the preset reverse oblique indexing rule adopt virtual zero-padding processing for the boundary position, including: if the data block position calculated according to the preset oblique indexing rule or the preset reverse oblique indexing rule exceeds the range of data stripes in multiple data disks, then the coordinate value of the data block position is mapped to the range through a modulo operation to obtain the mapping position; and the data block corresponding to the mapping position is determined to be a zero value.

[0057] When implementing boundary processing within geometric validation rules, range checks are performed on the data block locations calculated using the oblique and reverse oblique indexing rules. For example, assume that data stripes are distributed across three data disks, each with three data blocks, forming a 3×3 data matrix. If the calculated data block coordinate value exceeds the range of the data stripe across multiple data disks, a modulo calculation mechanism is automatically triggered. This mechanism performs a modulo calculation on the out-of-range coordinate value with the dimension parameters of the data disk array, mapping the coordinate value back into the range and generating a valid mapping position.

[0058] For example, suppose the oblique indexing rule calculates a data block position of (3,4), which is outside the range of the 3×3 matrix. Using a modulo operation, the coordinate value (3,4) is mapped back to the range of the 3×3 matrix, generating a valid mapping position. The data block corresponding to the mapped position is treated as a virtual zero-padding block and assigned a value of zero. This process is implemented through memory mapping. During data verification calculations, a virtual data block table is maintained. If it is detected that the data block at the mapped position does not actually exist, the zero-valued data is automatically retrieved from this table for exclusive OR operation.

[0059] Table 1

[0060]

[0061] As shown in Table 1, when calculating the check block Q1=D of the oblique check rule 11 ⊕D 22 ⊕D 33 When D 33 Out of bounds, mapped to the virtual zero-filling block through modulo operation, the actual calculation becomes Q1=D 11 ⊕D 22 ⊕0.

[0062] Specifically, the XOR relationship between the data block and the check block in Table 1 above is: 11 ⊕D 21 ⊕c1=P1;D 12 ⊕D 22 ⊕D 32 = P2;D 13 ⊕D 23 ⊕D 33 = P3;D 11 ⊕D 22 ⊕D 33 =Q1;D 12 ⊕D 23 = Q2;D 13 ⊕0⊕D 31 =Q3;D 21 ⊕D 32 =Q4;D 11 ⊕D 33 ⊕0=R1;D 23 ⊕D 32 ⊕0=R2;D 13 ⊕D 22 ⊕D 31 =R3;D 21 ⊕D 12 ⊕0=R4.

[0063] This processing method not only ensures the continuity of the geometric verification rules, but also avoids calculation errors caused by boundary violations, ensuring that the verification blocks can be correctly generated under various data distribution conditions, and effectively improving the system's data redundancy protection capabilities.

[0064] According to an embodiment of the present invention, the horizontal coordinate in the coordinate value of the data block position represents the index of the disk; for the preset oblique index rule, the vertical coordinate in the coordinate value is obtained by subtracting the index of the data stripe from the index of the disk; for the preset reverse oblique index rule, the vertical coordinate in the coordinate value is obtained by summing the index of the data stripe and the index of the disk.

[0065] When implementing coordinate calculations for geometric validation rules, each data block is assigned a two-dimensional coordinate, with the disk index as the horizontal coordinate and the result of a specific operation as the vertical coordinate. For diagonal validation rules, the vertical coordinate is calculated by subtracting the disk index i from the current data stripe's index j, forming a diagonal path with a specific slope. For example, if the current data stripe is being processed, the corresponding data block with disk index 2 has a vertical coordinate of 3-2=1, thus determining the data block's position in the diagonal path.

[0066] For the reverse-slope verification rule, the vertical coordinate is calculated by adding the data stripe index j to the disk index i, forming a reverse-slope verification path. For example, when processing the third data stripe, the vertical coordinate corresponding to the data block with disk index 2 is 3+2=5, and a reverse-slope data block sequence is constructed in this manner. These two coordinate calculation methods efficiently map data blocks to verification paths in different directions through simple addition and subtraction operations, providing a clear data location mechanism for parallel geometric verification.

[0067] According to an embodiment of the present invention, the coordinate value of the data block position is mapped to a range through a modulo operation to obtain a mapping position, including: comparing the vertical coordinate in the coordinate value with the range of the data stripe to obtain a comparison result; if the comparison result indicates that the vertical coordinate is out of range, a modulo operation is performed on the vertical coordinate to determine the mapping position based on the calculated vertical coordinate.

[0068] When implementing boundary processing for geometric validation rules, a range check is performed on the calculated data block coordinates. Specifically, after the vertical coordinate of the data block is determined using the oblique indexing rule or the reverse oblique indexing rule, the vertical coordinate value is compared with the valid range of the data stripe. This comparison is performed by determining whether the vertical coordinate is less than the lower limit of the range or greater than the upper limit of the range.

[0069] If the comparison result indicates that the vertical coordinate is outside the valid range, the modulo operation mechanism is immediately activated. For the out-of-range vertical coordinate value, the modulo operation is performed on it with the total number of data stripes. For example, if the total number of data stripes is N, and the calculated vertical coordinate is Y, when Y is greater than or equal to N, the modulo operation result is Y%N, ensuring that the result is positive and within the valid range.

[0070] The new ordinate value obtained through the modulo operation replaces the original out-of-range ordinate and, together with the original abscissa, forms the mapping position. The corresponding data block is accessed based on this mapping position. If no physical data block exists at that location, it is treated as zero for verification calculations according to preset rules. This processing method ensures that the geometric verification rules remain valid at data stripe boundaries. It implements virtual zero-padding logic through mathematical mapping, ensuring the correctness of the verification algorithm while avoiding complex boundary condition judgments, thereby improving the system's processing efficiency and reliability.

[0071] After performing the modulo operation, the oblique check rule and the reverse oblique check rule are used to check the written data to generate the corresponding check block. The specific calculation method of the oblique check rule is shown in formula (2), and the specific calculation method of the reverse oblique check rule is shown in formula (3).

[0072]

[0073] in, Indicates the parity block obtained by the Q parity disk. Indicates the check block obtained by checking the R check disk. Represents the modulo operation.

[0074] Figure 4 A flow chart of data verification according to an embodiment of the present invention is shown.

[0075] like Figure 4 As shown, the coding algorithm for verifying written data using the verification disks P, Q, and R includes operations S410 to S440.

[0076] In operation S410, it is determined whether the variable j satisfies "0≤j≤n-1". If not, the process ends. If so, operation S420 is executed.

[0077] In operation S420, the horizontal P check of row j is calculated and the result is stored in P[j]. This step involves performing a horizontal check operation on the j-th row of data to ensure the accuracy of the data in that dimension and storing the check result in the corresponding array element.

[0078] In operation S430, the j-th oblique Q check is calculated and the result is stored in Q[j]. This operation performs an oblique check on the j-th data block, calculates the check value using a specific algorithm, and stores it in the corresponding position of the Q array.

[0079] In operation S440, the jth reverse oblique direction R check is calculated and the result is stored in R[j]. This step is to perform a reverse oblique direction check on the jth data block and store the calculated check result in the corresponding position of the R array.

[0080] After completing the three verification calculation operations above, the entire verification process ends. This process clearly demonstrates that under certain conditions, multi-directional data verification and storage of verification results ensure the integrity of data verification.

[0081] According to an embodiment of the present invention, if a disk failure is detected, a recovery method is determined based on the type of the failed disk and the number of failed disks, including: if the type is a check disk, the recovery method is determined to be, based on the mapping relationship between data blocks and check blocks, forward calling the corresponding geometric check rules to recheck the data blocks stored in multiple data disks.

[0082] When a disk failure is detected, the system first identifies the type and number of failed disks to determine the appropriate recovery method. If the failed disk is identified as a parity disk, a forward recovery mechanism based on geometric verification rules is activated. This mechanism relies on the mapping relationship between data blocks and parity blocks, and recovers the data on the failed parity disk through reverification.

[0083] Based on the geometric check rule type (horizontal, diagonal, or reverse diagonal) corresponding to the parity disk, the relevant data blocks are located from multiple data disks. For horizontal check rules, the contents of all data blocks in the same data stripe are read and the horizontal check block is recalculated using the XOR operation. For diagonal and reverse diagonal check rules, the data blocks involved in the operation are selected from the set of data blocks across the data stripe according to the preset index rules. Boundary processing mechanisms (such as modulo operation and virtual zero padding) are used to ensure the accuracy of data block selection.

[0084] After acquiring all relevant data blocks, the corresponding geometric checksum calculations are performed in parallel to generate new checksum blocks. These recalculated checksum blocks are written to a spare checksum disk or a new device that replaces the failed checksum disk, thereby restoring the integrity of the checksum data. The entire recovery process fully utilizes the mathematical properties of geometric checksum rules. By using forward calculations rather than traditional data reconstruction methods, recovery efficiency is significantly improved, I / O (input / output) overhead is reduced during data recovery, and the system can quickly return to normal operation.

[0085] According to an embodiment of the present invention, a disk array allows r to have simultaneous failures in the number of disks, where r is an integer greater than or equal to 1 and less than or equal to n; if a disk failure is detected, a recovery method is determined based on the type of the failed disk and the number of failed disks, and the method also includes: if the type is a data disk, the recovery method is determined to be, based on the number of disks, reversely calling the geometric check rules of each check disk to perform back-replacement processing on the check blocks stored in the check disk.

[0086] When a disk array detects a data disk failure, it initiates a reverse recovery process based on the number of failed disks and the failure conditions, as the number of disks that can fail simultaneously is r (1≤r≤n). Based on the disk array architecture and the geometric verification rules for each parity disk, the parity blocks stored on the parity disks are used as key recovery data.

[0087] The geometric check rules of each check disk are reverse analyzed, and the mathematical relationship between the check blocks and the data blocks is used to construct the process group. According to the horizontal check rule, the check blocks in the same horizontal check group are combined with the data blocks of other normal data disks, and the original data of the faulty data disk is tried to be deduced through the inverse operation of the XOR operation. Assume that the check block P1 and the data block D of the normal data disk are 11 、D 21 Combined with the inverse operation of XOR operation, try to deduce the original data D of the faulty data disk 31 The specific calculation is D 31 =P1⊕D 11 ⊕D 21 .

[0088] For oblique and reverse oblique check rules, the relevant check blocks and normal data blocks are located according to the preset index rules, and the reverse logic of the check rules is also used to disassemble and reverse the information contained in the check blocks. For example, in the oblique check rule, the check block Q1 and the normal data block D 11 、D 22 Combined, the fault data block D is derived through reverse operation 33 , calculated as D 33 =Q1⊕D 11 ⊕D 22 .

[0089] During back-replacement, if the number of failed disks is less than the maximum allowed number of failures, r, the multi-dimensional parity blocks provided by multiple parity disks are used to accurately restore the data blocks on the failed data disk through a simultaneous reverse calculation method. For example, if the number of failed disks is 2 and r = 3, the parity blocks of the three parity disks, P, Q, and R, can be used to accurately restore the data blocks on the failed data disk through a system of simultaneous equations.

[0090] In extreme cases, if the number of failed disks approaches r, the validation rules are adaptively adjusted by incorporating extended logic using virtual zero padding. For example, if three data disks fail simultaneously, approaching the array's fault tolerance limit, virtual zero padding ensures that even in this situation, the data lost on the failed data disk can be recovered from the parity blocks on the parity disk by reversing the geometric validation rules, thus ensuring data integrity and availability.

[0091] According to an embodiment of the present invention, if a disk failure is detected, a recovery method is determined based on the type of the failed disk and the number of failed disks. It also includes: if the type includes data disks and check disks, based on the number of disks and the mapping relationship between data blocks and check blocks, a forward and reverse combination of calls are made to the geometric check rules to perform data recovery.

[0092] When simultaneous data and parity drive failures are detected in a disk array, a hybrid recovery mechanism is activated, combining forward and reverse invocation of geometric parity rules for data recovery. For example, suppose a disk array has three data drives (D1, D2, and D3) and three parity drives (P, Q, and R), and drive D2 and drive Q fail. In this case, the system first assesses whether the number of failed drives is within the array's fault tolerance (i.e., no more than r). The system then analyzes the mapping between data blocks and parity blocks to determine which data and parity blocks are affected.

[0093] The affected data blocks and parity blocks are divided into different recovery groups. For each recovery group, the corresponding recovery strategy is selected based on its mapping relationship. For data blocks that still have enough parity blocks available, the geometric check rule is called forward, and the data on the failed data disk is recalculated using the parity blocks in the normal parity disk. For example, if the P disk corresponding to the horizontal check rule is normal, and the data block D in the D2 disk is 21 If a failure occurs, the check block P1=D of the P disk can be used to check the 11 ⊕D 21 ⊕D 31 , use XOR operation to regenerate the fault data block D 21 The value of D 21 =P1⊕D 11 ⊕D31.

[0094] For data blocks that cannot be directly recovered due to a parity disk failure, the geometric check rule is called in reverse, and the remaining check blocks and normal data blocks are combined for back-calculation. For example, when the oblique check disk Q fails, the relevant check blocks R1 are obtained from the reverse oblique check rule. By reverse parsing the information contained in these check blocks, the original content of the failed data block can be gradually deduced. Assume that R1 = D 13 ⊕D 22 ⊕D 31 , and D22 If a fault occurs, you can 22 =R1⊕D 13 ⊕D 31 Derive D 22 value.

[0095] During the combined call process, a recovery matrix consisting of forward and inverse equations is constructed. Leveraging the mapping between data blocks and parity blocks, matrix operations are used to determine the data on all failed disks. In edge cases, a virtual zero-padding mechanism is automatically applied to ensure that geometric checksum rules function properly under all circumstances. This approach efficiently utilizes remaining parity and data information, minimizing recovery time and resource consumption while ensuring recovery accuracy.

[0096] According to an embodiment of the present invention, the data processing method further includes: if it is detected that the failed disk needs to be reconstructed, writing the recovered data into the reconstructed disk.

[0097] When a failed disk is detected and needs to be rebuilt, the reconstruction process is first initiated, and a new physical disk is allocated in the storage array or a spare disk is activated as the reconstruction target. For example, if data disk D2 in the disk array fails, a new spare disk D2' will be automatically allocated as the reconstruction target.

[0098] Then, based on the previously recovered data content, the recovered data is written to the reconstructed disk in an orderly manner according to a specific writing strategy. For example, if the recovered data block is D 21 、D 22 、D 23 , these data blocks will be written to the corresponding positions of the reconstructed disk D2 'according to the original distribution strategy.

[0099] During the write process, recovered data blocks are mapped to corresponding locations on the reconstructed disk according to the data stripe distribution rules. For striped data, this ensures that data blocks in each stripe are written to the correct sectors on the reconstructed disk according to the original distribution strategy. Furthermore, the mapping between data blocks and parity blocks is maintained to ensure that newly written data remains consistent with data on other healthy disks.

[0100] To ensure data integrity, the reconstructed disk undergoes an integrity check after the write is complete. This involves recalculating the check block of the written data and comparing it with the check block generated during the recovery process. If the comparison results are consistent, the reconstruction is confirmed to be successful. If there are any discrepancies, the recovery and write operations are repeated until the check passes. The entire reconstruction process is recorded in the system log for administrators to track and audit. This method allows for efficient reconstruction of failed disks without affecting business continuity, restoring the redundancy and data reliability of the storage array.

[0101] According to an embodiment of the present invention, based on the number of disks, the geometric check rules of each check disk are reversely called to perform back-replacement processing on the check blocks stored in the check disk, including: if the number of disks r is 1, a target disk is selected from n check disks to perform a reverse XOR operation on the check blocks in the target disk according to the geometric check rules corresponding to the target disk to obtain the data of the faulty disk.

[0102] When only one data disk fails in a disk array (i.e., r = 1), a target disk is intelligently selected from the n parity disks for data recovery. This selection process is based on the data mapping relationship between the parity disk and the failed data disk, prioritizing the parity disk with the optimal data distribution and a direct mapping relationship with the failed data disk as the target disk. For example, if data disk D2 fails, and among the parity disks P, Q, and R, disk P has a direct mapping relationship with disk D2 and the optimal data distribution, disk P will be selected as the target disk.

[0103] After selecting the target disk, the reverse XOR operation is initiated based on the target disk's corresponding geometric verification rule (horizontal, diagonal, or reverse diagonal). For horizontal verification, data blocks from all other healthy data disks in the same horizontal verification group are collected and continuously XORed with the contents of these blocks to ultimately obtain the original data of the faulty data disk. For diagonal or reverse diagonal verification, data blocks participating in the operation are selected from the set of data blocks across the data stripe according to the preset index rules. The faulty data is also restored through reverse XOR.

[0104] During the entire back-substitution process, edge cases are automatically handled. For example, when encountering virtual zero-padding locations, they are treated as zero values ​​for the calculation, ensuring the accuracy of the reverse calculation. This method allows efficient recovery of data from failed data disks by simply invoking the geometric verification rules of a single parity disk, significantly reducing I / O overhead and computational complexity during data recovery, and improving system recovery efficiency and availability.

[0105] According to an embodiment of the present invention, based on the number of disks, the geometric check rules of each check disk are reversely called to perform back-substitution processing on the check blocks stored in the check disk, and the method also includes: if the number of disks r is greater than 1 and less than n, n-1 target disks are selected from the n check disks; a group of equations is constructed according to the geometric check rules corresponding to the n-1 target disks respectively; and the data of the faulty disk is derived from the group of equations through a reverse XOR operation.

[0106] When the number of failed disks r in a disk array is greater than 1 and less than n, n-1 target disks are quickly selected from the n parity disks. This selection process comprehensively considers the data relevance between each parity disk and the failed data disk, as well as the data distribution balance, to ensure that the selected target disks provide the most effective information for data recovery. For example, if there are three parity disks (P, Q, and R) in the disk array and two data disks fail, two target disks (P and Q) are selected.

[0107] After selecting n-1 target disks, a data recovery equation system is constructed based on the geometric verification rules (horizontal, diagonal, or reverse diagonal) corresponding to each target disk. For horizontal verification rules, the relationship between the data blocks of normal data disks and the parity blocks of the target disks within the same horizontal parity group is converted into an equation. For diagonal and reverse diagonal verification rules, the logical relationship between data blocks and parity blocks across data stripes is abstracted into equation expressions based on pre-set indexing rules. These equations are linked through the exclusive-OR operation between data blocks and parity blocks to form a complete system of equations.

[0108] When deriving the data on the failed disk, the reverse XOR operation is used to logically decompose the parity block values ​​in the equation system from the data block values ​​on the normal data disk. By using the reversibility of the XOR operation, redundant terms are gradually eliminated from the simultaneous equations, ultimately solving for the data block contents on the failed disk. During the calculation process, if boundary data is involved, a virtual zero-padding mechanism is automatically enabled to ensure the integrity of the equation system and the accuracy of the calculation. This allows for efficient and accurate recovery of the failed disk data, ensuring rapid restoration of data integrity and availability of the disk array.

[0109] According to an embodiment of the present invention, based on the number of disks, the geometric check rules of each check disk are reversely called to perform back-substitution processing on the check blocks stored in the check disk, and also includes: if the number of disks r is n, then according to the geometric check rules corresponding to the n target disks, a group of equations is constructed; through the reverse XOR operation, the data of the faulty disk is derived from the group of equations.

[0110] When all n disks in a disk array fail simultaneously (r=n), full recovery is initiated, using the geometric verification rules of all parity disks to construct a complete set of equations for data recovery. At this point, all n parity disks are considered target disks. Using the geometric verification rules (horizontal, oblique, or reverse oblique) specific to each parity disk, combined with the mapping between data blocks and parity blocks, a set of n independent equations is constructed.

[0111] Each equation corresponds to the verification rule for a parity disk. The left side of the equation contains the parity block value, and the right side contains the XOR expression for the data blocks participating in the verification rule. Based on pre-set indexing rules, data blocks across data stripes are correctly incorporated into the corresponding equation, and virtual zero padding is performed on data blocks at the boundary to ensure the integrity of the equation system. Leveraging the reversibility of the inverse XOR operation, the equation system is converted into a solvable mathematical model. Using algorithms such as matrix operations or Gaussian elimination, redundant information is gradually eliminated to determine the contents of the data blocks on all failed disks.

[0112] During the solution process, the mapping relationship between verification rules and the association between data blocks are automatically processed to ensure the independence and solvability of each equation. By utilizing the geometric verification rules of all verification disks, the system of equations has sufficient constraints to uniquely determine the data on all failed disks. Finally, the solution results are written to the reconstructed disk and integrity checked to ensure the accuracy of the recovered data. This approach fully utilizes the redundancy of geometric verification rules, effectively recovering data even in extreme situations, and ensuring the high availability and data security of the storage system.

[0113] Figure 5 A flowchart of three-disk failure recovery according to an embodiment of the present invention is shown.

[0114] like Figure 5 As shown, when three disks in the disk array fail, operations S501 to S512 are performed.

[0115] In operation S501, it is determined whether the failed disk is three data disks. If so, operation S512 is executed. If not, operation S502 is executed.

[0116] In operation S502, it is determined whether the failed disk is one of the three parity disks. If so, operation S507 is executed. If not, operation S503 is executed.

[0117] In operation S503, it is determined whether the failed disk is two data disks and one parity disk. If so, operation S506 is executed. If not, operation S504 is executed.

[0118] In operation S504, it is determined whether the faulty disk is a data disk and two parity disks. If so, operation S505 is performed. If not, the faulty disk repair is terminated.

[0119] In operation S505, it is determined whether the failed check disk is an oblique check disk or a reverse oblique check disk. If not, operation S508 is performed. If yes, operation S509 is performed.

[0120] In operation S506, it is determined whether the faulty disk is a horizontal parity disk. If not, operation S510 is performed. If not, operation S510 is performed. If yes, operation S511 is performed.

[0121] In operation S507 , verification is re-performed based on the data disk.

[0122] In operation S508, the failed data disk and the transverse parity disk are first recovered based on the normally operating parity disk, and the oblique parity disk (or the reverse oblique parity disk) is finally calculated.

[0123] In operation S509, the failed data disk is recovered based on the transverse parity disk, and then the oblique parity disk and the reverse oblique parity disk are recovered based on all the recovered data disks.

[0124] In operation S510, a failed data disk is recovered based on a non-failed transverse check disk and a slant check disk (or reverse slant check disk), and then the failed slant check disk (or reverse slant check disk) is recovered based on the data disk and the transverse check disk.

[0125] In operation S511, recovery of a failed data disk is performed based on a non-failed oblique parity disk and a reverse oblique parity disk, and then recovery of a failed transverse parity disk is performed based on the data disk.

[0126] In operation S512, a data disk failure is processed first, and data is recovered on the remaining two parity disks through recovery.

[0127] The above decoding algorithm targets the situation where three disks in a disk array fail. It determines the type of the failed disk through a series of judgments and then performs corresponding operations, such as using different check disks to restore the data disk and recheck, to achieve data repair on the failed disk.

[0128] According to an embodiment of the present invention, the data processing method further includes: if the number of disks r is n, calling a decoding algorithm corresponding to the recovery method, and performing the following operations: calculating the relative spacing parameters between the faulty disk and the remaining disks based on the distribution position of the faulty disk in the disk array to construct a composite equation group including horizontal verification rules, oblique verification rules, and reverse oblique verification rules; based on the calculated minimum iteration step parameter, iteratively performing multiple rounds of XOR simplification on the composite equation group to gradually eliminate intermediate variables; and performing back substitution calculation on the target equation group after eliminating the intermediate variables to determine the missing data blocks in the faulty disk.

[0129] In the extreme case where all disks (r=n) in a disk array fail, a complex set of equations is constructed based on the mathematical properties of geometric check rules for data recovery. First, based on the distribution of the failed disks in the array, the relative spacing parameters between them and the remaining disks are calculated. These parameters, including lateral spacing, oblique slope, and reverse oblique slope, are used to determine the geometric relationship between data blocks. Based on these parameters, the lateral, oblique, and reverse oblique check rules are integrated to construct a complex set of equations. Each equation corresponds to a check rule, and the relationship between data blocks and check blocks is expressed through an exclusive-OR operation.

[0130] To simplify this complex system of equations, a minimum iteration step parameter is calculated. This parameter is derived by analyzing the relative positions of the faulty disks and the periodic nature of the verification rules. Based on this minimum iteration step parameter, multiple rounds of XOR simplification are performed on the complex system of equations. Each round of simplification selects a specific combination of equations for XOR operations, gradually eliminating intermediate variables and reducing the complexity of the system. This iterative simplification process fully utilizes the reversibility and commutativity of the XOR operation, ensuring that the solution to the system remains unchanged at each simplification step.

[0131] After multiple rounds of simplification, the complex system of equations is transformed into a target system of equations containing only the data blocks on the failed disk. Back-substitution is then used to solve these unknown data blocks. Starting from known boundary conditions or virtual zero-padding positions, the simplified equations are used to gradually derive the data block values ​​on each failed disk. During the calculation process, modular operations and edge cases are automatically handled to ensure that all calculations are within the valid range. This approach enables the system to recover complete data even in the extreme case of complete disk failure, ensuring high availability and data security for the storage system. The detailed calculation process is shown below.

[0132] Assume that the indexes of the failed disks are x, y, z, and x<y<z. After x, y, and z fail, the data blocks of the remaining disks Known. x, y, z The data block is unknown and needs to be restored. The data block of the jth stripe of disk, the total number of disks , a total of indivual. Corresponding to the virtual data stripe, assume that all values ​​are 0, and let the horizontal P check disk of the virtual row .

[0133] Step 1: Calculation : .

[0134] Step 2: Set up the equations and calculate ,in, .

[0135] Sub-step 2.1: For row, use Calculation of horizontal P check for rows: .

[0136] Sub-step 2.2: Use Calculation of horizontal P check for rows: .

[0137] Sub-step 2.3: Use Backslash R checksum calculation: OK R checksum for backslashes: .use Verification calculation: .

[0138] Sub-step 2.4: Use Slash Q check calculation: OK Participated in the slash Q check: . Calculate using Q check: .

[0139] Sub-step 2.5: Sub-step 2.1 to 2.4 above calculated Add: .storage .

[0140] Repeat the above sub-steps 2.1 to 2.5 until all , .

[0141] It should be noted that the addition in step 2 is performed on the data block and the check block, so it is a modulo 2 addition. , which is essentially an XOR operation on data. All the , .

[0142] Step 3: Calculation . They satisfy: Pick .

[0143] like or , then swap the values ​​of g and h, that is: .

[0144] like or ,but ;otherwise ;

[0145] Step 4: Perform pairwise sum simplification for each row and calculate . Calculate and store .

[0146] Steps 3 and 4 are to ensure the consistency of subsequent processing. During the simulation, it was found that there would be special cases where m=m1=m4 or m=m2=m3. That is, among m1, m2, m3, and m4, two values ​​are equal and all equal to the minimum value m. This special case needs to be handled carefully in steps 3 and 4.

[0147] Step 5: Back-substitute the solution to recover the failed disk y and iterate the pairwise summation.

[0148] Input: Calculated array ,p,flag. Output: Data on failed disk y is restored.

[0149] Specifically, for j=0 to (p-1), perform the following operations: r=(p-1)+(j\times2h), k=r+2h, calculate A[y,k]=B[r+flag\time2h]-A[y,r], and the loop ends.

[0150] In the above algorithm: Because in modulo 2 arithmetic, addition and subtraction are equivalent: .

[0151] At this point, the data in the fault data disk All solved.

[0152] According to an embodiment of the present invention, based on the number of disks and the mapping relationship between data blocks and check blocks, a combination of forward and reverse calls are made to the geometric check rules to perform data recovery, including: recovering the data of the type of data disk in the faulty disk by reversely calling the corresponding geometric check rules based on the check blocks and the number of disks in the normally operating check disk; if it is determined that the recovery of the data of the type data disk in the faulty disk is completed, re-checking the data blocks stored in multiple data disks to recover the data of the type check disk in the faulty disk.

[0153] When both a data disk and a check disk fail in a disk array, a combined forward and reverse check rule invocation mechanism is initiated based on the mapping relationship between the number of disks and the data block-check block. First, the check block data is extracted from the normally operating check disk, and based on the current number of failed disks, the mathematical relationship between each check rule and the data block is reversely analyzed. For example, for a failed data disk, based on the reverse logic of geometric check rules such as oblique and reverse oblique directions, the check block is used as a known quantity, and a set of equations is constructed in combination with the data blocks of the remaining normal data disks. The original data of the failed data disk is deduced through the inverse process of the XOR operation. During this process, the relative spacing parameters are calculated based on the distribution position of the failed disk to ensure that the set of equations can accurately cover all the data blocks to be recovered, and virtual zero padding is used at the boundary positions to ensure calculation integrity.

[0154] After confirming that all the data in the faulty data disk has been recovered, the forward verification process is started to restore the faulty verification disk. At this point, the recovered data disk and other normal data disks constitute a complete data set, and all data blocks are re-performed with XOR operations according to the horizontal, oblique, and reverse oblique geometric verification rules. For example, for the horizontal verification rule, parallel XOR calculations are performed on the data blocks in the same data stripe to generate new horizontal verification values. For the oblique verification rule and the reverse oblique verification rule, data blocks are selected across stripes for verification operations based on the preset index rules. The newly generated verification value will be written to the corresponding position of the faulty verification disk according to the mapping relationship between the data block and the verification block to complete the recovery of the verification data. During the entire combined call process, the recovery efficiency is improved through multi-threaded parallel processing, and the verification value comparison is performed at each step to ensure the accuracy and completeness of data recovery.

[0155] According to an embodiment of the present invention, the data processing method also includes: creating a structure memory space for storing the topological relationship of the disk array and the status information of the faulty disk, and a full memory page for writing recovery data; if data recovery is completed, the recovery data in the full memory page is written to the faulty disk.

[0156] First, a dedicated structure is allocated in memory to store the disk array topology and the status of the failed disk. This structure contains the physical layout parameters of the disk array (such as the number of disks, stripe size, and parity type), the logical mapping between each disk (such as the storage location index of data blocks and parity blocks), and the real-time status of the failed disk (such as the failure type, failure time, and recovery progress). Simultaneously, a contiguous full memory page is requested, sized to match the total capacity of the failed disk, to temporarily store the complete data generated during the recovery process.

[0157] Once the data recovery process is initiated, the recovered data blocks are written to the corresponding locations on the full memory page according to the disk array's topology, based on the forward or reverse calculation results of the geometric validation rules. For example, recovered data blocks are accurately placed at the specified offset address on the memory page, following the striping strategy and the coordinate mapping relationship specified in the geometric validation rules, ensuring that the physical layout of the data is consistent with the original array. During the recovery process, the fault status information in the structure is updated in real time, recording the range of recovered data blocks and the remaining tasks to be processed.

[0158] After the data of all failed disks has been calculated using geometric verification rules and integrated into the memory pages, the data write operation is triggered. At this point, the storage controller will sequentially write the recovery data in the full memory page to the newly replaced failed disk or the reconstructed target disk based on the disk topology information recorded in the structure. During the write process, cache pre-reading and batch write strategies are adopted to improve data write efficiency. After the write is completed, the disk integrity is verified. By comparing the verification value of the memory page data with the actual data stored on the disk, it is ensured that the recovery data is accurately written to the failed disk, ultimately restoring the normal operation of the disk array.

[0159] Based on the above data processing method, the present invention also provides a data processing device. Figure 6 The device is described in detail.

[0160] Figure 6 A structural block diagram of a data processing device according to an embodiment of the present invention is shown.

[0161] like Figure 6 As shown, the data processing device 600 of this embodiment includes a data verification module 610 , a mode determination module 620 and a data recovery module 630 .

[0162] Data verification module 610 is configured to, in response to a data write operation, verify the written data according to the geometric verification rules corresponding to each verification disk, and store the verified verification blocks on the corresponding verification disk. The verification blocks are mapped to the written data stored on the multiple data disks. In one embodiment, data verification module 610 can be used to perform operation S210 described above, and will not be further described here.

[0163] The method determination module 620 is used to determine the recovery method according to the type of the failed disk and the number of failed disks if a disk failure is detected. In one embodiment, the method determination module 620 can be used to perform the operation S220 described above, which will not be repeated here.

[0164] The data recovery module 630 is used to forwardly and / or reversely call the geometric check rule based on the mapping relationship and the recovery method to recover the data of the failed disk. In one embodiment, the data recovery module 630 can be used to perform the operation S230 described above, which will not be repeated here.

[0165] According to an embodiment of the present invention, any multiple modules among the data verification module 610, the method determination module 620, and the data recovery module 630 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to an embodiment of the present invention, at least one of the data verification module 610, the method determination module 620, and the data recovery module 630 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable method of circuit integration or packaging, or may be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of any of these. Alternatively, at least one of the data verification module 610, the method determination module 620, and the data recovery module 630 may be at least partially implemented as a computer program module that, when executed, performs the corresponding functionality.

[0166] It should be noted that the data processing device part in the embodiment of the present invention corresponds to the data processing method part in the embodiment of the present invention. The description of the data processing device part specifically refers to the data processing method part and will not be repeated here.

[0167] Figure 7 A block diagram of an electronic device suitable for implementing a data processing method according to an embodiment of the present invention is shown.

[0168] like Figure 7 As shown, an electronic device 700 according to an embodiment of the present invention includes a processor 701, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 702 or programs loaded from a storage unit 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or related chipsets and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include onboard memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0169] The RAM 703 stores various programs and data required for the operation of the electronic device 700. The processor 701, ROM 702, and RAM 703 are connected to each other via a bus 704. The processor 701 executes the programs in the ROM 702 and / or RAM 703 to perform the various operations of the method flow according to the embodiment of the present invention. It should be noted that the programs may also be stored in one or more memories other than the ROM 702 and RAM 703. The processor 701 may also execute the programs stored in one or more memories to perform the various operations of the method flow according to the embodiment of the present invention.

[0170] According to an embodiment of the present invention, electronic device 700 may further include an input / output (I / O) interface 705, which is also connected to bus 704. Electronic device 700 may also include one or more of the following components connected to I / O interface 705: an input section 706 including a keyboard, mouse, etc.; an output section 707 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 708 including a hard disk; and a communication section 709 including a network interface card such as a LAN card or modem. Communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to I / O interface 705 as needed. Removable media 711, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 710 as needed, so that computer programs read from the removable media can be installed into storage section 708 as needed.

[0171] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the present invention.

[0172] According to an embodiment of the present invention, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, a computer-readable storage medium may include the ROM 702 and / or RAM 703 described above, and / or one or more memories other than ROM 702 and RAM 703.

[0173] The embodiments of the present invention further include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is executed in a computer system, the program code is used to cause the computer system to implement the data processing method provided by the embodiments of the present invention.

[0174] The computer program executes the above functions defined in the system / device of the embodiment of the present invention when the computer program is executed by the processor 701. According to the embodiment of the present invention, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0175] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 709, and / or installed from a removable medium 711. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0176] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709 and / or installed from the removable medium 711. When the computer program is executed by the processor 701, the above-described functions defined in the system of the embodiment of the present invention are performed. According to the embodiment of the present invention, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.

[0177] According to an embodiment of the present invention, the program code for executing the computer program provided by the embodiment of the present invention can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).

[0178] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0179] It will be understood by those skilled in the art that the features described in the various embodiments of the present invention may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention may be combined and / or coupled in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or couplings fall within the scope of the present invention.

[0180] The above describes embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.

Claims

1. A data processing method, characterized in that: Applied to a disk array, the disk array being composed of a plurality of disks, the plurality of disks including a plurality of data disks and n parity disks, where n is an integer greater than 2, the method comprising: In response to a data write operation, the written data is verified according to a geometric verification rule corresponding to each of the verification disks, so that a verification block obtained by verification is stored in the corresponding verification disk. A mapping relationship exists between the verification block and the written data stored in the multiple data disks. The geometric verification rule includes a horizontal verification rule, an oblique verification rule, and a reverse oblique verification rule. The horizontal verification rule performs verification calculations on data blocks in the same row or the same horizontal dimension. The oblique verification rule performs combined verification on the data blocks in an anti-diagonal direction with a specific slope according to a path where the data blocks are obliquely distributed in the disk array. The reverse oblique verification rule performs combined verification on the data blocks in a diagonal direction with a slope opposite to the specific slope according to a path where the data blocks are reversely distributed in the disk array. If a failure of the disk is detected, a recovery method is determined based on the type of the failed disk and the number of failed disks; Based on the mapping relationship, the geometric check rule is called forward and / or reversely according to the recovery method to recover the data of the failed disk.

2. The method according to claim 1, characterized in that Verifying the written data according to the geometric verification rules corresponding to each of the verification disks includes: Determining that each data stripe in the written data is stored in a data block in the plurality of data disks, wherein the data stripes are obtained by dividing the written data according to a preset stripe size; Based on the distribution position of each data block in the data stripe in the multiple data disks, the data stripes are verified in parallel according to the geometric verification rules corresponding to each of the verification disks.

3. The method according to claim 2, characterized in that The verifying the data stripes in parallel based on the distribution position of each data block in the data stripe in the plurality of data disks and according to the geometric verification rules corresponding to each of the verification disks includes: According to the horizontal verification rule, performing an XOR operation on a plurality of data blocks distributed in the plurality of data disks in the same data stripe; With respect to the oblique check rule, based on the distribution position and according to the preset oblique index rule, selecting data blocks across data stripes to perform an exclusive OR operation; With respect to the reverse oblique check rule, based on the distribution position and according to the preset reverse oblique index rule, data blocks across data stripes are selected to perform an exclusive OR operation; Wherein, when selecting a data block, the preset oblique indexing rule and the preset reverse oblique indexing rule adopt virtual zero padding processing for the boundary position.

4. The method according to claim 3, characterized in that When selecting a data block, the preset oblique indexing rule and the preset reverse oblique indexing rule adopt virtual zero padding for boundary positions, including: If the data block position calculated according to the preset oblique indexing rule or the preset reverse oblique indexing rule exceeds the range of the data stripes in the multiple data disks, the coordinate value of the data block position is mapped into the range by a modulo operation to obtain a mapped position; The data block corresponding to the mapping position is determined to be zero value.

5. The method according to claim 4, characterized in that The horizontal coordinate of the coordinate value of the data block position represents the index of the disk; With respect to the preset oblique indexing rule, the vertical coordinate in the coordinate value is obtained by subtracting the index of the data stripe from the index of the disk; With respect to the preset reverse oblique indexing rule, the vertical coordinate in the coordinate value is obtained by summing the index of the data stripe and the index of the disk.

6. The method according to claim 5, characterized in that Mapping the coordinate value of the data block position to the range by a modulo operation to obtain a mapping position includes: Comparing the vertical coordinate in the coordinate value with the range of the data strip to obtain a comparison result; If the comparison result indicates that the ordinate exceeds the range, a modulo operation is performed on the ordinate to determine the mapping position according to the calculated ordinate.

7. The method according to claim 2, characterized in that If the disk failure is detected, a recovery method is determined based on the type of the failed disk and the number of failed disks, including: If the type is a check disk, the recovery method is determined to be, based on the mapping relationship between the data blocks and the check blocks, forward calling the corresponding geometric check rules to recheck the data blocks stored in the multiple data disks.

8. The method according to claim 2, characterized in that The disk array allows a number of disks r to fail simultaneously, where r is an integer greater than or equal to 1 and less than or equal to n. If a disk failure is detected, a recovery method is determined based on the type of the failed disk and the number of failed disks, including: If the type is a data disk, the recovery method is determined to be, based on the number of disks, reversely calling the geometric verification rules of each of the verification disks to perform back-replacement processing on the verification blocks stored in the verification disks.

9. The method according to claim 2, characterized in that If the disk failure is detected, a recovery method is determined based on the type of the failed disk and the number of failed disks, including: If the type includes the data disk and the check disk, based on the number of disks and the mapping relationship between the data blocks and the check blocks, the geometric check rule is called in a forward and reverse combination to perform data recovery.

10. The method according to claim 1, characterized in that The method further comprises: If it is detected that the failed disk needs to be reconstructed, the recovered data is written to the reconstructed disk.

11. The method according to claim 8, characterized in that The reverse calling of the geometric check rules of each of the check disks based on the number of disks to perform back-replacement processing on the check blocks stored in the check disks includes: If the number r of disks is 1, a target disk is selected from the n check disks, and a reverse XOR operation is performed on the check blocks in the target disk according to the geometric check rule corresponding to the target disk to obtain the data of the failed disk.

12. The method according to claim 8, characterized in that The reverse calling of the geometric check rules of each of the check disks based on the number of disks to perform back-replacement processing on the check blocks stored in the check disks includes: If the number of disks r is greater than 1 and less than n, n-1 target disks are selected from the n check disks; Constructing a set of equations according to the geometric verification rules corresponding to the n-1 target disks respectively; The data of the failed disk is derived from the equation group through an inverse XOR operation.

13. The method according to claim 8, characterized in that The reverse calling of the geometric check rules of each of the check disks based on the number of disks to perform back-replacement processing on the check blocks stored in the check disks includes: If the number r of the disks is n, then construct a set of equations according to the geometric verification rules corresponding to the n verification disks respectively; The data of the failed disk is derived from the equation group through an inverse XOR operation.

14. The method according to claim 13, characterized in that The method further comprises: If the number of disks r is n, the decoding algorithm corresponding to the recovery method is called to perform the following operations: Calculating a relative spacing parameter between the faulty disk and the remaining disks according to the distribution position of the faulty disk in the disk array to construct a composite equation group including the transverse verification rule, the oblique verification rule, and the reverse oblique verification rule; Iteratively performing multiple rounds of XOR simplification on the composite equations based on the calculated minimum iteration step parameter to gradually eliminate intermediate variables; The data blocks missing from the faulty disk are determined by performing back substitution calculation on the target equation group after eliminating the intermediate variables.

15. The method according to claim 9, characterized in that The method of performing a forward and reverse combined call on the geometric verification rule based on the number of disks and the mapping relationship between the data block and the verification block to perform data recovery includes: Recovering the data of the data disk in the faulty disk by reversely invoking the corresponding geometric check rule based on the check blocks in the normally operating check disk and the number of disks; If it is determined that the data recovery of the data disk type in the failed disk is completed, the data blocks stored in the multiple data disks are rechecked to recover the data of the check disk type in the failed disk.

16. The method according to claim 1, characterized in that The method further comprises: Creating a structure memory space for storing the topological relationship of the disk array and status information of the failed disk, and a full memory page for writing recovery data; If data recovery is complete, the recovered data in the full memory page is written to the failed disk.

17. A data processing device, characterized in that: Applied to a disk array, the disk array is composed of multiple disks, the multiple disks include multiple data disks and n check disks, where n is an integer greater than 2, the device comprises: a data verification module for, in response to a data write operation, verifying the written data according to geometric verification rules corresponding to each of the verification disks, so as to store a verification block obtained by verification in the corresponding verification disk, wherein a mapping relationship exists between the verification block and the written data stored in the multiple data disks, the geometric verification rules including a horizontal verification rule, an oblique verification rule, and a reverse oblique verification rule. The horizontal verification rule performs verification calculations on data blocks in the same row or the same horizontal dimension. The oblique verification rule performs a combined verification on the data blocks in an anti-diagonal direction having a specific slope according to a path along which the data blocks are obliquely distributed in the disk array. The reverse oblique verification rule performs a combined verification on the data blocks in a diagonal direction having a slope opposite to the specific slope according to a path along which the data blocks are reversely distributed in the disk array. a mode determination module, configured to determine a recovery mode based on the type of the failed disk and the number of failed disks if a failure of the disk is detected; A data recovery module is used to call the geometric verification rule forwardly and / or reversely based on the mapping relationship and the recovery method to recover the data of the failed disk.

18. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 16.

19. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instructions are executed by a processor, the steps of the method according to any one of claims 1 to 16 are implemented.

20. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Method and device for data backup and recovery in redundant array of inexpensive disks

    CN101770409A

  • Method for constructing disk array by horizontal grouping parallel concentrated verification

    CN101976175A