Abnormal data positioning method and device, equipment, program product and storage medium
By determining the verification chain calculation results and hierarchical traversal strategies of the target strip in the RAID system, accurately locate the disk where the abnormal data is located, and recover the abnormal data through the verification chain combination and real-time data reconstruction process, the problem of the inability to locate abnormal data in the existing technology is solved, and the user experience and system stability are improved.
Patent Information
- Application Number
- CN202510969579.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-07-15
AI Technical Summary
The disk where abnormal data is located cannot be accurately located in the existing RAID system, making it difficult to effectively recover and reduce the user experience.
By determining the verification chain calculation results of each verification chain in the target strip, if the disk corresponding to the verification chain with some or all zeros is the disk where the abnormal data is located, combined with the hierarchical traversal and verification chain combination strategy, the disk where the abnormal data is located is accurately located and recovered through the real-time data reconstruction process.
Efficiently and accurately locate and restore abnormal data, improve user experience, and ensure data security and system stability.
Smart Images

Figure CN120469849A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of independent redundant arrays of disks, and in particular to an abnormal data locating method, device, equipment, program product and storage medium. Background Art
[0002] In current RAID (Redundant Array of Independent Disks) systems, when abnormal data exists on a disk, stripe consistency verification can be used to determine the stripe to which the abnormal data belongs. However, a single stripe involves multiple disks, and related technologies cannot further determine which disk in the stripe the abnormal data is located on. This makes it difficult to recover the abnormal data in the stripe, reducing the user experience.
[0003] Therefore, how to provide a solution to the above technical problems is a problem that those skilled in the art need to solve at present. Summary of the Invention
[0004] The purpose of the present invention is to provide a method, device, equipment, program product and storage medium for locating abnormal data. In the present invention, for a target stripe with abnormal data, the check chain calculation results of each check chain can be determined. If the calculation results of each check chain are partially zero, the disk corresponding to the check chain with a non-zero calculation result can be used as the disk where the abnormal data is located, thereby efficiently and accurately locating the disk where the abnormal data is located, facilitating the recovery of the abnormal data and improving the user experience.
[0005] To solve the above technical problems, the present invention provides a method for locating abnormal data, comprising: For a target stripe with abnormal data in the storage system, the parity chain calculation results of each parity chain of the target stripe are determined based on the current data of each data block in the target stripe. If the calculation results of each check chain are partially zero, the disk corresponding to the check chain with a non-zero calculation result is regarded as the disk where the abnormal data is located; The data blocks include business data blocks and check data blocks. The disk corresponding to the check chain is the disk to which the check data in the check chain belongs in the target stripe. The storage system is a redundant array of independent disks system.
[0006] On the other hand, for a target stripe in the storage system having abnormal data, after determining the parity chain calculation results of each parity chain of the target stripe based on the current data of each data block in the target stripe, the abnormal data locating method further includes: If the calculation results of all check chains are non-zero, the check chains are divided into two check chain combinations, where the number of check chains in each check chain combination is a preset number, which is the total number of check chains minus one; For any business data block combination in the target stripe, business data recovery is performed on the business data block combination through each check chain combination to obtain business data recovery results of each check chain combination for the business data block combination; Determining whether the business data recovery results of the business data block combination are consistent; If they are consistent, the disk to which each business data block in the business data block combination belongs is used as the disk where the abnormal data is located.
[0007] On the other hand, the disk to which each business data block in the business data block combination belongs is used as the disk where the abnormal data is located, including: Determining whether there is a business data block in the business data block combination that meets a preset elimination condition, wherein the preset elimination condition includes that the restored data is consistent with the existing data; If so, the disk to which the business data blocks in the business data block combination, excluding the business data blocks that meet the preset elimination conditions, belong is used as the disk where the abnormal data is located.
[0008] In another aspect, the target stripe has at least three parity chains; If the calculation results of each check chain are all non-zero, the check chain is divided into two check chain combinations including: If the calculation results of each check chain are all non-zero, a check chain of the target stripe is selected as the initial check chain; For any service data block in the target stripe, recover the service data block using the initial check chain to obtain recovered data of the service data block; If there is a business data block in the target stripe where the restored data is consistent with the existing data, the disk where the business data block in the target stripe where the restored data is consistent with the existing data belongs is used as the disk where the abnormal data is located; If there is no business data block in the target stripe whose restored data is consistent with the existing data, the check chain is divided into two check chain combinations.
[0009] On the other hand, dividing the check chain into two check chain combinations includes: Randomly select a preset number of check chains from all check chains of the target stripe to form the first check chain combination; Randomly select a preset number of check chains from all check chains of the target stripe to form a second check chain combination; The first check chain combination is different from the second check chain combination.
[0010] On the other hand, the generation of the business data block combination adopts a hierarchical traversal strategy: Prioritize traversing the data block combinations on the disks that have recently experienced input and output errors; Secondly, it traverses the disk data block combinations marked as high-risk aging levels by the storage system; Finally, traverse the data block combinations of the remaining disks.
[0011] On the other hand, for any business data block combination in the target stripe, business data recovery is performed on the business data block combination through each check chain combination, and the business data recovery results of each check chain combination for the business data block combination include: For any combination of business data blocks in the target stripe, a first set of data recovery equations is constructed for the business data block combination by using a first check chain combination, and a second set of data recovery equations is constructed for the business data block combination by using a second check chain combination; Using the solution of the first set of data recovery equations as the business data recovery result of the first check chain combination for the business data block combination; The solution result of the second set of data recovery equations is used as the business data recovery result of the second check chain combination for the business data block combination.
[0012] On the other hand, the abnormal data locating method further includes: After the disk where the abnormal data is located is identified, the real-time data reconstruction process is triggered, including: Read data from normal business data blocks and check data blocks in the target stripe, and reversely calculate the target value of the abnormal data block based on the check chain equation; Overwrite the target value to the corresponding exception data block.
[0013] On the other hand, after overwriting the target value to the corresponding abnormal data block, the abnormal data locating method further includes: Verify whether the abnormal data blocks in the target stripe can be successfully recovered; If recovery is unsuccessful, the disk containing the abnormal data will be marked as untrustworthy and isolated.
[0014] On the other hand, verifying whether the abnormal data blocks in the target stripe can be successfully recovered includes: Calculate the check chain calculation result of any check chain of the target stripe; If the check chain calculation result is non-zero, the real-time data reconstruction process is triggered for the second time; Calculate the parity chain calculation result of any parity chain of the target stripe based on the abnormal data block recovered for the second time; If the checksum calculation result is non-zero, the physical disk is marked as untrusted and quarantined.
[0015] On the other hand, the abnormal data locating method further includes: After determining the disk where the abnormal data is located, before recovering the abnormal data, mark the disk where the abnormal data is located as abnormal, and adjust the read and write policy of the storage system to avoid further access to the abnormal disk.
[0016] On the other hand, the redundant array of independent disks system applies a double check technology or a triple check technology; The target stripe includes at least two check data blocks, and each check data block corresponds to a different check chain.
[0017] On the other hand, for a target stripe in the storage system having abnormal data, after determining the parity chain calculation results of each parity chain of the target stripe based on the current data of each data block in the target stripe, the abnormal data locating method further includes: If the calculation results of each check chain are all zero, it is determined that the data blocks in the target stripe are normal and there is no need to perform abnormal data location operations.
[0018] On the other hand, the abnormal data locating method further includes: Perform periodic inspections on multiple stripes in the storage system and identify target stripes with abnormal data by comparing the parity chain calculation results of each stripe with the preset benchmark value.
[0019] On the other hand, the abnormal data locating method further includes: When determining the parity chain calculation result, if non-zero results are detected for N consecutive stripes in the same parity chain, it is determined that the physical disk corresponding to the parity chain has a systematic fault, so as to trigger a preventive disk replacement process.
[0020] To solve the above technical problems, the present invention further provides an abnormal data locating device, comprising: A first determining module is configured to determine, for a target stripe having abnormal data in the storage system, a check chain calculation result of each check chain of the target stripe based on current data of each data block in the target stripe; The second determining module is configured to, if the calculation results of each check chain are partially zero, determine the disk corresponding to the check chain with a non-zero calculation result as the disk where the abnormal data is located; The data blocks include business data blocks and check data blocks. The disk corresponding to the check chain is the disk to which the check data in the check chain belongs in the target stripe. The storage system is a redundant array of independent disks system.
[0021] To solve the above technical problems, the present invention further provides an abnormal data locating device, comprising: memory for storing computer programs; A processor is configured to implement the steps of the abnormal data locating method described above when executing the computer program.
[0022] On the other hand, the processor is integrated into the field programmable gate array of the redundant array of independent disks controller card, and the check chain calculation is performed by the following hardware acceleration unit: Galois field multiplier array for parallel calculation of the coefficients of the check equation; XOR operation tree module, used for XOR chain calculation.
[0023] To solve the above technical problem, the present invention further provides a computer program product, including a computer program / instruction, which implements the steps of the above abnormal data locating method when executed by a processor.
[0024] To solve the above technical problems, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above abnormal data locating method are implemented.
[0025] Beneficial effect: The present invention provides a method for locating abnormal data. Considering that in a target stripe with abnormal data, if there is a check chain with a calculation result of zero, it indicates that all business data blocks are normal. At this time, the disk corresponding to the check chain with a non-zero calculation result is the disk where the abnormal data is located. Therefore, the present invention can determine the check chain calculation results of each check chain for the target stripe with abnormal data. If the calculation results of each check chain are partially zero, the disk corresponding to the check chain with a non-zero calculation result can be used as the disk where the abnormal data is located, thereby efficiently and accurately locating the disk where the abnormal data is located, facilitating the recovery of the abnormal data, and improving the user experience.
[0026] The present invention also provides an abnormal data locating device, equipment, program product and storage medium, which have the same beneficial effects as the above abnormal data locating method. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the relevant technologies and the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0028] Figure 1 A schematic diagram of a flow chart of an abnormal data locating method provided by the present invention; Figure 2A schematic diagram of the distribution of abnormal data in the first RAID6 stripe provided by the present invention; Figure 3 A schematic diagram of the distribution of abnormal data in the first RAID TP stripe provided by the present invention; Figure 4 A schematic diagram of the distribution of abnormal data in the second RAID6 stripe provided by the present invention; Figure 5 A schematic diagram showing the distribution of abnormal data in the second RAID TP stripe provided by the present invention; Figure 6 A schematic structural diagram of an abnormal data locating device provided by the present invention; Figure 7 A schematic structural diagram of an abnormal data locating device provided by the present invention; Figure 8 A schematic structural diagram of a computer-readable storage medium provided by the present invention. DETAILED DESCRIPTION
[0029] The core of the present invention is to provide a method, device, equipment, program product and storage medium for locating abnormal data. In the present invention, for the target stripe with abnormal data, the check chain calculation results of each check chain can be determined. If the calculation results of each check chain are partially zero, the disk corresponding to the check chain with a non-zero calculation result can be used as the disk where the abnormal data is located, thereby efficiently and accurately locating the disk where the abnormal data is located, facilitating the recovery of the abnormal data and improving the user experience.
[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0031] Please refer to Figure 1 , Figure 1 This is a flow chart of a method for locating abnormal data provided by the present invention, which includes: S101: For a target stripe having abnormal data in the storage system, determine the parity chain calculation results of each parity chain of the target stripe based on the current data of each data block in the target stripe; Specifically, taking into account the technical problems in the above background technology, and considering that in the target stripe where abnormal data exists, if there is a check chain with a calculation result of zero, it indicates that all business data blocks are normal. At this time, the disk corresponding to the check chain with a non-zero calculation result is the disk where the abnormal data is located. Therefore, in order to achieve accurate positioning of the disk where the abnormal data is located, the embodiment of the present invention is based on the correlation between the check chain calculation result and the disk abnormality, and sets up this scheme to screen abnormal disks through the check chain calculation result, thereby providing an accurate target for data repair.
[0032] Specifically, in this step, for the target stripe with abnormal data in the storage system, the check chain calculation results of each check chain of the target stripe can be determined through the current data of each data block in the target stripe, so as to serve as the data basis for subsequent steps.
[0033] S102: If the calculation results of each check chain are partially zero, the disk corresponding to the check chain with a non-zero calculation result is regarded as the disk where the abnormal data is located; The data blocks include business data blocks and check data blocks. The disk corresponding to the check chain is the disk to which the check data in the check chain belongs in the target stripe. The storage system is a redundant array of independent disks system.
[0034] Specifically, after determining the check chain calculation results of each check chain of the target stripe, based on the above concept, when the calculation results of some check chains are zero, the disk corresponding to the check chain with a non-zero check chain calculation result will be used as the disk where the abnormal data is located, thereby achieving accurate positioning of the abnormal data in the target stripe.
[0035] For example, in a RAID TP (Redundant Array of Independent Disks with Triple Parity) system consisting of seven disks, the target stripe contains business data blocks (such as d0-dn, where 0-n are the sequence numbers of the business data blocks and n is the total number of business data blocks, in 4KB units) and parity data blocks (P, Q, and R, corresponding to different parity chains). For the current data of each data block in the target stripe, calculate the parity chain calculation results for each parity chain: the P parity chain result is d0+d1+d2+d3+d4+d5+P, the Q parity chain result is a0×d0+a1×d1+a2×d2+a3×d3+a4×d4+a5×d5+Q, and the R parity chain result is (a0)²×d0+(a1)²×d1+(a2)²×d2+(a3)²×d3+(a4)²×d4+(a5)²×d5+R (where ai is a constant, i∈0-n). If the P parity chain calculation result is nonzero and the Q and R parity chain calculation results are zero, the disk corresponding to the P parity chain is determined to be the disk where the abnormal data is located. A stripe is a collection of data blocks with related locations on different partitions of the array. A data block is 4KB of data. The storage system is a redundant array of independent disks.
[0036] Specifically, to better illustrate the embodiments of the present invention, please refer to Figures 2 to 5 , Figure 2 This is a schematic diagram of the distribution of abnormal data in the first RAID6 stripe provided by the present invention. Figure 3 This is a schematic diagram of the distribution of abnormal data in the first RAID TP stripe provided by the present invention. Figure 4 This is a schematic diagram of the distribution of abnormal data in the second RAID6 stripe provided by the present invention. Figure 5 This is a schematic diagram of the distribution of abnormal data in the second RAID TP stripe provided by the present invention. Figures 2 to 5 The data blocks containing abnormal data are all represented by bold borders. Figure 2 and Figure 3 The abnormal data in are all distributed in the check data block. Figure 4 and Figure 5 The abnormal data in the distributed business data blocks.
[0037] The present invention provides a method for locating abnormal data. Considering that in a target stripe with abnormal data, if there is a check chain with a calculation result of zero, it indicates that each business data block is normal. At this time, the disk corresponding to the check chain with a non-zero calculation result is the disk where the abnormal data is located. Therefore, the present invention can determine the check chain calculation results of each check chain for the target stripe with abnormal data. If the calculation results of each check chain are partially zero, the disk corresponding to the check chain with a non-zero calculation result can be used as the disk where the abnormal data is located, thereby efficiently and accurately locating the disk where the abnormal data is located, facilitating the recovery of the abnormal data, and improving the user experience.
[0038] Based on the above embodiment: As an optional embodiment, for a target stripe in the storage system having abnormal data, after determining the parity chain calculation results of each parity chain of the target stripe based on the current data of each data block in the target stripe, the abnormal data locating method further includes: If the calculation results of all check chains are non-zero, the check chains are divided into two check chain combinations, where the number of check chains in each check chain combination is a preset number, which is the total number of check chains minus one; For any business data block combination in the target stripe, business data recovery is performed on the business data block combination through each check chain combination to obtain the business data recovery result of each check chain combination for the business data block combination; Determine whether the recovery results of each business data of the business data block combination are consistent; If they are consistent, the disk to which each business data block in the business data block combination belongs is used as the disk where the abnormal data is located.
[0039] Specifically, when all parity chain calculation results are non-zero, there may be multiple abnormal disks or complex abnormal situations, which cannot be located using a single parity chain alone. To address this problem, the present invention divides the parity chain into parity chain combinations. Business data block combinations are restored using different parity chain combinations. The consistency of the restored results is used to determine the abnormal disk, improving the location capability in complex scenarios.
[0040] For example, in a RAID 6 system (n data disks, 2 parity disks), if the P parity chain result X = d0 + d1 + … + dn + P ≠ 0 and the Q parity chain result Y = a0 × d0 + a1 × d1 + … + an × dn + Q ≠ 0, the parity chains are divided into two combinations: {P} and {Q} (each combination contains 1 parity chain, and the total number of parity chains is 2, with the preset number being the total number of parity chains minus one). For the business data block combination {d0}, d0 is recovered using the {P} combination: d0' = -d1 - d2 - … - dn - P; d0 is recovered using the {Q} combination: d0'' = (-a1 × d1 - … - an × dn - Q) / a0. Determine whether d0' and d0'' are consistent. If they are, the disk to which business data block d0 in {d0} belongs is identified as the disk where the abnormal data resides. Among them, RAID6 is a RAID array that includes two check disks and has error correction capabilities. When two or fewer disks fail, the data on the failed disk can be restored through other disks. A stripe is a collection of location-related data blocks on different partitions of the array.
[0041] As an optional embodiment, the disk to which each business data block in the business data block combination belongs is used as the disk where the abnormal data is located, including: Determine whether there is a business data block in the business data block combination that meets a preset elimination condition, wherein the preset elimination condition includes that the restored data is consistent with the existing data; If so, the disk to which the business data blocks in the business data block combination, except for the business data blocks that meet the preset elimination conditions, belong is regarded as the disk where the abnormal data is located.
[0042] Specifically, given that some business data blocks may be normal while others may be abnormal, direct location may lead to misjudgment. To accurately screen abnormal disks, preset exclusion conditions are set to eliminate interference from normal data blocks, ensuring that the location results only include truly abnormal disks.
[0043] Specifically, based on the previous embodiment, if the recovery result d0'=d0'' for the business data block combination {d0, d1} is further determined to determine whether the recovered data d0' of d0 is consistent with the existing data d0. The preset exclusion condition is that the recovered data is consistent with the existing data. If d0'=d0, then a business data block d0 that meets the preset exclusion condition exists. The disk containing the business data block d1 other than d0 in the business data block combination is then identified as the disk where the abnormal data resides.
[0044] Of course, in addition to this specific form, "using the disk to which each business data block in the business data block combination belongs as the disk where the abnormal data is located" can also be in other forms, which are not limited in this embodiment of the present invention.
[0045] As an optional embodiment, the target stripe has at least three check chains; If the calculation results of each check chain are all non-zero, the check chain is divided into two check chain combinations including: If the calculation results of each check chain are all non-zero, a check chain of the target stripe is selected as the initial check chain; For any business data block in the target stripe, data recovery is performed on the business data block through the initial check chain to obtain the recovered data of the business data block; If there is a business data block in the target stripe where the restored data is consistent with the existing data, the disk where the business data block in the target stripe where the restored data is consistent with the existing data belongs is used as the disk where the abnormal data is located; If there is no business data block in the target stripe whose restored data is consistent with the existing data, the check chain is divided into two check chain combinations.
[0046] Specifically, considering that the target stripe has at least three parity chains, single parity chain recovery may have limitations. Therefore, in this embodiment of the present invention, by first recovering data using the initial parity chain and determining whether consistent data blocks exist, the abnormal disk can be quickly located. If not, the parity chains are then divided into groups to improve location efficiency and accuracy.
[0047] Specifically, for example, in a RAID TP system (n data disks, 3 parity disks), the target stripe has three parity chains: P, Q, and R. If the calculated results of the P, Q, and R parity chains are all non-zero, the P parity chain is selected as the initial parity chain, and business data block d0 is recovered: d0' = -d1-d2-…-dn-P. If d0' in the target stripe is consistent with the existing data d0, the disk to which d0 belongs is considered the disk with the abnormal data. If there is no business data block in the target stripe where the recovered data is consistent with the existing data, the parity chain is divided into two combinations: {P,Q} and {P,R}. RAID TP is a RAID array with three parity disks and has error correction capabilities. If three or fewer disks fail, the data on the failed disk can be recovered using the remaining disks.
[0048] Of course, in addition to this specific form, “if all calculation results of each check chain are non-zero, then dividing the check chain into two check chain combinations” may also be other forms, which are not limited in this embodiment of the present invention.
[0049] As an optional embodiment, dividing the check chain into two check chain combinations includes: Randomly select a preset number of check chains from all check chains of the target stripe to form the first check chain combination; Randomly select a preset number of check chains from all check chains of the target stripe to form a second check chain combination; The first checksum chain combination is different from the second checksum chain combination.
[0050] Specifically, in order to ensure the randomness and comprehensiveness of the verification chain combination and avoid positioning deviation caused by fixed combinations, by randomly selecting verification chain combinations, more possible abnormal situations can be covered and the reliability of positioning can be improved.
[0051] For example, in a RAID TP system, two parity chains (the preset number is the total number of parity chains 3 minus one) are randomly selected from all parity chains P, Q, and R of the target stripe to form a first parity chain combination, such as {P, Q}; and two parity chains are randomly selected from all parity chains of the target stripe to form a second parity chain combination, such as {P, R}, and the first parity chain combination is different from the second parity chain combination.
[0052] Of course, in addition to this specific form, “dividing the check chain into two check chain combinations” may also be in other forms, which are not limited in this embodiment of the present invention.
[0053] As an optional embodiment, the generation of the business data block combination adopts a hierarchical traversal strategy: Prioritize traversing the data block combinations on the disks that have recently experienced input and output errors; Secondly, it traverses the disk data block combinations marked as high-risk aging levels by the storage system; Finally, traverse the data block combinations of the remaining disks.
[0054] Specifically, given the varying probability of abnormalities for different disks, disks that have recently experienced input / output errors or are marked as high-risk are more likely to have abnormalities. Therefore, the present invention employs a hierarchical traversal strategy, prioritizing high-risk disks. This improves the efficiency of abnormality location and allows for rapid identification of the disk containing abnormal data.
[0055] Specifically, for example, a hierarchical traversal strategy can be used to generate business data block combinations. In a RAID system, data block combinations on disks that recently experienced input / output (I / O) errors, such as disk A, are prioritized. Next, data block combinations on disks marked as high-risk by the storage system, such as disk B, are traversed. Finally, data block combinations on the remaining disks, such as disks C and D, are traversed. I / O, short for input and output, describes the transfer of data between storage devices and processing units.
[0056] Of course, in addition to this specific form, the hierarchical traversal strategy may also be in other forms, which are not limited in the embodiment of the present invention.
[0057] As an optional embodiment, for any business data block combination in the target stripe, business data recovery is performed on the business data block combination using each check chain combination, and the business data recovery results obtained by each check chain combination for the business data block combination include: For any combination of business data blocks in the target stripe, a first set of data recovery equations is constructed for the business data block combination using the first check chain combination, and a second set of data recovery equations is constructed for the business data block combination using the second check chain combination; The solution of the first set of data recovery equations is used as the business data recovery result of the first check chain combination for the business data block combination; The solution result of the second set of data recovery equations is used as the business data recovery result of the second check chain combination for the business data block combination.
[0058] Specifically, considering that by constructing a data recovery equation group and using the mathematical relationship of the check chain to solve the business data recovery result, the accuracy and logic of the recovery result can be ensured, and a reliable basis for judging abnormal disks can be provided, therefore, in an embodiment of the present invention, for any business data block combination in the target stripe, a first set of data recovery equation groups can be constructed for the business data block combination through the first check chain combination, and a second set of data recovery equation groups can be constructed for the business data block combination through the second check chain combination; the solution result of the first set of data recovery equation groups is used as the business data recovery result of the first check chain combination for the business data block combination; and the solution result of the second set of data recovery equation groups is used as the business data recovery result of the second check chain combination for the business data block combination.
[0059] Specifically, for the service data block combination {d0, d1} in the target stripe of the RAID6 system, a first set of data recovery equations is constructed for this service data block combination using the first parity chain combination {P}: d0+d1+d2+…+dn+P=0; a second set of data recovery equations is constructed for this service data block combination using the second parity chain combination {Q}: a0×d0+a1×d1+a2×d2+…+an×dn+Q=0. The solution of the first set of data recovery equations, d0'=-d1-d2-…-dn-P, is used as the service data recovery result of the first parity chain combination for this service data block combination; the solution of the second set of data recovery equations, d0''=(-a1×d1-a2×d2-…-an×dn-Q) / a0, is used as the service data recovery result of the second parity chain combination for this service data block combination. Here, ai is a constant, and RAID6 is a RAID array containing two check disks.
[0060] As an optional embodiment, the abnormal data locating method further includes: After the disk where the abnormal data is located is identified, the real-time data reconstruction process is triggered, including: Read data from normal business data blocks and check data blocks in the target stripe, and reversely calculate the target value of the abnormal data block based on the check chain equation; Overwrite the target value to the corresponding exception data block.
[0061] Specifically, after determining the disk where the abnormal data is located, it is necessary to promptly repair the abnormal data to prevent further data damage or affect the normal operation of the system. Therefore, in the embodiment of the present invention, a real-time data reconstruction process can be triggered to reversely calculate the target value of the abnormal data using normal data and verification data and overwrite it, thereby achieving data repair.
[0062] Specifically, for example, in a RAID6 system, after determining that the Q parity disk is abnormal, the real-time data reconstruction process is triggered, and data is read from the normal business data blocks d0, d1, etc. and the P parity data block in the target stripe. Based on the parity chain equation Q=a0×d0+a1×d1+…+an×dn (where ai is a constant), the target value Q'=a0×d0+a1×d1+…+an×dn of the abnormal data block Q is reversely calculated, and the target value Q' is overwritten to the abnormal data block Q to which it belongs.
[0063] As an optional embodiment, after overwriting the target value to the corresponding abnormal data block, the abnormal data locating method further includes: Verify whether the abnormal data blocks in the target stripe can be successfully recovered; If recovery is unsuccessful, the disk containing the abnormal data will be marked as untrustworthy and isolated.
[0064] Specifically, considering that data may fail after reconstruction, the repair results need to be verified to ensure data reliability. Therefore, in the embodiment of the present invention, if recovery is unsuccessful, the disk is marked and isolated to prevent the abnormal disk from further affecting the system and ensure data security.
[0065] In the previous embodiment, after overwriting the target value Q' to the corresponding abnormal data block Q, the abnormal data block in the target stripe is verified to be recoverable. The parity chain calculation result of the target stripe's P parity chain is calculated as X = d0 + d1 + ... + dn + Q'. If X is non-zero, indicating that the abnormal data block cannot be successfully recovered, the disk containing the abnormal data is marked as untrusted and isolated.
[0066] As an optional embodiment, verifying whether the abnormal data block in the target stripe can be successfully recovered includes: Calculate the check chain calculation result of any check chain of the target stripe; If the check chain calculation result is non-zero, the real-time data reconstruction process is triggered for the second time; Calculate the parity chain calculation result of any parity chain of the target stripe based on the abnormal data block recovered for the second time; If the checksum calculation result is non-zero, the physical disk is marked as untrusted and quarantined.
[0067] Specifically, considering that other factors may affect the data after a data reconstruction failure, by triggering the reconstruction process again and verifying it, it is possible to further confirm whether the disk is truly irreparable. If it still fails after two reconstructions, it means that the disk has a serious fault and needs to be marked for isolation. Therefore, in an embodiment of the present invention, it is possible to: calculate the check chain calculation result of any check chain of the target stripe; if the check chain calculation result is non-zero, trigger the real-time data reconstruction process for a second time; based on the abnormal data block recovered for the second time, calculate the check chain calculation result of any check chain of the target stripe; if the check chain calculation result is non-zero, mark the physical disk as untrustworthy and isolate it.
[0068] Specifically, for example, when verifying whether the abnormal data block in the target stripe can be successfully recovered, the check chain calculation result X=d0+d1+…+dn+Q' of the target stripe's P check chain is calculated. If X is non-zero, the real-time data reconstruction process is triggered a second time, and data is re-read from the normal business data blocks and check data blocks in the target stripe. The target value Q'' of the abnormal data block Q is calculated and overwritten to the abnormal data block Q. Based on the abnormal data block Q'' recovered for the second time, the check chain calculation result X'=d0+d1+…+dn+Q'' of the target stripe's P check chain is calculated. If X' is non-zero, the physical disk is marked as untrusted and isolated.
[0069] Of course, in addition to this specific form, "verifying whether the abnormal data blocks in the target stripe can be successfully recovered" can also be in other forms, which are not limited in this embodiment of the present invention.
[0070] As an optional embodiment, the abnormal data locating method further includes: After determining the disk where the abnormal data is located, before recovering the abnormal data, mark the disk where the abnormal data is located as abnormal, and adjust the read and write policy of the storage system to avoid further access to the abnormal disk. Specifically, considering that before recovering abnormal data, in order to avoid further access to the abnormal disk causing data corruption or system failure, the embodiment of the present invention first marks the abnormal state and adjusts the read and write strategy to ensure system stability and data security.
[0071] For example, in a RAID system, after determining that a data disk d0 is abnormal, before recovering the abnormal data, the disk where the abnormal data is located is marked as abnormal, and the read and write policy of the storage system is adjusted, such as prohibiting writing data to the disk, and giving priority to obtaining data from other normal disks when reading data to avoid further access to the abnormal disk.
[0072] As an optional embodiment, the redundant array of independent disks system applies a double check technology or a triple check technology; The target stripe includes at least two check data blocks, and each check data block corresponds to a different check chain.
[0073] Specifically, considering that redundant array of independent disks systems employ double parity (e.g., RAID 6) or triple parity (e.g., RAID TP) to improve system fault tolerance and data security, the target stripe in this embodiment of the present invention includes at least two parity data blocks, corresponding to different parity chains. This provides more verification evidence for locating abnormal data, improving the accuracy and reliability of locating data.
[0074] Specifically, the Redundant Array of Independent Disks (RAID6) system uses dual parity technology. The target stripe contains two parity blocks, P and Q, and each parity block corresponds to a different parity chain (P parity chain and Q parity chain). RAID6 is a RAID array that includes two parity disks and has error correction capabilities. If two or fewer disks fail, the data on the failed disk can be recovered using the remaining disks.
[0075] As an optional embodiment, for a target stripe in the storage system having abnormal data, after determining the parity chain calculation results of each parity chain of the target stripe based on the current data of each data block in the target stripe, the abnormal data locating method further includes: If the calculation results of each check chain are all zero, it is determined that the data blocks in the target stripe are normal and there is no need to perform abnormal data location operations.
[0076] Specifically, considering that when the calculation results of all check chains are all zero, it means that the data blocks in the target stripe are all normal and there is no need to perform abnormal data locating operations, therefore, in an embodiment of the present invention, if the calculation results of each check chain are all zero, it is determined that the data blocks in the target stripe are all normal and there is no need to perform abnormal data locating operations, thereby avoiding unnecessary calculations and resource waste and improving system efficiency.
[0077] For example, for a target stripe with abnormal data in a RAID6 system, the P parity chain calculation result X=d0+d1+…+dn+P and the Q parity chain calculation result Y=a0×d0+a1×d1+…+an×dn+Q of the target stripe are calculated using the current data of each data block in the target stripe. If both X and Y are zero, it is determined that the data blocks in the target stripe are normal and there is no need to perform the abnormal data locating operation.
[0078] As an optional embodiment, the abnormal data locating method further includes: Perform periodic inspections on multiple stripes in the storage system and identify target stripes with abnormal data by comparing the parity chain calculation results of each stripe with the preset benchmark value.
[0079] Specifically, considering that periodic inspections are performed on multiple stripes in a storage system, target stripes with abnormal data can be discovered in a timely manner so that abnormalities can be located and repaired as early as possible to ensure data security and system stability, in an embodiment of the present invention, periodic inspections can be performed on multiple stripes in a storage system, and target stripes with abnormal data can be identified by comparing the check chain calculation results of each stripe with a preset benchmark value.
[0080] Specifically, for example, multiple stripes in the storage system are periodically inspected by comparing the parity chain calculation results of each stripe with a preset reference value (the preset reference value is zero). For example, the P parity chain result X and the Q parity chain result Y of a certain stripe are calculated. If X or Y is non-zero, the stripe is identified as a target stripe with abnormal data.
[0081] As an optional embodiment, the abnormal data locating method further includes: When determining the parity chain calculation result, if non-zero results are detected for N consecutive stripes in the same parity chain, it is determined that the physical disk corresponding to the parity chain has a systematic fault, so as to trigger a preventive disk replacement process.
[0082] Specifically, considering that if non-zero results appear in the same parity chain for N consecutive stripes, it indicates that the physical disk corresponding to the parity chain may have a systemic fault, triggering a preventive disk replacement process, and replacing the faulty disk in advance to avoid data loss and system failure. Therefore, in an embodiment of the present invention, when determining the parity chain calculation result, if non-zero results are detected in the same parity chain for N consecutive stripes, it is determined that a systemic fault exists in the physical disk corresponding to the parity chain, so as to trigger the preventive disk replacement process.
[0083] For example, when determining the parity chain calculation result, if three consecutive stripes are detected to have non-zero results in the Q parity chain, it is determined that the physical disk corresponding to the Q parity chain has a systemic failure, so as to trigger the preventive disk replacement process and replace the physical disk in advance.
[0084] Please refer to Figure 6 , Figure 6 This is a structural diagram of an abnormal data locating device provided by the present invention, the abnormal data locating device comprising: A first determining module 61 is configured to determine, for a target stripe having abnormal data in the storage system, a parity chain calculation result of each parity chain of the target stripe based on current data of each data block in the target stripe; The second determining module 62 is configured to, if the calculation results of each check chain are partially zero, determine the disk corresponding to the check chain with a non-zero calculation result as the disk where the abnormal data is located; The data blocks include business data blocks and check data blocks. The disk corresponding to the check chain is the disk to which the check data in the check chain belongs in the target stripe. The storage system is a redundant array of independent disks system.
[0085] For an introduction to the abnormal data locating device provided by an embodiment of the present invention, please refer to the aforementioned embodiment of the abnormal data locating method, and the embodiment of the present invention will not be described in detail here.
[0086] Please refer to Figure 7 , Figure 7 This is a structural diagram of an abnormal data locating device provided by the present invention, the abnormal data locating device comprising: Memory 71, for storing computer programs; The processor 72 is configured to implement the steps of the abnormal data locating method in the aforementioned embodiment when executing the computer program.
[0087] As an optional embodiment, the processor is integrated into a field programmable gate array of a redundant array of independent disks controller card, and the check chain calculation is performed by the following hardware acceleration unit: Galois field multiplier array for parallel calculation of the coefficients of the check equation; XOR operation tree module, used for XOR chain calculation.
[0088] For an introduction to the abnormal data locating device provided by an embodiment of the present invention, please refer to the aforementioned embodiment of the abnormal data locating method, and the embodiment of the present invention will not be described in detail here.
[0089] Specifically, considering that the processor is integrated into the Field-Programmable Gate Array (FPGA) of the independent disk redundant array control card, the parallel computing capability of the FPGA is utilized, combined with hardware acceleration units such as the Galois Field multiplier array and the XOR operation tree module, to improve the speed and efficiency of the check chain calculation and meet the high performance requirements of the storage system. Therefore, the processor and hardware acceleration unit in the embodiment of the present invention are designed.
[0090] Specifically, the processor is integrated into the FPGA of the redundant array of independent disks controller card. Check chain calculations are performed by the following hardware acceleration units: a Galois Field multiplier array is used to parallelize the coefficients of check equations, such as ai × di (where ai is a constant and di is a data block); an XOR operation tree module is used for XOR chain calculations, such as d0 + d1 + … + dn + P (where "+" represents an XOR operation). An FPGA is a field-programmable integrated circuit used to implement customized hardware logic; a Galois Field is a finite field used for mathematical operations on data checksums and encoding.
[0091] The present invention also provides a computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the abnormal data locating method in the aforementioned embodiment.
[0092] For an introduction to the computer program product provided by the embodiment of the present invention, please refer to the aforementioned embodiment of the abnormal data locating method, and the embodiment of the present invention will not be described in detail here.
[0093] Please refer to Figure 8 , Figure 8 This is a structural diagram of a computer-readable storage medium provided by the present invention. A computer program 82 is stored on the computer-readable storage medium 81. When the computer program 82 is executed by a processor, the steps of the abnormal data locating method in the aforementioned embodiment are implemented.
[0094] For an introduction to the computer-readable storage medium provided in an embodiment of the present invention, please refer to the aforementioned embodiment of the abnormal data locating method, and the embodiment of the present invention will not be described in detail here.
[0095] In this specification, the various embodiments are described in a progressive manner, with each embodiment focusing on the differences from the other embodiments. Similar or identical parts between the various embodiments may be referred to in conjunction with each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and for relevant parts, reference may be made to the method description. It should also be noted that, in this specification, relational terms such as first and second, etc., are used solely to distinguish one entity or operation from another, and do not necessarily require or imply any actual relationship or order between these entities or operations. Furthermore, the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, article, or device comprising that element.
[0096] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for locating abnormal data, characterized in that: include: For a target stripe with abnormal data in the storage system, the parity chain calculation results of each parity chain of the target stripe are determined based on the current data of each data block in the target stripe. If the calculation results of each check chain are partially zero, the disk corresponding to the check chain with a non-zero calculation result is regarded as the disk where the abnormal data is located; The data blocks include business data blocks and check data blocks. The disk corresponding to the check chain is the disk to which the check data in the check chain belongs in the target stripe. The storage system is a redundant array of independent disks system.
2. The abnormal data locating method according to claim 1, characterized in that: For a target stripe in the storage system that has abnormal data, after determining the parity chain calculation results of each parity chain of the target stripe based on the current data of each data block in the target stripe, the abnormal data locating method further includes: If the calculation results of all check chains are non-zero, the check chains are divided into two check chain combinations, where the number of check chains in each check chain combination is a preset number, which is the total number of check chains minus one; For any business data block combination in the target stripe, business data recovery is performed on the business data block combination through each check chain combination to obtain business data recovery results of each check chain combination for the business data block combination; Determining whether the business data recovery results of the business data block combination are consistent; If they are consistent, the disk to which each business data block in the business data block combination belongs is used as the disk where the abnormal data is located.
3. The abnormal data locating method according to claim 2, characterized in that: The disk to which each business data block in the business data block combination belongs is used as the disk where the abnormal data is located, including: Determining whether there is a business data block in the business data block combination that meets a preset elimination condition, wherein the preset elimination condition includes that the restored data is consistent with the existing data; If so, the disk to which the business data blocks in the business data block combination, excluding the business data blocks that meet the preset elimination conditions, belong is used as the disk where the abnormal data is located.
4. The abnormal data locating method according to claim 2, characterized in that: The target stripe has at least three check chains; If the calculation results of each check chain are all non-zero, the check chain is divided into two check chain combinations including: If the calculation results of each check chain are all non-zero, a check chain of the target stripe is selected as the initial check chain; For any service data block in the target stripe, recover the service data block using the initial check chain to obtain recovered data of the service data block; If there is a business data block in the target stripe where the restored data is consistent with the existing data, the disk where the business data block in the target stripe where the restored data is consistent with the existing data belongs is used as the disk where the abnormal data is located; If there is no business data block in the target stripe whose restored data is consistent with the existing data, the check chain is divided into two check chain combinations.
5. The abnormal data locating method according to claim 2, characterized in that: Dividing the check chain into two check chain combinations includes: Randomly select a preset number of check chains from all check chains of the target stripe to form the first check chain combination; Randomly select a preset number of check chains from all check chains of the target stripe to form a second check chain combination; The first check chain combination is different from the second check chain combination.
6. The abnormal data locating method according to claim 2, characterized in that: The generation of the business data block combination adopts a hierarchical traversal strategy: Prioritize traversing the data block combinations on the disks that have recently experienced input and output errors; Secondly, it traverses the disk data block combinations marked as high-risk aging levels by the storage system; Finally, traverse the data block combinations of the remaining disks.
7. The abnormal data locating method according to claim 2, characterized in that: For any business data block combination in the target stripe, business data recovery is performed on the business data block combination through each check chain combination, and the business data recovery results of each check chain combination for the business data block combination include: For any combination of business data blocks in the target stripe, a first set of data recovery equations is constructed for the business data block combination by using a first check chain combination, and a second set of data recovery equations is constructed for the business data block combination by using a second check chain combination; Using the solution of the first set of data recovery equations as the business data recovery result of the first check chain combination for the business data block combination; The solution result of the second set of data recovery equations is used as the business data recovery result of the second check chain combination for the business data block combination.
8. The abnormal data locating method according to claim 1, characterized in that: The abnormal data locating method further includes: After the disk where the abnormal data is located is identified, the real-time data reconstruction process is triggered, including: Read data from normal business data blocks and check data blocks in the target stripe, and reversely calculate the target value of the abnormal data block based on the check chain equation; Overwrite the target value to the corresponding exception data block.
9. The abnormal data locating method according to claim 8, characterized in that: After overwriting the target value to the corresponding abnormal data block, the abnormal data locating method further includes: Verify whether the abnormal data blocks in the target stripe can be successfully recovered; If recovery is unsuccessful, the disk containing the abnormal data will be marked as untrustworthy and isolated.
10. The abnormal data locating method according to claim 9, characterized in that: Verifying whether the abnormal data blocks in the target stripe can be successfully recovered includes: Calculate the check chain calculation result of any check chain of the target stripe; If the check chain calculation result is non-zero, the real-time data reconstruction process is triggered for the second time; Calculate the parity chain calculation result of any parity chain of the target stripe based on the abnormal data block recovered for the second time; If the checksum calculation result is non-zero, the physical disk is marked as untrusted and quarantined.
11. The abnormal data locating method according to claim 1, characterized in that: The abnormal data locating method further includes: After determining the disk where the abnormal data is located, before recovering the abnormal data, mark the disk where the abnormal data is located as abnormal, and adjust the read and write policy of the storage system to avoid further access to the abnormal disk.
12. The abnormal data locating method according to claim 1, characterized in that: The redundant array of independent disks system uses a double check technology or a triple check technology; The target stripe includes at least two check data blocks, and each check data block corresponds to a different check chain.
13. The abnormal data locating method according to claim 1, characterized in that: For a target stripe in the storage system that has abnormal data, after determining the parity chain calculation results of each parity chain of the target stripe based on the current data of each data block in the target stripe, the abnormal data locating method further includes: If the calculation results of each check chain are all zero, it is determined that the data blocks in the target stripe are normal and there is no need to perform abnormal data location operations.
14. The abnormal data locating method according to claim 1, characterized in that: The abnormal data locating method further includes: Perform periodic inspections on multiple stripes in the storage system and identify target stripes with abnormal data by comparing the parity chain calculation results of each stripe with the preset benchmark value.
15. The abnormal data locating method according to any one of claims 1 to 14, characterized in that: The abnormal data locating method further includes: When determining the parity chain calculation result, if non-zero results are detected for N consecutive stripes in the same parity chain, it is determined that the physical disk corresponding to the parity chain has a systematic fault, so as to trigger a preventive disk replacement process.
16. An abnormal data locating device, characterized in that: include: A first determining module is configured to determine, for a target stripe having abnormal data in the storage system, a check chain calculation result of each check chain of the target stripe based on current data of each data block in the target stripe; The second determining module is configured to, if the calculation results of each check chain are partially zero, determine the disk corresponding to the check chain with a non-zero calculation result as the disk where the abnormal data is located; The data blocks include business data blocks and check data blocks. The disk corresponding to the check chain is the disk to which the check data in the check chain belongs in the target stripe. The storage system is a redundant array of independent disks system.
17. An abnormal data locating device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the abnormal data locating method according to any one of claims 1 to 15 when executing the computer program.
18. The abnormal data locating device according to claim 17, characterized in that: The processor is integrated into the field programmable gate array of the redundant array of independent disks controller card, and the check chain calculation is performed by the following hardware acceleration unit: Galois field multiplier array for parallel calculation of the coefficients of the check equation; XOR operation tree module, used for XOR chain calculation.
19. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the abnormal data locating method according to any one of claims 1 to 15 are implemented.
20. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the abnormal data locating method according to any one of claims 1 to 15.
Citation Information
Patent Citations
RAID (redundant array of independent disk) system and data recovery method thereof
CN102043685A
N-Code-based RAID6 disk array capacity expansion method and data filling method
CN112799604A
Coding method, decoding method and device of RAID6 (redundant array of independent disks 6) and medium
CN115080303A
Secure storage method and device, equipment and storage medium
CN115793985A
Disk array fault data recovery method and device and product
CN118708402A
Cited By
Abnormal data processing method and device, storage medium and electronic equipment
CN120723520A
Control method of disk array and electronic equipment
CN120929347A