Data processing method and device, computer device, and storage medium
By writing and reading data within the storage area according to preset rules, and performing matching and verification directly between storage units, the high cost and failure impact caused by relying on additional storage in existing technologies are solved, and efficient data reliability verification is achieved.
Patent Information
- Application Number
- CN202411762807.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-12-03
AI Technical Summary
Existing data reliability verification methods rely on additional storage devices or storage systems, resulting in high costs and the impact of failures on verification results.
By writing data to multiple storage units in the target storage area according to preset write rules, and reading and matching the data of adjacent storage units in sequence based on the data writing order, the verification is performed directly, avoiding the introduction of additional storage devices.
It saves on data reliability verification costs and avoids the impact of verification caused by the unavailability of additional storage devices in failure scenarios, thereby improving the availability of the storage system.
Smart Images

Figure CN119883108B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data access technology, and more specifically, to a data processing method, apparatus, computer equipment, and storage medium. Background Technology
[0002] In related technologies, current data reliability verification methods primarily rely on comparison. This involves comparing the data to be tested with verification data to assess the data's completeness and accuracy. However, verification data is often stored in additional storage devices or systems. This increases the cost of data reliability verification due to the need for additional storage devices or systems. Furthermore, disconnections or failures in the storage devices or systems storing verification data can also affect the results of the data reliability verification. Summary of the Invention
[0003] This application proposes a data processing method, apparatus, computer equipment, and storage medium to improve the above-mentioned deficiencies.
[0004] In a first aspect, embodiments of this application provide a data processing method, the method comprising: sequentially writing data to multiple storage units in a target storage area according to a preset writing rule, the preset writing rule including at least a data writing order and a data writing strategy between storage units with adjacent writing orders; sequentially reading data in the multiple storage units based on the data writing order; matching the data in the current storage unit with the data in the preceding storage unit, the current storage unit being the storage unit currently being read among the multiple storage units, and the preceding storage unit being the storage unit preceding the current storage unit that was read; if the data in the current storage unit conforms to the data writing strategy with the data in the preceding storage unit, then determining that the data in the current storage unit has passed the verification.
[0005] Secondly, embodiments of this application provide a data processing apparatus, comprising: a data writing module, a data reading module, and a data verification module. The data writing module is used to sequentially write data to multiple storage units in a target storage area according to preset writing rules, the preset writing rules including at least a data writing order and a data writing strategy between storage units with adjacent writing orders. The data reading module is used to sequentially read data from the multiple storage units based on the data writing order. The data verification module is used to match the data in the current storage unit with the data in the preceding storage unit, the current storage unit being the storage unit currently being read from among the multiple storage units, and the preceding storage unit being the storage unit preceding the current storage unit that was read from. If the data in the current storage unit conforms to the data writing strategy with the data in the preceding storage unit, then the data in the current storage unit is determined to have passed verification.
[0006] Thirdly, embodiments of this application also provide a computer device, including: one or more processors; a memory; one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to perform the methods described above.
[0007] Fourthly, embodiments of this application also provide a computer-readable storage medium storing program code that can be invoked by a processor to execute the above-described method.
[0008] The solution provided in this application sequentially writes data to multiple storage units in the target storage area according to a preset write rule. The preset write rule includes at least the data write order and the data write strategy between adjacent storage units. Based on the data write order, data in multiple storage units is read sequentially. The data in the current storage unit is matched with the data in the preceding storage unit. The current storage unit is the storage unit currently being read from among the multiple storage units, and the preceding storage unit is the storage unit preceding the current storage unit that was read. If the data in the current storage unit conforms to the data write strategy with the data in the preceding storage unit, then the data in the current storage unit is considered to have passed the verification. Thus, there is no need to introduce additional storage for storing verification data. Instead, by reading the data in the current storage unit and the preceding storage unit, and performing matching and verification, it is possible to determine whether the data in the current storage unit conforms to the preset write strategy, thereby completing the data reliability verification. This saves the cost of data reliability verification and avoids the impact of additional storage failures on data verification, as no additional storage is required. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 A schematic flowchart of a data processing method provided in an embodiment of this application is shown.
[0011] Figure 2 A schematic diagram of the data writing and data verification process provided in an embodiment of this application is shown.
[0012] Figure 3 This illustration shows a schematic diagram of the partitioning of a storage unit according to an embodiment of this application.
[0013] Figure 4 This illustration shows a schematic representation of a linked list for writing data, as provided in an embodiment of this application.
[0014] Figure 5 This illustration shows a schematic diagram of the contents of a storage unit provided in an embodiment of this application.
[0015] Figure 6 A flowchart illustrating a data processing method provided in another embodiment of this application is shown.
[0016] Figure 7 This illustration shows a data writing diagram of a storage unit provided in an embodiment of this application.
[0017] Figure 8 This illustration shows a data verification diagram of a storage unit provided in an embodiment of this application.
[0018] Figure 9 It shows Figure 6 A flowchart illustrating the sub-steps of step S250.
[0019] Figure 10 This is a block diagram of a data processing apparatus according to an embodiment of this application.
[0020] Figure 11 This is a block diagram of a computer device for performing a data processing method according to an embodiment of this application.
[0021] Figure 12 This is a storage unit in this application embodiment for storing or carrying program code that implements the data processing method according to this application embodiment. Detailed Implementation
[0022] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative effort are within the scope of protection of the present application.
[0023] It should be noted that some processes described in the specification, claims, and accompanying drawings of this application include multiple operations that appear in a specific order. These operations may not be performed in the order they appear herein, or they may be performed in parallel. Operation numbers such as S110, S120, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be performed sequentially or in parallel. Also, the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having”, and any variations thereof, are intended to cover non-exclusive inclusion, such that a process, method, system, product, or server that includes a series of steps or sub-modules is not necessarily limited to those steps or sub-modules that are explicitly listed, but may include other steps or sub-modules that are not explicitly listed or that are inherent to such process, method, product, or device.
[0024] In related technologies, the verification method for data reliability is mainly based on comparison. The integrity and accuracy of the data are evaluated by comparing the data to be tested with the verification data. Specific comparison methods generally include copy comparison, check value comparison, and idempotent data comparison.
[0025] In copy verification, in addition to writing the data to the test target, a complete copy of the data needs to be written to another dedicated storage device (which can be of various types or systems) for verification data. During verification, the reliability of the data is assessed by comparing the two stored data. Clearly, copy verification requires additional storage devices or systems, with a one-to-one storage overhead, increasing the cost of data reliability verification. Furthermore, disconnections or failures in the storage device or system storing the verification data will also affect the verification test results.
[0026] In the verification process, besides writing data to the test target, the calculated verification value is also written to a separate storage device specifically for verification data. During verification, the data is read, the recalculated verification value is compared with the original value to assess data reliability. Clearly, this verification method requires additional storage devices or systems, increasing the cost of data reliability verification. Furthermore, disconnections or failures in the storage device or system storing verification data can affect the verification test results. Additionally, extra computing power is needed to calculate the verification value.
[0027] Idempotent data comparison involves writing data to a test target using a "seed data" dataset. During the writing process, in addition to writing data to the test target, an update and its sequence number (metadata) are recorded in another storage location. Verification involves reading the data and regenerating the data using the "seed data," while simultaneously reading the metadata and comparing it to the content of the target storage location to assess data reliability. Clearly, idempotent data comparison requires additional storage devices or systems, increasing the cost of data reliability verification. Furthermore, disconnections or failures in the storage devices or systems storing the verification data can also affect the verification test results.
[0028] It is evident that the data reliability verification methods in related technologies all rely on additional storage devices or storage systems, which result in high costs; furthermore, disconnections or failures of additional storage devices or storage systems can also affect the verification test results.
[0029] To address the aforementioned problems, the inventors have proposed a data processing method, apparatus, computer device, and storage medium. The data processing method provided in the embodiments of this application will be described in detail below.
[0030] Please refer to Figure 1 , Figure 1 This is a schematic flowchart illustrating a data processing method provided in one embodiment of this application. The following will be combined with... Figure 1 The data processing method provided in the embodiments of this application will be described in detail. This data processing method may include the following steps:
[0031] Step S110: Write data sequentially to multiple storage units in the target storage area according to the preset write rules. The preset write rules include at least the data write order and the data write strategy between storage units with adjacent write orders.
[0032] In this embodiment, please refer to Figure 2 First, the metadata will be initialized, that is, the metadata parameters will be written to... Figure 3The target storage area shown corresponds to the metadata area. Metadata, also known as intermediary data or relay data, is data that describes data, mainly information describing data attributes, used to support functions such as indicating storage location, historical data, resource lookup, and file records. Metadata is a kind of electronic catalog; to achieve the purpose of cataloging, it is necessary to describe and collect the content or characteristics of the data, thereby assisting in data retrieval.
[0033] The aforementioned metadata parameters include at least a storage unit partitioning strategy. Then, based on the storage unit partitioning strategy, the target storage area can be partitioned to obtain multiple storage units, such as... Figure 3 Storage units 1 to 16 are shown.
[0034] Optionally, the storage unit partitioning strategy may include a preset unit size for the storage units. In this case, the size of each storage unit among the multiple storage units obtained by partitioning the target storage area is the preset unit size. For example, if the preset unit size of the storage unit is a hardware sector of 512 bytes, then the size of each storage unit among the multiple storage units obtained by partitioning the target storage area is 512 bytes. Here, a sector refers to a region on the disk. Each track on the disk is equally divided into several arc segments, which are the sectors of the disk. Hard disk read and write operations are performed using sectors as the basic unit.
[0035] Optionally, the storage unit partitioning strategy may also include a preset number of storage units. In this case, the target storage area is divided into the preset number of storage units. For example, if the target storage area is 1024 bytes in size and the preset number is 16, then the target storage area is divided into 16 storage units. Of course, this embodiment does not limit the specific method used to divide the target storage area into multiple storage units.
[0036] Furthermore, after the target storage area is partitioned, data can be written sequentially to multiple storage units within the target storage area based on the data write order and the data write strategy between adjacent storage units. Specifically, the aforementioned initialized metadata parameters may include not only the storage unit partitioning strategy but also a preset concatenation algorithm, a data write strategy, and a preset generation algorithm. The preset concatenation algorithm is used to indicate the data write order to be written to multiple storage units. It should be noted that after multiple storage units are partitioned, a unit identifier is assigned to each storage unit, which can be a numerical identifier.
[0037] In other words, based on a preset concatenation algorithm, the storage unit to be added to the write list is determined from multiple storage units as the target storage unit. The preset concatenation algorithm can either concatenate the storage units according to their numerical identifiers in ascending order, or in descending order, or concatenate the storage units according to their odd-numbered numerical identifiers in ascending order, and then concatenate them according to their even-numbered numerical identifiers in ascending order, forming a write list. Figure 4 The example shows a write-to-link list. Alternatively, the default concatenation algorithm can concatenate the storage units with odd-numbered identifiers in descending order, then concatenate them in descending order of even-numbered identifiers to form the write-to-link list. Another default concatenation algorithm can randomly combine multiple storage units to form the write-to-link list.
[0038] It should be noted that the head of the linked list is determined based on a preset concatenation algorithm in the metadata area, such as... Figure 5 As shown, after determining that storage unit 1 is the head of the linked list to be written based on the preset concatenation algorithm in the metadata area, the unit identifier of storage unit 1 is recorded in the metadata area. Furthermore, for storage unit 1, the preset generation algorithm in the metadata area is used to generate data for writing into storage unit 1, and then the generated data is written into storage unit 1.
[0039] Furthermore, if the target storage unit is not the head of the write list, a preset generation algorithm is used to generate the data to be written based on the data in the storage unit preceding the target storage unit that has been added to the write list. For example, if the determined target storage unit is... Figure 5 Storage unit 2 in the target storage unit is preceded by storage unit 1, which was added to the write list. In this case, a preset generation algorithm is used to generate the data to be written based on the data in storage unit 1. Finally, based on the data writing strategy between adjacent storage units, the data to be written is written to the target storage unit. The data writing strategy can be understood as follows:
[0040] In other words, this embodiment writes data to multiple storage units in the target storage area using a linked list until all storage units have completed the data writing process. The length of the linked list depends on the size of the metadata area. For example, if the storage unit number is stored in 4-byte format and the metadata area is 64 bytes, then the length of the linked list is 64 / 4 = 16. It should be noted that if the linked list has reached its maximum length, such as 16, when writing data, the next storage unit to be written is determined according to the data writing order, and the address of that storage unit is appended to the end of the linked list. Then, the unit closest to the head of the linked list is removed to ensure that the length of the linked list is fixed. Then, data is written to the storage unit to be written.
[0041] Step S120: Based on the data writing order, read the data in the plurality of storage units sequentially.
[0042] Furthermore, such as Figure 2 As shown, after writing data to multiple storage units in the target storage area, the metadata area can be read, and based on the header identifier in the metadata area's header, the storage unit that needs to be read at the current moment can be determined from the multiple storage units. For example, as... Figure 5 As shown, by reading the metadata area, it can be known that during the data writing phase, the head of the linked list is storage unit 1. Therefore, storage unit 1 is used as the current storage unit for data reading.
[0043] Step S130: Match the data in the current storage unit with the data in the previous storage unit. The current storage unit is the storage unit currently being read among the plurality of storage units, and the previous storage unit is the storage unit that was read before the current storage unit.
[0044] Step S140: If the data in the current storage unit conforms to the data writing strategy with the data in the previous storage unit, then it is determined that the data in the current storage unit has passed the verification.
[0045] During the data writing phase, data is written to adjacent storage units based on a strategy that prioritizes adjacent units. This means that the data within any two adjacent storage units exhibits a certain pattern. Therefore, to verify the reliability of the data within multiple storage units, after writing data to multiple units, the data in each unit is read sequentially based on the writing order. The data read from the current storage unit is then matched with the data from the preceding storage unit and the data from the adjacent preceding storage unit to determine if the data in the current storage unit conforms to the data writing strategy. Optionally, if the data in the current storage unit conforms to the data writing strategy, it indicates that the data in the current storage unit possesses integrity and reliability, meaning the data verification for the current storage unit has passed.
[0046] In this embodiment, there is no need to introduce additional storage for storing verification data. Instead, by reading the data of the current storage unit and the preceding storage unit and performing matching and verification, it can be determined whether the data in the current storage unit conforms to the preset writing strategy, thereby completing the data reliability verification. This saves the cost of data reliability verification. Furthermore, since there is no need to introduce additional storage, it avoids the situation where the storage system fails to cover the business branch as expected in the failure scenario due to the unavailability of the additional storage itself during the execution of failure scenario cases, which ultimately leads to the failure to detect the reliability problem of the storage system in a timely manner.
[0047] Please refer to Figure 6 , Figure 6 This is a flowchart illustrating a data processing method according to another embodiment of this application. The following will be combined with... Figure 6 The data processing method provided in the embodiments of this application will be described in detail. This data processing method may include the following steps:
[0048] Step S210: Write data sequentially to multiple storage units in the target storage area according to a preset write rule. The preset write rule includes at least the data write order and the data write strategy between storage units with adjacent write orders.
[0049] Step S220: Read the metadata area, which includes at least magic number information.
[0050] Step S230: If the magic number information in the metadata area meets the preset conditions, then based on the data writing order, the data in the plurality of storage units are read sequentially.
[0051] Magic number information is a parameter used to determine the file type. The first few bytes of a file are fixed (either intentionally padded or inherently so), and these bytes are called the magic number because the file type can be determined based on them. Obviously, to ensure the correctness of the read metadata area, magic number information can be pre-set in the metadata area. Based on this, during the data reliability verification of the target storage area, the first step is to read the metadata area corresponding to the target storage area and obtain the magic number information within it. Then, it is determined whether the content of the first few bytes of the metadata area (i.e., the magic number information) matches the preset information. If the magic number information matches the preset information, it is determined that the magic number information in the metadata area matches the preset information; if the magic number information does not match the preset information, it is determined that the magic number information in the metadata area does not match the preset information, indicating that the currently read storage area is not the metadata area. If the magic number information in the metadata area matches the preset information, it indicates that the data area read at this time is indeed the metadata area. Further, based on the data writing order, data in multiple storage units is read sequentially. If the magic number information in the metadata area matches the preset information, it indicates that the data is intact, and subsequent data reading operations can then be performed, thereby enhancing the security and reliability of data reading.
[0052] Step S240: Match the header data in the current storage unit with the tail data in the preceding storage unit. The header data is all data in the current storage unit except for the first data, and the tail data is all data in the preceding storage unit except for the last data.
[0053] In some implementations, the write strategy between sequentially adjacent memory cells can be to remove the first data from the preceding memory cell and append target data to the end of the data in the preceding memory cell after removing the first data. The target data can be generated using a preset algorithm based on the data in the preceding memory cell. Finally, the data in the preceding memory cell with the target data appended is written to the current memory cell. For an example, please refer to [link to example]. Figure 7 The data in the preceding storage unit can be {23456}. Based on the aforementioned writing strategy, the first data 2 in the preceding storage unit is removed, and the target data 7 is generated based on the data in the preceding storage unit using a preset generation algorithm. Finally, the target data 7 is added to the end of the data {3456} in the preceding storage unit after the first data is removed, and the data in the preceding storage unit with the target data added is written to the current storage unit. Thus, the final data in the current storage unit is {34567}.
[0054] In this method, during the data reading phase, the data read from the current storage unit is matched with the data read from the preceding storage unit. This can be achieved by matching the header data of the current storage unit with the tail data of the preceding storage unit. The header data consists of all data in the current storage unit except for the first data item, and the tail data consists of all data in the preceding storage unit except for the last data item. Obviously, if the header data in the current storage unit is the same as the tail data in the preceding storage unit, such as... Figure 8 As shown, if the header data in the current storage unit and the tail data in the previous storage unit are both {3456}, then it can be determined that the header data in the current storage unit matches the tail data in the previous storage unit. Optionally, if the header data in the current storage unit is different from the tail data in the previous storage unit, for example, if the header data in the current storage unit is {3456} and the tail data in the previous storage unit is {3458}, then it can be determined that the header data in the current storage unit does not match the tail data in the previous storage unit.
[0055] Step S250: If the header data in the current storage unit matches the tail data in the preceding storage unit, then the data verification in the current storage unit is confirmed to be successful.
[0056] Furthermore, if the header data in the current storage unit matches the tail data in the preceding storage unit, then the reliability verification of the data in the current storage unit can be considered successful. In this way, data reliability testing of the storage system is achieved directly through mutual verification of data within each storage unit of the target storage area. This eliminates the need for additional storage devices or systems to store verification data, improving the availability of the testing system itself and reducing testing costs.
[0057] Optionally, if the header data in the current storage unit does not match the tail data in the preceding storage unit, it is determined that the data verification in the current storage unit has failed, and a prompt message is output to alert the tester that the data in the current storage unit may have been tampered with or damaged.
[0058] In this embodiment, with Figure 4 Taking the writing order as an example, storage unit 1, storage unit 2, storage unit 5... storage unit 14 and storage unit 16 will be read sequentially, and the data verification process in steps S280 to S290 will be performed.
[0059] In some implementations, please refer to Figure 9 Step S250 may include the contents of steps S251 to S253:
[0060] Step S251: If the header data in the current storage unit matches the tail data in the preceding storage unit, then verification data is generated according to the preset generation algorithm and the data in the preceding storage unit.
[0061] After determining that the header data in the current storage unit matches the tail data in the preceding storage unit, in order to improve the accuracy of data verification, further verification data for the current storage unit can be generated based on the preset generation algorithm and the data in the preceding storage unit.
[0062] Optionally, if the preceding storage unit does not contain a preset generation algorithm, verification data is generated based on the preset generation algorithm in the metadata area and the data in the preceding storage unit. That is, the data written to all storage units is generated using the preset generation algorithm in the metadata area. Thus, during the data writing phase, a unified preset generation algorithm can be used to quickly generate data and write it to the current storage unit; similarly, during the verification phase after data reading, a unified preset generation algorithm can also be used to quickly generate verification data.
[0063] Optionally, if the preceding storage unit contains a preset generation algorithm, the verification data is generated based on the preset generation algorithm and the data in the preceding storage unit. That is, for different storage units in the target storage area, corresponding preset generation algorithms can be set; the preset generation algorithms for different storage units can be different or the same. The preset generation algorithm used by each storage unit can be stored in the storage unit preceding the adjacent write order. This reduces the probability of data collisions between storage units. Furthermore, in the verification stage after data reading, verification data is generated for each current storage unit using the preset generation algorithm in its preceding storage unit.
[0064] Step S252: Match the data in the current storage unit with the verification data.
[0065] Step S253: If the data in the current storage unit matches the verification data, then the data in the current storage unit is verified as passed.
[0066] Furthermore, the verification data generated for the current storage unit is matched with the data read from the current storage unit. Specifically, it is determined whether the verification data generated for the current storage unit is equal to the data in the current storage unit. If they are equal, the data in the current storage unit is considered to match the verification data, and the verification of the data in the current storage unit is deemed successful. If they are not equal, the data in the current storage unit is considered to be inconsistent with the verification data, and the verification of the data in the current storage unit is deemed to have failed. A prompt message is then output to indicate that the data in the current storage unit may have been tampered with or corrupted.
[0067] Optionally, the verification data generated for the current storage unit can be the last data in the current storage unit. That is, if the header data in the current storage unit matches the tail data in the preceding storage unit, the verification data generated for the current storage unit will be further matched with the last data read from the current storage unit. If the verification data is the same as the last data read from the current storage unit, the data verification is successful; if the verification data is different from the last data read from the current storage unit, the data verification is unsuccessful.
[0068] Optionally, if the current storage unit is the first storage unit in the data writing sequence, steps S240 and S250 are not required. Verification data is directly generated based on the preset generation algorithm and the data in the preceding storage unit. All data read from the current storage unit is matched with the verification data. If all data in the current storage unit is the same as the verification data, a match is determined, and the data in the current storage unit is verified as passed. If some data in the current storage unit is different from the verification data, a mismatch is determined, and the data in the current storage unit is verified as failed. A prompt message is output to alert the tester that the data in the current storage area is corrupted or tampered with.
[0069] In some implementations, the metadata area may also include a timestamp. If it is determined that the data verification in the current storage unit fails, the timestamp corresponding to the current storage unit can be obtained. If the timestamp is a preset timestamp, it is determined that the data in the current storage area is corrupted. If the timestamp is not a preset timestamp, and the time represented by the timestamp is closer to the current time than the time represented by the preset timestamp, it is determined that the data in the current storage unit has been tampered with.
[0070] In this embodiment, the data reliability test of the storage system is achieved by cross-verifying the data in adjacent storage units within the target storage area. This eliminates the need for additional storage to store the verification data, thereby improving the availability of the testing system and reducing testing costs.
[0071] Please refer to Figure 10 The diagram illustrates a structural block diagram of a data processing apparatus 300 according to an embodiment of this application. The apparatus 300 may include a data writing module 310, a data reading module 320, and a data verification module 330.
[0072] The data writing module 310 is used to write data sequentially to multiple storage units in the target storage area according to a preset writing rule. The preset writing rule includes at least the data writing order and the data writing strategy between storage units with adjacent writing orders.
[0073] The data reading module 320 is used to read data from the plurality of storage units sequentially based on the data writing order.
[0074] The data verification module 330 is used to match the data in the current storage unit with the data in the preceding storage unit. The current storage unit is the storage unit currently being read from among the plurality of storage units, and the preceding storage unit is the storage unit that was read before the current storage unit. If the data in the current storage unit conforms to the data writing strategy with the data in the preceding storage unit, then it is determined that the data in the current storage unit has passed the verification.
[0075] In some implementations, the data verification module 330 may be specifically used to: match the header data in the current storage unit with the tail data in the preceding storage unit, wherein the header data is all data in the current storage unit except for the first data, and the tail data is all data in the preceding storage unit except for the last data; if the header data in the current storage unit matches the tail data in the preceding storage unit, then it is determined that the data in the current storage unit has passed verification.
[0076] In some embodiments, the data verification module 330 may further include a verification data generation unit and a verification unit. The first verification unit may be used to generate verification data based on a preset generation algorithm and the data in the preceding storage unit if the header data in the current storage unit matches the tail data in the preceding storage unit. The verification unit may be used to match the data in the current storage unit with the verification data; if the data in the current storage unit matches the verification data, it is determined that the data in the current storage unit has passed verification.
[0077] In this approach, the metadata area corresponding to the target storage area includes a preset generation algorithm. The verification data generation unit can be specifically used to: if the preceding storage unit does not contain a preset generation algorithm, generate the verification data based on the preset generation algorithm in the metadata area and the data in the preceding storage unit; if the preceding storage unit contains a preset generation algorithm, generate the verification data based on the preset generation algorithm in the preceding storage unit and the data in the preceding storage unit.
[0078] In some embodiments, the data processing device 300 may include a metadata reading module and a magic number verification module. The metadata reading module can be used after data has been written sequentially to multiple storage units in the target storage area according to preset writing rules. The magic number verification module can be used to, if the magic number information in the metadata area matches preset information, then perform the step of sequentially reading data from the multiple storage units based on the data writing order.
[0079] In some embodiments, the data processing device 300 may include a metadata writing module and a storage partitioning module. The metadata writing module may be used to write metadata parameters to a metadata area corresponding to the target storage area before writing data sequentially to multiple storage units in the target storage area according to a preset writing rule. The metadata parameters may include at least a storage unit partitioning strategy. The storage partitioning module may be used to partition the target storage area according to the storage unit partitioning strategy to obtain the multiple storage units.
[0080] In this mode, the data writing module 310 can be specifically used to determine, based on the preset concatenation algorithm, the storage unit to be added to the write list as the target storage unit from the plurality of storage units; generate the data to be written using the preset generation algorithm and based on the data in the storage unit that was added to the write list before the target storage unit; and write the data to be written to the target storage unit based on the data writing strategy.
[0081] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0082] In the several embodiments provided in this application, the coupling between modules can be electrical, mechanical, or other forms of coupling.
[0083] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0084] In summary, according to preset write rules, data is sequentially written to multiple storage units in the target storage area. These preset write rules include at least the data write order and the data write strategy between adjacent storage units. Based on the data write order, data is sequentially read from multiple storage units. The data in the current storage unit is matched with the data in the preceding storage unit. The current storage unit is the storage unit currently being read, and the preceding storage unit is the storage unit preceding the current storage unit that was read. If the data in the current storage unit conforms to the data write strategy with the data in the preceding storage unit, then the data in the current storage unit is considered validated. This eliminates the need for additional storage for validation data. Instead, by reading and matching the data from the current and preceding storage units, the system can determine whether the data in the current storage unit conforms to the preset write strategy, thus completing the data reliability verification. This saves on the cost of data reliability verification. Furthermore, since no additional storage is required, it avoids the situation where, during failure scenario execution, the unavailability of additional storage causes the storage system to fail to cover business branches as expected in failure scenarios, ultimately leading to the failure to detect storage system reliability issues in a timely manner.
[0085] The following will combine Figure 11 This application describes a computer device.
[0086] Reference Figure 11 , Figure 11 This diagram illustrates a structural block diagram of a computer device 400 according to an embodiment of this application. The method described above in this embodiment can be executed by this computer device 400. The computer device can be an electronic terminal with data processing capabilities, including but not limited to tablet computers, laptops, and desktop computers. Alternatively, the computer device can be a server, which can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0087] The computer device 400 in this application embodiment may include one or more of the following components: processor 401, memory 402, and one or more application programs, wherein the one or more application programs may be stored in memory 402 and configured to be executed by one or more processors 401, and the one or more programs are configured to perform the methods as described in the foregoing method embodiments.
[0088] Processor 401 may include one or more processing cores. Processor 401 connects to various parts within the computer device 400 using various interfaces and lines, and performs various functions and processes data of the computer device 400 by running or executing instructions, programs, code sets, or instruction sets stored in memory 402, and by calling data stored in memory 402. Optionally, processor 401 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). Processor 401 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the aforementioned modem can also be integrated into processor 401 and implemented using a separate communication chip.
[0089] The memory 402 may include random access memory (RAM) or read-only memory (ROM). The memory 402 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 402 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), and instructions for implementing the various method embodiments described below. The data storage area may also store data created by the computer device 400 during use (such as the various correspondences described above).
[0090] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0091] In the several embodiments provided in this application, the coupling or direct coupling or communication connection between the modules shown or discussed may be an indirect coupling or communication connection through some interface, device or module, and may be electrical, mechanical or other forms.
[0092] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0093] Please refer to Figure 12 This diagram illustrates a structural block diagram of a computer-readable storage medium provided in an embodiment of this application. The computer-readable medium 500 stores program code that can be called by a processor to execute the methods described in the above method embodiments.
[0094] Computer-readable storage medium 500 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, computer-readable storage medium 500 includes non-transitory computer-readable storage medium. Computer-readable storage medium 500 has storage space for program code 510 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code 510 may be compressed, for example, in a suitable form.
[0095] In some embodiments, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and executes the computer instructions, causing the computer device to perform the steps in the above-described method embodiments.
[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A data processing method, characterized in that, The method includes: According to a preset write rule, data is sequentially written to multiple storage units in the target storage area. The preset write rule includes at least the data write order and the data write strategy between storage units with adjacent write orders. Based on the data writing order, the data in the plurality of storage units are read sequentially; The header data in the current storage unit is matched with the tail data in the preceding storage unit. The current storage unit is the storage unit that is currently being read among the plurality of storage units, and the preceding storage unit is the storage unit that was read before the current storage unit. The header data is all data in the current storage unit except for the first data, and the tail data is all data in the preceding storage unit except for the last data. If the header data in the current storage unit matches the tail data in the preceding storage unit, then the data verification in the current storage unit is deemed successful.
2. The method according to claim 1, characterized in that, If the tail data in the current storage unit matches the head data in the preceding storage unit, then the data verification in the current storage unit is determined to be successful, including: If the header data in the current storage unit matches the tail data in the preceding storage unit, then verification data is generated according to the preset generation algorithm and the data in the preceding storage unit. Match the data in the current storage unit with the verification data; If the data in the current storage unit matches the verification data, then the verification of the data in the current storage unit is deemed successful.
3. The method according to claim 2, characterized in that, The metadata area corresponding to the target storage area includes a preset generation algorithm. Generating verification data based on the preset generation algorithm and the data in the preceding storage unit includes: If the preceding storage unit does not contain a preset generation algorithm, then the verification data is generated according to the preset generation algorithm in the metadata area and the data in the preceding storage unit. If the preceding storage unit contains a preset generation algorithm, then the verification data is generated according to the preset generation algorithm in the preceding storage unit and the data in the preceding storage unit.
4. The method according to claim 1, characterized in that, After writing data sequentially to multiple storage units in the target storage area according to a preset write rule, the method further includes: Read the metadata area, which includes at least magic number information; If the magic number information in the metadata area matches the preset information, then the data in the plurality of storage units is read sequentially based on the data writing order.
5. The method according to any one of claims 1-4, characterized in that, Before writing data sequentially to multiple storage units in the target storage area according to a preset write rule, the method further includes: Write metadata parameters to the metadata area corresponding to the target storage area, wherein the metadata parameters include at least the storage unit partitioning strategy; According to the storage unit partitioning strategy, the target storage area is divided to obtain the multiple storage units.
6. The method according to claim 5, characterized in that, The metadata parameters also include a preset concatenation algorithm, the data writing strategy, and a preset generation algorithm. The preset concatenation algorithm is used to indicate the data writing order in which data is written to the plurality of storage units. The step of writing data sequentially to multiple storage units in the target storage area according to a preset writing rule includes: Based on the preset concatenation algorithm, the storage unit to be added to the write list is determined from the plurality of storage units as the target storage unit; Using the preset generation algorithm, and based on the data in the storage unit preceding the target storage unit that has been added to the write linked list, data to be written is generated; Based on the data writing strategy, the data to be written is written to the target storage unit.
7. A data processing apparatus, characterized in that, The device includes: The data writing module is used to write data sequentially to multiple storage units in the target storage area according to a preset writing rule. The preset writing rule includes at least the data writing order and the data writing strategy between storage units with adjacent writing orders. The data reading module is used to read data from the plurality of storage units sequentially based on the data writing order; The data verification module is used to match the header data in the current storage unit with the tail data in the preceding storage unit. The current storage unit is the storage unit currently being read from among the plurality of storage units, and the preceding storage unit is the storage unit that was read before the current storage unit. The header data consists of all data in the current storage unit except for the first data, and the tail data consists of all data in the preceding storage unit except for the last data. If the header data in the current storage unit matches the tail data in the preceding storage unit, then the data in the current storage unit is determined to have passed the verification.
8. A computer device, characterized in that, The computer device includes: One or more processors; Memory; One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs being configured to perform the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program code that can be invoked by a processor to execute the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Data writing method and device and verification method and device
CN107844273A
Data processing method, system and device and medium
CN117331852A