Novel distributed data verification mode and algorithm
Through a multi-step distributed data verification algorithm, the problem of data backup loss in large-scale Internet architecture is solved, the accuracy and integrity of data in multi-center and multi-copy backup are achieved, service availability is ensured, and economic and resource waste is reduced.
Patent Information
- Application Number
- CN202410250736.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-05
- Publication Date
- 2025-09-05
AI Technical Summary
In existing technologies, in large-scale Internet architectures, data is easily lost when backed up in multiple centers and multiple copies, resulting in service unavailability and economic and resource waste.
A new distributed data verification method and algorithm is adopted to ensure the accuracy and integrity of data backup in multiple locations and multiple centers through multi-step calculation and verification data plane processing, including multi-directional verification calculation from left to right, right to left, top to bottom, bottom to top, etc., and the verification data is copied and compared in memory and on disk.
Effectively reduce data loss, improve the reliability and integrity of data backup, avoid service unavailability, and reduce economic and human resource waste.
Smart Images

Figure CN120596570A_ABST
Abstract
Description
[0001] The present invention belongs to the field of computer data verification, and in particular relates to a distributed network data verification method and algorithm. Background Art
[0002] Currently, there are various data verification methods available, such as MD5 for grouped verification, the SM4 symmetric checksum required by the National Security Agency, and hash digests such as SHA and SM3. There are also verification methods with error correction, such as the Hamming checksum. Simple verification methods include parity checks. In an era when single machines and disks were prohibitively expensive, simple verification methods were particularly cost-effective and significantly reduced costs and increased efficiency. my country has made significant progress in disk research, with capacity continuously expanding, size shrinking, and read and write speeds increasing. However, large-scale internet architectures are becoming increasingly complex, with increasingly demanding storage requirements. The internet often requires high data availability, high access availability, multiple data centers, and multiple replicas to ensure data loss prevention. However, software design is inherently flawed. For example, when two machines are connected via the internet, excessive data access can cause network jitter, leading to data loss. In severe cases, this can lead to data congestion and even software crashes. Based on the current low disk prices, mature technology, and the complexity of large-scale software architectures, this patent proposes a novel distributed data step-by-step verification method and algorithm. Summary of the Invention
[0003] Organize the entire data into a data plane, such as Figure 1 As shown. Through a certain algorithm, for example, first uniformly add each row of data from left to right to obtain a set of verification data A, B, C, D, the result is as follows Figure 2 As shown. From right to left, using the same algorithm, we will inevitably get the same set of data A, B, C, D. The result is as follows Figure 3 As shown. The third step, use different algorithms such as subtracting each column of data from top to bottom and taking the absolute value of the result to obtain a set of data H, I, J, K. The result is as follows Figure 4 As shown. In the fourth step, the algorithm of the third step is used to calculate from bottom to top, and another set of data is obtained, which must be H, I, J, K. The result is as follows Figure 5 After the calculation is completed Figure 5 After that, the fifth step can continue with a similar algorithm, such as obtaining X through the algorithm for the first row and the first column, and the result is as follows Figure 6 As shown, the same algorithm as in step 5 is used to process the sixth row and sixth column, and X is also obtained. The result is as follows Figure 7 As shown. Similarly, calculate the first column and sixth row to get Z, as shown Figure 8 As shown. Calculate the first row and sixth column to get Z, and the final result is as follows Figure 9 shown. Specific implementation
[0004] This invention aims to fully utilize the advantages of domestic storage technology to design a step-by-step verification method, so that data can be backed up in multiple locations and centers, reducing data loss and service unavailability, which would otherwise cause huge economic, resource and manpower waste.
[0005] The present invention provides a new data verification method and algorithm, including: In step S1, the data is organized into a data plane in memory. This data is then copied to disk to ensure local persistence. The data is then distributed to other machines via the network for storage. Step S2: Use an algorithm to calculate the entire data from left to right to obtain a set of verification data. Copy the verification data to the disk as a local verification data. Step S3: Use the same algorithm as in S2 to calculate the entire data from right to left to obtain a set of verification data. This reorganized verification data is compared with the verification data obtained in S2 in memory. If they are the same, the data is correct. If they are different, the data is incorrect, and the error is reported. Step S2 is re-executed to verify whether the data has any problems. If there are problems, the data is discarded and an alarm is issued to indicate the data is incorrect. If the results are the same, a copy of the verification data is saved locally. Step S4: Use an algorithm to calculate the entire data from top to bottom to obtain a set of verification data, and copy the verification data to the disk as a local verification data; Step S5: Use the same algorithm as in step S4 to calculate the entire data from right to left to obtain a set of verification data. This reorganized verification data is compared with the verification data obtained in step S4 in memory. If they are the same, the data is correct. If they are different, the data is incorrect, an error is reported, and step S4 is re-executed to verify whether the data has any problems. If there are problems, the data is discarded and an alarm is issued to indicate the data is incorrect. If the results are the same, a copy of the verification data is saved locally. Step S6: Use an algorithm to calculate the check data of S3 and S5 to obtain a check number, and copy it to the disk; Step S7: Use the same algorithm as in S6 to calculate the checksum from S2 and S4 to obtain a checksum. This checksum is compared with the checksum from S6. If they are the same, the data is correct. If they are different, the data is incorrect and an error needs to be reported. Step S6 is then repeated to verify the data. If there is a problem, the data is discarded and an alarm is issued to indicate the data is incorrect. If they are the same, a copy is saved to disk. Step S8: Use an algorithm to calculate the check data of S3 and S4 to obtain a check number, and copy it to the disk; In step S9, the same algorithm as in S8 is used to calculate the verification data of S2 and S5 to obtain a verification number. This verification number is compared with the verification number obtained in S6. If they are the same, the data is correct. If they are different, the data is incorrect and the error needs to be reported. Step S6 is then re-executed to verify whether there is a problem with the data. If there is a problem, the data is discarded and an alarm is issued to indicate that the data is incorrect. If they are the same, a copy is saved to disk. After executing S2-S9, other machines also execute steps S2-S9 accordingly. Finally, the obtained verification values are compared. If they are the same, the data synchronization is completed correctly. If not, the data is incorrect and the verification error needs to be reported separately. If necessary, manual intervention is required to solve the problem or the data is discarded directly. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Figure 1 This is a flow chart of a data verification method provided by the present invention; Figure 2 It is a planar schematic diagram of data arrangement provided during the implementation of the present invention; Figure 3 This is a schematic diagram of the result obtained after the algorithm performs verification from left to right; Figure 4 This is a schematic diagram of the result obtained after the algorithm performs a right-to-left check; Figure 5 This is a schematic diagram of the results obtained after the algorithm performs verification from top to bottom; Figure 6 This is a schematic diagram of the results obtained after the algorithm performs verification from bottom to top; Figure 7 This is a schematic diagram of the results obtained after the algorithm executes the first row and first column of verification data; Figure 8 This is a schematic diagram of the results obtained after the algorithm executes the sixth row and sixth column verification data; Figure 9 This is a schematic diagram of the results obtained after the algorithm executes the verification data in the first column and the sixth row; Figure 10 This is a schematic diagram of the results obtained after the algorithm executes the verification data in the first row and the sixth column.
Claims
1. A new data verification method and algorithm, characterized in that: The method comprises: In step S1, the data is organized into a data plane in memory. This data is then copied to disk to ensure local persistence. The data is then distributed to other machines via the network for storage. Step S2: Use an algorithm to calculate the entire data from left to right to obtain a set of verification data. Copy the verification data to the disk as a local verification data. Step S3: Use the same algorithm as in S2 to calculate the entire data from right to left to obtain a set of verification data. This reorganized verification data is compared with the verification data obtained in S2 in memory. If they are the same, the data is correct. If they are different, the data is incorrect, and the error is reported. Step S2 is re-executed to verify whether the data has any problems. If there are problems, the data is discarded and an alarm is issued to indicate the data is incorrect. If the results are the same, a copy of the verification data is saved locally. Step S4: Use an algorithm to calculate the entire data from top to bottom to obtain a set of verification data, and copy the verification data to the disk as a local verification data; Step S5: Use the same algorithm as in step S4 to calculate the entire data from right to left to obtain a set of verification data. This reorganized verification data is compared with the verification data obtained in step S4 in memory. If they are the same, the data is correct. If they are different, the data is incorrect, an error is reported, and step S4 is re-executed to verify whether the data has any problems. If there are problems, the data is discarded and an alarm is issued to indicate the data is incorrect. If the results are the same, a copy of the verification data is saved locally. Step S6: Use an algorithm to calculate the check data of S3 and S5 to obtain a check number, and copy it to the disk; Step S7: Use the same algorithm as in S6 to calculate the checksum from S2 and S4 to obtain a checksum. This checksum is compared with the checksum from S6. If they are the same, the data is correct. If they are different, the data is incorrect and an error needs to be reported. Step S6 is then repeated to verify the data. If there is a problem, the data is discarded and an alarm is issued to indicate the data is incorrect. If they are the same, a copy is saved to disk. Step S8: Use an algorithm to calculate the check data of S3 and S4 to obtain a check number, and copy it to the disk; In step S9, the same algorithm as in S8 is used to calculate the verification data of S2 and S5 to obtain a verification number. This verification number is compared with the verification number obtained in S6. If they are the same, the data is correct. If they are different, the data is incorrect and the error needs to be reported. Step S6 is then re-executed to verify whether there is a problem with the data. If there is a problem, the data is discarded and an alarm is issued to indicate that the data is incorrect. If they are the same, a copy is saved to disk. After executing S2-S9, other machines also execute steps S2-S9 accordingly. Finally, the obtained verification values are compared. If they are the same, the data synchronization is completed correctly. If not, the data is incorrect and the verification error needs to be reported separately. If necessary, manual intervention is required to solve the problem or the data is discarded directly.
2. As claimed in claim 1, it is characterized in that Each verification can basically be calculated using a different algorithm, so that each data verification can be different. After the data is encrypted and processed into blocks of data, each data block can be protected more securely and reliably. Even if part of the data is leaked, the user's privacy will not be completely leaked, so local verification, encryption and other operations can be performed to ensure the security of user data.
3. As stated in claim 1, in distributed data backup, due to various reasons such as unreliable hardware, unreliable signal transmission, and unreliable storage, each data synchronization may cause data loss. However, through this step-by-step synchronization, step-by-step calculation, and then comparison through checksum values, the calculation can be effectively dispersed, the impact of factors such as machine failure and signal instability can be reduced, and the reliability of data transmission can be improved.
4. As claimed in claim 1, it is characterized in that After data verification is completed, the data can be localized. If data is lost in different data centers, the verification value is used to determine the lost part of the data, and then the corresponding data blocks are synchronized. In this way, with multiple data backups, data can be better restored off-site to ensure data security.
5. As described in claim 1, it is characterized in that After data verification is completed, the data can be localized and accessed by different users in different data centers, thereby accelerating the process, reducing user waiting time, and improving user experience.
6. As described in claim 1, it is characterized in that When edge data is collected, the data can be directly transmitted to the data center, and then the data is verified to ensure data security.
7. As described in claim 1, it is characterized in that When everything is connected, the data of other connected devices can be verified through a large-capacity data center, so that data can be restored on demand anytime and anywhere.