Data processing method and device and electronic equipment
By partitioning data to be backed up data and encapsulating meta-information into blockchain transactions, the problem of data consistency verification in the prior art depends on the integrity of storage hash value, and the integrity and verifiability guarantee of backup data are achieved.
Patent Information
- Application Number
- CN202510112225.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-27
AI Technical Summary
Existing data consistency verification methods rely on the integrity of files or records that store hash values and are easily tampered with by attackers, resulting in the integrity and verifiability of backup data that cannot be effectively guaranteed.
By obtaining the preprocessed data to be backed up, the data to be backed up is divided according to the data logical correlation and the preset data block size, the meta information of each data block is extracted, and it is encapsulated into blockchain transactions and recorded in the blockchain ledger, and the immutability of blockchain is used to ensure the secure storage of data block meta information.
It effectively protects the meta information of the data block, ensures the integrity and verifiability of the backup data, and avoids the risk of hash value being tampered with.
Smart Images

Figure CN120045623A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a data processing method, apparatus, and electronic device. Background Art
[0002] In the era of big data, Tencent Distributed SQL (TDSQL) stores a large amount of critical business data. To ensure the security of critical business data, it is usually backed up to the Hadoop Distributed File System (HDFS). To ensure that the backup data is exactly the same as the original data, data consistency verification is required.
[0003] The traditional data consistency verification method is to calculate a fixed hash value for the entire backup file after completing the TDSQL database backup and store it in a separate file or database record. During data recovery, the hash value of the backup file is recalculated and compared with the stored hash value to determine the consistency of the backup data.
[0004] However, the consistency verification process of this method mainly depends on the integrity of the file or record storing the hash value. Since the hash value can be easily obtained and modified by an attacker, if the file or record storing the hash value is tampered with and the tampered hash value matches the tampered backup data, the entire verification mechanism will fail, making the integrity and verifiability of the backup data unable to be effectively guaranteed. Summary of the Invention
[0005] In view of this, this application provides a data processing method, apparatus, and electronic device, mainly aiming to solve the technical problem that the current consistency verification process mainly depends on the integrity of the file or record storing the hash value. Since the hash value can be easily obtained and modified by an attacker, if the file or record storing the hash value is tampered with and the tampered hash value matches the tampered backup data, the entire verification mechanism will fail, making the integrity and consistency of the backup data unable to be effectively guaranteed.
[0006] According to the first aspect of this application, a data processing method is provided, including:
[0007] Obtain the preprocessed data to be backed up;
[0008] Divide the data to be backed up into data blocks according to the data logical relevance and the preset data block size;
[0009] Extract the meta-information of each data block, encapsulate the meta-information of each data block into a blockchain transaction according to a custom format, and record the blockchain transaction in a blockchain ledger.
[0010] According to a second aspect of the present application, there is provided a data processing apparatus, including:
[0011] An acquisition module, configured to acquire preprocessed data to be backed up;
[0012] A partitioning module, configured to partition the data to be backed up into data blocks according to data logical relevance and a preset data block size;
[0013] A recording module, configured to extract the meta-information of each data block, encapsulate the meta-information of each data block into a blockchain transaction according to a custom format, and record the blockchain transaction in a blockchain ledger.
[0014] According to a third aspect of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the data processing method described in the first aspect is implemented.
[0015] According to a fourth aspect of the present application, there is provided an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, and when the processor executes the computer program, the data processing method described in the first aspect is implemented.
[0016] By means of the above technical solutions, a data processing method, apparatus, and electronic device provided by the present application can acquire preprocessed data to be backed up; partition the data to be backed up into data blocks according to data logical relevance and a preset data block size; extract the meta-information of each data block, encapsulate the meta-information of each data block into a blockchain transaction according to a custom format, and record the blockchain transaction in a blockchain ledger. Compared with the current existing technologies, by encapsulating the meta-information of each data block into a blockchain transaction according to a custom format and recording the blockchain transaction in a blockchain ledger, the present application can ensure the secure storage of this information by using the immutability of the blockchain, thereby effectively protecting the meta-information of the data block and ensuring the integrity and verifiability of the backup data.
[0017] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the description. And in order to make the above and other purposes, features, and advantages of the present application more obvious and understandable, the specific embodiments of the present application are hereinafter specifically exemplified. Description of the Drawings
[0018] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0019] Figure 1 Flow diagram of a data processing method provided by an embodiment of the present disclosure;
[0020] Figure 2 Flow diagram of a data backup provided by an embodiment of the present disclosure;
[0021] Figure 3 Flow diagram of a data processing method provided by an embodiment of the present disclosure;
[0022] Figure 4 Flow diagram of a data consistency check provided by an embodiment of the present disclosure;
[0023] Figure 5 Structural diagram of a data processing device provided by an embodiment of the present disclosure. Detailed implementation manners
[0024] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details included in the embodiments of the present disclosure are helpful for understanding and should be regarded as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.
[0025] The following describes a data processing method, device, and electronic device according to embodiments of the present disclosure with reference to the accompanying drawings.
[0026] In order to solve the technical problem that the current consistency verification process mainly relies on the integrity of files or records storing hash values. Since hash values are easily obtained and modified by attackers, if the files or records storing hash values are tampered with and the tampered hash values are made to match the tampered backup data, then the entire verification mechanism will fail, making it impossible to effectively guarantee the integrity and consistency of the backup data. This embodiment provides a data processing method, as Figure 1 shown, the method includes:
[0027] Step 101, obtain the data to be backed up after preprocessing.
[0028] For the embodiments of the present disclosure, as Figure 2 shown, the data preprocessing program can be started first, and the data preprocessing program is used to perform a comprehensive format check and optimization on the data to be backed up to ensure that the data meets the requirements of backup and storage, improving the data quality and backup efficiency. Among them, the data to be backed up can be a data set to be backed up, and can include database content, system files, application program data, metadata, etc.
[0029] The preprocessing at least includes cleaning invalid, obsolete or redundant data in the data to be backed up, compressing or deduplicating the data to be backed up, performing data integrity checks on the data to be backed up, evaluating and optimizing and / or reconstructing the target database index of the data to be backed up, and optimizing the storage layout of the data to be backed up.
[0030] Among them, cleaning invalid, obsolete or redundant data in the data to be backed up can be deleting or marking invalid, obsolete or redundant data to be backed up to reduce the total amount of backup data and improve the backup efficiency. At the same time, the data to be backed up can be compressed or deduplicated to save storage space and backup time. For some data columns with a high repetition rate, specific algorithms can be used for deduplication, only storing unique data values, and recording the number of repetitions or associated information so that the original state of the data can be restored during data recovery.
[0031] To ensure that the data to be backed up is not lost or damaged, data integrity checks can be performed on the data to be backed up. The data integrity checks can include checking the integrity of data pages and records, and using a checksum verification mechanism to detect whether the data has been damaged during storage.
[0032] To ensure that the index structure of the backup data is reasonable, the target database index can be evaluated and optimized. In some cases, to improve data access efficiency and backup performance, the target database index of the data to be backed up can be reconstructed. For example, if the index becomes fragmented or the performance degrades, the target database index can be reconstructed during the backup process so that data queries and operations can be more efficiently supported after data recovery. For the indexes of large data tables, they can be optimized and adjusted according to the data distribution and access patterns to ensure that the indexes can effectively accelerate data retrieval. Among them, the target database can be a TDSQL distributed database.
[0033] To improve the read and write performance of data during backup and recovery, the storage layout of the data can be optimized according to the characteristics of the target storage system, and related data blocks can be stored continuously. Among them, the target storage system can be an HDFS storage system.
[0034] The present disclosure reduces the amount of backup data through data preprocessing, optimizes the data storage layout and index structure, and improves the data read and write performance.
[0035] Step 102: Divide the data to be backed up into data blocks according to the data logical relevance and the preset data block size.
[0036] For the embodiments of the present disclosure, such as Figure 2As shown, the preset data block size can be determined based on the basic processing capabilities of the target storage system and the blockchain node, and according to the preset data block size, the data to be backed up with a logical relevance greater than the preset logical relevance can be divided into the same data block.
[0037] Specifically, for the embodiments of the present disclosure, after the preprocessing is completed, in order to maintain the business logic relationship of the data during the backup, transmission, processing, and recovery processes, improve efficiency and availability, and enable the data block to better adapt to the storage and verification requirements of blockchain technology, the data to be backed up can be divided into data blocks according to the data logic relevance principle.
[0038] When dividing data blocks, the internal logical structure of the data in the TDSQL distributed database can be deeply analyzed to identify data subsets that are logically closely related. For example, in a relational database, the data rows belonging to the same transaction, having foreign key associations, or being frequently co-queryed can be divided into the same data block to maintain the semantic integrity and business coherence of the data during subsequent backup, verification, and recovery processes.
[0039] When performing data recovery, the data relationship can be quickly reconstructed according to the data blocks divided based on the data logic relevance principle, reducing the problem of scattered associated data caused by unreasonable data block division, thereby improving the recovery efficiency and data availability.
[0040] During the blockchain storage and operation processes, maintaining the integrity of the data is crucial for performing consistency checks and rapid recovery. For example, in a financial transaction database, the account information, transaction details, and other related data of the same transaction can be grouped together to ensure that its integrity is maintained during the blockchain storage and operation processes, facilitating consistency checks and rapid recovery.
[0041] For the embodiments of the present disclosure, considering that the basic storage unit of the HDFS storage system is a block, and its block storage characteristic is usually 128MB, which is a relatively reasonable basic block size set based on its distributed storage architecture and data read / write mechanism. When dividing data blocks, the block storage characteristics of the HDFS storage system can be referred to. If the data block is divided too large, it may exceed the processing limit of a single block of the HDFS storage system, resulting in complex splitting and merging operations during data storage and reading, increasing system overhead and processing time.
[0042] At the same time, the processing power of blockchain nodes also has an important impact on the data block size. When blockchain nodes perform data verification, consensus mechanism operations, and data storage, their computing resources and storage resources are limited. If the data block is too small, it will lead to a sharp increase in the number of blockchain transactions. Each data block needs to be independently verified and recorded in the blockchain network, which will greatly increase the computing burden on blockchain nodes, cause network congestion, and reduce the processing speed of the entire system.
[0043] Therefore, in the initial stage of the backup process, a preset data block size can be set based on the basic processing capabilities of the HDFS storage system and blockchain nodes. Among them, the preset data block size can be a relatively moderate starting value set based on the basic processing capabilities of the HDFS storage system and blockchain nodes, such as 32MB.
[0044] Step 103: Extract the meta-information of each data block, encapsulate the meta-information of each data block into a blockchain transaction in a custom format, and record the blockchain transaction in the blockchain ledger.
[0045] The meta-information includes at least the hash value, timestamp, and identification number of each data block.
[0046] Among them, for each data block, the hash value can be a set of multiple hash values calculated using multiple hash algorithms. For example, for a 128MB data to be backed up, it can be divided into two 64MB data blocks, and the hash values of the SHA-256 and SHA-3 encryption hash algorithms can be calculated respectively to obtain two hash values. This data block hash calculation method combining multiple hash algorithms can significantly enhance the reliability of data integrity verification compared with the traditional single hash algorithm.
[0047] At the same time, a unique identification number can be assigned to each data block, and the numbers increase sequentially starting from 1. The identification number can be used to ensure the sequential consistency of data blocks in subsequent data processing and recovery processes.
[0048] At the same time, the system can call a high-precision clock module to obtain the accurate timestamp when the data backup operation starts, and the accuracy of the timestamp reaches the millisecond level. Among them, the format of the timestamp can follow the international standard time format (such as ISO 8601) and include time zone information to ensure consistency and accuracy globally, so as to accurately track the time sequence of data backup subsequently.
[0049] After obtaining the timestamp, the encryption verification technology can be immediately used to verify the validity of the timestamp. The encryption verification can include checking whether the encryption signature of the timestamp is correct, whether the timestamp is within a reasonable time range, and whether it conforms to the time series logic of the backup operation, so as to prevent the timestamp from being illegally tampered with or incorrect time records caused by system failures.
[0050] For the embodiments of the present disclosure, as Figure 2 shown, detailed meta-information such as the hash value, identification number, and timestamp of each data block can be combined, and the meta-information of each data block can be encapsulated in a custom structured data format. For example, a binary-based compact format can be designed, which sets fields of fixed length at the head of the data to identify the types and lengths of various meta-information. Then, in a specified order, the encrypted hash value, the signed sequence number, the data block size with a checksum, and the time information including a time authentication token and a time interval can be stored in the data block information unit in sequence. Among them, encryption and signature can ensure the security and integrity of information during transmission and storage.
[0051] As Figure 2 shown, the constructed data block information unit can be encapsulated into a standard blockchain transaction, and through the blockchain network, the encapsulated blockchain transaction can be recorded in the blockchain ledger. The distributed storage feature of the blockchain can be utilized to ensure that the data block information is stored on multiple nodes globally, increasing the availability and fault tolerance of the data. The encryption feature and the immutability feature of the blockchain can be utilized to effectively prevent the hash value from being tampered with, and ensure the security of data consistency verification. Through the combination of multiple hash algorithms and encryption verification technologies, the integrity and reliability of the data are further enhanced, and the security of the backup data during transmission and storage is ensured.
[0052] In summary, according to a data processing method of the present disclosure, preprocessed data to be backed up can be obtained; the data to be backed up can be divided into data blocks according to the data logical relevance and a preset data block size; the meta-information of each data block can be extracted, the meta-information of each data block can be encapsulated into a blockchain transaction according to a custom format, and the blockchain transaction can be recorded in the blockchain ledger. Compared with the current existing technologies, in this application, by encapsulating the meta-information of each data block into a blockchain transaction according to a custom format and recording the blockchain transaction in the blockchain ledger, the technical features of the blockchain can be utilized to ensure the secure storage and immutability of this information, thereby effectively protecting the meta-information of the data block and ensuring the integrity and verifiability of the backup data.
[0053] Further, as a refinement and extension of the above embodiments, in order to fully illustrate the specific implementation process of the method of the present disclosure, the present disclosure provides a specific method as Figure 3 shown, and this method includes:
[0054] Step 201, monitor the first key performance indicator of the target storage system in real time, and monitor the second key performance indicator of the blockchain node in real time, and dynamically adjust the preset data block size according to the monitored first key performance indicator and second key performance indicator.
[0055] For the embodiments of the present disclosure, as the backup process progresses, through a dynamic monitoring and feedback mechanism, the data block size can be dynamically adjusted according to the system operating state, ensuring the high efficiency and stability of the backup and verification processes and achieving all-round real-time monitoring of the backup process.
[0056] Specifically, the preset data block size can be dynamically adjusted according to the first key performance indicators of the target storage system monitored in real time and the second key performance indicators of the blockchain nodes monitored in real time. Among them, the first key performance indicators can include storage load, data read and write speed, etc., and the second key performance indicators can include CPU usage rate, memory occupancy rate, network bandwidth, etc.
[0057] Exemplarily, if it is monitored that the HDFS storage system has more storage free space, faster data read and write speed, and the computing resources and network bandwidth of the blockchain nodes are relatively abundant, the size of the data block can be appropriately increased (for example, increased to 48MB or 64MB) to reduce the number of data blocks and improve the overall processing efficiency.
[0058] On the contrary, if it is monitored that the storage pressure of the HDFS storage system increases and the read and write speed slows down, or if it is monitored that the resources of the blockchain nodes are tense, the data block size can be reduced (for example, reduced to 24MB or 16MB) to ensure that the data can be successfully stored and processed on the HDFS storage system and the blockchain. Through this dynamic adjustment, it is possible to always maintain an optimal balance point of storage and processing efficiency of the data block between the HDFS storage system and the blockchain nodes, while maintaining the logical integrity and coherence of the data, and ensuring the stability and efficiency of the entire backup and verification process.
[0059] Step 202: Continuously monitor the process of writing data blocks to the target storage system; when the data block is written to the target storage system, read the data block from the target storage system and recalculate the hash value of the data block; compare the recalculated hash value with the hash value of the data block stored in the blockchain ledger; if the two hash values are the same, it is determined that the data block consistency verification passes; if the two hash values are different, trigger the in-depth verification program.
[0060] For the embodiments of the present disclosure, as Figure 4 shown, during the backup process, the data blocks can be verified in real time to promptly discover data inconsistency problems and reduce the risk of data loss. Specifically, the transmission and storage conditions of the data blocks can be continuously monitored. When the data block is written to the HDFS storage system, the data block can be immediately read from the HDFS storage system and its hash value can be recalculated, and the recalculated hash value can be compared with the original hash value recorded in the blockchain ledger. If the two hash values are the same, the data block can be marked as passing the consistency verification; if the two hash values are different, the in-depth verification program can be triggered.
[0061] Among them, the in-depth verification program can utilize the chain structure and timestamp information of the blockchain to trace back all relevant blockchain transaction records of the data block during the entire backup process, including node information during data extraction and transmission, storage operations, etc., to determine the specific link and reason for data inconsistency.
[0062] For example, if it is found that the hash value of the data block is inconsistent, by querying the blockchain transaction records, it can be found that an abnormal network fluctuation occurred in a certain network node through which the data block passed during transmission, which may cause partial loss or damage of the data, thus accurately locating the root cause of the problem. During the backtracking process, the smart contract of the blockchain can be used to automatically execute the verification logic to improve the verification efficiency and accuracy. For example, the smart contract can define verification rules and processes, and automatically find the problem nodes and relevant information according to the timestamp and transaction records.
[0063] Correspondingly, triggering the in-depth verification program specifically may include:
[0064] Utilize the chain structure and timestamp information of the blockchain to trace back all relevant blockchain transaction records of the data block during the backup process; based on preset verification rules and processes, determine the specific link and reason for data inconsistency according to the timestamp information and relevant blockchain transaction records.
[0065] For the embodiments of the present disclosure, as Figure 2 shown, the automatic execution of verification rules and exception handling can be realized based on the blockchain smart contract. Among them, the smart contract can pre-define rules such as the standards, frequencies, and exception handling processes for data consistency verification. For example, when a certain proportion (such as 10%) of the data blocks have consistency problems, the smart contract can automatically start the data repair program.
[0066] The repair program can re-obtain the correct data block from the TDSQL distributed database or backup redundant copy according to the original data information recorded on the blockchain, and rewrite the correct data block into the HDFS storage system to replace the damaged or inconsistent data block, and update the relevant records in the blockchain ledger to ensure the consistency and integrity of the data.
[0067] Among them, the distributed architecture of the blockchain avoids the problem that the verification mechanism fails due to the failure of a single storage point. Even if some blockchain nodes or storage devices fail, the system can still verify and recover through other nodes and backup data, ensuring the continuity and reliability of data consistency verification.
[0068] For example, if a data block is found to be damaged, the smart contract can extract the corresponding data block from the TDSQL distributed database according to the metadata of the data block stored on the blockchain, rewrite it into the HDFS storage system to replace the damaged data block, and update the hash value and relevant records on the blockchain. In addition, the smart contract can also notify the administrator or relevant monitoring system of the verification results and repair status in a timely manner to achieve transparent management and real-time monitoring of the data backup process. For example, by sending emails or push notifications, the verification results and repair reports are sent to the administrator so that the administrator can understand the backup situation in a timely manner and take further measures.
[0069] For the embodiments of the present disclosure, by adopting multiple hash algorithms and in-depth verification programs, combined with the traceability function of the blockchain, the root cause of data problems can be determined more accurately, and the accuracy and reliability of verification can be improved. Through the automated verification and repair mechanism driven by the smart contract, manual intervention is reduced, and the processing efficiency and the success rate of data recovery are improved.
[0070] In summary, according to a data processing method of the present disclosure, the to-be-backup data after preprocessing can be obtained; the to-be-backup data can be divided into data blocks according to the data logical relevance and the preset data block size; the metadata of each data block is extracted, and the metadata of each data block is encapsulated into a blockchain transaction in a custom format, and the blockchain transaction is recorded in the blockchain ledger. Compared with the current existing technologies, in this application, by encapsulating the metadata of each data block into a blockchain transaction in a custom format and recording the blockchain transaction in the blockchain ledger, the technical characteristics of the blockchain can be used to ensure the secure storage and immutability of this information, thereby effectively protecting the metadata of the data block and ensuring the integrity and verifiability of the backup data.
[0071] Based on the above Figure 1 and Figure 3 specific implementation of the method shown, this embodiment provides a data processing device, as Figure 5 shown, the device includes: an acquisition module 31, a division module 32, and a recording module 33;
[0072] The acquisition module 31 is used to acquire the to-be-backup data after preprocessing;
[0073] The division module 32 is used to divide the to-be-backup data into data blocks according to the data logical relevance and the preset data block size;
[0074] The recording module 33 is used to extract the metadata of each data block, encapsulate the metadata of each data block into a blockchain transaction in a custom format, and record the blockchain transaction in the blockchain ledger.
[0075] In a specific application scenario, the partitioning module 32 can be used to determine the preset data block size based on the basic processing capabilities of the target storage system and the blockchain node;
[0076] According to the preset data block size, the data to be backed up with a logical relevance greater than the preset logical relevance is partitioned into the same data block.
[0077] In a specific application scenario, the device further includes: a first monitoring module 34 and an adjustment module 35;
[0078] The first monitoring module 34 is used to monitor the first key performance indicators of the target storage system in real time and monitor the second key performance indicators of the blockchain node in real time. The first key performance indicators at least include storage load and data read / write speed, and the second key performance indicators at least include CPU usage rate, memory occupancy rate, and network bandwidth;
[0079] The adjustment module 35 is used to dynamically adjust the preset data block size according to the monitored first key performance indicators and the second key performance indicators.
[0080] In a specific application scenario, the meta-information at least includes the hash value, timestamp, and identification number of each data block; the hash value is the multiple hash values of each data block calculated using multiple hash algorithms; the timestamp is the start time of the data backup operation that has passed validity verification; the identification number is the unique identification number assigned to each data block.
[0081] In a specific application scenario, the device further includes: a second monitoring module 36, a reading module 37, a comparison module 38, a determination module 39, and a triggering module 40;
[0082] The second monitoring module 36 is used to continuously monitor the process of writing the data block to the target storage system;
[0083] The reading module 37 is used to, after the data block is written to the target storage system, read the data block from the target storage system and recalculate the hash value of the data block;
[0084] The comparison module 38 is used to compare the recalculated hash value with the hash value of the data block stored in the blockchain ledger;
[0085] The determination module 39 is used to determine that the data block consistency verification passes if the two hash values are the same;
[0086] The triggering module 40 is used to trigger a deep verification program if the two hash values are different.
[0087] In a specific application scenario, the triggering module 40 can be used to trace all relevant blockchain transaction records of the data block during the backup process by utilizing the chain structure and timestamp information of the blockchain;
[0088] Based on preset verification rules and processes, determine the specific links and reasons for data inconsistency according to the timestamp information and the relevant blockchain transaction records.
[0089] In a specific application scenario, the device further includes: a processing module 41;
[0090] The processing module 41 is used to preprocess the data to be backed up. The preprocessing at least includes cleaning invalid, outdated or redundant data in the data to be backed up, compressing or deduplicating the data to be backed up, performing a data integrity check on the data to be backed up, evaluating and optimizing and / or reconstructing the target database index of the data to be backed up, and optimizing the storage layout of the data to be backed up.
[0091] It should be noted that for other corresponding descriptions of each functional unit involved in the data processing device provided in this embodiment, reference can be made to Figure 1 and Figure 3 the corresponding descriptions in the method, which will not be elaborated here.
[0092] Based on the above as Figure 1 and Figure 3 shown in the method, correspondingly, the present disclosure also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method as shown in the above as Figure 1 and Figure 3 shown is implemented.
[0093] Based on such an understanding, the technical solution of the present disclosure can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods in various implementation scenarios of the present disclosure.
[0094] Based on the above as Figure 1 and Figure 3 shown in the method, as well as Figure 5 shown in the virtual device embodiment, in order to achieve the above object, the embodiments of the present disclosure also provide an electronic device. The device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the method as shown in the above as Figure 1 and Figure 3 shown.
[0095] Optionally, the above-mentioned physical device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, and so on. The user interface may include a display, an input unit such as a keyboard, etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), etc.
[0096] Those skilled in the art can understand that the above-mentioned physical device structure provided by the present disclosure does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0097] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing the hardware and software resources of the above-mentioned physical device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement communication between components inside the storage medium, as well as communication between other hardware and software in the information processing physical device.
[0098] Through the description of the above embodiments, those skilled in the art can clearly understand that the present disclosure can be implemented by means of software plus a necessary general hardware platform, or can also be implemented by hardware. The data processing method, device, and electronic device provided by the present disclosure, compared with the prior art, the present disclosure obtains the data to be backed up after preprocessing; divides the data to be backed up into data blocks according to the data logical relevance and the preset data block size; extracts the meta-information of each data block, encapsulates the meta-information of each data block into a blockchain transaction in a custom format, and records the blockchain transaction in the blockchain ledger. Compared with the current prior art, the present application encapsulates the meta-information of each data block into a blockchain transaction in a custom format, and records the blockchain transaction in the blockchain ledger. By using the technical characteristics of the blockchain, the security storage and immutability of this information can be ensured, thereby effectively protecting the meta-information of the data block and ensuring the integrity and verifiability of the backup data.
[0099] It should be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0100] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.
Claims
1. A data processing method, characterized in that: include: Obtain the pre-processed data to be backed up; Dividing the data to be backed up into data blocks according to data logical relevance and preset data block sizes; Extract the meta information of each data block after division, encapsulate the meta information of each data block into a blockchain transaction according to a custom format, and record the blockchain transaction in the blockchain account book.
2. The method according to claim 1, characterized in that The step of dividing the data to be backed up into data blocks according to the data logical association and the preset data block size includes: Determine the preset data block size based on the basic processing capabilities of the target storage system and the blockchain node; According to the preset data block size, the to-be-backed-up data with a logical correlation greater than the preset logical correlation are divided into the same data block.
3. The method according to claim 2, characterized in that The method further comprises: Monitor the first key performance indicator of the target storage system in real time, and monitor the second key performance indicator of the blockchain node in real time, wherein the first key performance indicator includes at least storage load and data read and write speed, and the second key performance indicator includes at least CPU usage, memory occupancy, and network bandwidth; The preset data block size is dynamically adjusted according to the monitored first key performance indicator and the second key performance indicator.
4. The method according to claim 1, characterized in that: The meta-information includes at least a hash value, a timestamp, and an identification number for each data block; the hash value is a plurality of hash values for each data block calculated using a plurality of hash algorithms; the timestamp is a time for characterizing the start of a data backup operation that has been verified for validity; and the identification number is a unique identification number assigned to each data block.
5. The method according to claim 4, characterized in that The method further comprises: Continuously monitoring the process of writing data blocks to the target storage system; After the data block is written to the target storage system, the data block is read from the target storage system and the hash value of the data block is recalculated; Compare the recalculated hash value with the hash value of the data block stored in the blockchain ledger; If the two hash values are consistent, it is determined that the data block consistency verification has passed; If the two hash values do not match, the deep verification procedure is triggered.
6. The method according to claim 5, characterized in that The triggering deep verification procedure includes: Using the chain structure and timestamp information of the blockchain, all relevant blockchain transaction records of the data block during the backup process are traced back; Based on the preset verification rules and processes, the specific links and causes of the data inconsistency are determined according to the timestamp information and the relevant blockchain transaction records.
7. The method according to claim 1, characterized in that The method further comprises: The data to be backed up is preprocessed, and the preprocessing at least includes cleaning up invalid, obsolete or redundant data in the data to be backed up, compressing or deduplicating the data to be backed up, performing data integrity check on the data to be backed up, evaluating, optimizing and / or rebuilding the target database index of the data to be backed up, and optimizing the storage layout of the data to be backed up.
8. A data processing device, characterized in that: include: An acquisition module is used to acquire the pre-processed data to be backed up; A partitioning module, used for partitioning the data to be backed up into data blocks according to data logical relevance and preset data block sizes; The recording module is used to extract the meta information of each data block, encapsulate the meta information of each data block into a blockchain transaction according to a custom format, and record the blockchain transaction in a blockchain account book.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. An electronic device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Code auditing system based on block chain
CN121211463A