Big data block chain evidence storage method

By using technical means such as multi-dimensional hashing calculation, timestamp binding and digital signature in big data proof storage, data fingerprints are generated and blockchain proof storage is solved, and data fingerprints are not unique enough and data processing in the existing technology is ineffective, achieving efficient and secure data proof storage.

CN120067205APending Publication Date: 2025-05-30SINKO (SHAANXI) INFORMATION TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411698666.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art may use a single hashing algorithm or a hashing process that is not complex enough in big data proofing, resulting in the data fingerprint not having enough uniqueness, reducing the verifiability of the data, and inefficient data processing process.

Method used

Multi-dimensional hashing calculation, timestamp binding, digital signature application and other steps are used to generate data fingerprints, and verified through blockchain technology to ensure the uniqueness and verifiability of the data.

Benefits of technology

It improves the efficiency and security of data processing, ensures the uniqueness and immutability of data fingerprints, and ensures the integrity and reliability of proof-stored data.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention discloses a big data block chain evidence storage method. According to the method, the uniqueness and verifiability of the data are ensured through the generation of the data fingerprints. By adopting Hash algorithms such as SHA-256 and the like, each data fragment is converted into a Hash value with a fixed length, the Hash value is unique, and it is almost impossible that two different data fragments generate the same Hash value. The uniqueness enables the data fingerprint to become a reliable identifier of block chain evidence storage. In the data storage process, the data fingerprints serve as digital fingerprints of the data, and the integrity and authenticity of the data can be quickly verified. Whether the data is tampered or damaged or not can be confirmed by comparing the data fingerprints in the data uploading, storage or subsequent verification process, so that the integrity and reliability of the evidence storage data are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of blockchain evidence storage, and specifically relates to a method for storing big data on the blockchain. Background Art

[0002] With the rapid development of information technology, big data has become an important part of various fields of society. However, big data faces various challenges in the processes of storage, transmission, and processing, such as data security, data integrity, and data credibility. Traditional data storage and verification methods can no longer meet the requirements of the big data era. Therefore, blockchain technology, as a decentralized and immutable distributed ledger technology, has gradually become a key technology for solving the problem of big data evidence storage.

[0003] However, existing technologies may only use a single hashing algorithm or an insufficiently complex hashing process, which may result in the data fingerprints not having sufficient uniqueness, such that different data shards may generate the same hash value, thereby reducing the verifiability of the data. At the same time, the data processing process may not be efficient enough, especially in the generation of data fingerprints, lacking steps such as multi-dimensional hashing calculation and timestamp binding, resulting in low efficiency in the processes of data uploading, storage, and verification. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for storing big data on the blockchain in order to solve the above-mentioned problems.

[0005] The technical solution adopted by the present invention is as follows: A method for storing big data on the blockchain, characterized in that it includes the following steps:

[0006] S1: Perform data preprocessing. First, clean, deduplicate, and format the original big data to ensure data quality; then encrypt the data to ensure data security;

[0007] S2: Perform data sharding. Divide the preprocessed data into several data shards to facilitate storage and verification on the blockchain; ensure that the size of the data shards is suitable for the transaction limits of the blockchain network;

[0008] S3: Generate data fingerprints. Generate unique data fingerprints for each data shard; the data fingerprints will be used as the unique identifier for blockchain evidence storage.

[0009] S4: Construct an evidence storage transaction. Package the data fingerprints and information related to the data shards into an evidence storage transaction; ensure that the format of the evidence storage transaction meets the requirements of the blockchain network during this process.

[0010] S5: Conduct an upload and deposit transaction, send the deposit transaction to the blockchain network, and wait for network consensus confirmation; then ensure that the deposit transaction is successfully recorded on the chain and obtain the corresponding transaction hash;

[0011] S6: Conduct deposit verification. Query the deposit transaction hash through a blockchain browser to verify whether the data is successfully recorded on the chain; then compare the data fingerprints on the chain with the local data fingerprints to ensure data consistency;

[0012] S7: Conduct deposit information registration. Create a deposit information record for each data shard on the blockchain, including data fingerprints, transaction hashes, and deposit times; the deposit information records can be used for subsequent data query, verification, and auditing;

[0013] S8: Construct a data index. Based on the keywords and tag information of the data shards, construct a data index; the data index facilitates the rapid retrieval and positioning of the deposit data on the blockchain;

[0014] S9: Back up the deposit data. Back up the original data and the blockchain deposit information to a secure storage medium or cloud service to ensure the security and recoverability of the data backup;

[0015] S10: Conduct regular auditing and maintenance. Regularly audit the blockchain deposit data to check data integrity and security; maintain the blockchain network to ensure its stable operation, and update and optimize the deposit method in a timely manner.

[0016] In a preferred embodiment, in step S1, the Kafka real-time data processing framework is used for the preprocessing of the data stream; at the same time, to ensure data security, the AES encryption algorithm is used to encrypt the data to ensure the security of the data during transmission and storage.

[0017] In a preferred embodiment, in step S2, the consistent hashing algorithm is used to implement data sharding. This algorithm can shard data according to the content of the data rather than its location, ensuring the uniformity of data distribution; first, define a hash ring, then map the data to the hash ring through a hash function, and finally allocate the data to different shards according to the mapping results.

[0018] In a preferred embodiment, step S3 specifically includes the following steps:

[0019] S3-1. Enhanced processing of data shards: Perform secondary encryption on each data shard using different encryption algorithms to increase data security; introduce a random number generator during the encryption process to ensure that each encryption result is different and improve the complexity of the data fingerprint;

[0020] S3-2. Multi-dimensional Hash Calculation: Perform hash calculations on each data shard using multiple hash algorithms to generate multiple hash values; combine the multiple hash values to form a multi-dimensional hash fingerprint of the data shard;

[0021] S3-3. Timestamp Binding: Bind the current timestamp to the multi-dimensional hash fingerprint of the data shard to ensure the non-tamperability of the data fingerprint in terms of time; use the timestamp service provided by the blockchain network to further verify the authenticity of the timestamp;

[0022] S3-4. Digital Signature Application: Use the private key to perform digital signatures on the multi-dimensional hash fingerprint and the timestamp to ensure the integrity of the data fingerprint and the verifiability of the source; attach the digital signature to the data fingerprint to form the final data fingerprint package;

[0023] S3-5. Data Fingerprint Verification: Verify the generated data fingerprint package locally to ensure it has not been tampered with; the verification content includes the correctness of the hash value, the accuracy of the timestamp, and the validity of the digital signature;

[0024] S3-6. Data Fingerprint Storage Preparation: Store the successfully verified data fingerprint package in a secure temporary storage area pending upload to the blockchain network; ensure access control of the temporary storage area to prevent unauthorized access and data leakage;

[0025] Among them, the SHA-256 multi-dimensional hash calculation formula specifically includes the following:

[0026] Initial hash value: H 0 = h 0 H 1 = h 1 ... H 7 = h 7 ;

[0027] Pad data block: P = M||1||0 k ||L;

[0028] Process data block: W[t] = M[t for t = 0 to 15

[0029] W[t] = σ 1 (W[t-2]) + W[t-7] + σ 0 (W[t-15]) + W[t-16] for t = 16 to 63;

[0030] Compression function: a = H 0 J, b = H1,... h = H 7 for t = 0 to 63

[0031] T 1 = h + ∈ 1(a) + CH(a, b, c) + K t + W[t]T 2 ∈ 0 (e) + MAJ(e, f, g) h = g g = f f = e e = d + T 1

[0032] d = c c = b b = a a = T 1 + T 2 end for H 0 = H 0 + a, H 1 = H 1 + b,..., H 7 = H 7 + h

[0033] Final hash value: H = H 0 J || H 1 ||...|| H 7 ;

[0034] Where:

[0035] H 1 = h 1 ,..., H 7 refers to the initial hash value, which is a fixed value preset by the SHA-256 algorithm

[0036] M refers to the original data shard;

[0037] 1 refers to an additional 1-bit;

[0038] O k refers to the additional 0-bits for padding;

[0039] L refers to the 64-bit length field indicating the length of the original data;

[0040] W[t] refers to the 32-bit word generated by message expansion;

[0041] σ 1 , σ 0 refers to the logical function for message expansion;

[0042] CH, MAJ refer to the logical functions for the compression function;

[0043] K t refers to 64 constants for the compression function;

[0044] T 1 , T 2 refers to the intermediate variable for the calculation of the compression function;

[0045] a, b, c, d, e, f, g, h denote working variables for the calculation of the compression function.

[0046] In a preferred embodiment, in step S4, it is necessary to organize relevant information into a transaction structure in accordance with the regulations of the blockchain network; the transaction structure includes input, output, and witness data parts to ensure that the deposit transaction can be correctly processed by the blockchain network.

[0047] In a preferred embodiment, in step S5, through the node client of the blockchain network, the transaction is broadcast to the network; the system in the network will verify the transaction and confirm the validity of the transaction through a consensus algorithm; once the transaction is confirmed, a transaction hash will be obtained as a certificate of successful deposit.

[0048] In a preferred embodiment, in step S6, first use a blockchain browser to query the transaction hash to confirm whether the transaction is accepted by the network; then, compare the data fingerprint on the chain with the local data fingerprint to ensure that the data has not been tampered with during storage, thereby verifying the data consistency;

[0049] Step S7 is implemented through a smart contract. The smart contract will provide an interface to allow users to create records on the blockchain; first call the function of the smart contract, passing in the data fingerprint and transaction hash parameters, and then the contract will create a new record entry on the blockchain. This record can be queried by anyone but cannot be tampered with.

[0050] In a preferred embodiment, in step S8, use the Elasticsearch search engine to build an index; the specific process is as follows: extract the keywords and tag information of the data shards, and then create an inverted index to associate the keywords with the deposit records on the blockchain.

[0051] In a preferred embodiment, in step S9, adopt a multi-copy backup strategy and combine cloud services for data storage; first encrypt the data, and then upload it to multiple data centers of the cloud service to ensure redundant backup of the data.

[0052] In a preferred embodiment, in step S10, first review the transaction records on the blockchain to verify the consistency of the data fingerprint, secondly check the version and configuration of the node software, and monitor the network performance; to regularly check the integrity and security of the blockchain deposit data; at the same time, perform necessary maintenance work according to the audit results.

[0053] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are:

[0054] 1. In the present invention, the generation of data fingerprints ensures the uniqueness and verifiability of data. By using hash algorithms such as SHA-256, each data shard is converted into a hash value of a fixed length, which is unique. It is almost impossible for two different data shards to produce the same hash value. This uniqueness makes the data fingerprint a reliable identifier for blockchain evidence storage. During the process of data evidence storage, the data fingerprint, as the "digital fingerprint" of the data, can quickly verify the integrity and authenticity of the data. Whether during data upload, storage, or subsequent verification processes, the integrity and reliability of the evidence-stored data can be ensured by comparing the data fingerprints to confirm whether the data has been tampered with or damaged.

[0055] 2. In the present invention, the generation of data fingerprints improves the efficiency and security of data processing. In the blockchain evidence storage system, the generation of data fingerprints is not just a simple hash calculation process. It also includes multiple steps such as multi-dimensional hash calculation, timestamp binding, and digital signature application. These steps together enhance the complexity and security of the data fingerprint. Multi-dimensional hash calculation increases the complexity of the data fingerprint by using multiple hash algorithms, making it difficult for attackers to reverse-engineer and crack the original data. The application of timestamp binding and digital signature adds a time dimension and identity verification function to the data fingerprint, ensuring the immutability of the data fingerprint during the generation process and the traceability of its source. These measures together improve the security of the entire evidence storage system and also provide an efficient way for subsequent data retrieval and verification. Because the data fingerprint can directly locate the evidence storage records on the blockchain without traversing the entire blockchain network, greatly improving the processing efficiency. Generally speaking, the steps of generating data fingerprints not only provide strong data security protection for big data blockchain evidence storage but also lay a foundation for the reliability and practicality of the entire evidence storage system by improving data processing efficiency. Detailed implementation manners

[0056] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following further elaborates on the present invention in combination with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0057] Embodiment:

[0058] A big data blockchain evidence storage method includes the following steps:

[0059] S1: Perform data preprocessing. First, perform operations such as cleaning, deduplication, and formatting on the original big data to ensure data quality. Then encrypt the data to ensure data security;

[0060] S2: Perform data sharding, divide the preprocessed data into several data shards to facilitate storage and verification on the blockchain. Ensure that the size of the data shards is suitable for the transaction limits of the blockchain network;

[0061] S3: Generate data fingerprints, generate unique data fingerprints (hash values) for each data shard. The data fingerprints will be used as the unique identifiers for blockchain evidence storage;

[0062] S4: Construct evidence storage transactions, package the data fingerprints, data shard-related information (such as data source, timestamp) into evidence storage transactions. Ensure that the format of the evidence storage transactions meets the requirements of the blockchain network during this process;

[0063] S5: Upload evidence storage transactions, send the evidence storage transactions to the blockchain network and wait for network consensus confirmation. Then ensure that the evidence storage transactions are successfully recorded on the chain and obtain the corresponding transaction hashes;

[0064] S6: Perform evidence storage verification, query the evidence storage transaction hashes through the blockchain browser to verify whether the data is successfully recorded on the chain. Then compare the data fingerprints on the chain with the local data fingerprints to ensure data consistency;

[0065] S7: Record evidence storage information, create an evidence storage information record for each data shard on the blockchain, including data fingerprints, transaction hashes, evidence storage time, etc. The evidence storage information records can be used for subsequent data query, verification and auditing;

[0066] S8: Construct data indexes, construct data indexes based on information such as keywords and tags of data shards. The data indexes facilitate the quick retrieval and positioning of evidence storage data on the blockchain;

[0067] S9: Back up evidence storage data, back up the original data and blockchain evidence storage information to a secure storage medium or cloud service, which can ensure the security and recoverability of data backup;

[0068] S10: Perform regular auditing and maintenance, regularly audit the blockchain evidence storage data to check data integrity and security. Maintain the blockchain network to ensure its stable operation, and update and optimize the evidence storage method in a timely manner.

[0069] In step S1, a real-time data processing framework such as Kafka is used for the preprocessing of data streams. At the same time, to ensure data security, the AES encryption algorithm is used to encrypt the data to ensure the security of data during transmission and storage.

[0070] In step S2, the consistent hashing algorithm is adopted to implement data sharding. This algorithm can perform sharding based on the content of the data rather than its location, ensuring the uniformity of data distribution. First, a hash ring is defined, then the data is mapped to the hash ring through a hash function, and finally, the data is allocated to different shards according to the mapping result.

[0071] Step S3 specifically includes the following steps:

[0072] S3-1. Data shard enhancement processing: Each data shard is encrypted twice using different encryption algorithms to increase data security. A random number generator is introduced during the encryption process to ensure that each encryption result is different, improving the complexity of the data fingerprint.

[0073] S3-2. Multi-dimensional hash calculation: Multiple hash algorithms (such as SHA-256, SHA-3, BLAKE2) are used to perform hash calculations on each data shard to generate multiple hash values. The multiple hash values are combined to form the multi-dimensional hash fingerprint of the data shard.

[0074] S3-3. Timestamp binding: The current timestamp is bound to the multi-dimensional hash fingerprint of the data shard to ensure the non-tamperability of the data fingerprint over time. The timestamp service provided by the blockchain network is used to further verify the authenticity of the timestamp.

[0075] S3-4. Digital signature application: The private key is used to digitally sign the multi-dimensional hash fingerprint and the timestamp to ensure the integrity of the data fingerprint and the verifiability of its source. The digital signature is attached to the data fingerprint to form the final data fingerprint package.

[0076] S3-5. Data fingerprint verification: The generated data fingerprint package is verified locally to ensure that it has not been tampered with. The verification content includes the correctness of the hash value, the accuracy of the timestamp, and the validity of the digital signature.

[0077] S3-6. Data fingerprint storage preparation: The data fingerprint package that passes the verification is stored in a secure temporary storage area, waiting to be uploaded to the blockchain network. Ensure access control for the temporary storage area to prevent unauthorized access and data leakage.

[0078] The specific SHA-256 multi-dimensional hash calculation formula is as follows:

[0079] Initialize the hash value (the initial hash value is fixed): H 0 = h 0 , H 1 = h 1 ,..., H 7 = h 7 ;

[0080] Padding data block: P = M||1||0 k ||L;

[0081] Processing data block (taking a 512-bit block as an example): W[t] = M[t] for t = 0 to 15, W[t] = σ 1 (W[t - 2]) + W[t - 7] + σ 0 (W[t - 15]) + W[t - 16] for t = 16 to 63;

[0082] Compression function (for each block): a = H 0 , b = H 1 ,..., h = H 7 for t = 0 to 63

[0083] T 1 = h + ∈ 1 (a) + CH(a, b, c) + K t + W[t][T 2 ∈ 0 (e) + MAJ(e, f, g), h = g, g = f, f = e, e = d + T 1

[0084] d = c, c = b, b = a, a = T 1 + T 2 end for, H 0 = H 0 + a, H 1 = H 1 + b,..., H 7 = H 7 + h

[0085] Final hash value: H = H 0 |||H 1 ||...||H 7 .

[0086] Where:

[0087] H 1 = h 1 ,..., H 7 refers to the initial hash value, which is a fixed value preset by the SHA-256 algorithm

[0088] M refers to the original data shard.

[0089] 1 refers to an additional 1-bit.

[0090] 0 k refers to the additional 0-bit for padding.

[0091] L refers to the 64-bit length field indicating the length of the original data.

[0092] W[t] represents the 32-bit word generated by message expansion.

[0093] σ 1 , σ 0 refers to the logical function used for message expansion.

[0094] CH and MAJ refer to the logical functions used for the compression function.

[0095] K t refers to 64 constants used for the compression function.

[0096] T 1 , T 2 refers to the intermediate variable used for the calculation of the compression function.

[0097] a, b, c, d, e, f, g, h refer to the working variables used for the calculation of the compression function.

[0098] In step S4, it is necessary to follow the regulations of the blockchain network, such as the scripting language, and organize the relevant information into a transaction structure. The transaction structure includes parts such as inputs, outputs, and witness data, ensuring that the deposit transaction can be correctly processed by the blockchain network.

[0099] In step S5, through the node client of the blockchain network, the transaction is broadcast to the network. The system in the network will verify the transaction and confirm the validity of the transaction through a consensus algorithm (such as Proof of Work - PoW). Once the transaction is confirmed, a transaction hash will be obtained as a certificate of successful deposit.

[0100] In step S6, first use the blockchain browser to query the transaction hash to confirm whether the transaction is accepted by the network. Then, compare the data fingerprint on the chain with the local data fingerprint to ensure that the data has not been tampered with during storage, thereby verifying the data consistency.

[0101] In step S7, it is implemented through a smart contract. The smart contract will provide an interface that allows users to create records on the blockchain. First, call the function of the smart contract, passing in parameters such as the data fingerprint and the transaction hash. Then, the contract will create a new record entry on the blockchain. This record can be queried by anyone but cannot be tampered with.

[0102] In step S8, use a search engine such as Elasticsearch to build an index. The specific process is as follows: Extract the keyword, label, and other information of the data shards, and then create an inverted index to associate the keywords with the deposit records on the blockchain.

[0103] In step S9, a multi-copy backup strategy is adopted, and data is stored in combination with cloud services such as Amazon S3. First, the data is encrypted, and then it is uploaded to multiple data centers of the cloud service to ensure redundant backup of the data.

[0104] In step S10, first, the transaction records on the blockchain are reviewed to verify the consistency of the data fingerprints. Secondly, the versions and configurations of the node software are checked, and the network performance is monitored. To regularly check the integrity and security of the blockchain evidence data. At the same time, necessary maintenance work is carried out according to the audit results, such as upgrading the software.

[0105] In the present invention, the generation of data fingerprints ensures the uniqueness and verifiability of the data. By adopting hash algorithms such as SHA-256, each data shard is converted into a hash value of a fixed length, and this hash value is unique. It is almost impossible for two different data shards to generate the same hash value. This uniqueness makes the data fingerprint a reliable identifier for blockchain evidence. During the data evidence process, the data fingerprint, as the "digital fingerprint" of the data, can quickly verify the integrity and authenticity of the data. Whether during data upload, storage, or subsequent verification, the integrity and reliability of the evidence data can be ensured by comparing the data fingerprints to confirm whether the data has been tampered with or damaged.

[0106] In the present invention, the generation of data fingerprints improves the efficiency and security of data processing. In the blockchain evidence system, the generation of data fingerprints is not just a simple hash calculation process. It also includes multiple steps such as multi-dimensional hash calculation, timestamp binding, and digital signature application. These steps together enhance the complexity and security of the data fingerprint. Multi-dimensional hash calculation increases the complexity of the data fingerprint by using multiple hash algorithms, making it difficult for attackers to reverse engineer and crack the original data. The application of timestamp binding and digital signature adds a time dimension and an identity authentication function to the data fingerprint, ensuring the immutability of the data fingerprint during the generation process and the traceability of the source. These measures together improve the security of the entire evidence system, and at the same time provide an efficient way for subsequent data retrieval and verification. Because the data fingerprint can directly locate the evidence records on the blockchain without traversing the entire blockchain network, greatly improving the processing efficiency. Generally speaking, the steps of generating data fingerprints not only provide strong data security protection for big data blockchain evidence, but also lay the foundation for the reliability and practicality of the entire evidence system by improving data processing efficiency.

[0107] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0108] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A big data blockchain evidence storage method, characterized by: The steps include: S1: Perform data preprocessing. First, clean, deduplicate, and format the original big data to ensure data quality. Then encrypt the data to ensure data security. S2: Perform data sharding, dividing the preprocessed data into several data shards for storage and verification on the blockchain; Ensure that the size of the data shard is appropriate for the transaction limit of the blockchain network; S3: Generate a data fingerprint, and generate a unique data fingerprint for each data shard; the data fingerprint will serve as the unique identifier of the blockchain evidence; S4: Construct a proof transaction, and package the data fingerprint and data sharding related information into a proof transaction; in this process, ensure that the format of the proof transaction meets the requirements of the blockchain network; S5: Upload the evidence transaction and send it to the blockchain network for network consensus confirmation; then ensure that the evidence transaction is successfully uploaded to the chain and obtain the corresponding transaction hash; S6: Perform evidence verification, query the evidence transaction hash through the blockchain browser to verify whether the data is successfully uploaded to the chain; then compare the on-chain data fingerprint with the local data fingerprint to ensure data consistency; S7: Register the evidence information and create an evidence information record for each data shard on the blockchain, including data fingerprint, transaction hash, and evidence storage time; the evidence information record can be used for subsequent data query, verification, and audit; S8: Build a data index based on the keywords and tag information of the data shards; Data indexing facilitates quick retrieval and location of evidence data on the blockchain; S9: Perform a backup of the evidence data, back up the original data and blockchain evidence information to a secure storage medium or cloud service to ensure the security and recoverability of the data backup; S10: Conduct regular audits and maintenance, regularly audit blockchain evidence data, check data integrity and security; maintain the blockchain network to ensure its stable operation, and timely update and optimize evidence storage methods.

2. A big data blockchain evidence storage method as claimed in claim 1, characterized in that: In step S1, the Kafka real-time data processing framework is used to pre-process the data stream; at the same time, in order to ensure data security, the AES encryption algorithm is used to encrypt the data to ensure the security of the data during transmission and storage.

3. A big data blockchain evidence storage method as claimed in claim 1, characterized in that: In step S2, a consistent hashing algorithm is used to implement data sharding. The algorithm can shard data based on the content rather than the location, thereby ensuring the uniformity of data distribution. First, a hash ring is defined, and then the data is mapped to the hash ring through a hash function. Finally, the data is distributed to different shards based on the mapping result.

4. A big data blockchain evidence storage method as claimed in claim 1, characterized in that: The step S3 specifically includes the following steps: S3-1. Data sharding enhancement processing: Perform secondary encryption on each data shard, use different encryption algorithms to increase data security; introduce a random number generator during the encryption process to ensure that each encryption result is different, and increase the complexity of the data fingerprint; S3-2. Multi-dimensional hash calculation: Use multiple hash algorithms to perform hash calculation on each data shard to generate multiple hash values; combine multiple hash values ​​to form a multi-dimensional hash fingerprint of the data shard; S3-3. Timestamp binding: Bind the current timestamp to the multi-dimensional hash fingerprint of the data shard to ensure that the time of the data fingerprint cannot be tampered with; use the timestamp service provided by the blockchain network to further verify the authenticity of the timestamp; S3-4. Digital signature application: Use the private key to digitally sign the multi-dimensional hash fingerprint and timestamp to ensure the integrity of the data fingerprint and the verifiability of the source; attach the digital signature to the data fingerprint to form the final data fingerprint package; S3-5. Data fingerprint verification: Verify the generated data fingerprint package locally to ensure that it has not been tampered with; the verification content includes the correctness of the hash value, the accuracy of the timestamp, and the validity of the digital signature; S3-6. Data fingerprint storage preparation: store the verified data fingerprint package in a secure temporary storage area for uploading to the blockchain network; ensure access control of the temporary storage area to prevent unauthorized access and data leakage; The SHA-256 multi-dimensional hash calculation formula specifically includes the following: Initialize hash values: H0=h0, H1=h1, ..., H7=h7; Filling data block: P = M || 1 || 0 k ||L; Processing data block: W[t] = M[t for t = 0 to 15 W[t]=σ1(W[t-2])+W[t-7]+σ0(W[t-15])+W[t-16] for t=16 to 63; Compression function: a=H0, b=H1, ..., h=H7fort=0to63 T1=h+∈1(a)+CH(a,b,c)+K t +W[t]T2=e0(e)+MAJ(e,f,g)h=g g=f f=e e=d+T1 d=cc=bb=aa=T1+T2end for H0=H0+α,H1=H1+b,...,H7=H7+hFinal hash value: H=H0||H1|||...||H7; in: H1=h1,...,H7 refers to the initial hash value, which is a fixed value preset by the SHA-256 algorithm M refers to the original data shard; 1 refers to an additional 1 bit; 0 k Refers to additional 0 bits for padding; L refers to the 64-bit length field, which indicates the length of the original data; W[t] refers to the 32-bit word generated by message expansion; σ1, σ0 refer to logic functions, which are used for message expansion; CH,MAJ refers to the logical function, which is used for compression function; K t Refers to 64 constants used for compression functions; T1 and T2 refer to intermediate variables, which are used for the calculation of compression functions; a, b, c, d, e, f, g, h refer to working variables and are used in the calculation of the compression function.

5. A big data blockchain evidence storage method as claimed in claim 1, characterized in that: In step S4, the relevant information must be organized into a transaction structure in accordance with the regulations of the blockchain network; the transaction structure includes input, output and witness data parts to ensure that the evidence transaction can be correctly processed by the blockchain network.

6. A big data blockchain evidence storage method as claimed in claim 1, characterized in that: In step S5, the transaction is broadcast to the network through the node client of the blockchain network; the system in the network verifies the transaction and confirms the validity of the transaction through a consensus algorithm; once the transaction is confirmed, a transaction hash will be obtained as proof of successful storage.

7. A big data blockchain evidence storage method as claimed in claim 1, characterized in that: In step S6, the transaction hash is first queried using a blockchain browser to confirm whether the transaction is accepted by the network; then, the data fingerprint on the chain is compared with the local data fingerprint to ensure that the data has not been tampered with during the storage process, thereby verifying the consistency of the data; In the step S7, it is implemented through a smart contract, which provides an interface allowing users to create records on the blockchain. First, the function of the smart contract is called, and the data fingerprint and transaction hash parameters are passed in. Then the contract creates a new record entry on the blockchain. This record can be queried by anyone, but cannot be tampered with.

8. A big data blockchain evidence storage method as claimed in claim 1, characterized in that: In step S8, the Elasticsearch search engine is used to build an index; the specific process is: extracting the keywords and tag information of the data shards, and then creating an inverted index to associate the keywords with the evidence records on the blockchain.

9. A big data blockchain evidence storage method as claimed in claim 1, characterized in that: In step S9, a multi-copy backup strategy is adopted to store data in combination with cloud services; the data is first encrypted and then uploaded to multiple data centers of the cloud service to ensure redundant backup of the data.

10. A big data blockchain evidence storage method as claimed in claim 1, characterized in that: In step S10, the transaction records on the blockchain are first reviewed to verify the consistency of the data fingerprint, and then the version and configuration of the node software are checked, and the network performance is monitored; the integrity and security of the blockchain evidence data are regularly checked; at the same time, necessary maintenance work is performed based on the audit results.

Citation Information

Cited By

  • Data element circulation full-link credible evidence storage method

    CN120281486A

  • Air environment monitoring data identification system and method

    CN120372700A

  • Data storage evidence preservation processing method and system based on distributed storage technology

    CN122348814A