Hadoop Big Data Security With Integrated Blockchain Layers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing Big Data Systems based on Hadoop lack sufficient data security, as MD5CRC algorithms are vulnerable to attacks, and integrating a standalone Blockchain network for data storage results in slow performance and limited data manipulation.
Innovation Solution
Integrate Blockchain technologies into the Hadoop-based Big Data System architecture without forming a standalone Blockchain network, adding layers of data security and processing capabilities, including data encryption, masking, and automatic correction, while maintaining near-real-time data processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If MD5CRC algorithms are used for data security in HDFS, then data integrity checking is enabled, but the system becomes vulnerable to preimage attacks and brute force attacks
Solution Approach 1:
The patent changes the cryptographic parameter from MD5CRC to SHA-3 hash algorithm. SHA-3 provides stronger security against preimage attacks and brute force attacks while maintaining compatibility with the HDFS architecture. This parameter change directly addresses the vulnerability issue while preserving data integrity checking capabilities.
Solution Approach 2:
The patent creates a composite security mechanism by integrating Blockchain technology with HDFS. The Blockchain ledger stores cryptographic hashes of data blocks, creating a layered security structure where SHA-3 hashing and Blockchain verification work together to provide robust protection against various attack vectors.
2Reliability
If a standalone Blockchain network is integrated for data storage, then data security is enhanced, but system performance slows down and data manipulation becomes limited
Solution Approach 1:
The patent extracts only the essential security components of Blockchain technology (hashing, ledger, consensus mechanism) and integrates them into HDFS, rather than implementing a complete standalone Blockchain network. This selective extraction maintains security benefits while avoiding the performance overhead of full Blockchain functionality.
Solution Approach 2:
The patent merges Blockchain security mechanisms with HDFS storage architecture into a unified system. The Blockchain ledger is integrated within HDFS to store cryptographic hashes, allowing both systems to function together seamlessly without the performance penalties of a separate Blockchain network.
3Reliability
If SHA-3 hash algorithm is applied to HDFS files, then cryptographic security is improved, but user programming intervention in Hadoop code is required
Solution Approach 1:
The patent implements self-service by making the HDFS system automatically perform SHA-3 hashing and Blockchain verification operations without requiring user programming. The system autonomously manages the cryptographic operations, eliminating the need for users to modify Hadoop code while still providing enhanced security.
4Reliability
If Blockchain ledger is used to store all big data transactions, then data security is enhanced, but processing speed becomes very slow especially for large data volumes
Solution Approach 1:
The patent applies partial Blockchain functionality by storing only cryptographic hashes of data blocks in the Blockchain ledger rather than complete transaction records. This partial implementation provides sufficient security verification while dramatically reducing storage requirements and processing time compared to storing all data in Blockchain.
Data Source
Figure 1

AI summary
The patent is in the field of information technology, enhancing the security of Big Data Systems using Blockchain technologies and it consists of the following blocks: Block 1 - Big Data System, Block 2 - Data transmission networks, Block 3 - Data Collector, Data Cleaning and Enrichment, Block 4 - Data Allocator, Block 5 - Management of Primary Data Repository for real-time processing, Block 6 - Application Transaction Creator, Block 7 - Blockchain ledger block generator with encryption, masking and editing, Block 8 - Blockchain Ledger Verifier Part 1, Block 9 - Blockchain Ledger Verifier Part 2, Block 10 - Blockchain Ledger Verifier part 3, Block 11 - Obtaining Consent, Block 12 - Recording in the User Audit Register, Block 13 - Recording in Blockchain Ledger, Block 14 - Managing a User Blockchain ledger for data analytics, quality and security, Block 15 - Defect Recorder in the Blockchain Ledger, Block 16 - Creating a correction for the Blockchain ledger, Block 17- User system, Block 18 - Integrating Blockchain Authentication and Kerberos/Hadoop Authentication, Implementing User Keys and Roles, Block 19 - Data sharing between authenticated users, Block 20 - Business Rules Processor for Blockchain Ledgers, Block 21 - Managing the life cycle of a User Blockchain ledger. Basically, in Big Data Systems based on Hadoop, data security. The patent proposes a System and Method for enhancing the security of Big Data Systems using Blockchain technologies, adding to their architecture blocks implementing individual Blockchain technologies, expanding the Big Data System architecture, without integrating with a standalone Blockchain network (system), and through these proposed System and Method, data security in Big Data Systems is extended with 2 advanced levels of security, with the data at the first advanced level of security also having the possibility of processing in "near real time", and at the second advanced level, having higher security, the data is presented as a Blockchain ledger, operating predominantly not in "near real-time mode", but with faster processing compared to processing the Blockchain ledger in a Blockchain network, and also with automatic correction when detected mistake. The proposed System and Method can extend the security of any Hadoop-based Big Data System. The system consists of 21 blocks and the Method consists of 5 security functions. The first extended security level creates data with an increased level of security capable of "near real-time" processing applied to Block 1 "Big Data System" whose security functions are extended using Block 17 "User System " through which the user is authenticated and authorized to work with the extended Big Data System, Block 18 "Integration of Blockchain Authentication and Kerberos/Hadoop Authentication, Application of User Keys and Roles" performing authentication and authorization of users and user systems with their user roles for general interaction with Kerberos the authentication of Big Data Systems using the created cryptographic keys and set potential data sources, Block 19 "Data sharing between authenticated users" for data exchange between users and/or user systems using the authorized rules and roles , Block 2 "Networks" for receiving data, the security of which is to be extended by the proposed System and Method, Block 3 "Data Collector, Cleaning and Enrichment" where data is collected and routed for storage, undergoing cleaning by content point of view and are enriched by adding new data or by modifying the existing data with other already arrived data of authenticated same or other users or user systems, stored through the Collector and having a logical and business connection with the incoming data, Block 4 "Distributor of data" redirecting the data to Block 5 "Management of Primary Data Storage for real-time processing", where data storage will be organized for which first advanced level security will be implemented. In the case of integrated authentication, the activity of the Kerberos function in Big Data Systems is used, authentication and authorization parameters are created for each user: a pair of cryptographic keys for authentication and for encryption; a list of authenticated users whose data is authorized for access, and these authentication and authorization parameters are integrated into Kerberos as encrypted data encrypted by the system administrative keys, and the authorization parameters are recorded in the Access Control List of the Big Data System. The processing of the data arriving from the networks also consists of: Removing duplicate records, correcting inconsistencies, filling in missing values and converting data into a certain format; Adding new data or by modifying existing data with other data already received; Modifying sensitive data in such a way that it has no or negligible value to unauthorized users ensuring anonymization or tokenization of the data; Edit by changing the content of the data. The created data with advanced security features on the one hand based on the primary data format in Big Data Systems - HDFS files, and on the other hand offers both increased security features and advanced forms of big data operation in "close to real time' such as in Spark, Hbase and Cassandra data formats. In the second extended level of security, a Blockchain ledger is used, in which level of security is operated with Block 2 "Networks" for receiving data, the security of which is to be extended by the proposed System and Method, Block 3 "Data Collector, Clearing and enrichment" where the data is collected and sent to storage, being cleaned from a content point of view and enriched by adding new data or by modifying the existing data with other already arrived data of authenticated same or other users or user systems, stored through the Collector and having a logical and business relationship with the incoming data, Block 4 "Data Distributor" forwarding the data to Block 6 "Applied Transaction Creator" to create a transaction to be used as the basis for creating a block of data in the Blockchain ledger from Block 7 "Blockchain ledger block generator with encryption, masking and editing", as the current state of the Blockchain ledger from Block 14 "Management of User Blockchain ledger for data analytics, quality and security" is checked by Block 8 "Verifier of Blockchain ledger Part 1, Block 9 "Verifier of Blockchain ledger Part 2" and Block 10 "Verifier of Blockchain ledger Part 3" and upon positive verification from Block 11 "Obtaining consent" go to record this block in the Blockchain ledger with Block 14 "Management of a User Blockchain ledger for data analytics, quality and security" using Block 12 "Entry in a User Audit Register", Block 13 "Entry in the Blockchain book", in which the term "transaction" is used in the Blockchain book, it represents a certain amount of user data (data in the Big Data System) serving to create a "block of data" in the Blockchain book - a unit of record in the Blockchain book , while the term "Blockchain ledger" is a structured system of data-blocks, where each new record (Blockchain ledger block) contains the parameters "time-of-creation", "data source/sources", "data content (the transaction)", "authentication parameters" and "hash function of the previous record (block)", as the Blockchain ledger exists in several replicas, with the previous block-next block validation performed in 3 parallel parts (each verifier about 1/3 of the blocks in the Blockchain book), and the connection of the next block is made through the hash function of the previous block, determined by the principles of the Blockchain book, where the first 2 blocks of each Blockchain book contain the system administrative keys related to decryption of the authentication parameters of the individual user. When reading, writing or modifying data in the Blockchain ledger, the current content of the Blockchain ledger is checked through the specified checks. When a data impropriety is detected in a Blockchain ledger, the data is corrected (using the multiple replicas of the Blockchain ledger) and the corrected ledger is saved with its corrected content automatically, without user intervention.