Big Data Processing with In-Flight Masking and Dynamic Resource Rebating
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing big data handling systems, such as Hadoop, face challenges with resource management, security, and data integrity when processing large datasets, including memory constraints, data security vulnerabilities, and performance issues due to failed tasks and data loss during distributed processing.
Innovation Solution
Implementing in-flight data masking and on-demand encryption, resource allocation and rebalancing, fault handling, and fallback control mechanisms to securely manage and process big data across networks, ensuring efficient data transfer, obfuscation, and recovery of failed tasks while optimizing resource usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is kept in memory to ensure cost-efficient processing of big data, then processing efficiency is improved, but resource consumption increases and can result in starvation
Solution Approach 1:
The system dynamically allocates and adjusts memory resources based on workload demands. The configurable memory management allows the system to adapt memory allocation to specific dataset requirements, enabling efficient processing while preventing resource exhaustion through dynamic scaling rather than static allocation.
Solution Approach 2:
The patent implements configurable memory management parameters that can be adjusted based on the specific characteristics of the dataset being processed. By changing memory allocation parameters dynamically according to workload characteristics, the system achieves cost-efficient processing without excessive resource consumption or starvation conditions.
2Speed
If data flows from disparate sources are processed without obfuscation, then processing speed is improved, but security vulnerabilities increase due to insider and external threats
Solution Approach 1:
The system performs data obfuscation as a preliminary action before data flows from disparate sources are processed. By masking non-public personal information data upfront, the system eliminates security vulnerabilities related to insider and external threats while maintaining processing speed, as the obfuscation is integrated into the data ingestion pipeline rather than being a post-processing step.
3Device complexity
If classic perimeter defenses are used for data security, then system complexity is reduced, but they become obsolete and vulnerable to malicious insiders with access to data
Solution Approach 1:
Instead of relying on uniform perimeter defenses, the system applies data obfuscation at the local level where data resides and flows. By masking non-public personal information data specifically at the point of processing and transfer, the system provides targeted security protection that remains effective even when insiders have access to the system, without requiring complex perimeter defense architectures.
4Device complexity
If tasks are not monitored for completion in distributed processing, then system complexity is reduced, but performance drastically impacts when tasks fail or cached datasets are lost
Solution Approach 1:
The system implements feedback mechanisms that monitor task completion status in distributed processing. By tracking whether tasks have successfully completed and verifying data integrity, the system can detect failures early and trigger appropriate recovery actions. This feedback loop maintains high performance by preventing cascading failures and enabling rapid remediation without requiring overly complex monitoring infrastructure.
Data Source
AI summary
Aspects of the disclosure relate to resource allocation and rebating during in-flight data masking and on-demand encryption of big data on a network. Computer machine(s), cluster managers, nodes, and/or multilevel platforms can request, receive, and/or authenticate requests for a big data dataset, containing sensitive and non-sensitive data. Profiles can be auto provisioned, and access rights can be assigned. Server configuration and data connection properties can be defined. Secure connection(s) to the data store can be established. Sensitive information can be redacted into a sanitized dataset based on one or more data obfuscation types. RAM requirements and current RAM allocation can be diagnosed. Portion(s) of the current RAM allocation exceeding the RAM requirements can be rebated. The encrypted data can be transmitted, in response to the request, to a source, a target, and/or another computer machine and can be decrypted back into the sanitized dataset.


