Web-Scale Data Fabric Using Commodity Hardware for Big Data Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current database management systems are inadequate for processing, storing, and analyzing large and complex data sets known as 'big data,' requiring extensive resources and hardware, which is costly and inefficient, and fails to provide instantaneous access to information.
Innovation Solution
A system utilizing a software-defined network (SDN) with commodity hardware, including multiple computer servers, direct attached storage, random access memory, and a stream processor with message brokers and complex event processors, configured for resilient high-throughput-low-latency data processing, enabling economical large-scale computation and storage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If existing database management systems are used to process big data, then data processing capability is maintained at current levels, but hardware costs and resource requirements become excessively high
Solution Approach 1:
The patent applies this principle by using commodity hardware components that can be easily replaced and upgraded. The system is designed to run on standard, off-the-shelf hardware rather than specialized expensive equipment, allowing organizations to process big data using economical, readily available components that can be discarded or replaced as needed.
Solution Approach 2:
The patent segments the big data processing system into multiple independent components including data ingestion layer, processing layer, storage layer, and analytics layer. Each layer can be independently scaled and optimized, allowing the system to handle large data volumes without requiring a monolithic expensive hardware infrastructure.
2Speed
If traditional data processing systems are used, then system complexity is kept manageable, but instantaneous access to information cannot be provided
Solution Approach 1:
The patent implements preliminary action by pre-processing and pre-organizing data in advance using multiple processing layers. Data is ingested, cleaned, transformed, and indexed before being stored, so that when queries are executed, the system can retrieve information instantly without performing heavy processing at query time.
Solution Approach 2:
The patent introduces an intermediary layer between data storage and data retrieval operations. This includes caching mechanisms, index structures, and query optimization layers that mediate between the user's information needs and the underlying data storage, enabling fast access without exposing the full system complexity to users.
3Ease of manufacture
If commodity hardware is used for economy, then cost efficiency is improved, but system reliability and resilience become more challenging to maintain
Solution Approach 1:
The patent applies local quality by implementing different reliability mechanisms at different levels of the system architecture. Critical components have enhanced error handling and redundancy, while non-critical components use simpler approaches. This allows the system to achieve overall high reliability while maintaining cost efficiency through selective application of reliability measures where they are most needed.
Solution Approach 2:
The patent implements beforehand cushioning by incorporating error handling, data validation, and recovery mechanisms at each processing layer before failures can propagate. Checksum validation, transaction logging, and automated recovery procedures are built into the architecture to cushion against hardware failures inherent in commodity systems.
Data Source
AI summary
Methods and systems for processing machine accelerated and augmented customer data using a Web-Scale Data Fabric (WSDF). According to embodiments, the data may be received as data transfer objects from a set of business operations client applications. The data transfer objects may be analyzed using complex event processing (CEP) and, based on the analyzing, rules specific to the business operations client application may be applied. The methods and systems may semantically classify text specific to the business operations client application. A federated database (FD) may archive the receive data transfer objects as well as analysis data specific to the business operations client application.


