Big Data Management System for Dynamic Keyword Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems lack the ability to effectively leverage big data sources due to the challenges of volume, velocity, and variety, making it difficult to process and analyze the increasing amounts of data from diverse sources, which hinders real-time situational awareness and automated issue resolution.
Innovation Solution
A big data management system that dynamically identifies keywords from multiple data sources, classifies them as related to technical problems or solutions, weights the associated data based on attributes, and stores it in a unified database for analysis, enabling reactive, proactive, and preemptive generation of solutions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional database and software techniques are used to process data, then processing capabilities are maintained within existing organizational capacities, but the volume, velocity, and variety of big data exceed these capabilities and cannot be effectively processed
Solution Approach 1:
The patent segments the monolithic data processing approach into a distributed architecture where data is divided into manageable chunks processed by multiple worker nodes. The master node coordinates these workers, allowing the system to handle large volumes of data by processing segments in parallel, thereby scaling productivity with data quantity.
Solution Approach 2:
The patent transitions from traditional single-dimension processing to multi-dimensional processing by implementing a distributed computing framework that processes data across multiple dimensions: spatial distribution across multiple nodes, temporal parallelism through concurrent processing, and hierarchical organization with master-worker relationships. This dimensional expansion enables the system to handle big data volumes that exceed traditional processing capabilities.
2Loss of information
If data is collected from increasing numbers of sensors and devices at higher velocities, then data availability and real-time awareness are improved, but the velocity and volume of data streaming exceed traditional processing speeds
Solution Approach 1:
The patent implements preliminary action through data preprocessing and filtering at the source before data enters the main processing pipeline. Worker nodes perform initial data validation, formatting, and filtering operations, preparing data in advance for more efficient processing. This preliminary action reduces the burden on the master node and enables faster handling of high-velocity data streams.
Solution Approach 2:
The patent ensures continuous processing of data streams through persistent worker nodes that continuously receive and process incoming data without interruption. The system maintains continuous connections with data sources, continuously ingests data streams, and continuously processes information, ensuring no loss of real-time situational awareness even at high velocities.
3Adaptability or versatility
If diverse data formats from multiple sources are collected, then data variety and comprehensiveness are improved, but the variety and complexity of data formats exceed traditional processing capabilities
Solution Approach 1:
The patent implements universality through a standardized data interface and common data model that all worker nodes adhere to. Despite receiving diverse data formats from different sources, the system uses a universal processing framework that can handle multiple formats through standardized methods. The master node coordinates these universal operations, enabling the system to process varied data types without proportionally increasing complexity.
4Loss of time
If data is processed and analyzed to generate solutions in real-time, then customer satisfaction and mean time to repair are improved, but the computational resources and processing time required increase significantly
Solution Approach 1:
The patent applies partial action by processing only the most critical and relevant data subsets first, rather than analyzing all available data uniformly. The system identifies and prioritizes high-impact data for immediate processing, generating solutions for the most urgent issues while deferring less critical analyses. This approach reduces computational resource consumption while maintaining improved mean time to repair for critical issues.
Data Source
AI summary
Techniques are presented herein to monitor a plurality of big data sources in order to dynamically identify keywords. The big data sources are analyzed to classify the keywords as related to either a technical problem or to a solution to the technical problem. In addition, data associated with the keywords is weighted based on one or more attributes of the data and stored in a database in a problem-solution format.


