Parallel Filter Graph Optimization for Real-Time Data Stream Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for monitoring social media fail to effectively extract actionable intelligence from vast amounts of data, often producing more data than companies can react to and missing critical information.
Innovation Solution
A method and system that utilize hierarchical, parallel models to identify high-value information in real-time data streams by combining classification models based on content and metadata, distributed across multiple processors for efficient execution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If conventional tools and methods are used for monitoring social media, then data collection is performed, but the ability to identify valuable information is poor and the data volume overwhelms companies
Solution Approach 1:
The patent segments the data stream processing by dividing it into multiple parallel processing paths, each handling different aspects of data analysis. Classification models are distributed across multiple processors, with each processor executing specific models on portions of the data stream, enabling efficient filtering and identification of valuable information from large volumes of social media data
Solution Approach 2:
The patent changes the parameter of information value by using multiple classification models with different criteria to evaluate data packets. Each model assesses packets based on different parameters (content, metadata, engagement metrics), and packets are identified as high-value when they meet specific thresholds across multiple models, transforming raw data into classified information with varying value parameters
2Productivity
If multiple classification models are executed in parallel on multiple processors, then the extraction of high-value information is improved, but the system complexity increases
Solution Approach 1:
The patent creates a universal processing framework where multiple classification models can be executed in parallel across multiple processors using a common architecture. The system uses standardized data packet structures, uniform distribution mechanisms, and consistent result aggregation methods, allowing different classification models to work together efficiently without requiring separate processing systems for each model
Solution Approach 2:
The patent introduces intermediary components including a data distribution mechanism that routes packets to appropriate processors, and a result aggregation system that consolidates findings from multiple parallel executions. These intermediaries manage the complexity of parallel processing by providing standardized interfaces between the diverse classification models and the overall system
3Loss of time
If data is processed in real-time with parallel execution, then the response time is reduced, but the computational resources required increase
Solution Approach 1:
The patent applies partial action by executing multiple classification models in parallel rather than sequentially, processing data packets through selected models simultaneously across multiple processors. This parallel partial execution reduces overall processing time while optimizing resource usage by not requiring all possible models to analyze every single packet, thereby achieving real-time performance with manageable computational resources
Data Source
AI summary
A computer system identifies high-value information in data streams. The computer system receives a filter graph definition. The filter graph definition includes a plurality of filter nodes, each filter node including one or more filters that accept or reject packets. Each respective filter is categorized by a number of operations, and the one or more filters are arranged in a general graph. The computer system performs one or more optimization operations, including: determining if a closed circuit exists within the graph, and when the closed circuit exists within the graph, removing the closed circuit; reordering the filters based at least in part on the number of operations; and parallelizing the general graph such that the one or more filters are configured to be executed on one or more processors.


