Numerical Data Model for Semi-Structured DPI Traffic Parsing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large businesses face challenges in timely processing and analyzing high-volume data related to customer interactions, leading to delayed business-relevant offers and increased resource wastage, as existing data management systems struggle to handle real-time data insights effectively.
Innovation Solution
A system and method utilizing a numerical data model for semi-structured Deep Packet Inspection (DPI) data, incorporating algebraic techniques, fast tags, fixed-width counters, semi-flexible grains, business datatype aware pool-based memory management, and an embedded rule-based event trip framework to efficiently parse and label DPI traffic, enabling real-time contextual targeting and marketing offers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If traditional data management systems are used to store and process customer data, then data can be stored in databases, but the processing speed is too slow to provide real-time insights and offers
Solution Approach 1:
The patent replaces traditional mechanical database query systems with a streaming data processing architecture that continuously processes data in motion. The system uses stream processing engines to analyze customer interaction data as it arrives, eliminating the batch processing delays inherent in traditional database systems and enabling real-time offer generation.
Solution Approach 2:
The system performs preliminary actions by pre-processing and enriching customer data in real-time as it streams through the system. Data is validated, normalized, and tagged with contextual information before being stored, enabling faster query response times and eliminating the need for complex real-time data preparation during offer generation.
2Productivity
If more processing power and bandwidth are increased to handle high-volume data, then data processing capability improves, but equipment expense and operating costs increase
Solution Approach 1:
The patent segments the data processing workload across multiple distributed processing nodes rather than using a single powerful system. Each node handles a portion of the data stream independently, allowing the system to scale horizontally by adding more modest-capacity machines rather than investing in expensive high-performance hardware.
Solution Approach 2:
The system dynamically adjusts processing parameters such as batch size, parallelism degree, and memory allocation based on incoming data volume and system load. This allows the system to optimize resource utilization and avoid over-provisioning hardware, reducing both equipment expenses and operating costs while maintaining high processing capability.
3Ease of operation
If data is stored in relational databases for structured queries, then data organization is improved, but the system cannot timely process relevant data for real-time offers
Solution Approach 1:
The patent introduces an intermediary layer between data ingestion and storage/query operations. A streaming data processing engine sits between the data sources and the database systems, performing real-time filtering, aggregation, and enrichment of data before it reaches the databases. This intermediary enables both real-time processing capabilities and maintains organized data storage for structured queries.
Data Source
AI summary
A system for efficiently parsing semi-structured deep packet inspection traffic data tied to a telecommunications entity. The system is capable of parsing such records at million-records-per-second scale through use of a numerical data model, leverage on proven fundamental algebraic techniques, and shortcuts to label streaming traffic on the fly. In some embodiments, the system may perform parallel accumulation of data traffic into business grade counters using elementary techniques and subsequently identify subscribers exhibiting specific data patterns in real time for contextual targeting of promotional offers. A method of efficiently parsing the traffic data via the system of the disclosure.


