Data Processing Platform Reducing Dataset Volume by 99 Percent

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional big data platforms face challenges in efficiently processing and analyzing large datasets due to the need to identify user identities and causative characteristics for every data point, which consumes significant computing resources and may not guarantee the identification of useful insights.

Innovation Solution

The approach disregards user identity and causative characteristics initially, aggregating and normalizing data before computation, reducing the dataset by up to 99% through predetermined rules and algorithms to identify critical intelligence information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional big data platforms analyze each unit of datum to identify useful insights, then measurement precision is improved, but productivity deteriorates due to significant computing resource consumption and processing delays

Engineering Contradiction:
Improveidentification accuracy of useful insightsVSAvoiddata processing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs preliminary actions by applying predetermined rules and algorithms to aggregate and normalize data before full analysis. This preprocessing step reduces the data volume to a manageable subset that is then analyzed in detail, achieving both processing efficiency and identification accuracy without analyzing every single datum point

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The data processing is segmented into distinct phases: initial data aggregation using predetermined rules, normalization of the reduced dataset, and then detailed analysis of the normalized data. This segmentation allows the system to handle large volumes of data efficiently while maintaining precision in the final analysis stage

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If software applications apply substantial analysis against each unit of datum, then manufacturing precision is improved, but loss of energy increases due to overuse of computer processors, memory, and technical computing elements

Engineering Contradiction:
Improvedata analysis accuracyVSAvoidcomputing resource consumption
Core Design Contradiction:
Manufacturing precisionVSLoss of energy

Solution Approach 1:

The system extracts only the essential data elements that are likely to contain useful insights by applying predetermined filtering rules and algorithms. Instead of analyzing all data units with substantial computational power, the system extracts a reduced subset of critical data points for detailed analysis, significantly reducing computing resource consumption while maintaining analysis accuracy

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system applies partial action by performing substantial analysis only on a portion of the data (the normalized subset) rather than all data units. The predetermined rules and algorithms enable the system to skip unnecessary analysis on data points that are unlikely to yield useful insights, reducing energy consumption while preserving the ability to identify critical intelligence information

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If the system identifies user identity and causative characteristics for every data point, then reliability is improved, but device complexity increases due to the need to process and store extensive metadata

Engineering Contradiction:
Improvedata traceabilityVSAvoidsystem architecture complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by capturing essential metadata (user identity and causative characteristics) during the initial data aggregation phase using predetermined rules. This metadata is stored alongside the reduced dataset rather than for every individual data point, maintaining data traceability and reliability while avoiding the complexity of managing extensive metadata for all原始数据 points

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11276006B2System, apparatus, and method to identify intelligence using a data processing platform
Publication Date: 2022.03.15 OUTLIER AI INC
  • US11276006B2 patent drawing
  • US11276006B2 patent drawing
  • US11276006B2 patent drawing

AI summary

A system and method includes: implementing an intelligence and insights service; identifying an anomalous observation output by a data processing pipeline based on streams of data sourced from a subscriber to the intelligence and insights service; recursively inputting into a subset of the data processing pipeline of the intelligence and insights service a plurality of dimensions of the streams of data based on attributes of the anomalous observation; automatically identifying one or more driving factors causing the output of the anomalous observation based on an analysis within the subset of the data processing pipeline of plurality of dimensions of the streams of data; generating a story component based on a conversion of the one or more driving factors; and augmenting the story component to a pre-existing story relating to the anomalous observation that is provided to the subscriber via a user interface.