Data Classification Engines for Metadata Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Businesses face challenges in accessing and analyzing vast amounts of data from multiple sources, including metadata, due to its complexity and the need for integrated processing and storage across various systems, which existing technologies struggle to efficiently manage.

Innovation Solution

A system comprising a priori and a posteriori classification engines, along with a heuristics engine, that classifies data using classification rules, probabilistic algorithms, and heuristic analysis to extract metadata and trends from internal and external data sources, enabling comprehensive data aggregation and analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If data from multiple sources is aggregated and analyzed, then valuable insights and trends can be extracted, but the complexity of accessing and processing the data increases

Engineering Contradiction:
Improveextraction of metadata and trendsVSAvoidsystem complexity for data aggregation
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system divides data aggregation into multiple specialized engines: a priori classification engine for initial categorization, a posteriori classification engine for refined analysis, and heuristics engine for pattern recognition. Each engine handles specific aspects of data processing, making the complex task of extracting metadata and trends from multiple sources more manageable and efficient

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces classification engines as intermediary components between raw data sources and business analysis systems. These engines serve as mediators that automatically process, classify, and structure data from diverse sources (sales databases, email systems, social media, etc.), reducing the complexity burden on end-user systems while enabling comprehensive data aggregation

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If comprehensive data classification is performed across multiple systems, then unified data analysis is achieved, but the processing time and computational resources increase

Engineering Contradiction:
Improveunified data classification capabilityVSAvoiddata processing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The a priori classification engine performs preliminary classification of data as it is collected from multiple sources, organizing data into categories before detailed analysis is required. This advance categorization reduces the time needed for subsequent a posteriori classification and heuristic analysis, enabling faster unified data analysis across systems

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system employs dynamic classification where the three engines (a priori, a posteriori, and heuristics) operate in flexible sequences and can be activated based on data characteristics and analysis requirements. This dynamic approach optimizes processing time by applying the appropriate level of classification intensity to different data types and sources

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS9501744B1System and method for classifying data
Publication Date: 2016.11.22 QUEST SOFTWARE INC
  • US9501744B1 patent drawing
  • US9501744B1 patent drawing
  • US9501744B1 patent drawing

AI summary

In one embodiment, a method includes providing an a priori classification engine, an a posteriori classification engine, and a heuristics engine. The a priori classification engine is operable to perform an a priori classification. The a posteriori classification engine is operable to perform an a posteriori classification. The heuristics engine is operable to perform a heuristics classification. In addition, the method includes accessing data from at least one source. The method further includes, responsive to an indication that the a priori classification should be performed, performing the a priori classification on the data. The method also includes, responsive to an indication that the a posteriori classification should be performed, performing the a posteriori classification on the data. Further, the method includes, responsive to an indication that the heuristics classification should be performed, performing the heuristics classification on the data.