Classification Tool for Structured Data Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Processing large datasets from online user activities to extract useful information is challenging due to varying representation techniques and lack of methods to identify significant information across different online entities.

Innovation Solution

A classification tool is developed to identify semantically similar portions of structured data lacking a common ontology, using a training system that groups raw data into cases, presents them to trainers for labeling, and applies machine-learning algorithms to automatically classify target classes within the data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual processing techniques are used to identify significant information in large datasets from online user activities, then processing accuracy can be maintained, but processing time and labor requirements increase significantly

Engineering Contradiction:
Improveidentification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-processing raw data into structured formats with standardized field names and data types before classification. This preliminary structuring enables faster subsequent processing while maintaining accuracy, as the data is already organized and ready for automated classification algorithms.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary classification tool that acts as a bridge between raw unstructured data and meaningful insights. This classification tool uses machine learning algorithms to automatically identify and categorize significant information, reducing both processing time and manual labor while maintaining identification accuracy through trained classification models.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If automated classification algorithms are applied directly to raw unstructured data, then processing speed increases, but identification accuracy decreases due to varying representation techniques and lack of common ontology

Engineering Contradiction:
Improveprocessing speedVSAvoididentification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system applies preliminary action by structuring raw unstructured data into standardized formats with consistent field names, data types, and schemas before applying automated classification. This pre-structuring ensures that automated algorithms can process data quickly while maintaining accuracy, as the varying representation techniques are normalized in advance.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes parameters by transforming raw data into structured data with standardized field names, data types, and formats. This parameter transformation creates a common ontology framework that enables automated classification algorithms to process diverse data sources accurately and efficiently, resolving the conflict between speed and precision.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If data from multiple online entities with different representation techniques is processed using a single methodology, then system complexity is reduced, but processing accuracy deteriorates due to lack of common ontology

Engineering Contradiction:
Improvesystem complexityVSAvoidclassification accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The classification tool achieves universality by designing a standardized data structure and classification framework that can handle multiple data sources with different representation techniques. The system uses universal field names, data types, and classification categories that work across diverse online entities, maintaining both low system complexity and high classification accuracy through a unified approach.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent applies parameter changes by transforming diverse data representations into a standardized format with common field names, data types, and structures. This parameter normalization creates a common ontology that enables accurate classification across multiple online entities while keeping the system simple through consistent processing methodology.

Inventive Principle:
Principle #35Parameter changes

4Measurement precision

If extensive manual labeling of training data is performed to improve classifier accuracy, then classification precision improves, but time and labor requirements increase

Engineering Contradiction:
Improveclassification precisionVSAvoidtraining efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs preliminary action by automatically pre-labeling training data using initial classification rules and heuristics before human trainers review it. This pre-labeling reduces the amount of manual work required while maintaining high classification precision, as trainers only need to review and correct a small subset of pre-labeled data rather than labeling everything from scratch.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The classification system applies self-service by using its own initial classifications and feedback to automatically improve its performance over time. The system learns from trainer corrections and automatically refines its classification algorithms, reducing the need for continuous manual labeling while maintaining or improving classification precision through automated self-improvement.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS8355997B2Method and system for developing a classification tool
Publication Date: 2013.01.15 MICRO FOCUS LLC
  • US8355997B2 patent drawing
  • US8355997B2 patent drawing
  • US8355997B2 patent drawing

AI summary

An exemplary embodiment of the present invention provides a computer implemented method of developing a classifier. The method includes receiving input for a case, the case comprising a plurality of instances and an example, the example comprising a plurality of data fields each corresponding to one of the plurality of instances, wherein the input indicates which, if any, of the instances includes a data field belonging to a target class. The method also includes training the classifier based, at least in part, on the input from the trainer.