Classification Tool for Structured Data Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Processing large datasets from online user activities to extract useful information is challenging due to varying representation techniques and lack of methods to identify significant information across different online entities.
Innovation Solution
A classification tool is developed to identify semantically similar portions of structured data lacking a common ontology, using a training system that groups raw data into cases, presents them to trainers for labeling, and applies machine-learning algorithms to automatically classify target classes within the data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual processing techniques are used to identify significant information in large datasets from online user activities, then processing accuracy can be maintained, but processing time and labor requirements increase significantly
Solution Approach 1:
The system performs preliminary actions by pre-processing raw data into structured formats with standardized field names and data types before classification. This preliminary structuring enables faster subsequent processing while maintaining accuracy, as the data is already organized and ready for automated classification algorithms.
Solution Approach 2:
The patent introduces an intermediary classification tool that acts as a bridge between raw unstructured data and meaningful insights. This classification tool uses machine learning algorithms to automatically identify and categorize significant information, reducing both processing time and manual labor while maintaining identification accuracy through trained classification models.
2Productivity
If automated classification algorithms are applied directly to raw unstructured data, then processing speed increases, but identification accuracy decreases due to varying representation techniques and lack of common ontology
Solution Approach 1:
The system applies preliminary action by structuring raw unstructured data into standardized formats with consistent field names, data types, and schemas before applying automated classification. This pre-structuring ensures that automated algorithms can process data quickly while maintaining accuracy, as the varying representation techniques are normalized in advance.
Solution Approach 2:
The patent changes parameters by transforming raw data into structured data with standardized field names, data types, and formats. This parameter transformation creates a common ontology framework that enables automated classification algorithms to process diverse data sources accurately and efficiently, resolving the conflict between speed and precision.
3Device complexity
If data from multiple online entities with different representation techniques is processed using a single methodology, then system complexity is reduced, but processing accuracy deteriorates due to lack of common ontology
Solution Approach 1:
The classification tool achieves universality by designing a standardized data structure and classification framework that can handle multiple data sources with different representation techniques. The system uses universal field names, data types, and classification categories that work across diverse online entities, maintaining both low system complexity and high classification accuracy through a unified approach.
Solution Approach 2:
The patent applies parameter changes by transforming diverse data representations into a standardized format with common field names, data types, and structures. This parameter normalization creates a common ontology that enables accurate classification across multiple online entities while keeping the system simple through consistent processing methodology.
4Measurement precision
If extensive manual labeling of training data is performed to improve classifier accuracy, then classification precision improves, but time and labor requirements increase
Solution Approach 1:
The system performs preliminary action by automatically pre-labeling training data using initial classification rules and heuristics before human trainers review it. This pre-labeling reduces the amount of manual work required while maintaining high classification precision, as trainers only need to review and correct a small subset of pre-labeled data rather than labeling everything from scratch.
Solution Approach 2:
The classification system applies self-service by using its own initial classifications and feedback to automatically improve its performance over time. The system learns from trainer corrections and automatically refines its classification algorithms, reducing the need for continuous manual labeling while maintaining or improving classification precision through automated self-improvement.
Data Source
AI summary
An exemplary embodiment of the present invention provides a computer implemented method of developing a classifier. The method includes receiving input for a case, the case comprising a plurality of instances and an example, the example comprising a plurality of data fields each corresponding to one of the plurality of instances, wherein the input indicates which, if any, of the instances includes a data field belonging to a target class. The method also includes training the classifier based, at least in part, on the input from the trainer.


