Automated Taxonomy Engine for Classifier Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current taxonomy creation methods rely heavily on human scoring, which is time-consuming and prone to errors, requiring significant effort and time to identify classifiers from large collections of documents, making it inefficient for real-time analysis and action.

Innovation Solution

A computer-implemented taxonomy automation engine that generates classifiers by parsing documents to identify topic and sentiment terms, creating a colocation matrix, and using distance metrics to select relevant classifiers, thereby automating the taxonomy development process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If human scoring is used for taxonomy creation, then accuracy and reliability of classifier identification is improved, but time consumption and labor effort increase significantly

Engineering Contradiction:
Improveaccuracy of classifier identificationVSAvoidtime required for taxonomy development
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent replaces the manual human scoring process with an automated computer-implemented system that uses computational methods to identify classifiers. The system processes documents automatically to extract topic terms, sentiment terms, and candidate classifiers, eliminating the need for human annotators while maintaining classification accuracy through algorithmic analysis.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service taxonomy development by automatically generating classifiers without requiring human intervention. The computational method independently analyzes documents, identifies patterns, and produces taxonomies that can be directly applied, allowing the system to serve itself rather than requiring human scoring operations.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If manual taxonomy development is used, then precision and control over classifier selection is improved, but productivity and speed of analysis decrease

Engineering Contradiction:
Improveprecision of classifier selectionVSAvoidspeed of document analysis
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent substitutes manual precision judgment with automated computational methods that use defined algorithms for identifying topic terms, sentiment terms, and candidate classifiers. The system applies consistent computational rules to ensure precise classifier selection while processing documents at machine speed, achieving both precision and high productivity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs preliminary processing by automatically pre-identifying topic terms and sentiment terms before final classifier selection. This preliminary action organizes and pre-sorts document content, making the subsequent classifier identification process faster and more efficient while maintaining precision through structured analysis.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If automated classification methods are used, then speed of analysis and productivity is improved, but complexity of the system and difficulty of implementation increase

Engineering Contradiction:
Improvespeed of taxonomy developmentVSAvoidcomplexity of classification system
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the complex taxonomy development process into distinct manageable stages: topic term identification, sentiment term identification, candidate classifier generation, and final classifier selection. Each stage has specific computational operations and criteria, breaking down the overall complexity into sequential steps that are easier to implement and maintain.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system manages complexity by allowing adjustment of key parameters such as topic threshold distance, sentiment threshold distance, and colocation matrix thresholds. These parameter changes enable the system to adapt to different document collections and requirements without fundamentally altering the core algorithmic structure, simplifying implementation while maintaining high productivity.

Inventive Principle:
Principle #35Parameter changes

4Reliability

If human scoring is used for taxonomy creation, then accuracy and reliability of classifier identification is improved, but ease of operation and scalability decrease

Engineering Contradiction:
Improveaccuracy of classifier identificationVSAvoidscalability of taxonomy development
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent replaces manual human scoring with an automated system that scales efficiently. The computational method can process large volumes of documents automatically, enabling the system to scale from small to large datasets without requiring proportional increases in human resources, thereby achieving both accuracy and scalability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system provides universal applicability by using general computational methods that can be applied across different document collections, industries, and topics. The same core algorithmic framework handles various scenarios, making the system versatile and scalable without requiring topic-specific manual scoring procedures.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS9116985B2Computer-implemented systems and methods for taxonomy development
Publication Date: 2015.08.25 SAS INSTITUTE INC
  • US9116985B2 patent drawing
  • US9116985B2 patent drawing
  • US9116985B2 patent drawing

AI summary

Systems and methods are provided for generating a set of classifiers. A location is determined for each instance of a topic term in a collection of documents. One or more topic term phrases are identified, and one or more sentiment terms within each topic term phrase. Candidate classifiers are identified by parsing words in the one or more topic term phrases, and a colocation matrix is generated. A seed row of the colocation associated with a particular attribute is identified, and distance metrics are determined by comparing each row of the colocation matrix to the seed row. A set of classifiers are generated for the particular attribute, where classifiers in the set of classifiers are selected using the distance metrics.