Automated Taxonomy Engine for Classifier Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current taxonomy creation methods rely heavily on human scoring, which is time-consuming and prone to errors, requiring significant effort and time to identify classifiers from large collections of documents, making it inefficient for real-time analysis and action.
Innovation Solution
A computer-implemented taxonomy automation engine that generates classifiers by parsing documents to identify topic and sentiment terms, creating a colocation matrix, and using distance metrics to select relevant classifiers, thereby automating the taxonomy development process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If human scoring is used for taxonomy creation, then accuracy and reliability of classifier identification is improved, but time consumption and labor effort increase significantly
Solution Approach 1:
The patent replaces the manual human scoring process with an automated computer-implemented system that uses computational methods to identify classifiers. The system processes documents automatically to extract topic terms, sentiment terms, and candidate classifiers, eliminating the need for human annotators while maintaining classification accuracy through algorithmic analysis.
Solution Approach 2:
The system enables self-service taxonomy development by automatically generating classifiers without requiring human intervention. The computational method independently analyzes documents, identifies patterns, and produces taxonomies that can be directly applied, allowing the system to serve itself rather than requiring human scoring operations.
2Measurement precision
If manual taxonomy development is used, then precision and control over classifier selection is improved, but productivity and speed of analysis decrease
Solution Approach 1:
The patent substitutes manual precision judgment with automated computational methods that use defined algorithms for identifying topic terms, sentiment terms, and candidate classifiers. The system applies consistent computational rules to ensure precise classifier selection while processing documents at machine speed, achieving both precision and high productivity.
Solution Approach 2:
The system performs preliminary processing by automatically pre-identifying topic terms and sentiment terms before final classifier selection. This preliminary action organizes and pre-sorts document content, making the subsequent classifier identification process faster and more efficient while maintaining precision through structured analysis.
3Productivity
If automated classification methods are used, then speed of analysis and productivity is improved, but complexity of the system and difficulty of implementation increase
Solution Approach 1:
The patent segments the complex taxonomy development process into distinct manageable stages: topic term identification, sentiment term identification, candidate classifier generation, and final classifier selection. Each stage has specific computational operations and criteria, breaking down the overall complexity into sequential steps that are easier to implement and maintain.
Solution Approach 2:
The system manages complexity by allowing adjustment of key parameters such as topic threshold distance, sentiment threshold distance, and colocation matrix thresholds. These parameter changes enable the system to adapt to different document collections and requirements without fundamentally altering the core algorithmic structure, simplifying implementation while maintaining high productivity.
4Reliability
If human scoring is used for taxonomy creation, then accuracy and reliability of classifier identification is improved, but ease of operation and scalability decrease
Solution Approach 1:
The patent replaces manual human scoring with an automated system that scales efficiently. The computational method can process large volumes of documents automatically, enabling the system to scale from small to large datasets without requiring proportional increases in human resources, thereby achieving both accuracy and scalability.
Solution Approach 2:
The system provides universal applicability by using general computational methods that can be applied across different document collections, industries, and topics. The same core algorithmic framework handles various scenarios, making the system versatile and scalable without requiring topic-specific manual scoring procedures.
Data Source
AI summary
Systems and methods are provided for generating a set of classifiers. A location is determined for each instance of a topic term in a collection of documents. One or more topic term phrases are identified, and one or more sentiment terms within each topic term phrase. Candidate classifiers are identified by parsing words in the one or more topic term phrases, and a colocation matrix is generated. A seed row of the colocation associated with a particular attribute is identified, and distance metrics are determined by comparing each row of the colocation matrix to the seed row. A set of classifiers are generated for the particular attribute, where classifiers in the set of classifiers are selected using the distance metrics.


