Automated Web Taxonomy Update via Structured Content Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current taxonomic classification methods are hindered by the need for human intervention to label documents, leading to delays and increased costs, as well as inefficiencies in updating taxonomies to reflect changing content.
Innovation Solution
A computer-implemented method that extracts structured content from websites, determines a recent taxonomy by applying category rules, and updates a stored taxonomy by adding new categories, thereby automating the classification process without human intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If human readers are employed to label documents and apply taxonomy categories, then classification accuracy is improved, but time consumption and cost increase significantly
Solution Approach 1:
The system enables automatic self-classification of documents by extracting structured content and applying category rules without human intervention. The taxonomy updater automatically processes documents, determines categories, and updates the taxonomy structure, eliminating the need for human readers to manually label documents while maintaining classification accuracy through automated rule-based processing
Solution Approach 2:
The patent replaces the mechanical human labeling process with an automated computer-based system that extracts structured content from documents, applies predefined category rules, and updates taxonomies automatically. This substitution of manual human operation with automated computational processes reduces time consumption while preserving classification precision through systematic rule application
2Measurement precision
If human readers manually label documents, then taxonomy accuracy is maintained, but the process becomes expensive and slow
Solution Approach 1:
The system performs self-service by automatically extracting structured content from documents and applying category rules without requiring human readers. The taxonomy updater independently processes documents, determines appropriate categories, and updates the taxonomy structure, thereby maintaining taxonomy accuracy while dramatically improving processing speed and eliminating the expensive manual labeling process
Solution Approach 2:
The system performs preliminary actions by pre-extracting structured content from documents and pre-applying category rules before final taxonomy determination. This preliminary processing enables rapid automated taxonomy updates without requiring slow manual review, thereby maintaining accuracy while increasing productivity through preparatory computational steps
3Measurement precision
If taxonomy is updated manually based on new content, then taxonomy remains accurate, but updating process is time-consuming and costly
Solution Approach 1:
The taxonomy updater performs self-service by automatically extracting structured content from websites and applying category rules to determine new categories. The system independently updates the taxonomy structure without human intervention, maintaining accuracy through systematic rule-based analysis while dramatically reducing updating time and eliminating the need for costly manual taxonomy maintenance
Solution Approach 2:
The system incorporates feedback mechanisms by continuously monitoring website content, extracting structured information, and using category rules to determine when taxonomy updates are needed. This feedback loop enables automatic, real-time taxonomy adjustments based on actual content changes, maintaining accuracy while eliminating time-consuming manual review processes
Data Source
AI summary
According to an example implementation, a computer-implemented method may include extracting, by a computing device, structured content from a website, determining a recent taxonomy by applying category rules to the structured content, the recent taxonomy including multiple categories and a new category, and updating a stored taxonomy based on the determined recent taxonomy by adding the new category to the stored taxonomy.


