Unsupervised Taxonomic Tree Generation via Vector Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Unsupervised computing systems struggle to accurately understand human language due to the lack of human-provided meaning, leading to decreased performance in data searches, product recommendations, and other computerized services.
Innovation Solution
A taxonomic tree is generated in an unsupervised manner by extracting hierarchical structures from documents, embedding categories as multidimensional vectors, and grouping them to create category clusters, which are then used to refine domain-specific topics without human supervision.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If human-supervised techniques are used to provide underlying meaning, then understanding accuracy is improved, but cost and scalability deteriorate
Solution Approach 1:
The system performs self-service by automatically generating taxonomic trees and category structures without human intervention. The unsupervised learning algorithms autonomously extract categories from documents, embed them as vectors, group them into clusters, and generate taxonomic trees, eliminating the need for expensive human supervision while maintaining understanding accuracy
Solution Approach 2:
The patent replaces the mechanical human supervision process with an automated computational system. Human experts manually creating category structures is substituted with unsupervised learning algorithms that automatically perform category extraction, embedding, clustering, and taxonomic tree generation, achieving both accuracy and scalability
2Productivity
If unsupervised techniques are used for topic extraction, then scalability is improved, but understanding accuracy deteriorates
Solution Approach 1:
The unsupervised process is segmented into four distinct stages: category extraction from documents, embedding categories as multidimensional vectors, grouping vectors into clusters using similarity conditions, and generating taxonomic trees from clusters. This segmentation allows each stage to be optimized independently while maintaining overall accuracy
Solution Approach 2:
The patent introduces multidimensional category vectors as an intermediary representation between raw text categories and final taxonomic structures. These vectors capture semantic relationships and enable accurate clustering, bridging the gap between unsupervised processing and human-level understanding accuracy
3Measurement precision
If human experts provide category structures, then topic extraction accuracy is improved, but bias from single supervisors is introduced
Solution Approach 1:
The system eliminates supervisor bias by performing self-service through unsupervised learning. Multiple algorithms work together autonomously to extract categories, create vector representations, perform clustering, and generate taxonomic trees, ensuring the process is free from individual human biases while maintaining high accuracy
Solution Approach 2:
The unsupervised system achieves universality by integrating multiple functions: category extraction from diverse document sources, multidimensional embedding to capture semantic relationships, similarity-based clustering, and taxonomic tree generation. This multi-functional approach processes various document types uniformly without human bias
Data Source
AI summary
A computing system generates a taxonomic tree for a domain in an unsupervised manner (e.g., without human intervention). Hierarchical structures of documents of the domain are collected from a document index. A category for each node of each of the hierarchical structures is extracted. The extracted categories are embedded as multidimensional category vectors in a multidimensional vector space. The multidimensional category vectors are grouped into multiple groups. The multidimensional category vectors of a first group satisfy a similarity condition for the first group better than the multidimensional category vectors of a second group. Each group of the multidimensional category vectors constitutes a category cluster. Each category cluster includes multidimensional category vectors for extracted categories from different hierarchical levels of the hierarchical structures. The taxonomic tree is generated with each category cluster inserted as a category node of the taxonomic tree.


