Unsupervised Taxonomic Tree Generation via Vector Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Unsupervised computing systems struggle to accurately understand human language due to the lack of human-provided meaning, leading to decreased performance in data searches, product recommendations, and other computerized services.

Innovation Solution

A taxonomic tree is generated in an unsupervised manner by extracting hierarchical structures from documents, embedding categories as multidimensional vectors, and grouping them to create category clusters, which are then used to refine domain-specific topics without human supervision.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If human-supervised techniques are used to provide underlying meaning, then understanding accuracy is improved, but cost and scalability deteriorate

Engineering Contradiction:
Improveunderstanding accuracyVSAvoidscalability
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs self-service by automatically generating taxonomic trees and category structures without human intervention. The unsupervised learning algorithms autonomously extract categories from documents, embed them as vectors, group them into clusters, and generate taxonomic trees, eliminating the need for expensive human supervision while maintaining understanding accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical human supervision process with an automated computational system. Human experts manually creating category structures is substituted with unsupervised learning algorithms that automatically perform category extraction, embedding, clustering, and taxonomic tree generation, achieving both accuracy and scalability

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If unsupervised techniques are used for topic extraction, then scalability is improved, but understanding accuracy deteriorates

Engineering Contradiction:
ImprovescalabilityVSAvoidunderstanding accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The unsupervised process is segmented into four distinct stages: category extraction from documents, embedding categories as multidimensional vectors, grouping vectors into clusters using similarity conditions, and generating taxonomic trees from clusters. This segmentation allows each stage to be optimized independently while maintaining overall accuracy

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces multidimensional category vectors as an intermediary representation between raw text categories and final taxonomic structures. These vectors capture semantic relationships and enable accurate clustering, bridging the gap between unsupervised processing and human-level understanding accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If human experts provide category structures, then topic extraction accuracy is improved, but bias from single supervisors is introduced

Engineering Contradiction:
Improvetopic extraction accuracyVSAvoidsupervisor bias
Core Design Contradiction:
Measurement precisionVSObject-generated harmful factors

Solution Approach 1:

The system eliminates supervisor bias by performing self-service through unsupervised learning. Multiple algorithms work together autonomously to extract categories, create vector representations, perform clustering, and generate taxonomic trees, ensuring the process is free from individual human biases while maintaining high accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The unsupervised system achieves universality by integrating multiple functions: category extraction from diverse document sources, multidimensional embedding to capture semantic relationships, similarity-based clustering, and taxonomic tree generation. This multi-functional approach processes various document types uniformly without human bias

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10929439B2Taxonomic tree generation
Publication Date: 2021.02.23 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10929439B2 patent drawing
  • US10929439B2 patent drawing
  • US10929439B2 patent drawing

AI summary

A computing system generates a taxonomic tree for a domain in an unsupervised manner (e.g., without human intervention). Hierarchical structures of documents of the domain are collected from a document index. A category for each node of each of the hierarchical structures is extracted. The extracted categories are embedded as multidimensional category vectors in a multidimensional vector space. The multidimensional category vectors are grouped into multiple groups. The multidimensional category vectors of a first group satisfy a similarity condition for the first group better than the multidimensional category vectors of a second group. Each group of the multidimensional category vectors constitutes a category cluster. Each category cluster includes multidimensional category vectors for extracted categories from different hierarchical levels of the hierarchical structures. The taxonomic tree is generated with each category cluster inserted as a category node of the taxonomic tree.