Hierarchical Taxonomy Generation via Density-Based Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current cluster analysis procedures struggle to automatically organize terms into complex and informative structures, often requiring human intervention for labeling and specifying the number of output clusters.

Innovation Solution

A computerized system and method that automatically cluster entities by calculating generality scores, selecting exemplars, and clustering unselected nodes to produce a hierarchical taxonomy, while handling outlier points and categorizing interactions among remotely connected computers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If current cluster analysis procedures are used to group terms according to similarity measures, then terms can be grouped into clusters, but the clusters require human intervention for labeling and specifying the number of output clusters

Engineering Contradiction:
Improveautomatic clusteringVSAvoidhuman intervention required
Core Design Contradiction:
Extent of automationVSEase of operation

Solution Approach 1:

The patent implements self-service by enabling the clustering system to automatically determine the number of clusters through density-based algorithms and autonomously generate meaningful labels using representative entity selection, eliminating the need for human operators to specify cluster count or provide manual labeling

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces intermediary mechanisms including density estimation functions and representative entity selection algorithms that mediate between raw similarity measurements and final cluster structures, automatically translating similarity data into labeled hierarchical clusters without human intervention

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If simple clustering algorithms are used to group entities, then the process is computationally efficient, but the resulting structures lack complexity and informativeness

Engineering Contradiction:
Improveclustering speedVSAvoidsemantic relationships
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent applies segmentation by dividing the clustering process into distinct computational stages: initial cluster formation using efficient similarity measures, subsequent hierarchical refinement through density-based operations, and final labeling through representative entity selection, allowing each stage to optimize for its specific computational task

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces additional dimensional information by computing density estimates and hierarchical relationships alongside traditional similarity measures, enabling the system to capture complex semantic structures without sacrificing computational efficiency through multi-dimensional feature integration

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Loss of information

If the system attempts to automatically generate complex hierarchical taxonomies, then more informative structures are produced, but the computational complexity and processing time increase

Engineering Contradiction:
Improvesemantic hierarchyVSAvoidalgorithm complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent implements preliminary action by pre-computing similarity matrices and density estimates before the actual clustering process, and by pre-selecting representative entities for potential labeling roles, thereby reducing the computational burden during the hierarchical taxonomy generation phase

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies local quality by applying different computational strategies to different parts of the clustering process: using efficient distance-based methods for initial cluster formation, density-based operations for hierarchical refinement, and localized representative entity selection for labeling, optimizing computational resources for each specific task

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20250028752A1Systems and methods for agglomerative clustering
Publication Date: 2025.01.23 NICE LTD
  • US20250028752A1 patent drawing
  • US20250028752A1 patent drawing
  • US20250028752A1 patent drawing

AI summary

A computerized system and method may provide a robust, automated clustering procedure, including handling of outlier points, which may involve measuring and/or quantifying degrees of relevance and/or generality for a plurality of input entities. In some embodiments, a clustering procedure may be used, e.g., to generate a hierarchical, multi-tiered taxonomy of such entities. In some embodiments, a computerized system comprising a processor, and a memory, may be used for calculating a distance between nodes for each of a plurality of pairs of nodes, where the pairs may comprise a plurality of input entities and/or initial clusters; selecting one or more of the pairs based on the calculated distances; and merging one or more of the selected pairs, which may include a common node, into one or more final clusters. Some embodiments of the invention may allow routing interactions between remotely connected computer systems based on an automatically generated taxonomy.