Learning Device Semantic Classification Merging Clusters
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional distributed learning techniques face accuracy degradation when the number of input documents is small, particularly in determining semantic classifications, leading to subdivided words and reduced accuracy.
Innovation Solution
A learning device that clusters documents based on common labels assigned to similar clusters, using a processor to acquire and re-cluster documents, and employing a context storage unit to generate and update contexts for each word, thereby ensuring sufficient input documents for accurate classification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional distributed learning techniques are used with a small number of input documents, then the processing speed is maintained, but the accuracy of semantic classification deteriorates
Solution Approach 1:
The patent merges clusters that have similar distribution characteristics together. By calculating distribution differences between clusters and grouping similar clusters, the system effectively combines the document sets associated with each cluster. This merging allows the system to achieve accurate semantic classification even with a small number of input documents, as the combined document sets provide sufficient statistical basis for learning.
2Manufacturing precision
If clusters are subdivided based on concept names, then the semantic classification becomes more detailed, but the number of input documents for each cluster decreases leading to reduced accuracy
Solution Approach 1:
The patent addresses this contradiction by merging clusters with similar distribution characteristics. When clusters are subdivided into detailed semantic categories, the system calculates the distribution difference between these clusters and merges those with small differences. This ensures that each merged cluster contains sufficient input documents while still maintaining detailed semantic classification through the hierarchical structure of merged and unmerged clusters.
3Measurement precision
If the number of clusters is increased for more fine-grained classification, then the classification precision improves, but the accuracy of distributed learning deteriorates due to insufficient documents per cluster
Solution Approach 1:
The patent resolves this contradiction by dynamically merging clusters based on distribution similarity. The system calculates distribution differences between all clusters and merges those with small differences, creating a hierarchical structure where fine-grained clusters are merged into coarser clusters when their distributions are similar. This ensures that each leaf cluster in the hierarchy contains sufficient documents for reliable learning while maintaining high classification precision through the detailed hierarchical structure.
Data Source
AI summary
A learning device includes a memory and a processor configured to execute a process including acquiring a plurality of documents, clustering the plurality of documents with respect to each of a first plurality of words, the first plurality of words being included in the plurality of documents, assigning a common label to a first word and a second word among the first plurality of words in a case where a cluster relating to the first word and a cluster relating to the second word resemble each other, and re-clustering, on the basis of the common label, the plurality of documents including the first word and the second word after the assigning the common label.


