Learning Device Semantic Classification Merging Clusters

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional distributed learning techniques face accuracy degradation when the number of input documents is small, particularly in determining semantic classifications, leading to subdivided words and reduced accuracy.

Innovation Solution

A learning device that clusters documents based on common labels assigned to similar clusters, using a processor to acquire and re-cluster documents, and employing a context storage unit to generate and update contexts for each word, thereby ensuring sufficient input documents for accurate classification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional distributed learning techniques are used with a small number of input documents, then the processing speed is maintained, but the accuracy of semantic classification deteriorates

Engineering Contradiction:
Improveaccuracy of semantic classificationVSAvoidnumber of input documents
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent merges clusters that have similar distribution characteristics together. By calculating distribution differences between clusters and grouping similar clusters, the system effectively combines the document sets associated with each cluster. This merging allows the system to achieve accurate semantic classification even with a small number of input documents, as the combined document sets provide sufficient statistical basis for learning.

Inventive Principle:
Principle #5Merging (Combining)

2Manufacturing precision

If clusters are subdivided based on concept names, then the semantic classification becomes more detailed, but the number of input documents for each cluster decreases leading to reduced accuracy

Engineering Contradiction:
Improvesemantic classification detailVSAvoidnumber of input documents per cluster
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent addresses this contradiction by merging clusters with similar distribution characteristics. When clusters are subdivided into detailed semantic categories, the system calculates the distribution difference between these clusters and merges those with small differences. This ensures that each merged cluster contains sufficient input documents while still maintaining detailed semantic classification through the hierarchical structure of merged and unmerged clusters.

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If the number of clusters is increased for more fine-grained classification, then the classification precision improves, but the accuracy of distributed learning deteriorates due to insufficient documents per cluster

Engineering Contradiction:
Improveclassification precisionVSAvoidaccuracy of distributed learning
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent resolves this contradiction by dynamically merging clusters based on distribution similarity. The system calculates distribution differences between all clusters and merges those with small differences, creating a hierarchical structure where fine-grained clusters are merged into coarser clusters when their distributions are similar. This ensures that each leaf cluster in the hierarchy contains sufficient documents for reliable learning while maintaining high classification precision through the detailed hierarchical structure.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS10747955B2Learning device and learning method
Publication Date: 2020.08.18 FUJITSU LTD
  • US10747955B2 patent drawing
  • US10747955B2 patent drawing
  • US10747955B2 patent drawing

AI summary

A learning device includes a memory and a processor configured to execute a process including acquiring a plurality of documents, clustering the plurality of documents with respect to each of a first plurality of words, the first plurality of words being included in the plurality of documents, assigning a common label to a first word and a second word among the first plurality of words in a case where a cluster relating to the first word and a cluster relating to the second word resemble each other, and re-clustering, on the basis of the common label, the plurality of documents including the first word and the second word after the assigning the common label.