Semantic Theme Identification via Stable Marriage Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional topic modeling methods struggle to effectively identify and name semantic themes in electronic documents, as they rely on top keywords that may not accurately represent the themes' semantic content, leading to confusion and difficulty in user interpretation.

Innovation Solution

A computerized method that performs topic modeling to yield a set of topics, then uses a processor to match and name themes by selecting a subset of words from each topic's output list, ensuring each word appears in no more than a predetermined number of subsets, employing algorithms like Stable Marriage and Hospital-Residents matching to optimize theme naming.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional topic modeling methods use top keywords to name themes, then the theme naming process is simple, but the semantic accuracy of theme names deteriorates

Engineering Contradiction:
Improvesemantic accuracy of theme namesVSAvoidtheme naming process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary matching process between words and themes using stable marriage algorithms. Instead of directly using top keywords as theme names, the system creates a matching layer that optimizes the assignment of words to themes based on semantic coherence and user preference, thereby improving semantic accuracy without excessive complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system enables themes to self-select their representative words through the stable marriage matching mechanism. Each theme evaluates its candidate words based on semantic relevance, and the matching algorithm automatically resolves conflicts and assignments, reducing the need for manual intervention while improving naming accuracy

Inventive Principle:
Principle #25Self-service

2Productivity

If multiple words are assigned to multiple themes, then word utilization is maximized, but theme identification accuracy deteriorates

Engineering Contradiction:
Improveword utilization efficiencyVSAvoidtheme identification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent changes the parameter of word-theme assignment from fixed multi-assignment to flexible single-or-multi-assignment based on stable marriage matching results. The system adjusts the assignment parameters dynamically, allowing words to be assigned to one or multiple themes only when the matching algorithm determines this improves overall semantic coherence, thus maintaining both productivity and accuracy

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If greedy algorithms are used for theme naming, then the processing speed is fast, but the quality of theme names deteriorates

Engineering Contradiction:
Improvetheme naming qualityVSAvoidprocessing time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-computing word-theme compatibility scores and preference rankings before the actual matching process. This preliminary preparation enables the stable marriage algorithm to operate more efficiently, reducing the time penalty associated with using more sophisticated matching methods compared to simple greedy approaches

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10002182B2System and method for computerized identification and effective presentation of semantic themes occurring in a set of electronic documents
Publication Date: 2018.06.19 MICROSOFT ISRAEL RES & DEV 2002 LTD
  • US10002182B2 patent drawing
  • US10002182B2 patent drawing
  • US10002182B2 patent drawing

AI summary

System and method for computerized identification and presentation of semantic themes occurring in a set of electronic documents, comprising performing topic modeling on the set of documents thereby to yield a set of topics and for each topic, a topic-modeling output list of words; and using a processor performing a matching algorithm to match only a subset of each topic-modeling output list of words, to the output list's corresponding topic, such that each word appears in no more than a predetermined number of subsets from among said subsets.