Semantic Text Tagging via Gibbs Sampling and Association Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional methods for classifying semantic text data, such as user comments, are inefficient due to reliance on manual marking and struggle with short, colloquial, and scattered information, making it challenging to extract meaningful features and reduce operational costs.

Innovation Solution

A method and device for matching semantic text data with tags by pre-processing the data, determining association and theme based on reproduction relationships, and using Gibbs iterative sampling to establish mapping probability relationships, allowing for automated tagging and classification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual marking is used for text classification, then classification accuracy can be maintained, but processing efficiency deteriorates significantly

Engineering Contradiction:
Improveclassification accuracyVSAvoidprocessing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system enables automatic text classification through self-learning mechanisms. The model automatically processes user comments, extracts features, and performs classification without requiring manual marking for each new text, thereby maintaining accuracy while dramatically improving processing efficiency

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual marking process with an automated machine learning system. The system uses feature extraction, probabilistic modeling (Gibbs sampling), and automated classification algorithms to substitute human manual classification work, achieving both accuracy and efficiency

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Device complexity

If traditional feature extraction methods are used for short text, then processing simplicity is maintained, but feature extraction precision deteriorates due to colloquial language and scattered information

Engineering Contradiction:
Improveprocessing simplicityVSAvoidfeature extraction precision
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The system transforms the feature extraction process by changing parameters from traditional TF-IDF weighting to a probabilistic framework based on Gibbs sampling. This allows the system to capture semantic relationships and contextual information in short, colloquial text more effectively, improving feature extraction precision while managing complexity through systematic probabilistic modeling

Inventive Principle:
Principle #35Parameter changes

3Stability of the object's composition

If a fixed sample classification tag system is used, then classification consistency is improved, but adaptability to new topics and expressions deteriorates

Engineering Contradiction:
Improveclassification consistencyVSAvoidadaptability to new topics
Core Design Contradiction:
Stability of the object's compositionVSAdaptability or versatility

Solution Approach 1:

The system implements dynamic classification through probabilistic modeling. Instead of rigid fixed tags, the system uses probability distributions to represent topic assignments, allowing flexible adaptation to new topics and expressions while maintaining overall classification consistency through the structured probabilistic framework

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent creates a universal classification system that can handle diverse text types and topics. The probabilistic topic model framework is adaptable to various domains and can incorporate new topics without requiring complete system redesign, achieving both consistency and versatility

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11586658B2Method and device for matching semantic text data with a tag, and computer-readable storage medium having stored instructions
Publication Date: 2023.02.21 CHINA UNIONPAY
  • US11586658B2 patent drawing
  • US11586658B2 patent drawing
  • US11586658B2 patent drawing

AI summary

A method for matching semantic text data with tags. The method includes: pre-processing multiple semantic text data to obtain original corpus data comprising multiple semantic independent members; determining the degree of association between any two of the multiple semantic independent members according to a reproduction relationship of the multiple semantic independent members in a natural text, determining a theme corresponding to the association according to the degree of association between any two, and thus determining a mapping probability relationship between the multiple semantic text data and the theme; selecting one of the multiple semantic independent members corresponding to the association as a tag of the theme, and mapping the multiple semantic text data to the tag according to the determined mapping probability relationship between the multiple semantic text data and the theme; and taking the determined mapping relationship between the multiple semantic text data and the tag as a supervision material, and matching the unmapped semantic text data with the tag according to the supervision material.