Keyword Library Construction for Short Text Topic Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional topic clustering analysis on short text user feedback is inefficient due to sparse feature words and high manpower requirements, leading to inaccurate subject term extraction and low work efficiency.
Innovation Solution
A data processing method that constructs a keyword library using a comment corpus, extracts partial comment corpora, and performs topic clustering to obtain subject terms, reducing manual extraction and improving efficiency by applying the topic clustering algorithm directly on short text corpora.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If topic clustering analysis is performed on short text user feedback using conventional methods, then the analysis can be conducted, but the extraction accuracy of subject terms is low and multiple subject terms are obtained
Solution Approach 1:
The patent segments the short text comment corpus into multiple partial comment corpora based on different keywords. Each partial corpus focuses on a specific keyword, allowing topic clustering to be performed on more targeted, less sparse data. This segmentation resolves the contradiction by transforming the original short text into structured partial corpora that maintain information while improving clustering accuracy.
Solution Approach 2:
The patent transforms the analysis from direct topic clustering on short texts to a multi-dimensional approach: first extracting keywords, then creating partial corpora around each keyword, and finally performing topic clustering on these enriched partial corpora. This dimensional transformation allows the system to overcome the sparsity limitation of short texts while maintaining extraction reliability.
2Measurement precision
If traditional manual method is used to extract subject terms from user feedback, then extraction can be performed, but it consumes a lot of manpower and resources with low efficiency
Solution Approach 1:
The patent implements an automated system where the computer automatically performs keyword extraction, partial corpus generation, and topic clustering without manual intervention. The system serves itself by using algorithms to process the comment corpus, replacing manual extraction methods and dramatically improving productivity while maintaining or enhancing extraction accuracy through systematic processing.
Solution Approach 2:
The patent replaces the mechanical manual extraction process with an automated computational system. Instead of human analysts manually reading and extracting subject terms, the system uses keyword extraction algorithms, corpus segmentation, and topic clustering algorithms to automatically identify and extract subject terms, thereby substituting mechanical human labor with automated processing that is both efficient and accurate.
3Quantity of substance
If long text such as news reports is used for topic clustering, then multiple subject terms can be obtained, but it is difficult to determine the users' focus
Solution Approach 1:
The patent applies local quality by creating different partial comment corpora for different keywords, where each partial corpus has specialized focus on a specific keyword. This allows the system to obtain multiple subject terms (one for each keyword) while maintaining precision because each partial corpus is locally optimized for its specific keyword context, making it easier to identify user focus areas.
Solution Approach 2:
The patent segments the comment corpus into keyword-specific partial corpora, transforming the original approach of analyzing entire long texts into analyzing focused segments. This segmentation enables the system to obtain multiple subject terms from different segments while improving user focus identification because each segment is dedicated to a specific keyword, reducing noise and improving relevance.
Data Source
AI summary
The present disclosure provides a data processing method. The method includes the steps of constructing a keyword library of an obtained comment corpus associated with a target object, the keyword library comprising a plurality of keywords; extracting a plurality of partial comment corpora from the comment corpus, each partial comment corpus of the plurality of partial comment corpora comprising a plurality of comment words including at least one of the plurality of keywords in the keyword library; combining the plurality of partial comment corpora to produce a candidate corpus; and performing a topic clustering process on the candidate corpus of each keyword to obtain a subject term for the target object.


