Keyword Extraction Using Concept Dictionary Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing keyword extraction techniques from electronic documents often favor technical terms over basic terms, making it difficult to obtain a comprehensive overview, especially in small-scale document sets, and struggle to balance statistical and conceptual correlations effectively.
Innovation Solution
A keyword presentation apparatus that extracts basic and technical terms using a general concept dictionary, evaluates relevancies, clusters terms based on weighted correlations, and selects representative keywords to present a clear overview, suitable for both large and small-scale document sets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If keyword extraction is based on statistical features such as frequencies of occurrence, then keywords can be extracted from large-scale electronic document sets, but basic terms are not extracted easier than technical terms and the method is not suited to small-scale document sets
Solution Approach 1:
The patent segments the keyword extraction process into two distinct phases: first extracting basic terms using statistical features (frequency, document coverage), then extracting technical terms using co-occurrence relationships with those basic terms. This segmentation allows each extraction method to be optimized for its appropriate purpose and document scale.
Solution Approach 2:
The patent performs preliminary extraction of basic terms before extracting technical terms. The basic terms serve as a foundation or scaffold that guides subsequent technical term extraction through co-occurrence analysis, ensuring that even in small-scale document sets, basic terms are reliably identified first.
2Ease of operation
If keywords are grouped based on co-occurrence relationships between keywords, then keyword groups can be presented to ascertain an overview, but co-occurrence relationships between basic terms having higher frequencies of occurrence are easily determined while technical terms are overlooked
Solution Approach 1:
The patent segments keyword grouping into two separate processes: first grouping basic terms based on their co-occurrence relationships to create basic term groups, then separately grouping technical terms based on their co-occurrence relationships with basic terms. This ensures both basic and technical terms are properly represented in the final keyword groups.
Solution Approach 2:
The patent uses basic terms as an intermediary to connect technical terms to keyword groups. Technical terms are grouped based on their co-occurrence relationships with basic terms, which act as mediators that bridge the gap between statistical frequency-based grouping and technical term identification.
3Productivity
If only extracted keywords are presented, then the user can recognize an overview and perform refined search, but the user cannot easily distinguish between basic terms and technical terms
Solution Approach 1:
The patent applies visual differentiation (analogous to color changes) by displaying basic terms and technical terms in different visual styles or formats. Basic terms may be shown in one style while technical terms are shown in another, allowing users to easily distinguish between the two types of keywords while maintaining search functionality.
Solution Approach 2:
The patent segments the keyword presentation by clearly separating and labeling basic terms from technical terms in the output display. This segmentation helps users understand the nature of each keyword type and how to use them effectively for different search purposes.
Data Source
AI summary
According to one embodiment, a keyword presentation apparatus includes an extraction unit, a selection unit and a clustering unit. The extraction unit is configured to extract, as technical terms, morpheme strings, which are not defined in a general concept dictionary, from a document set. The selection unit is configured to evaluate relevancies between each of basic term candidates and the technical terms, and to preferentially select basic term candidates having high relevancies as basic terms. The clustering unit is configured to calculate weighted sums of statistical degrees of correlation between the basic terms based on the document set, to calculate conceptual degrees of correlation between the basic terms based on the general concept dictionary, and to cluster the basic terms based on the weighted sums.


