Short text clustering and hotspot theme extraction method based on TF-IDF characteristics
A TF-IDF and extraction method technology, applied in the field of digital text mining, can solve problems such as complexity, unbalanced samples, and high complexity of clustering algorithms, and achieve the effect of supporting decision-making
- Summary
- Abstract
- Description
- Claims
- Application Information
AI Technical Summary
Problems solved by technology
Method used
Image
Examples
Embodiment Construction
[0044] To make the purpose, technical solution and advantages of the present invention more clear and understandable, the embodiments of the present invention will be further described in detail below in conjunction with the accompanying drawings.
[0045] like figure 1 Shown, the overall flow process of the present invention is described in detail as follows:
[0046] Step 1: Use the forward maximum matching method to perform Chinese word segmentation on all samples, and then sum the frequency of occurrence of all words to find the total word frequency of all words, and then divide all words according to their frequency of occurrence from large to small Sorting starts from the word with the largest word frequency and selects words in the order of decreasing word frequency until the ratio of the word frequency of the selected word to the total word frequency reaches 9:10, then stop. At this point, high-frequency words with higher frequency are screened out.
[0047] Step 2: U...
PUM
Abstract
Description
Claims
Application Information
- R&D Engineer
- R&D Manager
- IP Professional
- Industry Leading Data Capabilities
- Powerful AI technology
- Patent DNA Extraction
Browse by: Latest US Patents, China's latest patents, Technical Efficacy Thesaurus, Application Domain, Technology Topic, Popular Technical Reports.
© 2024 PatSnap. All rights reserved.Legal|Privacy policy|Modern Slavery Act Transparency Statement|Sitemap|About US| Contact US: help@patsnap.com