Valuable Word Net Extraction via Weighted Semantic Linking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for extracting valuable words from text, such as keyword extraction, often miss non-keywords but useful words and can be manipulated by factors like buzzwords or mixed languages, and fail to systematically combine keywords for effective application.
Innovation Solution
A system that autonomously collects and analyzes various texts using machine learning to extract valuable words, weights them based on usage metrics, and links them to form a valuable word net, adjusting weights by time, region, and field for improved relevance and application.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional keyword extraction methods are used, then the extraction process is simple, but valuable non-keyword words are missed and results can be manipulated by buzzwords
Solution Approach 1:
The patent segments the extraction process into multiple independent modules: a valuable word extraction module that identifies candidate words, a weighting module that calculates relevance scores, and a filtering module that removes manipulated terms. This segmentation allows each module to specialize in one aspect of extraction accuracy while maintaining manageable system complexity through modular architecture.
Solution Approach 2:
The patent introduces intermediary components including a pre-trained language model that acts as a mediator between raw text and extraction rules, and a weighting mechanism that mediates between word frequency and actual relevance. These intermediaries enable the system to capture nuanced semantic value beyond simple keyword matching without requiring complete redesign of the extraction pipeline.
2Extent of automation
If only machine learning is used to extract keywords, then the process is automated, but non-keyword useful words are missed and manipulation resistance is reduced
Solution Approach 1:
The patent merges multiple extraction strategies including statistical methods, rule-based filtering, and machine learning models into a unified hybrid system. This combination allows the system to automatically process text while maintaining resistance to manipulation through cross-validation of multiple approaches, where each method compensates for the weaknesses of others.
Solution Approach 2:
The extraction system uses a composite approach combining different algorithmic 'materials': frequency-based word scoring, semantic similarity metrics, and manipulation detection rules. This composite structure creates a robust extraction system that maintains reliability under automated operation by integrating diverse detection mechanisms that collectively resist various forms of text manipulation.
3Productivity
If valuable words are extracted without systematic sorting, then extraction is fast, but keywords cannot be effectively combined for application
Solution Approach 1:
The patent performs preliminary sorting and organization of extracted valuable words during the extraction phase itself, assigning categories, relevance scores, and relationship tags to each word. This preliminary action ensures that words are ready for effective combination and application without requiring additional processing steps, thus maintaining extraction speed while enabling versatile downstream use.
Solution Approach 2:
The patent adds dimensional organization to the extracted words by categorizing them across multiple dimensions such as semantic category, relevance level, and contextual relationship. This multi-dimensional sorting structure enables effective combination of keywords for various applications while maintaining extraction efficiency through parallel processing of different dimensional assignments.
Data Source
AI summary
Method and system for extracting valuable words and forming a valuable word net, wherein a server is employed to collect contents of articles from internet sources, EDM texts (email direct marketing), product descriptions and other texts, and extract valuable words by machine learning. Each valuable word performs connection weight based on the number of times the text has been read, the number of times the text has been clicked, the number of times the text has been cited, the correlation of external sites, the conversion of expert knowledge, the probability space, the Shannon entropy, the spatial distribution and other values and their conversion by machine learning; then the linked valuable words are integrated to form a valuable word net. When the valuable words are needed to be used, the valuable words and the valuable word net can be retrieved from the database for subsequent applications.


