Automated Keyword Classification via Semantic Similarity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing content analysis systems require manual entry of search keywords and synonyms, making it costly and inefficient to generate databases, as they lack automated methods to determine keyword associations with categories based on user input.
Innovation Solution
An analyzing apparatus that calculates semantic similarities between words and applies them to template sentences to classify words into categories, using a model trained on large datasets to automate the process of determining keyword associations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual entry of search keywords and synonyms is used, then keyword coverage can be achieved, but the cost and time required to generate the database increases significantly
Solution Approach 1:
The system automatically generates search keywords and synonyms by analyzing user search queries and content data without requiring manual developer intervention. The keyword generation apparatus autonomously processes search logs, extracts meaningful keywords, generates synonyms using NLP techniques, and updates the search database automatically, enabling the system to serve itself in maintaining keyword coverage.
Solution Approach 2:
The patent replaces the manual mechanical process of keyword entry with an automated computational system. Instead of developers manually thinking about and entering search keywords, the system uses natural language processing, machine learning models, and automated text analysis to generate keywords and synonyms programmatically, substituting human cognitive and manual labor with automated algorithms.
2Reliability
If manual entry of search keywords is used, then keyword associations can be established, but the complexity and cost of the process increases
Solution Approach 1:
The system introduces an intermediary automated keyword generation apparatus that mediates between user search queries and the search database. This intermediary component analyzes search patterns, extracts keywords, generates synonyms, and establishes associations automatically, serving as a bridge that eliminates the need for direct manual intervention while maintaining association accuracy through systematic computational processes.
Solution Approach 2:
The patent utilizes parameter changes in the form of adjusting similarity thresholds, frequency cutoffs, and weighting parameters in the keyword extraction and synonym generation algorithms. By dynamically adjusting these parameters based on analysis of search data and content characteristics, the system optimizes keyword association accuracy without requiring manual tuning or complex configuration procedures.
3Adaptability or versatility
If automated similarity calculation is used, then keyword expansion can be improved, but the need for training data and model computation increases
Solution Approach 1:
The system performs preliminary action by pre-processing and storing search query data, user behavior patterns, and content metadata in structured formats before keyword generation is needed. This preliminary organization of data enables the automated keyword expansion to proceed efficiently using pre-computed statistics and patterns, reducing the computational burden during actual keyword generation operations.
Solution Approach 2:
The patent applies partial action by generating keywords and synonyms only for specific categories or subsets of content based on priority, frequency, or relevance criteria. Instead of attempting to expand all keywords across the entire database simultaneously, the system focuses computational resources on high-priority areas, achieving effective keyword expansion with reduced data processing requirements.
Data Source
AI summary
An analyzing apparatus according to an embodiment includes a memory and a hardware processor coupled to the memory. The hardware processor is configured to: calculate a first similarity between a first word representing a category and a second word; apply, to one or more template sentences, each of one or more second words having the first similarity higher than a first threshold; analyze the one or more template sentences each including the second word and classify the second word into one or more first categories; and display, for each of the first categories, the second word used in the classification into the first categories on a display device.


