Co-occurrence Word Data Generation for Freshness and Relevance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing feature word automatic learning systems, such as JP 2010-9307 A, are unable to create a feature word database considering feature words from past web texts, resulting in a lack of freshness and relevance in the generated data.
Innovation Solution
A related data generating apparatus that generates co-occurrence word data from posted data across various periods, using a co-occurrence word data generating unit to identify words with stable appearance frequencies and high relevance to a predetermined keyword, and a related data generating unit to determine the freshness and relevance of these words.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If feature word automatic learning system stores feature words from current web texts only, then the system can maintain simplicity in data collection, but the generated feature word database lacks freshness and relevance
Solution Approach 1:
The patent applies preliminary action by pre-processing and storing feature words from historical web texts before they are needed. The system collects, processes, and stores feature words from multiple time periods in advance, creating a ready-to-use database that can quickly provide fresh and relevant feature words without requiring real-time processing when queries are made.
Solution Approach 2:
The patent implements dynamics by making the feature word database time-aware and adaptive. The system dynamically selects and weights feature words based on their temporal characteristics, adjusting the contribution of different time periods' feature words to optimize both freshness and relevance. This allows the database to adapt to changing requirements over time.
2Reliability
If the system analyzes web texts from multiple time periods, then the freshness and relevance of generated data improve, but the processing complexity and computational resources increase
Solution Approach 1:
The patent applies segmentation by dividing the historical web text data into distinct time periods or epochs. Each segment is processed independently to extract feature words specific to that period, making the overall processing task more manageable. This segmentation allows the system to handle multiple time periods without overwhelming computational complexity.
Solution Approach 2:
The patent uses copying by creating structured representations of feature words from different time periods that can be stored and reused. Instead of re-processing original web texts whenever needed, the system copies and stores extracted feature words in an organized manner, allowing quick retrieval and combination without repeating the expensive text processing operations.
3Loss of information
If the system processes and stores feature words from historical data, then the comprehensiveness of the feature word database improves, but the storage requirements and data management complexity increase
Solution Approach 1:
The patent applies extraction by selectively removing and storing only the essential feature words from historical web texts, rather than storing the complete original texts or all possible word combinations. This extraction process identifies and retains only the most relevant and representative feature words for each time period, significantly reducing storage requirements while maintaining database comprehensiveness.
Solution Approach 2:
The patent implements discarding and recovering by selectively discarding redundant or less important feature words while retaining and recovering the most valuable ones. The system evaluates feature words based on their contribution to freshness and relevance, discarding duplicates or low-value entries while preserving high-quality feature words that provide maximum information value per storage unit.
Data Source
AI summary
Provided are a co-occurrence word data generating unit configured to generate co-occurrence word data including a co-occurrence word that is a vocabulary used along with a predetermined keyword in posted data in all periods and an appearance frequency of the co-occurrence word, from among pieces of posted data posed in a plurality of different periods; and a related data generating unit configured to generate related data including the co-occurrence word as a usual related word when a temporal variation in the appearance frequency of the co-occurrence word is smaller than a first threshold value and the appearance frequency is higher than a second threshold value.


