Dynamic Word Amplification for Accurate Phrase Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for calculating relationships between words are inadequate when dealing with limited data, leading to biased associations and the need for extensive phrase preparation across multiple fields.
Innovation Solution
An inter-word score calculation apparatus and method that adjusts the degree of relatedness calculation based on the amount of accumulated data, using a word combination unit to amplify candidate words and calculate scores, thereby extracting suitable related phrases.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If amplification of keywords from term lists is used to make associations, then associations between highly related keywords can be made in specific fields, but an enormous amount of phrases must be prepared and associations become biased when data amount increases
Solution Approach 1:
The patent dynamically adjusts the degree of amplification based on the amount of accumulated data. When data amount is small, amplification is strong to ensure reliable associations. When data amount increases, amplification is reduced to prevent bias. This dynamic adjustment resolves the contradiction by adapting the system behavior to data availability without requiring manual phrase preparation.
Solution Approach 2:
The patent changes the amplification parameter based on data amount. By monitoring the number of accumulated documents and adjusting the amplification degree accordingly, the system maintains association accuracy across different data volumes without requiring extensive phrase preparation or manual maintenance.
2Reliability
If amplification of keywords is used to make associations, then associations can be made with limited data, but associations become biased towards list of terms when data amount increases
Solution Approach 1:
The patent dynamically reduces amplification strength as data amount increases. This prevents the system from becoming overly dependent on predefined term lists when sufficient data is available, thereby reducing bias while maintaining reliability in the early stages when data is limited.
Solution Approach 2:
The system uses feedback from the actual data to adjust amplification. By monitoring data accumulation and responding by reducing amplification strength, the system automatically corrects the harmful bias effect without requiring manual intervention to maintain term list accuracy.
3Reliability
If term list data is continuously maintained to prevent bias, then associations remain accurate, but the system becomes complex and maintenance-intensive
Solution Approach 1:
The patent implements self-service by automatically adjusting amplification based on data amount without requiring manual maintenance of term lists. The system monitors its own data accumulation and self-regulates its behavior, eliminating the need for continuous human intervention while maintaining association accuracy.
Data Source
AI summary
An inter-word score calculation apparatus calculates a degree of relatedness between words included in an amount of data from at least one document. The inter-word score calculation apparatus includes a memory storing document data from the documents, term list data wherein predetermined terms are written and a processor. The processor performs a combination process of amplifying an amplification candidate word, which is a word corresponding to a term in the term list data and included in the document data, creating an amplified word, and adding the amplified word to the document data creating processed document data, calculate the degree of relatedness between words included in the processed document data using a predetermined calculation method, and when an amount of documents accumulated in the document data is smaller than a first predetermined amount, add the amplification candidate word to the processed document data.


