Automated Term Similarity Determination via Semantic Vector Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for determining synonyms and polysemes in documents are inaccurate due to reliance on human intervention, which introduces variations, omissions, and errors, and fail to effectively classify argument-predicate pairs as synonymous or antonymous.
Innovation Solution
A computer-based method and apparatus that performs semantic analysis on sentences to generate semantic structures, extracts characteristic vectors from these structures, and uses machine learning to determine similarity between terms, improving the accuracy of synonym and polyseme detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If human intervention is used to determine synonyms and polysemes, then flexibility in handling complex cases is maintained, but accuracy deteriorates due to variations, omissions, and errors
Solution Approach 1:
The patent replaces human manual determination with an automated information processing system that uses semantic analysis and machine learning algorithms to identify synonyms and polysemes, eliminating human variability and improving consistency
Solution Approach 2:
The system performs self-learning through training on labeled data, automatically improving its determination accuracy without requiring continuous human intervention, thereby maintaining high reliability while processing large volumes of text
2Productivity
If manual checking is performed for synonym and polyseme determination, then attention to contextual nuances is maintained, but productivity deteriorates due to time-consuming processes
Solution Approach 1:
The system continuously processes text data through automated semantic analysis and comparison, maintaining constant operation to identify synonyms and polysemes without the interruptions and time limits inherent in manual checking
Solution Approach 2:
The patent introduces semantic structures and characteristic vectors as intermediary representations that bridge the gap between raw text and determination results, enabling fast automated comparison while preserving contextual accuracy
3Measurement precision
If existing classification methods are used for argument-predicate pairs, then processing simplicity is maintained, but measurement precision deteriorates in classifying synonymous and antonymous relationships
Solution Approach 1:
The patent segments the semantic analysis into distinct components: semantic structure generation, characteristic vector extraction, and machine learning classification, allowing each component to be optimized independently while improving overall precision
Solution Approach 2:
The system transforms textual semantic information into numerical characteristic vectors with specific dimensions and weightings, enabling precise machine learning classification of argument-predicate relationships that was not possible with traditional methods
Data Source
AI summary
A determination method executed by a computer including a memory and a processor coupled to the memory, includes receiving a plurality of sentences and designation of terms included in the plurality of sentences, generating, for each term for which the designation is received, information indicating a relation between the term and each of other terms included in one of the plurality of sentences containing the term, extracting, for each term for which the designation is received, information indicating a specific relation from the generated information indicating the relation, generating characteristic information that uses the extracted information as a feature, and determining similarity between a plurality of the terms based on the generated characteristic information for each term.


