Glossary Term Identification System for Text Ambiguity Resolution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Text documents often contain terms with multiple meanings, leading to misinterpretation by different readers, which can result in incorrect system design or costly mistakes, especially in system requirements documents.
Innovation Solution
A device or method that analyzes text to identify linguistic units, resolves ambiguities, and determines a set of glossary terms by applying linguistic unit analysis and glossary term analysis techniques, including semantic relatedness scoring to clarify intended meanings.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If terms with multiple meanings are used in text documents, then the document can be more concise and expressive, but misinterpretation by readers increases leading to incorrect system design
Solution Approach 1:
The patent introduces an intermediary system (glossary term identification system) that mediates between the ambiguous terms in the document and the readers. The system automatically identifies potentially ambiguous terms and creates a glossary with disambiguated definitions, serving as an intermediate layer that clarifies meanings without changing the original document text.
Solution Approach 2:
The system performs preliminary analysis of the document text to identify ambiguous terms before the reading or processing occurs. By pre-identifying terms that may have multiple meanings and preparing disambiguated versions in advance, the system prevents misinterpretation from occurring in the first place.
2Reliability
If a glossary is manually created to clarify terms, then interpretation accuracy improves, but time consumption and labor requirements increase
Solution Approach 1:
The system enables self-service by automatically identifying ambiguous terms and generating glossary entries without requiring manual intervention. The automated process analyzes the document context, identifies terms with multiple possible meanings, and creates appropriate definitions based on the document's content and structure.
Solution Approach 2:
The patent replaces the mechanical manual process of creating glossaries with an automated computational system. Instead of requiring human analysts to manually review and create glossary entries, the system uses algorithms to automatically process the document text, identify ambiguous terms, and generate glossary content.
3Quantity of substance
If all possible terms are analyzed to ensure complete glossary coverage, then comprehensiveness improves, but processing complexity and computational resources increase
Solution Approach 1:
The system applies local quality by focusing the analysis only on specific portions of the text that are likely to contain ambiguous terms, rather than uniformly processing the entire document. It identifies and analyzes terms based on their contextual characteristics and potential for ambiguity, concentrating computational resources where needed.
Solution Approach 2:
The system uses partial action by analyzing only the necessary portions of the text to identify ambiguous terms, rather than exhaustively processing every possible term. It applies filtering and selection criteria to focus on terms that are most likely to be ambiguous based on their usage patterns and contextual information.
Data Source
AI summary
A device may obtain text to be analyzed to identify glossary terms. The device may analyze a linguistic unit to generate multiple linguistic units related to the linguistic unit. The device may analyze the multiple linguistic units to generate potential glossary terms. The device may perform a glossary term analysis on the potential glossary terms to generate glossary terms that include a subset of the potential glossary terms. The device may identify included terms that are included in the glossary terms. The device may identify excluded terms that are excluded from the glossary terms. The device may determine a semantic relatedness score between at least one excluded term and at least one included term. The device may selectively add the excluded linguistic term to the glossary terms to form a final set of glossary terms based on the semantic relatedness score, and may output the final set of glossary terms.


