Automated Glossary Creation via Syntactic Structure Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Creating large glossaries manually is impractical due to the need to account for numerous variations in language, typographical errors, and interdependencies between parts, especially in complex domains like automotive warranty claims, which requires significant human effort and intervention.
Innovation Solution
An automated method for creating glossaries by identifying canonical and variant forms of glossary items, defining syntactic structures, and searching information sources to extract semantic classes such as symptoms, causes, and actions, using syntactic analysis and rule-based systems to convert unstructured text into structured data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If manual glossary creation is used, then human intervention and control are maintained, but the process becomes impractical and time-consuming for large glossaries with numerous variations
Solution Approach 1:
The glossary creation process is segmented into distinct automated stages: extracting candidate phrases from unstructured text, clustering similar phrases into canonical forms, and generating syntactic structures. This segmentation allows complex variation handling to be broken down into manageable automated steps, resolving the contradiction between automation extent and complexity management.
Solution Approach 2:
The system performs self-service by automatically extracting glossary items from unstructured text sources, clustering variations into canonical forms, and generating syntactic structures without requiring manual intervention. This self-service capability enables the system to handle large glossaries with numerous variations independently, achieving high automation while managing complexity through automated algorithms.
2Reliability
If comprehensive glossary coverage is achieved through manual creation, then all language variations and typographical errors can be accounted for, but the time and effort required become excessive
Solution Approach 1:
The mechanical manual process of reviewing and creating glossary entries is replaced with automated computational processes. The system uses text extraction algorithms, clustering algorithms, and syntactic analysis to automatically generate comprehensive glossary coverage, achieving high reliability in capturing language variations and typographical errors without the time investment required for manual creation.
Solution Approach 2:
The system changes the parameters of glossary creation from manual review to automated processing. By transforming the creation process into computational operations that can process large volumes of text rapidly, the system achieves comprehensive coverage of language variations and typographical errors while significantly reducing the time required compared to manual methods.
3Productivity
If automated phrase extraction is used, then glossary creation efficiency improves, but human intervention remains necessary for administration and rule management
Solution Approach 1:
The system achieves self-service by automatically extracting phrases from unstructured text, clustering them into canonical forms, and generating syntactic structures without requiring human intervention in administration or rule management. This complete automation eliminates the need for human involvement in glossary administration while maintaining high productivity in glossary creation.
Data Source
AI summary
A method and device for creating a glossary includes a processor operable for executing computer instructions for identifying, in at least one information source, at least one glossary item identifying a part or a component, determining at least one glossary item form as a canonical form, defining, by using the canonical form, at least one syntactic structure, that includes one of the at least one identified glossary items, for each of at least one semantic classes, and searching a second information source for the at least one syntactic structure of the semantic class.


