Korean Word Embedding Library Generation via Morpheme Tagging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current word embedding tools based on English fail to accurately embed Korean due to its complex grammatical structure, resulting in low accuracy.
Innovation Solution
A method and apparatus for generating a word embedding library that segments Korean text by morpheme, combines them according to preset rules, and matches tags based on morphological and syntactic attributes to classify and vectorize words, considering the unique structure of Korean language.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If word embedding tools based on English structure are used for Korean, then the tool can be applied, but the embedding accuracy is very low
Solution Approach 1:
The patent changes the structural parameters of word embedding from English-based sequential arrangement to Korean-based hierarchical aggregation. It modifies how words are vectorially combined by considering Korean's matrix sentence structure, embedded sentences, and suffix-based part of speech changes, thereby improving embedding accuracy for Korean language while maintaining tool applicability
Solution Approach 2:
Instead of adapting Korean to English embedding structure, the patent inverts the approach by designing embedding structure that adapts to Korean's unique grammatical characteristics. It reverses the conventional wisdom by making the embedding tool conform to Korean language structure rather than forcing Korean to conform to English structure
2Measurement precision
If Korean text is segmented and processed by morpheme combination, then the grammatical structure can be captured, but the processing complexity increases
Solution Approach 1:
The patent segments Korean text into morphemes and processes them through step-by-step combination according to preset rules. This segmentation allows the system to capture grammatical structures by analyzing how morphemes combine to form words with different parts of speech and syntactic roles, thereby improving structural capture accuracy
Solution Approach 2:
The patent applies preliminary action by pre-defining combination rules and tag matching criteria before processing Korean text. This allows the complex morpheme combination process to follow systematic predefined patterns, reducing actual processing complexity during implementation
Data Source
AI summary
The present invention relates to a method of generating a word embedding library, including: receiving, by a processor, original text composed of Hangul through an input interface; segmenting, by the processor, the original text by morpheme, combining segmented morphemes step by step according to a preset rule, and matching a tag to a combination of step-by-step morphemes according to a morphological attribute or a syntactic attribute of the combination of step-by-step morphemes; and generating, by the processor, a word embedding library by classifying the morphemes included in the original text based on the tag matched to the combination of step-by-step morphemes.


