Korean Word Embedding Library Generation via Morpheme Tagging

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current word embedding tools based on English fail to accurately embed Korean due to its complex grammatical structure, resulting in low accuracy.

Innovation Solution

A method and apparatus for generating a word embedding library that segments Korean text by morpheme, combines them according to preset rules, and matches tags based on morphological and syntactic attributes to classify and vectorize words, considering the unique structure of Korean language.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If word embedding tools based on English structure are used for Korean, then the tool can be applied, but the embedding accuracy is very low

Engineering Contradiction:
Improveapplicability of word embedding toolVSAvoidembedding accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent changes the structural parameters of word embedding from English-based sequential arrangement to Korean-based hierarchical aggregation. It modifies how words are vectorially combined by considering Korean's matrix sentence structure, embedded sentences, and suffix-based part of speech changes, thereby improving embedding accuracy for Korean language while maintaining tool applicability

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

Instead of adapting Korean to English embedding structure, the patent inverts the approach by designing embedding structure that adapts to Korean's unique grammatical characteristics. It reverses the conventional wisdom by making the embedding tool conform to Korean language structure rather than forcing Korean to conform to English structure

Inventive Principle:
Principle #13The other way round (Inversion)

2Measurement precision

If Korean text is segmented and processed by morpheme combination, then the grammatical structure can be captured, but the processing complexity increases

Engineering Contradiction:
Improvegrammatical structure capture accuracyVSAvoidtext processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments Korean text into morphemes and processes them through step-by-step combination according to preset rules. This segmentation allows the system to capture grammatical structures by analyzing how morphemes combine to form words with different parts of speech and syntactic roles, thereby improving structural capture accuracy

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by pre-defining combination rules and tag matching criteria before processing Korean text. This allows the complex morpheme combination process to follow systematic predefined patterns, reducing actual processing complexity during implementation

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12112128B2Apparatus and method for generating word embedding library
Publication Date: 2024.10.08 KOREA ELECTRIC POWER CORP
  • US12112128B2 patent drawing
  • US12112128B2 patent drawing
  • US12112128B2 patent drawing

AI summary

The present invention relates to a method of generating a word embedding library, including: receiving, by a processor, original text composed of Hangul through an input interface; segmenting, by the processor, the original text by morpheme, combining segmented morphemes step by step according to a preset rule, and matching a tag to a combination of step-by-step morphemes according to a morphological attribute or a syntactic attribute of the combination of step-by-step morphemes; and generating, by the processor, a word embedding library by classifying the morphemes included in the original text based on the tag matched to the combination of step-by-step morphemes.