Named Entity Label Enrichment for Similarity-Based Word Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for named entity recognition face challenges in accuracy due to limited information in label names and the time-consuming process of creating high-quality label descriptions, leading to inefficiencies in improving recognition accuracy.
Innovation Solution
An information processing apparatus and method that extracts phrases associated with classification names from training data, derives similarities between word representations and classification names, and selects the most appropriate classification names based on these similarities, using trained models to enhance accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If label name is used for machine learning, then the process is simple and quick, but the accuracy of machine learning and named entity recognition is insufficient due to limited information
Solution Approach 1:
The patent introduces a label description generation unit that creates intermediate descriptions bridging the gap between simple label names and detailed entity information. These generated descriptions serve as mediators that enrich the training data without requiring manual preparation, thereby improving accuracy while maintaining automation speed.
Solution Approach 2:
The system performs self-service by automatically generating label descriptions using the label name and related information from the training data itself. This self-generation mechanism eliminates the need for external manual description preparation while enhancing the information available for accurate named entity recognition.
2Measurement precision
If high-quality label description is prepared manually, then the accuracy of machine learning is improved, but the time and effort required increases significantly
Solution Approach 1:
The label description generation unit automatically creates label descriptions by processing the label name and extracting relevant information from the training data. This self-service approach eliminates manual preparation time while generating high-quality descriptions that improve machine learning accuracy.
Solution Approach 2:
The system performs preliminary action by pre-generating label descriptions before the machine learning process begins. The label description generation unit creates enriched descriptions in advance using the training data, so that when the learning process starts, the necessary information is already prepared and available.
3Measurement precision
If more information is added to label data, then the accuracy of named entity recognition improves, but the complexity of data preparation increases
Solution Approach 1:
The system automatically enriches the label data by generating descriptions itself rather than requiring complex manual data preparation processes. The label description generation unit processes the training data to create enriched labels, simplifying the overall data preparation complexity while improving information quality.
Data Source
AI summary
An information processing apparatus extracts a phrase corresponding to a named entity included in an input sentence used for training a natural language processing model from data in which a phrase and a classification name of a named entity are associated with each other, adds the extracted phrase to the classification name corresponding to the extracted phrase among a plurality of different classification names which are set in advance, derives a similarity between a distributed representation of each word included in the input sentence and a distributed representation of each of the plurality of classification names to which the phrase is added, and select a classification name of each word included in the input sentence from among the plurality of classification names based on the derived similarity.


