Named Entity Vector Retrieval for Unseen Question Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The accuracy of information retrieval using machine learning is low, particularly in responding to unseen questions, necessitating improved methods for enhancing the precision of information retrieval tasks.
Innovation Solution
An information processing apparatus and method utilizing deep learning to extract named entities from questions and documents, generate vectors based on these entities, calculate similarity, and retrieve relevant documents using a retrieval system that includes an extraction module, encoders, and a similarity calculation unit to enhance the accuracy of document retrieval.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If machine learning is used for information retrieval, then automation is improved, but measurement precision deteriorates
Solution Approach 1:
The patent segments the information retrieval process into distinct components: named entity extraction from queries and documents, vector generation for extracted entities, similarity calculation between vectors, and document ranking based on similarity scores. This segmentation allows each component to be optimized independently, improving overall retrieval accuracy while maintaining automation.
Solution Approach 2:
The patent introduces named entities as intermediary elements between the query and document matching process. By extracting named entities from both the query and documents, and calculating similarity between these entities, the system creates an intermediate representation that bridges the gap between automated processing and precise matching, thereby improving retrieval accuracy.
2Measurement precision
If named entity extraction and vector generation are implemented, then measurement precision is improved, but device complexity increases
Solution Approach 1:
The patent employs a universal named entity extraction module that can identify multiple types of entities (persons, organizations, locations, etc.) across different documents and queries using the same mechanism. The vector generation component also serves multiple functions by creating representations for both query entities and document entities, which are then used in similarity calculation. This multi-functionality reduces the need for separate specialized components, thereby managing complexity while improving precision.
Data Source
AI summary
According to one embodiment, an apparatus includes: an interface circuit configured to receive first data items respectively relating to documents and a second data item relating to a question; and a processor configured to process the first and second data items, wherein the processor is configured to: extract first named entities respectively from the first data items and extract a second named entity from the second data item; generate first vectors respectively relating to the first data items and the corresponding first named entities; generate a second vector relating to the second data item and the second named entity; calculate a similarity between each of the first vectors and the second vector; and acquire a third data item relating to an answer retrieved from the first data items based on a result of calculating the similarity.


