Wikipedia Question Answering via Multi-Document Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Wikipedia-based question answering systems face ambiguity and information loss when converting natural language to structured knowledge, and struggle to extract precise answers due to the complexity of mapping natural language queries to standardized knowledge base elements.

Innovation Solution

An apparatus and method that extracts and indexes various types of information from Wikipedia, including fulltext, section titles, info-boxes, and category documents, to enable precise text searches and answer extraction, using POS-based index terms and specific question patterns to generate accurate answers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If Wikipedia contents are converted to a standardized knowledge base, then ambiguity of knowledge is reduced, but information loss and distortion occur during the conversion process

Engineering Contradiction:
Improveambiguity reductionVSAvoidinformation loss
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent segments Wikipedia content into multiple document types (fulltext documents, section title documents, info-box documents, category documents, definition statement documents) rather than converting everything to a single standardized knowledge base format. This segmentation allows preserving the original information structure while enabling targeted access to different types of information, thus reducing information loss while still providing ambiguity reduction through structured organization.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension by creating multiple parallel document representations instead of a single knowledge base conversion path. Each document type serves a different function and preserves different aspects of the original Wikipedia content, allowing the system to access information from multiple dimensions simultaneously, thereby avoiding the information loss inherent in single-path conversion.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If natural language questions are converted to standardized knowledge base expressions, then precise answer extraction is enabled, but ambiguity is introduced in the conversion process

Engineering Contradiction:
Improveanswer extraction precisionVSAvoidambiguity
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent segments the question processing into multiple parallel paths, each handling different document types (fulltext, section titles, info-boxes, categories, definition statements). Each path extracts answers independently based on the characteristics of that document type, avoiding the need to convert the entire natural language question into a single standardized knowledge base expression, thus reducing conversion ambiguity while maintaining precision.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by extracting answers from only the relevant document types for each specific question, rather than converting and searching all Wikipedia content. For example, definition statement documents are used for definition-type questions, while info-box documents are used for factual attribute questions, reducing the ambiguity of full conversion while maintaining precise answer extraction for each question type.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If only structured knowledge base information is used for question answering, then answer precision is improved, but information loss occurs from excluding unstructured information

Engineering Contradiction:
Improveanswer precisionVSAvoidinformation loss
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent merges multiple document types with different structures into a unified question answering system. It combines fulltext documents (unstructured), section title documents (semi-structured), info-box documents (structured), category documents (hierarchical), and definition statement documents (structured) into a single system that processes questions through multiple parallel paths, thereby preserving both the precision of structured information and the completeness of unstructured information.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a universal question answering apparatus that handles multiple document types and question types through a unified framework. The system can process different kinds of questions (definition questions, factual questions, descriptive questions) using appropriate document types, making the system multi-functional while avoiding information loss from excluding any document type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Reliability

If Wikipedia content is fully converted to knowledge base format, then knowledge standardization is achieved, but costs and complexity of building the knowledge base increase

Engineering Contradiction:
Improveknowledge standardizationVSAvoidknowledge base building complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent creates simplified copies of Wikipedia content in multiple document types rather than performing a complete and complex conversion to a single standardized knowledge base. Each document type is a targeted copy that preserves specific aspects of the original content needed for different question types, reducing the complexity of knowledge base construction while maintaining standardization where needed.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent segments the knowledge representation into multiple specialized document types rather than creating one comprehensive standardized knowledge base. This segmentation reduces the complexity of building and maintaining the knowledge representation system, as each document type can be processed and maintained independently with simpler rules, while collectively providing comprehensive coverage.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10037381B2Apparatus and method for searching information based on Wikipedia's contents
Publication Date: 2018.07.31 ELECTRONICS & TELECOMM RES INST
  • US10037381B2 patent drawing
  • US10037381B2 patent drawing
  • US10037381B2 patent drawing

AI summary

The present invention is to provide an apparatus for searching information based on Wikipedia's contents comprising: a document converting part extracting fulltext documents, section title documents, info-box documents, category documents and definition statement documents from Wikipedia original documents and generating at least one of Wikipedia documents for questions and answers; a document indexing part analyzing the Wikipedia document for questions and answers, extracting POS-based index terms from the Wikipedia document for questions and answers, and generating a Wikipedia document index for questions and answers; a question analyzing part receiving a natural language question, analyzing a question pattern, an answer pattern and a question focus from the natural language question, and extracting document search keywords; a document searching part performing document search by using the document search keywords from the Wikipedia document index for questions and answers and generating document search result from each Wikipedia document index for questions and answers; an answer extracting part extracting first answers by using information about the question pattern, the answer pattern and the question focus from the document search result; and an answer integrating part integrating and prioritizing the first answer and generating a second answer.