Selective Fact Generation from Table Data in Cognitive Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current QA systems face inefficiencies when ingesting unstructured natural language documents with embedded structured data, such as tables, as they generate numerous irrelevant facts from these data structures, wasting processing resources and failing to focus on facts relevant to the document's content.

Innovation Solution

The proposed solution involves an ingestion engine that identifies and extracts table signatures from unstructured documents, evaluates references to these tables in the natural language text, and generates a prioritization plan to selectively ingest only the most relevant facts based on their importance and frequency of reference.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If the system ingests all facts from structured data portions in unstructured documents, then the quantity of ingested information increases, but processing efficiency decreases and irrelevant facts are generated

Engineering Contradiction:
Improvequantity of ingested factsVSAvoidprocessing efficiency
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent extracts and identifies only the relevant structured data portions (tables) from unstructured documents by analyzing references in the natural language text. The ingestion engine selectively extracts facts from table portions that are actually referenced in the document content, rather than ingesting all facts from all tables. This extraction approach resolves the contradiction by taking out only the necessary facts, maintaining processing efficiency while obtaining sufficient information quantity for QA tasks.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies local quality by differentiating between relevant and irrelevant table portions based on their reference patterns in the natural language text. The system evaluates each table portion's significance by analyzing how it is referenced in the document, assigning different ingestion priorities to different table regions. This selective approach ensures that processing resources are concentrated on locally relevant facts rather than uniformly processing all table data, thus maintaining efficiency while capturing necessary information.

Inventive Principle:
Principle #3Local quality

2Reliability

If the system generates facts from all structured data portions, then comprehensive coverage is achieved, but processing resources are wasted on irrelevant facts

Engineering Contradiction:
Improvecomprehensive coverageVSAvoidprocessing resource waste
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent performs preliminary analysis of the natural language text to identify references to structured data portions before generating facts. The ingestion engine first scans the document content to detect which tables are referenced and how they are used, then selectively generates facts only from those identified table portions. This preliminary action ensures comprehensive coverage of relevant information while avoiding the energy waste of generating facts from unreferenced tables, as the system knows in advance which tables are important.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses feedback from the natural language text references to guide fact generation from structured data. By analyzing how tables are referenced in the document content, the system receives feedback about which table portions are relevant to the document's information gaps and questions. This feedback mechanism ensures that fact generation is directed toward producing relevant facts that address actual information needs, maintaining comprehensive coverage while minimizing resource waste on irrelevant fact generation.

Inventive Principle:
Principle #23Feedback

3Loss of information

If the system processes all table data structures, then no information is missed, but the accuracy of relevant fact identification decreases due to noise

Engineering Contradiction:
Improveinformation completenessVSAvoidfact relevance accuracy
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The patent extracts only the table portions that are referenced in the natural language text, separating relevant structured data from irrelevant table data. By taking out only the referenced table portions for fact generation, the system maintains information completeness regarding relevant facts while removing the noise of unreferenced table data. This extraction process ensures that the fact identification process focuses solely on relevant information, improving accuracy without losing important information.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies local quality by evaluating each table portion's reference patterns to determine its relevance, then processing only those portions with significant reference patterns. The system analyzes the local context of each table reference in the natural language text to assess importance, and selectively processes only those table regions that demonstrate meaningful engagement in the document. This approach maintains completeness of relevant information while improving fact relevance accuracy by excluding irrelevant table data from the fact generation process.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10095740B2Selective fact generation from table data in a cognitive system
Publication Date: 2018.10.09 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10095740B2 patent drawing
  • US10095740B2 patent drawing
  • US10095740B2 patent drawing

AI summary

Mechanisms are provided for ingesting natural language textual content. Ingestion of natural language textual content is initiated and an embedded structured data portion within the natural language textual content is identified. A signature of the structured data portion is generated which comprises one or more metadata elements describing the configuration or content of the structured data portion. References to the structured data portion are identified in natural language text portions of the natural language textual content and evaluated based on the signature. An ingestion prioritization plan for ingesting a set of facts associated with a set of elements of the structured data portion is generated based on results of the evaluation. The ingestion prioritization plan is applied to generate the set of facts and store the set of facts in an ingested representation of the natural language textual content.