Table-Based Groundtruth Generation Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current question-answer (QA) systems are ineffective in generating significant QA pairs from structured data, such as tables, leading to challenges in providing groundtruth for training, especially when quantitative information is involved, as they often produce numerous irrelevant questions.

Innovation Solution

A method and system for automating the generation of table-based groundtruth by receiving a document with unstructured text and a table, applying templates to generate questions, and scoring QA pairs based on user interest using unstructured text and table content, to prioritize and filter relevant questions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If all possible questions are generated from repeated-structure content in tables, then the quantity of QA pairs increases, but the quality and relevance of QA pairs deteriorates due to generation of irrelevant questions

Engineering Contradiction:
Improvequantity of QA pairsVSAvoidrelevance of QA pairs
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent changes the parameter of question generation by introducing a scoring mechanism that evaluates each generated QA pair based on multiple criteria including user interest, table structure, and content importance. This filtering approach transforms the raw quantity of QA pairs into a quality-filtered set, resolving the contradiction between generating comprehensive QA pairs and maintaining their relevance.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediary scoring system that acts as a mediator between question generation and final QA pair selection. This scoring mechanism evaluates generated questions against multiple criteria and filters out irrelevant ones, allowing the system to maintain both high quantity and high quality of QA pairs simultaneously.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Extent of automation

If templates are applied to generate questions from table contents, then the automation level increases, but the precision and user interest alignment of generated questions deteriorates

Engineering Contradiction:
Improveautomation of QA pair generationVSAvoidprecision of user interest alignment
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent implements feedback mechanisms where generated QA pairs are scored based on their alignment with user interest, table structure, and content importance. This scoring feedback allows the automated system to evaluate and prioritize questions, improving the precision of user interest alignment while maintaining high automation levels.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent makes the question generation process dynamic by adjusting the scoring and filtering criteria based on table characteristics and user interest patterns. This dynamic approach allows the automated system to adapt to different contexts, improving precision while maintaining automation efficiency.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10387560B2Automating table-based groundtruth generation
Publication Date: 2019.08.20 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10387560B2 patent drawing
  • US10387560B2 patent drawing
  • US10387560B2 patent drawing

AI summary

A method, system and computer-usable medium are disclosed for automating the generation of table-based groundtruth, comprising: receiving a document comprising unstructured text and a table; generating questions by applying a template the contents of the table; performing QA pair generation operations on the table to generate QA pairs, each QA pair comprising a question generated by applying the template; and, assigning a score to each QA pair, the score providing an indicator of user interest to each QA pair, the score being based on a score generation methodology using the unstructured text and the table.