Data Acquisition Device for Tables with Missing Headlines

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data acquisition systems face challenges in extracting necessary data from tables when headlines are missing or insufficient, making it difficult to recognize attribute relationships and retrieve relevant information.

Innovation Solution

A data acquisition device that uses a processor to access correspondence information between attribute and non-attribute keywords, specifying and assigning annotations to extract specific tables, select relevant rows and columns, and acquire data from cells, even when headlines are missing or unclear.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional data acquisition methods are used, then data extraction from tables with clear headlines is effective, but data acquisition fails when headlines are missing or insufficient

Engineering Contradiction:
Improvedata acquisition reliabilityVSAvoidadaptability to tables with missing headlines
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary action by pre-defining correspondence information between attribute keywords and non-attribute keywords before data acquisition. This allows the system to proactively identify and establish attribute relationships even when table headlines are missing, enabling reliable data extraction from tables with insufficient headlines.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The correspondence information between attribute keywords and non-attribute keywords serves as an intermediary that bridges the gap between search keywords and table data. This intermediary mechanism enables the system to infer attribute relationships without relying on explicit table headlines, thereby improving adaptability to tables with missing or insufficient headlines.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If the system relies on table headlines for data extraction, then processing is simple when headlines are present, but data acquisition becomes impossible when headlines are missing

Engineering Contradiction:
Improvedata extraction efficiencyVSAvoiddata acquisition success rate
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system performs preliminary action by pre-defining correspondence information between attribute keywords and non-attribute keywords before data acquisition. This allows the system to proactively identify and establish attribute relationships even when table headlines are missing, enabling reliable data extraction from tables with insufficient headlines.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The correspondence information between attribute keywords and non-attribute keywords serves as an intermediary that bridges the gap between search keywords and table data. This intermediary mechanism enables the system to infer attribute relationships without relying on explicit table headlines, thereby improving adaptability to tables with missing or insufficient headlines.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Speed

If the system uses keyword matching without attribute classification, then processing is fast, but accurate data retrieval cannot be achieved

Engineering Contradiction:
Improvekeyword matching speedVSAvoiddata retrieval accuracy
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The system segments keywords into attribute keywords and non-attribute keywords based on pre-defined correspondence information. This segmentation allows the system to quickly identify which keywords represent attributes that can be matched with table data, maintaining processing speed while improving retrieval accuracy through structured keyword classification.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes the parameter of keyword representation by transforming plain search keywords into annotated keywords with attribute information. This parameter change enables the system to maintain fast processing through efficient keyword matching while achieving accurate data retrieval through attribute-based classification and matching.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11100099B2Data acquisition device, data acquisition method, and recording medium
Publication Date: 2021.08.24 HITACHI LTD
  • US11100099B2 patent drawing
  • US11100099B2 patent drawing
  • US11100099B2 patent drawing

AI summary

A data acquisition device is accessible to correspondence information that defines correspondence between an attribute keyword indicating an attribute and a non-attribute keyword that does not indicate the attribute, and is configured to execute: specifying the attribute keyword corresponding to the non-attribute keyword when the search keyword is the non-attribute keyword with respect to each of a plurality of search keywords; assigning the search keyword to a character string in a retrieval target document corresponding to the search keyword; extracting a specific table assigned with the annotation from one or more tables; selecting at least one of a specific row and a specific column relevant to each of the plurality of search keywords from rows and columns that constitute the specific table extracted on the basis of the annotation; and acquiring a cell in the specific table specified by a first selection result.