Data Acquisition Device for Tables with Missing Headlines
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data acquisition systems face challenges in extracting necessary data from tables when headlines are missing or insufficient, making it difficult to recognize attribute relationships and retrieve relevant information.
Innovation Solution
A data acquisition device that uses a processor to access correspondence information between attribute and non-attribute keywords, specifying and assigning annotations to extract specific tables, select relevant rows and columns, and acquire data from cells, even when headlines are missing or unclear.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional data acquisition methods are used, then data extraction from tables with clear headlines is effective, but data acquisition fails when headlines are missing or insufficient
Solution Approach 1:
The system performs preliminary action by pre-defining correspondence information between attribute keywords and non-attribute keywords before data acquisition. This allows the system to proactively identify and establish attribute relationships even when table headlines are missing, enabling reliable data extraction from tables with insufficient headlines.
Solution Approach 2:
The correspondence information between attribute keywords and non-attribute keywords serves as an intermediary that bridges the gap between search keywords and table data. This intermediary mechanism enables the system to infer attribute relationships without relying on explicit table headlines, thereby improving adaptability to tables with missing or insufficient headlines.
2Productivity
If the system relies on table headlines for data extraction, then processing is simple when headlines are present, but data acquisition becomes impossible when headlines are missing
Solution Approach 1:
The system performs preliminary action by pre-defining correspondence information between attribute keywords and non-attribute keywords before data acquisition. This allows the system to proactively identify and establish attribute relationships even when table headlines are missing, enabling reliable data extraction from tables with insufficient headlines.
Solution Approach 2:
The correspondence information between attribute keywords and non-attribute keywords serves as an intermediary that bridges the gap between search keywords and table data. This intermediary mechanism enables the system to infer attribute relationships without relying on explicit table headlines, thereby improving adaptability to tables with missing or insufficient headlines.
3Speed
If the system uses keyword matching without attribute classification, then processing is fast, but accurate data retrieval cannot be achieved
Solution Approach 1:
The system segments keywords into attribute keywords and non-attribute keywords based on pre-defined correspondence information. This segmentation allows the system to quickly identify which keywords represent attributes that can be matched with table data, maintaining processing speed while improving retrieval accuracy through structured keyword classification.
Solution Approach 2:
The system changes the parameter of keyword representation by transforming plain search keywords into annotated keywords with attribute information. This parameter change enables the system to maintain fast processing through efficient keyword matching while achieving accurate data retrieval through attribute-based classification and matching.
Data Source
AI summary
A data acquisition device is accessible to correspondence information that defines correspondence between an attribute keyword indicating an attribute and a non-attribute keyword that does not indicate the attribute, and is configured to execute: specifying the attribute keyword corresponding to the non-attribute keyword when the search keyword is the non-attribute keyword with respect to each of a plurality of search keywords; assigning the search keyword to a character string in a retrieval target document corresponding to the search keyword; extracting a specific table assigned with the annotation from one or more tables; selecting at least one of a specific row and a specific column relevant to each of the plurality of search keywords from rows and columns that constitute the specific table extracted on the basis of the annotation; and acquiring a cell in the specific table specified by a first selection result.


