Spreadsheet Population via Multi-Dimensional Analogy Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for populating spreadsheets require manual effort to extract and place data, which is time-consuming and costly, and fail to effectively utilize implicit relationships between data items.

Innovation Solution

A method that associates text with a spreadsheet to build a multi-dimensional analogy model, capturing implicit relationships between data items within a context window, allowing for automatic population of spreadsheet cells based on these relationships.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual extraction and placement of data is used to populate spreadsheets, then data accuracy and conformity to user semantics can be achieved, but time consumption and human resource costs increase

Engineering Contradiction:
Improvedata accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent introduces an intermediary system consisting of a processing engine with machine learning models that act as a mediator between raw document data and spreadsheet population. This intermediary automatically extracts entities, relationships, and attributes from unstructured documents and transforms them into structured spreadsheet data, eliminating the need for manual extraction while maintaining data accuracy through learned patterns and validation mechanisms.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical manual process of data extraction and spreadsheet population with an automated computational system using natural language processing, entity recognition, and machine learning algorithms. This substitution transforms the manual mechanical task into an automated intelligent process that can handle large volumes of data without proportional increases in time or resources.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of manufacture

If conventional techniques with pre-determined categories are used, then specific types of information can be extracted, but the system cannot work with example relationships between cells

Engineering Contradiction:
Improveextraction capabilityVSAvoidrelationship recognition
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic relationship modeling where the system learns and adapts to various types of relationships between spreadsheet cells through training on example data. Rather than using fixed pre-determined categories, the system dynamically identifies and extracts relationships such as temporal, spatial, causal, and hierarchical connections between entities, allowing it to handle diverse and evolving data structures without requiring explicit programming for each relationship type.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameter space from fixed categorical values to continuous relationship vectors that can represent multiple types of relationships simultaneously. By transforming rigid category-based extraction into flexible parameter-based relationship modeling, the system can capture nuanced connections between cells while maintaining the ability to extract specific information types as needed.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If automated methods are used to populate spreadsheets, then time consumption decreases, but the ability to conform to user semantics and implicit relationships may be compromised

Engineering Contradiction:
Improvepopulation speedVSAvoidsemantic conformity
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent incorporates feedback mechanisms where the system continuously learns from user corrections, validations, and interactions with the populated spreadsheets. This feedback loop allows the automated system to refine its understanding of user semantics and implicit relationships over time, improving semantic conformity while maintaining high productivity. The system adjusts its extraction and population strategies based on learned patterns from feedback.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent performs preliminary actions by pre-training the machine learning models on large corpora of documents and example relationships before actual spreadsheet population. This preliminary training establishes a strong foundation for understanding semantics and relationships, enabling the system to maintain high accuracy during automated population without requiring extensive real-time validation for each data point.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10699069B2Populating spreadsheets using relational information from documents
Publication Date: 2020.06.30 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10699069B2 patent drawing
  • US10699069B2 patent drawing
  • US10699069B2 patent drawing

AI summary

A spreadsheet population method, system, and computer program product include associating text with a spreadsheet, the text including candidate data items for populating the spreadsheet, building a multi-dimensional analogy model where each dimension comprises a unique pair of data items where the data items co-occur within a same context window, accepting example data items in the spreadsheet where the data items form tuples in a same implicit relationship according to a spatial configuration, and performing an assistance operation on the spreadsheet using the data item tuples retrieved using the analogy model from the example data items.