Spreadsheet Document Classification via Cell Value Pattern Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge lies in automatically classifying design documents from various vendors without common formats, as existing systems lack guidance and templates, leading to inefficiencies in inventory and classification processes, especially when dealing with large numbers of documents in spreadsheet form.

Innovation Solution

A computer-implemented method that reads and processes spreadsheet documents by identifying common cell values, calculating distances between candidate header labels, and generating classification data to categorize documents based on header combinations, allowing for the classification of documents without predefined templates.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual classification methods are used for documents from different vendors, then classification accuracy can be maintained through human judgment, but productivity decreases due to the large volume of documents requiring manual review

Engineering Contradiction:
Improveclassification accuracyVSAvoiddocument processing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system enables automatic self-classification of documents by analyzing spreadsheet structures and content without human intervention. The classification module automatically identifies document types, vendors, and categories by processing spreadsheet data structures, cell values, and formatting patterns, allowing the system to serve itself rather than requiring manual classification for each document

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical classification processes with an automated computational system. The classification module uses algorithmic analysis of spreadsheet structures, cell value patterns, and document metadata to substitute human judgment with automated decision-making, thereby increasing processing speed while maintaining classification accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If standardized document formats are imposed on vendor documents, then ease of operation improves through uniform processing, but adaptability decreases because vendor-specific formats cannot be accommodated

Engineering Contradiction:
Improveprocessing uniformityVSAvoidformat flexibility
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The system applies different processing rules and classification criteria to different regions or sections of spreadsheet documents based on their specific characteristics. The classification module identifies local patterns in cell structures, column arrangements, and data formats specific to each vendor or document type, allowing customized processing for each local context while maintaining overall system operation

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The classification system dynamically adapts its processing approach based on the input document characteristics. Rather than using fixed rigid rules, the system adjusts its analysis methods, threshold values, and classification criteria according to the specific format, structure, and content patterns detected in each vendor's spreadsheet documents, enabling flexible handling of diverse formats

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If comprehensive analysis of all document details is performed, then measurement precision improves for accurate classification, but loss of time increases due to the extensive processing required

Engineering Contradiction:
Improveclassification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The classification process is divided into multiple independent stages: initial screening based on spreadsheet structure, intermediate analysis of cell values and patterns, and final classification decision-making. Each segment processes specific aspects of the document independently, allowing parallel processing and reducing the time required for comprehensive analysis while maintaining overall classification accuracy

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs analysis at multiple levels of depth, applying comprehensive detailed analysis only when necessary. For clearly identifiable document types, the system uses expedited partial analysis based on key identifying features, while reserving full comprehensive analysis for ambiguous or complex cases, thereby reducing average processing time while maintaining accuracy for critical classifications

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10318568B2Generation of classification data used for classifying documents
Publication Date: 2019.06.11 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10318568B2 patent drawing
  • US10318568B2 patent drawing
  • US10318568B2 patent drawing

AI summary

Systems and methods are provided for generating classification data which is used for classifying documents. The method includes reading documents in a form of a spreadsheet; collecting cell values in each of the documents; finding one or more common cell values among the collected values; counting, for each of the common cell values, a number of the documents having the common cell value; storing, if the number of the documents is equal to or larger than a predetermined number, the common cell value as a candidate header label in a memory; calculating a distance between cell locations of the candidate header labels in each of the documents; choosing, according to the calculated distance, two or more candidate header labels among the candidate header labels for each of the documents; and storing one or more combinations of the chosen two or more candidate header labels as the classification data.