Document Classification System Using Hybrid Image and Human Review

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for organizing documents, particularly non-text content like images and videos, face challenges in identifying relevant information due to insufficient use of keywords, especially for subjective and dynamic content, leading to information overload and difficulty in sorting through content like photographs of celebrities or models.

Innovation Solution

A computer system that uses image-processing software to determine editing instructions and classification information for documents, allowing for subjective comments and corrections, and organizes images in a hierarchical data structure based on these instructions, enabling flexible identification and access to relevant content through a combination of human and computer-based analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If keywords are used to organize and retrieve documents, then text-based document classification is improved, but identification of non-text content (images, videos) and subjective content remains insufficient

Engineering Contradiction:
Improvedocument classification accuracyVSAvoidcontent type coverage
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system employs multiple classification approaches (keyword-based text classification, image processing software for visual content, and human reviewer classification) that can handle different content types including text, images, videos, and subjective content. Each classification method is applied based on the document type, creating a universal classification system that adapts to various content formats.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

Human reviewers act as intermediaries between automated classification systems and the final classification result. The computer system generates initial classification information, which is then reviewed, corrected, and supplemented by human experts who provide subjective assessment and refinement, particularly for images and subjective content.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If automated image processing software is used to determine classification information, then processing speed and productivity are improved, but accuracy for subjective content may be insufficient

Engineering Contradiction:
Improvedocument processing speedVSAvoidclassification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The computer system performs preliminary classification of documents using automated image processing software before human review. This preliminary classification handles the bulk of routine documentation quickly, and only requires human intervention for refinement and subjective content, thus maintaining high productivity while improving accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Human reviewers provide feedback on the automated classification results by correcting errors and adding subjective assessments. This feedback loop allows the system to maintain high processing speed through automation while continuously improving classification accuracy through human expertise.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If human reviewers are used to provide subjective comments and corrections, then classification accuracy for subjective content is improved, but processing time and complexity increase

Engineering Contradiction:
Improvesubjective content classification accuracyVSAvoiddocument processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

Human reviewers are applied partially - only to documents that require subjective assessment or correction of automated classification errors, rather than reviewing every document. This selective application of human expertise maintains high accuracy for subjective content while minimizing the time loss associated with human review.

Inventive Principle:
Principle #16Partial or excessive action

4Adaptability or versatility

If multiple classification methods are combined, then content coverage and versatility are improved, but system complexity increases

Engineering Contradiction:
Improvecontent type coverageVSAvoidsystem structure complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The classification system is segmented into distinct modules: keyword-based text classification, image processing software for visual content, and human reviewer classification for subjective content. Each module handles specific content types independently, and the results are integrated to produce the final classification. This segmentation manages complexity by dividing the system into specialized, manageable components.

Inventive Principle:
Principle #1Segmentation

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Facilitates effective organization and retrieval of relevant content by dynamically adjusting classification information and editing instructions, improving the identification of subjective and time-varying content, and enabling efficient sorting and display of images based on user queries.

Implementation Method 1

determines a first set of editing instructions and classification information associated with the documents using image-processing software

Methodology Applied
Scientific EffectImage processing: Image Processing

Data Source

PatentUS9158793B2System and technique for editing and classifying documents
Publication Date: 2015.10.13 ASCENTIAL INC
  • US9158793B2 patent drawing
  • US9158793B2 patent drawing
  • US9158793B2 patent drawing

AI summary

Embodiments of a computer system which determines information associated with documents are described. During operation, this computer system receives documents (such as images). Then, the computer system determines a first set of editing instructions and classification information associated with the documents using data-processing software. Next, the computer system receives a second set of editing instructions and classification information associated with the documents. Note that the second set of editing instructions and classification information are generated by a group of individuals and include modifications and additions to the first set of editing instructions and classification information.