Grammatical Parsing for Visual Recognition Tasks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional OCR techniques face difficulties in processing complex documents and two-dimensional recognition tasks due to computational complexity and the challenge of storing and retrieving vast amounts of possible patterns, especially with elements like equations and documents with enhanced structures.

Innovation Solution

The use of grammatical parsing with image recognition to score parse trees and subtrees, combined with geometric constraints and integral images, facilitates efficient analysis and recognition of documents by rendering images and employing machine learning for classification, thereby improving parsing speed and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional OCR techniques are used to process complex documents with equations and enhanced structures, then recognition capability is improved, but computational complexity increases exponentially

Engineering Contradiction:
Improverecognition capabilityVSAvoidcomputational complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the complex recognition task into multiple hierarchical levels: first segmenting the document into regions (headers, footers, body text), then further segmenting regions into lines and words. This hierarchical segmentation reduces the computational complexity by breaking down the exponential problem into manageable polynomial sub-problems at each level.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from traditional one-dimensional linear scanning to two-dimensional spatial analysis by introducing spatial relationships and geometric constraints. This dimensional change allows the system to leverage the natural two-dimensional layout of documents, reducing search space and computational complexity while improving recognition accuracy for complex structures.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If traditional OCR techniques are used to process complex documents with equations and enhanced structures, then recognition capability is improved, but processing speed decreases

Engineering Contradiction:
Improverecognition capabilityVSAvoidprocessing speed
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent performs preliminary actions by first segmenting the document into regions and establishing spatial relationships before actual character recognition. Geometric constraints and spatial models are pre-computed and used to guide subsequent recognition steps, avoiding exhaustive search and significantly improving processing speed for complex documents.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

By dividing the document into hierarchical segments (page → regions → lines → words → characters), the system processes only relevant portions at each level using geometric constraints, rather than processing the entire document uniformly. This selective segmentation dramatically improves processing speed while maintaining recognition capability.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If database oriented recognition systems are used to store equation elements, then recognition accuracy is improved, but database size and retrieval time increase

Engineering Contradiction:
Improverecognition accuracyVSAvoiddatabase size
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent changes the approach from storing complete equation element images in a database to using parameter-based geometric constraints and spatial relationships. Instead of retrieving stored images, the system defines equations through mathematical parameters and geometric properties, dramatically reducing database size while maintaining recognition accuracy through parametric matching.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS7639881B2Application of grammatical parsing to visual recognition tasks
Publication Date: 2009.12.29 MICROSOFT TECHNOLOGY LICENSING LLC
  • US7639881B2 patent drawing
  • US7639881B2 patent drawing
  • US7639881B2 patent drawing

AI summary

Image recognition is utilized to facilitate in scoring parse trees for two-dimensional recognition tasks. Trees and subtrees are rendered as images and then utilized to determine parsing scores. Other instances of the subject invention can incorporate additional features such as stroke curvature and/or nearby white space as rendered images as well. Geometric constraints can also be employed to increase performance of a parsing process, substantially improving parsing speed, some even resolvable in polynomial time. Additional performance enhancements can be achieved in yet other instances of the subject invention by employing constellations of integral images and/or integral images of document features.