Text Line Curvature Correction for Optical Character Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Text-containing images captured by hand-held devices or dedicated scanners often suffer from noise, optical blur, and curvature distortions due to curved-page-surface-induced and perspective-induced effects, which degrade the performance of optical-character recognition systems.

Innovation Solution

A method and system that identify the outline of a text-containing page, generate contours for text lines, determine centroids and inclination angles, and use a model to assign local displacements to pixels, thereby straightening text lines through an inclination-angle map or pixel-displacement map to correct for curvature.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If hand-held devices or dedicated scanners are used to capture text-containing images, then portability and ease of use are improved, but curvature distortions and perspective-induced effects worsen the quality of text images

Engineering Contradiction:
ImproveportabilityVSAvoidimage quality
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent applies preliminary image processing operations including noise reduction, optical deblurring, and curvature correction before optical character recognition. By performing these corrective actions in advance, the system compensates for the distortions introduced by hand-held device portability, thereby maintaining both ease of operation and image quality.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If curved-page-surface-induced and perspective-induced curvature effects occur during image capture, then portability and flexibility are improved, but text line straightness and character recognition accuracy worsen

Engineering Contradiction:
ImproveflexibilityVSAvoidcharacter recognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent transforms the curved text lines into straight lines by applying geometric transformations and parameter adjustments. The system detects curvature parameters and applies corresponding correction transformations to straighten text lines, thereby maintaining adaptability to different capture conditions while improving character recognition accuracy.

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If noise and optical blur are present in captured images, then ease of capture is improved, but optical character recognition performance worsens

Engineering Contradiction:
Improveease of captureVSAvoidrecognition performance
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent extracts and removes noise and blur components from the captured images through dedicated signal processing operations. By separating and eliminating these detrimental elements while preserving the actual text content, the system maintains ease of capture while significantly improving recognition performance.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10366469B2Method and system that efficiently prepares text images for optical-character recognition
Publication Date: 2019.07.30 ABBYY DEVELOPMENT INC
  • US10366469B2 patent drawing
  • US10366469B2 patent drawing
  • US10366469B2 patent drawing

AI summary

The current document is directed to methods and systems that straighten curvature in the text lines of text-containing digital images, including text-containing digital images generated from the two pages of an open book. Initial processing of a text-containing image identifies the outline of a text-containing page. Next, contours are generated to represent each text line. The midpoints and inclination angles of the links or vectors that comprise the contour lines are determined. A model is constructed for the perspective-induced curvature within the text image. In one implementation, the model, essentially an inclination-angle map, allows for assigning local displacements to pixels within the page image which are then used to straighten the text lines in the text image. In another implementation, the model is essentially a pixel-displacement map which is used to straighten the text lines in the text image.