Multi-Language OCR Character Recognition Using Unified Dictionary

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing optical character recognition (OCR) technologies face challenges in accurately recognizing characters from multiple languages within an image, leading to decreased recognition accuracy and search functionality when characters from different languages are present, as they typically rely on a single language environment for processing.

Innovation Solution

An electronic apparatus that uses multiple dictionary data sets for different language environments to recognize characters in images, converting characters into corresponding codes and storing recognition results, allowing for accurate keyword searches across languages.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a single language environment dictionary is used for character recognition, then the recognition process is simple and fast, but characters from unexpected languages cannot be correctly recognized

Engineering Contradiction:
Improvecharacter recognition accuracyVSAvoiddictionary data structure
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies universality by creating a dictionary data structure that serves multiple language environments simultaneously. The dictionary includes character codes from both a first language environment and a second language environment, allowing the same OCR system to accurately recognize characters from either language without requiring separate recognition processes.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent segments the dictionary data into distinct language environment sections (first language environment characters and second language environment characters). This segmentation allows the system to organize and manage multiple language dictionaries efficiently while maintaining a unified recognition process that can access both language sets.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If multiple dictionary data sets for different languages are used, then character recognition accuracy for multiple languages improves, but the complexity of the recognition process increases

Engineering Contradiction:
Improvemulti-language character recognition accuracyVSAvoidrecognition process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges multiple language environment dictionaries into a single unified dictionary data structure. By combining the first language environment characters and second language environment characters into one dictionary, the system achieves multi-language recognition capability without requiring separate recognition processes or multiple independent dictionary loading operations.

Inventive Principle:
Principle #5Merging (Combining)

3Adaptability or versatility

If OCR system uses multiple language dictionaries, then search functionality across multiple languages is enabled, but processing time and computational resources increase

Engineering Contradiction:
Improvemulti-language search capabilityVSAvoidcharacter recognition processing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-organizing multiple language environment dictionaries into a unified structure before the recognition process. The dictionary data is prepared in advance with character codes from both language environments already integrated, so that during actual OCR processing, the system can quickly access the appropriate character codes without performing complex language detection or switching operations.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10127478B2Electronic apparatus and method
Publication Date: 2018.11.13 DYNABOOK INC
  • US10127478B2 patent drawing
  • US10127478B2 patent drawing
  • US10127478B2 patent drawing

AI summary

According to one embodiment, an electronic apparatus includes a hardware processor. The hardware processor converts a first character in a first image of images in which characters of languages are rendered, into a first character code by using dictionary data for a first language environment, converts the first character into a second character code by using dictionary data for a second language environment, causes a memory to store a pair of the first character code and a first area in the first image corresponding to the first character code, and causes the memory to store a pair of the second character code and a second area in the first image corresponding to the second character code.