Single Character Model for Mixed-Language Text Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current text recognition technologies require multiple dedicated models for different text orientations and languages, increasing complexity and resource requirements, and may struggle with accuracy in mixed-language or deformed images.
Innovation Solution
A single character model is trained on text line areas with varying orientations and languages, allowing for unified text recognition without the need to determine orientation or language, simplifying the recognition process and reducing resource demands.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple dedicated models are trained for different text orientations and languages, then recognition accuracy for specific text types is improved, but device complexity and resource requirements increase
Solution Approach 1:
The patent applies universality by designing a single character recognition model that can handle multiple text orientations (horizontal, vertical, rotated) and multiple languages (Latin, Eastern) simultaneously. The model is trained on diverse training data encompassing various orientations and languages, enabling it to perform multiple recognition functions without requiring separate dedicated models for each text type, thus reducing device complexity while maintaining recognition accuracy
Solution Approach 2:
The patent employs parameter changes by modifying the training data parameters to include diverse orientations and languages. The model learns to recognize characters by adapting to variations in text orientation parameters and language characteristics during training, allowing a single model to adjust its recognition parameters dynamically based on the input text characteristics rather than requiring separate models for each parameter combination
2Measurement precision
If multiple dedicated models are trained for different text orientations and languages, then recognition accuracy for specific text types is improved, but resource requirements increase
Solution Approach 1:
The single character recognition model performs multiple recognition functions for different orientations and languages simultaneously, eliminating the need to load and execute multiple separate models. This reduces computational resources and energy consumption by using one unified model instead of multiple dedicated models, while still achieving accurate recognition across various text types
Solution Approach 2:
The patent merges the functionality of multiple dedicated models into a single character recognition model. By combining the recognition capabilities for different orientations and languages into one model, the system reduces the computational overhead of managing multiple models and decreases resource requirements while maintaining the accuracy benefits of specialized recognition
3Measurement precision
If text orientation and language must be determined during recognition, then recognition accuracy can be optimized, but processing time increases
Solution Approach 1:
The patent applies preliminary action by pre-training the character recognition model on diverse training data that includes various text orientations and languages. This preliminary training enables the model to automatically adapt to different text characteristics during recognition without requiring separate steps to determine orientation or language, thus reducing processing time while maintaining recognition accuracy
Data Source
AI summary
According to implementations of the subject matter described herein, there is provided a solution for text recognition in an image. In this solution, a target text line area, which is expected to include a text to be recognized, is determined from an image. Probability distribution information of a character model element(s) present in the target text line area is determined using a single character model. The single character model is trained based on training text line areas and respective ground-truth texts in the training text line areas. Texts in the training text line areas are arranged in different orientations, and/or the ground-truth texts comprise texts are related to various languages (e.g., texts related to a Latin and an Eastern languages). The text in the target text line area can be determined based on the determined probability distribution information. The single character model enables more efficient and convenient text recognition.


