End-to-End Text Recognition Model Reducing Computational Load

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional text recognition methods are cumbersome and computationally intensive, requiring complex preprocessing and multiple neural networks for effective recognition, especially in cases of non-uniform lighting and fuzzy images, leading to high computation costs.

Innovation Solution

A method involving a sample text dataset preprocessing, diffusion annotation, and training a text recognition model using convolutional layers, pooling, and up-sampling to generate clear-scale prediction maps for accurate text sequence extraction, optimizing the loss function with weights for text and character regions and boundaries.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional text recognition methods are used with complex preprocessing and multiple neural networks, then text recognition accuracy is improved, but computational load and process complexity increase significantly

Engineering Contradiction:
Improvetext recognition accuracyVSAvoidprocess complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges text region positioning and text recognition into a single end-to-end deep learning model. Instead of using separate neural networks for positioning and recognition, the invention integrates both functions into one unified model that directly outputs recognition results, thereby reducing process complexity while maintaining accuracy.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The unified deep learning model performs multiple functions simultaneously: it identifies text regions, positions characters, and recognizes text content all in one process. This multi-functional approach eliminates the need for separate preprocessing and positioning steps, simplifying the overall process while achieving accurate recognition.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If complex preprocessing measures are applied to handle non-uniform lighting and fuzzy images, then text recognition accuracy is improved, but computational load increases

Engineering Contradiction:
Improvetext recognition accuracyVSAvoidcomputational load
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent replaces traditional mechanical preprocessing operations (such as manual image enhancement, filtering, and adjustment steps) with a deep learning-based end-to-end recognition system. The neural network automatically adapts to various image conditions including non-uniform lighting and fuzziness, eliminating the need for computationally intensive preprocessing while maintaining recognition accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If two neural networks are trained separately for text positioning and recognition, then recognition accuracy is improved, but training time and computational resources increase

Engineering Contradiction:
Improverecognition accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent combines two separate neural networks into a single end-to-end model that performs both text positioning and recognition simultaneously. This unified model is trained in one go rather than training two separate networks, significantly reducing training time and computational resource requirements while maintaining the accuracy benefits of having both positioning and recognition capabilities.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20230154217A1Method for Recognizing Text, Apparatus and Terminal Device
Publication Date: 2023.05.18 TP-LINK SYSTEMS INC
  • US20230154217A1 patent drawing
  • US20230154217A1 patent drawing
  • US20230154217A1 patent drawing

AI summary

The present disclosure discloses a method for recognizing text, an apparatus and a terminal device. The method for recognizing text includes: acquiring a sample text dataset, preprocessing text image in the sample text dataset, and generating a label image; inputting the label image into a text recognition model for training, extracting image features, performing down-sampling, restoring an image resolution, normalizing an output probability for the last layer using a sigmoid layer to output a multiple prediction maps with different scales, and optimizing a loss function of the text recognition model to obtain a trained text recognition model; preprocessing a text image to be recognized, inputting the text image to be recognized which is preprocessed into the trained text recognition model, and outputting a clear-scale prediction map; and analyzing the clear-scale prediction map to obtain a text sequence of the text image to be recognized.