Text Image Super-Resolution Using Text-Guided Feature Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing text recognition algorithms struggle with low-resolution text images, leading to significant drops in recognition accuracy.

Innovation Solution

A text image super-resolution method using a pre-trained model with Gated Text Detection Blocks (GTDBs) and Convolutional Block Attention Module (CBAM) to enhance feature extraction, combined with a text-assisted loss function, to improve text readability and resolution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Length of stationary object

If traditional image super-resolution algorithms are used, then the resolution of text images can be improved, but the text recognition accuracy drops significantly in low-resolution images

Engineering Contradiction:
Improveimage resolutionVSAvoidtext recognition accuracy
Core Design Contradiction:
Length of stationary objectVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by performing super-resolution enhancement on text images before text recognition is performed. The model pre-processes low-resolution images to restore text details and clarity, ensuring that by the time recognition occurs, the image quality is sufficient for accurate reading. This is evident in the technical effect where the model reconstructs text images to improve readability and recognition accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies local quality by focusing the super-resolution process specifically on text regions rather than treating the entire image uniformly. The model identifies and enhances text areas with higher detail restoration, while applying different processing to non-text regions. This is reflected in the technical effect where text features are specifically enhanced to improve recognition accuracy.

Inventive Principle:
Principle #3Local quality

2Stability of the object's composition

If deep residual network with skip connection is used, then the network stability and convergence are improved, but the model complexity increases

Engineering Contradiction:
Improvenetwork stabilityVSAvoidmodel complexity
Core Design Contradiction:
Stability of the object's compositionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the deep residual network into multiple residual modules with skip connections. Each module is a self-contained unit that processes features independently while maintaining connections to preserve gradient flow. This modular structure improves network stability and convergence while making the complex model more manageable and trainability, as evidenced by the successful training of the super-resolution model.

Inventive Principle:
Principle #1Segmentation

3Manufacturing precision

If high-resolution image acquisition device is used, then clearer image details are obtained, but the cost increases

Engineering Contradiction:
Improveimage detail qualityVSAvoidacquisition cost
Core Design Contradiction:
Manufacturing precisionVSEase of manufacture

Solution Approach 1:

The patent applies copying by creating a high-resolution copy of the low-resolution input image through computational super-resolution. Instead of requiring expensive high-resolution acquisition devices, the model generates a synthetic high-resolution version of the existing low-resolution image, restoring text details and clarity. This is directly reflected in the technical effect where the model reconstructs high-resolution text images from low-resolution inputs, providing clear details without additional hardware costs.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12499697B2Text image super-resolution method based on text assistance
Publication Date: 2025.12.16 NANJING UNIV OF POSTS & TELECOMM
  • US12499697B2 patent drawing
  • US12499697B2 patent drawing

AI summary

The present application discloses a text image super-resolution method based on text assistance, including: obtaining a low-resolution text image to be reconstructed; inputting the low-resolution text image into a pre-trained text image super-resolution model, and determining a reconstructed text image based on an output of the text image super-resolution model; a method of constructing and training the text image super-resolution model includes: obtaining a text image dataset; and training the pre-constructed text image super-resolution model by using the text image dataset to obtain the trained text image super-resolution model. Compared to other ordinary super-resolution models, this text image super-resolution model fuses the text sequence features with the image texture features, and fully exploits and utilizes the text information in the low-resolution image, which can help to improve the quality of reconstructed text image.