Conditional Text Detection in Images for Flexible OCR Output

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional OCR methods are limited in their ability to recognize or detect characters in images, as they often require a predetermined format and do not adapt to user-specific requirements or variations in character detection performance.

Innovation Solution

A method and system for detecting text in images using a text detection model that receives an image and a text detection condition, generating a sequence indicating the detection result based on user-defined criteria, allowing for flexible and efficient text detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional OCR methods are used to detect text in images, then text recognition can be performed, but the detection format is limited and cannot adapt to user-specific requirements

Engineering Contradiction:
Improvedetection format adaptabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The text detection system dynamically adjusts detection parameters and output formats based on user instructions. The model receives text detection conditions as input and generates detection sequences according to specified criteria such as detection location, area, or shape, enabling flexible adaptation to different user requirements without requiring multiple fixed systems

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes detection parameters by receiving text detection conditions as input instructions. These conditions specify various parameters such as detection location, area, shape, and other criteria, which the model uses to generate appropriate detection sequences, allowing the same system to adapt to different detection requirements by modifying input parameters

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If conventional sequence generation-based text detection models are used, then text detection can be performed, but the output information amount is limited regardless of image size or text length

Engineering Contradiction:
Improveoutput information amountVSAvoiddetection efficiency
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The system incorporates feedback mechanisms where the text detection model processes text detection conditions as input instructions and generates detection sequences based on these conditions. This feedback loop allows the model to continuously adapt its output based on user requirements, enabling comprehensive text detection that scales with image size and text length while maintaining high detection efficiency

Inventive Principle:
Principle #23Feedback

3Ease of operation

If text detection is performed without user-defined conditions, then detection can be executed, but the results cannot be visualized or output at desired locations, areas, or shapes

Engineering Contradiction:
Improveuser-friendly operationVSAvoiddetection precision
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The system performs preliminary action by receiving text detection conditions as input instructions before executing the detection. These conditions specify the desired detection location, area, shape, and other criteria, allowing the model to prepare and execute detection tasks with precise control over output characteristics, thereby achieving both ease of operation and high precision

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250363815A1Method and device for detecting text in image
Publication Date: 2025.11.27 NAVER CORP
  • US20250363815A1 patent drawing
  • US20250363815A1 patent drawing
  • US20250363815A1 patent drawing

AI summary

A method for detecting text in an image includes receiving an image including text; receiving a command including a text detection condition; and inputting the image and the command into a text detection model so as to generate a sequence indicating a detection result of a text instance included in the image according to the text detection condition.