Document Image Instruction Conversion for Handwritten AI Prompts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems require users to manually input instruction sentences for generative AI to recognize handwritten portions, imposing a heavy burden due to the variety of handwriting styles and user-specific markings.

Innovation Solution

An information processing system that automatically identifies and converts user input instruction sentences into a format recognizable by generative AI, using document image analysis and natural language processing to extract and convert handwritten instructions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If users manually input instruction sentences for generative AI to recognize handwritten portions, then the system can process user instructions, but the user burden increases due to the variety of handwriting styles and user-specific markings

Engineering Contradiction:
Improveinstruction recognition accuracyVSAvoiduser input burden
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system automatically analyzes the document image to identify handwritten portions and converts them into instruction sentences without requiring users to manually input instructions. The system serves itself by autonomously extracting and interpreting handwritten content, thereby eliminating the user burden while maintaining reliable instruction recognition.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual input process with an automated image processing system. By using optical character recognition and natural language processing techniques, the system converts handwritten markings into digital instruction sentences automatically, substituting the manual mechanical action with an automated computational process.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If the system automatically identifies handwritten portions, then user burden is reduced, but the complexity of the system increases

Engineering Contradiction:
Improveuser input burdenVSAvoidsystem processing complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system divides the complex task of handwritten portion identification into separate processing stages: first detecting handwritten portions in the document image, then converting them into instruction sentences. This segmentation allows each stage to be optimized independently, managing overall system complexity through modular processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer that bridges the document image and the final instruction sentence. This intermediary stage automatically analyzes handwritten portions and translates them into standardized instruction formats, simplifying the interaction between the user and the generative AI system while managing processing complexity through a dedicated conversion module.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20260064784A1Information processing apparatus, information processing method, and storage medium
Publication Date: 2026.03.05 CANON KK
  • US20260064784A1 patent drawing
  • US20260064784A1 patent drawing
  • US20260064784A1 patent drawing

AI summary

A non-transitory computer-readable storage medium stores an application program which, when executed by one or more processors, causes an information processing apparatus to perform a control method, the control method including acquiring a document image including areas indicated by a plurality of handwritten portions on the document, acquiring an instruction sentence input by a user, identifying, from among the plurality of handwritten portions, an instruction portion that causes generative artificial intelligence (AI) to perform processing, converting the acquired instruction sentence input by the user into an instruction sentence enabling the generative AI to identify the instruction portion, and outputting the instruction sentence obtained by conversion and the acquired document image to the generative AI.