Document Image Conversion Using Sample Layout and Generative AI
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing generative AI systems struggle to convert scanned document images into application files that match user intentions without requiring burdensome manual input of layout and design instructions.
Innovation Solution
A system that inputs a pair of image and text to a generative AI, using a user instruction to rearrange information from an original document image based on a sample file layout, allowing automatic conversion to a desired format.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If a scanned document image is submitted to generative AI with a text instruction statement describing layout and design, then the system can generate converted data, but the user burden of inputting such instruction statements increases
Solution Approach 1:
The system enables self-service by automatically extracting layout and design information from the scanned document image itself, eliminating the need for users to manually input instruction statements. The generative AI process automatically generates converted data by analyzing the original document's structure, achieving both high automation and ease of operation.
Solution Approach 2:
The system performs preliminary action by pre-processing the scanned document image to extract layout and design information before submitting it to the generative AI. This preliminary extraction of structural information eliminates the need for users to provide detailed layout instructions, reducing user burden while maintaining conversion accuracy.
2Manufacturing precision
If manual input of layout and design instructions is required, then the converted data can match user intentions, but the conversion process becomes less efficient
Solution Approach 1:
The system replaces the mechanical process of manual instruction input with an automated image analysis process. The generative AI automatically extracts layout and design information from the scanned document image, substituting manual user input with automated computational analysis, thereby achieving both high accuracy and efficiency.
Solution Approach 2:
The system introduces an intermediary process of automatic image analysis and information extraction between the scanned document and the generative AI conversion process. This intermediary automatically bridges the gap between the original document structure and the desired converted format, eliminating manual instruction input while maintaining layout accuracy.
Data Source
AI summary
An information processing method includes receiving an instruction that causes a generative artificial intelligence (AI) to generate document data, the document data being generated by rearranging the information included in an original document image according to layout of a sample file, and the instruction being generated based on a user operation, acquiring the original document image, acquiring the sample file, generating an instruction statement corresponding to the received instruction, and transmitting the original document image, the sample file, and the instruction statement to a server of the generative AI.


