Swipe Gesture Region Definition for OCR Schedule Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing information processing apparatuses equipped with cameras and display devices face challenges in accurately extracting and associating event titles and dates from images, often resulting in improper recognition regions and incomplete schedule data generation.
Innovation Solution
An information processing apparatus with a camera, display device, touch panel, and control device that uses swipe gestures to define strip-shaped regions for text recognition, allowing the extraction and association of event titles and dates through OCR processing, ensuring accurate generation of schedule data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If an arbitrary location is specified and a recognition region is automatically defined, then the operation is simple, but the recognition region may be improper and text extraction is inaccurate
Solution Approach 1:
The recognition region definition method is made dynamic by allowing users to switch between automatic definition (simple operation) and manual definition (precise extraction). The system adapts the region definition approach based on user input, combining the benefits of both automated convenience and manual precision.
Solution Approach 2:
Instead of automatically defining the recognition region and hoping it captures the correct text, the invention inverts the approach by allowing users to manually specify the region when automatic definition fails. This reverses the control from system-automatic to user-controlled, ensuring accuracy when needed.
2Loss of information
If text recognition is performed on the entire image, then all text is captured, but it becomes difficult to associate specific text with correct events
Solution Approach 1:
The image is segmented into multiple recognition regions based on swipe gesture trajectories. Each region contains specific text relevant to particular events, allowing the system to process and associate text with events more easily by working with divided, focused regions rather than the entire image at once.
Solution Approach 2:
Different regions of the image are treated differently by creating specific recognition zones. Each swipe-defined region has the quality of being text-focused and event-specific, allowing targeted text extraction and association without the complexity of processing the entire image uniformly.
3Measurement precision
If manual region specification is required, then text extraction accuracy improves, but the operation becomes complex
Solution Approach 1:
The system dynamically adjusts the operation complexity based on user needs. Users can perform simple swipe gestures for common cases where automatic region definition suffices, while having the option to perform more detailed manual specifications only when higher precision is required, making the overall operation adaptable and user-friendly.
Data Source
AI summary
An information processing apparatus includes a camera, a display device, a touch panel, and a control device. The control device functions as a controller that, upon specification of, based on a trajectory of a swipe gesture accepted by the touch panel on an image captured by the camera and being displayed on the display device, a strip-shaped region containing start and end points of the swipe gesture and having a predetermined constant width perpendicular to a direction of the swipe gesture, recognizes a text in the specified region, extracts from the recognized text a title of an event and a date of the event, associates the extracted title of the event and date with each other, and sets the associated title of the event and date as schedule data.


