Machine-Learned Object Detection for Digital Graphic Novel Navigation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional ereaders, designed for text-based ebooks, fail to provide an optimal user experience for navigating digital graphic novels due to limitations in screen size and resolution, making it difficult for users to read and examine objects of interest like speech bubbles without frequent zooming.
Innovation Solution
A method and system that utilize a machine-learned model to identify and prioritize objects of interest in digital graphic novels, creating presentation metadata to enable expanded views of these objects on ereaders, allowing for improved navigation without manual zooming.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If graphic novels are displayed in conventional manner on ereaders, then the device complexity remains low, but the ease of operation deteriorates due to difficulty in navigating and reading text without frequent zooming
Solution Approach 1:
The system performs preliminary action by automatically detecting and identifying objects of interest (speech bubbles, text regions, characters) before the user needs to read them. The machine-learned model analyzes the graphic novel content in advance, creates a structured representation of objects with their locations and reading orders, and prepares navigation metadata that enables seamless access without requiring manual zooming or page-by-page flipping.
Solution Approach 2:
The patent introduces an intermediary system consisting of a machine-learned model and navigation metadata layer between the graphic novel content and the user interface. This intermediary automatically identifies objects of interest, determines their reading order, and creates structured navigation instructions that mediate between the visual content and the user's reading needs, enabling intelligent navigation without increasing device complexity.
2Productivity
If users manually zoom in and out to read text and examine objects, then the manufacturing precision of the display remains unchanged, but the loss of time increases due to repeated zooming operations
Solution Approach 1:
The system implements self-service by automatically detecting, identifying, and organizing objects of interest without requiring user intervention for navigation. The machine-learned model autonomously analyzes the graphic novel, identifies speech bubbles and text regions, determines their reading order based on visual flow and narrative structure, and creates navigation metadata that enables direct access to any object without manual zooming or page flipping, thereby eliminating time loss.
Solution Approach 2:
The patent replaces the mechanical manual zooming and page-flipping system with an automated computational system. Instead of users physically interacting with the display to zoom in on objects, the machine-learned model computationally identifies objects and their locations, creates a structured representation, and enables direct navigation to any object of interest through programmed metadata, substituting mechanical user actions with automated computational processes.
3Loss of information
If conventional ereader navigation is used for graphic novels, then the device complexity remains low, but the loss of information increases as users cannot easily access detailed views of objects of interest
Solution Approach 1:
The system performs preliminary action by automatically detecting and identifying all objects of interest (speech bubbles, text regions, characters, important visual elements) before the user needs to access them. The machine-learned model creates a comprehensive structured representation including locations, types, and reading orders of all significant objects, ensuring that detailed information about any object is prepared and accessible without requiring the user to manually navigate through entire pages or use zoom functions.
Solution Approach 2:
The patent introduces an intermediary metadata layer that mediates between the graphic novel content and the user's information needs. This intermediary system automatically identifies objects of interest, extracts their detailed information, organizes it according to reading flow and narrative structure, and provides direct access to detailed views without requiring manual navigation, thereby preventing information loss while adding minimal device complexity.
Data Source
Figure 1
Figure 2
Figure 3~4
AI summary
Locations and presentation orders of objects of interest (e.g., speech bubbles) in digital graphic novel content are identified such that expanded versions of the objects of interest can be presented to a reader. Specifically, digital graphic novel content is received and locations of interest regions (e.g., rectangular text regions of speech bubbles) in the content are identified by applying a machine-learned model to the content. Locations and presentation orders of objects of interest in the digital graphic novel content are identified based on the identified locations of the interest regions. The digital graphic novel content and presentation metadata including the locations and presentation orders of the objects of interest are provided to a reading device such that expanded versions of the objects of interest are presented to the user in accordance with the presentation metadata.