Electronic Apparatus for AI-Assisted UI Location Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic apparatuses face challenges in efficiently extracting and updating user interface (UI) location information from various content providing apparatuses due to software upgrades and diverse UI layouts, leading to processing delays and legal issues with captured content images.
Innovation Solution
An electronic apparatus utilizing artificial intelligence models to process combined images, extracting UI location information by removing noise, and acquiring reliable content information through a server, with the ability to request and update UI location information as needed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If OCR or ACR functions are used to extract content information from the entire screen, then content information can be extracted, but processing time becomes excessively long
Solution Approach 1:
The patent divides the screen into multiple regions and creates separate OCR/ACR processing tasks for each region. The processor identifies specific areas where content information is likely to appear (such as title areas, metadata regions) and applies text recognition only to those segmented portions rather than the entire screen, thereby reducing processing time while maintaining extraction accuracy.
Solution Approach 2:
The patent applies different processing qualities to different screen regions. High-precision OCR/ACR processing is applied only to regions identified as containing content information (such as overlay text areas, metadata regions), while other regions receive minimal or no processing. This local quality approach optimizes resource allocation and reduces overall processing time.
2Productivity
If UI display areas are stored in advance for all content providing apparatuses, then content information can be extracted efficiently, but it becomes difficult to adapt to new apparatuses and software upgrades
Solution Approach 1:
The patent implements a dynamic system where UI display area information is not fixed but can be updated and adapted. The electronic apparatus stores UI area information for known content providing apparatuses and uses this stored information to efficiently locate and extract content information. When new apparatuses or software versions are encountered, the system can learn and adapt to their UI layouts, balancing efficiency with adaptability.
Solution Approach 2:
The patent performs preliminary storage of UI display area information for various content providing apparatuses during setup or initial operation. This pre-stored information enables rapid content extraction when those specific apparatuses are used. The system prepares ahead by memorizing UI layouts, so when the same apparatus is used again, extraction is highly efficient without requiring real-time analysis.
3Measurement precision
If the entire screen is captured and transmitted to an external server, then content information can be extracted remotely, but legal problems related to content rights occur
Solution Approach 1:
The patent extracts only the specific text or information elements needed from the screen content without capturing or transmitting the entire visual content. Instead of sending full screen images that contain protected content, the system extracts only metadata such as titles, channel names, or other textual information elements, thereby avoiding content right infringement while still achieving the goal of content identification.
Solution Approach 2:
The patent converts the potential harm of capturing copyrighted content into a benefit by using OCR/ACR to extract only textual metadata. The system recognizes that capturing full images creates legal risks, so it transforms the approach to extract only non-copyrighted textual information (titles, descriptions, metadata) that serves the same functional purpose without infringing content rights.
Data Source
AI summary
An electronic apparatus comprises a memory configured to store identification information of a content providing apparatus; a display; and at least one processor configured to: control the display to display a first content received from the content providing apparatus, based on a determination that a predetermined event has occurred, transmit the identification information of the content providing apparatus to the server, receive, from the server, user interface (UI) location information corresponding to the identification information, based on reception of a user instruction for changing the displayed first content to a second content, control the display to display the second content, and acquire content information corresponding to the second content based on the received UI location information, wherein the UI location information is acquired from a combined image that includes a plurality of images overlapped and merged into the combined image.


