Dynamic Closed Caption Positioning via Neural Network ROI Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Display devices struggle to effectively display closed captions without obscuring important information on the screen, leading to reduced information delivery and readability, especially when the captions' position is not adjusted in real-time.
Innovation Solution
A display device equipped with a neural network that detects regions of interest in an image, generates integrated regions by grouping adjacent ROIs, and determines a closed caption output region that avoids overlapping with important information, allowing the captions to be displayed in a non-obstructive area, with adjustments made to color and transparency as needed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If closed captions are displayed at a fixed position according to received attribute information, then the display device can show captions with proper positioning, but important information on the broadcast screen may be overlapped and obscured by the captions
Solution Approach 1:
The patent implements dynamic closed caption positioning by continuously detecting regions of interest in the broadcast screen and adjusting caption positions in real-time based on detected ROI locations. This transforms the static caption positioning system into a dynamic one that adapts to changing screen content, preventing overlap with important information while maintaining proper caption display.
Solution Approach 2:
The system employs feedback mechanisms by detecting regions of interest in the broadcast screen and using this information to adjust closed caption positions. The detected ROI data feeds back into the caption positioning algorithm, creating a closed-loop system that continuously optimizes caption placement to avoid obscuring important screen content.
2Loss of information
If the closed caption display position is adjusted in real-time to avoid important information, then information delivery and readability improve, but the device complexity increases due to additional processing requirements
Solution Approach 1:
The patent extracts the region of interest detection function as a separate processing module that identifies important screen areas independently. By isolating this detection task, the system can focus computational resources on analyzing only the relevant portions of the screen rather than processing the entire broadcast content, thereby reducing overall system complexity.
Solution Approach 2:
The broadcast screen is segmented into regions of interest and non-interest areas through neural network detection. This segmentation allows the system to apply different processing rules to different screen areas, simplifying the caption positioning logic by only needing to avoid specific detected regions rather than analyzing the entire screen continuously.
3Measurement precision
If neural network processing is used to detect regions of interest, then caption positioning accuracy improves, but the use of energy increases due to computational requirements
Solution Approach 1:
The system applies partial neural network processing by detecting only the essential regions of interest needed for caption positioning rather than performing comprehensive analysis of all screen elements. This selective detection approach maintains sufficient accuracy for avoiding caption overlap while reducing computational energy consumption compared to full-screen analysis.
Data Source
AI summary
An embodiment of the present disclosure relates to a display device including a display, a memory storing at least one instruction, and a processor configured to execute the at least one instruction stored in the memory to receive an image and a closed caption corresponding to the image, detect at least one region of interest (ROI) included in the image by using a neural network, generate at least one integrated region by grouping the at least one ROI into at least one group of adjacent ROIs, determine a closed caption output region among at least one preset candidate closed caption region, based on whether the at least one preset candidate closed caption region overlaps at least one of the at least one ROI and the at least one integrated region, and control the display to display the closed caption in the closed caption output region.


