Video Text Region Detection and Contrast Enhancement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video processing technologies face challenges in enhancing the readability and accessibility of on-screen captions and subtitles due to factors like small display size, poor eyesight, language difficulties, rapid text changes, and background colors, which hinder viewer understanding.
Innovation Solution
The method involves detecting text regions in video data, applying enhancements such as magnification, contrast adjustment, and optical character recognition (OCR) to improve readability, allowing user control over subtitle advancement, and displaying text in a secondary window or scrolling format, with options for translation and voice synthesis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If subtitles are displayed in the original video format, then the video content is preserved, but the readability of the text deteriorates due to small display size, poor contrast, and rapid text changes
Solution Approach 1:
The patent extracts text regions from the video signal using detection algorithms that identify subtitle areas based on motion vector analysis and histogram processing. Once extracted, these text regions are processed separately through enhancement filters and can be displayed in magnified form in dedicated overlay regions, separating the text processing from the main video content while maintaining both qualities
Solution Approach 2:
The patent applies different processing qualities to different regions of the video signal. Text regions detected in the video are subjected to local enhancement processing including contrast adjustment, sharpness filtering, and magnification, while the rest of the video content remains unchanged. This allows high-quality text display without compromising overall video quality
2Productivity
If text is displayed quickly to match dialogue pace, then the video timing is maintained, but the viewer's ability to read and understand the text deteriorates
Solution Approach 1:
The patent performs preliminary detection and enhancement of text regions before they are displayed. By pre-processing the text regions with enhancement filters and preparing magnified versions in advance, the system ensures that when text appears on screen, it is already optimized for readability, giving viewers adequate time to process the information without disrupting the dialogue timing
Solution Approach 2:
The patent adds a temporal dimension to text display by extending the duration that enhanced text regions remain visible on screen. Text can be held longer in magnified overlay regions, allowing viewers more time to read and comprehend the content while the main video continues to play at its original pace
3Reliability
If text is displayed in the original font and size, then the video fidelity is maintained, but the accessibility for viewers with poor eyesight or language difficulties deteriorates
Solution Approach 1:
The patent creates a multi-functional text display system that can serve multiple viewer needs simultaneously. The same detected text region can be displayed in its original form for fidelity, while also providing an enhanced magnified version for accessibility. The system can adapt between these modes based on viewer requirements, making it universally applicable to different viewing conditions and abilities
Solution Approach 2:
The patent applies parameter changes to text regions including magnification scaling, contrast ratio adjustment, and sharpness enhancement. These parameter modifications create an enhanced version of the text that maintains the original content fidelity while improving visual characteristics for viewers with accessibility needs
Data Source
AI summary
The present disclosure relates to methods and apparatus for detecting text information in a video signal that includes subtitles, captions, credits, or other text, and also for applying enhancements to the display of text areas in video. The sharpness and/or contrast ratio of subtitles of detected text areas may be improved. Text areas may be displayed in a magnified form in a separate window on a display, or on a secondary display. Further disclosed are methods and apparatus for extending the duration for which subtitles appear on the display, for organizing subtitles to be displayed in a scrolling format, for allowing the user to control when a subtitle advances to the next subtitle using a remote control, and for allowing a user to scroll back to a past subtitle in cases where the user has not finished reading a subtitle. Additionally, optical character recognition (OCR) technology may be applied to detected areas of a video signal that include text, and the text may then be displayed in a more readable font, displayed in a translated language, or rendered using voice synthesis technology.


