Context-Based Dynamic Video Zooming for Visual Accessibility

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Individuals with visual impairments face difficulties in fully enjoying television and movie experiences due to missing contextually relevant visual details, despite efforts like straining vision or using glasses, as existing technologies do not effectively enhance video content accessibility.

Innovation Solution

A context-based dynamic zooming method that uses natural language processing to analyze text blocks from scripts or transcripts, correlate with video frames, and dynamically magnify the most relevant regions of interest through image segmentation, allowing for improved focus on key elements in video content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If individuals with visual impairments strain their vision or view from closer distance to see video content, then they can see more visual details, but they experience visual strain and fatigue

Engineering Contradiction:
Improvevisual detail detectionVSAvoidvisual strain
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The video content is segmented into multiple regions of interest based on contextual analysis of text blocks. Instead of requiring the viewer to see the entire screen, the system divides the video into relevant segments and presents only those portions that contain contextually important visual information, reducing the visual burden while maintaining comprehension.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies different quality levels to different regions of the video content. Regions identified as contextually relevant receive enhanced focus and magnification, while less important regions are minimized or excluded. This local quality enhancement allows visually impaired users to concentrate their visual effort on the most important areas without the strain of processing the entire frame.

Inventive Principle:
Principle #3Local quality

2Loss of information

If the entire video frame is displayed at standard resolution, then the overall context is preserved, but contextually relevant visual details are missed by visually impaired users

Engineering Contradiction:
Improvecontextual visual detailsVSAvoidvideo viewing accessibility
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The system performs preliminary analysis of the video content and text blocks before playback to identify contextually relevant regions. This advance preparation allows the system to pre-calculate and store region-of-interest data, enabling real-time magnification and focus enhancement without requiring complex processing during playback, thus improving accessibility while preserving contextual details.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces an intermediary processing layer between the video content and the user interface. This intermediary analyzes the relationship between text blocks and video frames, identifies relevant regions, and transforms the standard video output into an enhanced format that highlights contextually important visual details, making the content more accessible without losing overall context.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If the video content is magnified to help visually impaired users see details, then visual accessibility is improved, but the overall context and composition of the scene is lost

Engineering Contradiction:
Improvevisual accessibilityVSAvoidscene context
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The magnification level and region of interest are made dynamic rather than static. The system continuously adjusts which regions are magnified based on the current text block being processed and the temporal context in the video. This dynamic approach allows the magnified view to adapt to different scenes and moments, preserving overall context by only magnifying relevant portions at relevant times rather than continuously magnifying the entire scene.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20240305840A1Context-based dynamic zooming
Publication Date: 2024.09.12 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20240305840A1 patent drawing
  • US20240305840A1 patent drawing
  • US20240305840A1 patent drawing

AI summary

A method, computer system, and a computer program product for context-based dynamic zooming is provided. The present invention may include determining a context of a text block using natural language processing (NLP). The present invention may include identifying a clip of a video content synced with the text block, wherein the clip includes at least one frame of the video content. The present invention may include selecting a most relevant region of the at least one frame based on the context of the text block. The present invention may include storing the most relevant region of the at least one frame of the video content as a zoom region of the at least one frame. The present invention may include in response to displaying the at least one frame of the video content, dynamically magnifying the zoom region of the at least one frame of the video content.