Screen-to-Camera Communication via Temporal Domain Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current screen-to-camera communication systems face challenges such as noisy channels, poor image quality, perspective distortion, motion blur, and low signal-to-noise ratios, making it difficult to embed and retrieve information effectively without disrupting the viewing experience, especially in public display settings like digital billboards and TVs.
Innovation Solution
A computer-implemented system that converts information into data symbols, pilot symbols, and scannable barcodes, embedding them in media content frames using adaptive modulation techniques across color and luminance channels, allowing for unobtrusive and personalized information transmission and retrieval using a hand-held device, while addressing issues like perspective distortion and low SNR ratios.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If classical steganography or transform domain techniques are used to embed information, then information can be hidden in images, but the screen-camera channel noise and low-pass filtering effects destroy the transmitted information
Solution Approach 1:
The patent changes the frequency domain parameters by transmitting information in the temporal domain rather than spatial frequency domain. This avoids the low-pass filtering effects that destroy high-frequency information while being robust to the specific noise characteristics of the screen-camera channel.
Solution Approach 2:
The patent replaces classical steganography methods with a temporal domain transmission approach that uses synchronized frame manipulation. Instead of embedding information in spatial frequencies, it uses temporal patterns across consecutive frames that are resilient to channel noise and filtering.
2Loss of information
If QR codes are used for screen-to-camera communication, then information can be transmitted, but the aesthetic appearance is disrupted and user experience is interfered with
Solution Approach 1:
The patent extracts the information transmission function from visible visual elements like QR codes. By using imperceptible temporal patterns in the video frames, it separates the communication function from the visual content, allowing information transmission without aesthetic disruption.
Solution Approach 2:
The patent introduces temporal patterns as an intermediary carrier of information. These patterns are embedded in the temporal domain across consecutive frames and are imperceptible to human viewers, yet can be detected and decoded by a camera-based receiver system.
3Loss of information
If deep neural networks are used to embed and recover information, then information can be hidden in images and videos, but it is challenging to model the screen-camera channel accurately and networks are trained on specific datasets making them unsuitable for real-world scenarios
Solution Approach 1:
The patent creates a self-contained temporal domain transmission system that does not rely on external training data or complex neural network models. The method uses deterministic temporal patterns that can be generated and decoded without requiring adaptation to specific datasets or channel conditions.
Solution Approach 2:
The patent segments the information transmission into temporal components across multiple frames rather than relying on a monolithic neural network model. This allows the system to handle diverse real-world scenarios by processing temporal patterns frame-by-frame without requiring retraining.
4Loss of information
If intensity modulation is used for screen-to-camera communication, then information can be transmitted, but flicker occurs at low refresh rates and precise screen detection is required
Solution Approach 1:
The patent moves the information transmission from the spatial/intensity domain to the temporal domain. By encoding information in temporal patterns across consecutive frames rather than intensity variations within single frames, it eliminates flicker issues and removes the need for precise screen detection.
Data Source
AI summary
A system and method for managing encoded information in a real-time screen-to-camera communication environment are disclosed. The system converts information into a pre-defined number of characters and generates data symbols in shapes and pilot symbols corresponding to the characters. Further, the system embeds the data symbols in media content frames and modulates pixels and boundaries for display of display device, based on luminance, and adaptively displays frames as temporal-complementary frames. Furthermore, the system detects frames from recorded content, extracts data symbols based on grid and fixed pattern, and detects bit values by analyzing color differences. Additionally, the system generates information based on the detected bit values and outputs the information on an user device display, including products, recommendations, services, and relevant information related to the media content.


