Video Bitstream Text Prompt Signaling via SEI Messages
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video coding standards lack a protocol for carrying imagery along with corresponding supplemental information to neural network models that generate new or similar image-based results, necessitating direct referencing of neural networks from applications rather than inclusion in the coded video stream.
Innovation Solution
A method and apparatus for video processing that involves receiving a text prompt from a user device to instruct an AI generative process, encoding a video bitstream with images generated using the AI process, signaling the text prompt in a supplemental enhancement information (SEI) message within the bitstream, and signaling the encoded video containing these images.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If neural network models are carried or referenced within the coded video stream, then AI processing can be performed on decoded video, but the device complexity and bandwidth requirements increase significantly
Solution Approach 1:
The patent extracts the neural network model references from the video stream by using SEI messages that point to external model locations rather than embedding the models directly in the stream. This allows AI processing capability while reducing the complexity and bandwidth requirements of the video stream itself.
Solution Approach 2:
The patent introduces SEI messages as an intermediary mechanism between the video stream and neural network models. These messages serve as references that point to external model locations, enabling AI processing without requiring the models to be carried within the coded stream, thus resolving the contradiction between automation capability and device complexity.
2Manufacturing precision
If uncompressed video is transmitted, then image quality is preserved, but bandwidth and storage requirements become excessively high
Solution Approach 1:
The patent segments the video data into compressed video streams with supplemental enhancement information. This allows the video to be transmitted in a compressed form (reducing bandwidth) while maintaining sufficient quality through the compression standards and selective enhancement where needed.
Solution Approach 2:
The patent changes the parameter of video compression from uncompressed to compressed formats (such as H.264, H.265, AV1), which significantly reduces bandwidth and storage requirements while maintaining acceptable image quality for the intended applications.
3Quantity of substance
If lossy compression is applied to video, then bandwidth and storage requirements are reduced, but image quality deteriorates
Solution Approach 1:
The patent applies local quality enhancement by using SEI messages to provide supplemental information only where needed, rather than uniformly compressing the entire video stream. This allows lossy compression to reduce bandwidth while maintaining acceptable quality in most regions, with targeted enhancement applied locally when necessary.
4Adaptability or versatility
If AI-generated images are integrated into video streams, then content versatility is improved, but the complexity of video processing increases
Solution Approach 1:
The patent makes the video processing system universal by enabling it to handle both traditional video content and AI-generated content through the same SEI message framework. The system can process conventional video streams and integrate AI-generated images using the same compression and enhancement mechanisms, thus improving versatility without proportionally increasing complexity.
Data Source
AI summary
A method and apparatus comprising computer code for video processing, the method including receiving a text prompt from a user device, the text prompt being used for instructing an artificial intelligence (AI) generative process to generate one or more images; encoding a video in a bitstream, the video comprising the one or more images generated using the AI generative process and the text prompt; signaling the text prompt in the bitstream in a supplemental enhancement information (SEI) message; and signaling the encoded video comprising the one or more images.


