Multi-Layer Video Encoding for Bandwidth-Constrained AI Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing AI systems face challenges in efficiently processing and analyzing high-resolution video and audio signals captured remotely, as they require significant bandwidth and processing resources, making it difficult to perform advanced neural network analysis on-site without transferring large amounts of data to data centers or sacrificing analysis capabilities by moving processing resources to edge devices.
Innovation Solution
Implementing a multi-layer encoding scheme that breaks down signals into hierarchical tiers of quality, allowing for efficient transmission of lower-resolution data for initial analysis and on-demand transmission of higher-resolution details, using standards like SMPTE VC-6 and LCEVC, which enables cloud-based AI analysis and reduces bandwidth requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If high-resolution video signals are transmitted from remote locations to data centers for AI analysis, then analysis accuracy is improved, but bandwidth requirements and transmission time increase significantly
Solution Approach 1:
The video signal is segmented into multiple quality layers (base layer and enhancement layers) with different resolutions. The base layer provides low-resolution data for initial analysis, while enhancement layers provide additional high-resolution details only when needed, reducing overall bandwidth requirements while maintaining analysis accuracy when full resolution is required.
Solution Approach 2:
The system dynamically adjusts the transmission quality and resolution of video signals based on real-time bandwidth availability and analysis needs. During bandwidth constraints, only base layer data is transmitted; when bandwidth is available, enhancement layers are added to improve analysis accuracy, making the system adaptable to varying network conditions.
2Loss of time
If advanced neural network analysis is performed on-site at remote locations, then transmission time is reduced, but processing resource requirements and device complexity increase
Solution Approach 1:
The neural network analysis is segmented into multiple stages: initial analysis is performed on-site using lightweight models with base layer data, while more complex analysis is performed at the data center with enhancement layer data. This distribution reduces transmission time for critical decisions while maintaining overall analysis capability.
Solution Approach 2:
Instead of performing complete high-resolution analysis on-site, the system performs partial analysis using lower-resolution base layer data for immediate decisions, and supplements with full-resolution analysis at the data center when needed, achieving a balance between speed and accuracy.
3Measurement precision
If full-resolution video data is transmitted continuously, then analysis quality is maintained, but transmission bandwidth and energy consumption increase
Solution Approach 1:
The system transmits video data periodically at different quality levels. Base layer data is transmitted continuously at low resolution, while enhancement layer data is transmitted periodically or on-demand at high resolution, reducing overall energy consumption while maintaining analysis quality when needed.
Solution Approach 2:
The transmission parameters (resolution, bitrate, quality level) are dynamically changed based on bandwidth availability and analysis requirements. The system switches between transmitting only base layer data and transmitting full multi-layer data, optimizing energy consumption while maintaining analysis quality.
Data Source
AI summary
The present disclosure relates to a method of analysing a plurality of video camera feeds, the method comprising: encoding, at a first location, the plurality of video camera feeds using a layer-based encoding, including generating encoded data streams for each of a plurality of layers within the layer-based encoding, wherein different layers in the plurality of layers correspond to different spatial resolutions, higher layers representing higher spatial resolutions; transmitting, to a second location remote from the first location, encoded data streams for one or more lowest layers for the plurality of video camera feeds; decoding, at the second location, the encoded data streams to generate a set of reconstructions of the plurality of video camera feeds at a first spatial resolution; applying one or more video analysis functions to the set of reconstructions to identify one or more video camera feeds for further analysis; sending, to the first location for the identified one or more video camera feeds for further analysis, a request for further encoded data streams for one or more layers above the one or more lowest layers; responsive to the request, transmitting, to the second location, the further encoded data streams for one or more layers above the one or more lowest layers; decoding, at the second location, the further encoded data streams to generate a set of reconstructions for the identified one or more video camera feeds at a second spatial resolution; and applying one or more video analysis functions to the set of reconstructions at the second spatial resolution.


