Unified VLSI Engine for HEVC SAO Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing video encoding standards, particularly H.264, face challenges in achieving high performance and area efficiency for advanced filtering techniques like Sample Adaptive Offset (SAO) in HEVC, which is crucial for next-generation Ultra HDTV at 4K resolution and 60 frames per second, while maintaining compliance with bit-rate reduction and video quality.
Innovation Solution
A unified processing engine is designed to collect statistics on original and encoded/decoded pixels, determine optimal SAO parameters, and operate on a three-stage pipeline for efficient SAO filtering, incorporating programmable look-up tables and override mechanisms to enhance video quality and bit-rate savings.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If sophisticated SAO filtering is implemented in HEVC, then video quality and bit-rate efficiency are improved, but device complexity and processing requirements increase
Solution Approach 1:
The SAO filtering process is divided into distinct stages: statistics collection, offset calculation, and filtering application. Each stage processes specific data independently, allowing parallel execution and reducing overall complexity while maintaining video quality improvements
Solution Approach 2:
The filtering parameters and offset values are dynamically adjusted based on local image characteristics and statistics collected from the video content. This adaptive approach improves video quality by tailoring filtering to specific regions while the modular architecture manages the resulting complexity
2Productivity
If SAO filtering processes 4K video at 60 fps, then productivity and processing speed are improved, but use of energy and computational resources increase
Solution Approach 1:
Statistics for SAO filtering are collected and prepared in advance during the encoding process, allowing the actual filtering operation to execute efficiently during decoding. This preliminary preparation enables 4K@60fps processing by pre-computing offset values and categorization data that would otherwise require intensive real-time computation
Solution Approach 2:
The three-stage pipeline architecture ensures continuous processing of video data through overlapping operations: statistics collection, offset calculation, and filtering application occur in parallel across different data blocks. This continuous throughput achieves high frame rates while distributing computational load to manage energy consumption
3Area of stationary object
If area efficient VLSI architecture is used, then device area is reduced, but processing capability and productivity may be limited
Solution Approach 1:
The VLSI architecture employs unified processing units that handle multiple SAO filtering operations including luminance and chrominance processing, different filtering types (edge offset and band offset), and both encoding and decoding functions. This multi-functionality reduces overall device area while maintaining processing capability for 4K video at 60 fps through resource sharing and parallel operation modes
Data Source
AI summary
An apparatus for sample adaptive offset (SAO) filtering in video encoding. A unified processing engine collects statistics on a block of pixels, determines a minimum RD cost (J) for each category of band offsets and edge offsets; determines a RD cost to find the optimal SAO type and determines a cost for each of the left SAO parameters and the up SAO parameters. The unified processing engine operates for three iterations: once for luminance once for each chrominance. A SAO merge decision unit determines an optimal mode and generates current LCU Parameters. The RD offset unit determination includes determining whether the sign of the minimum offset is proper for the category of edge offset. The RD offset is determined using a programmable look-up table indexed by the offset to estimate a rate. The unified processing engine operates on a three stage pipeline: loading blocks; processing; and updating blocks.


