NPU Video Decoding for Machine Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies lack an effective method for machine-based image analysis, particularly in handling high-resolution video data and supporting advanced AI applications with efficient encoding and decoding techniques.
Innovation Solution
A neural processing unit (NPU) is developed to process and decode video and feature maps, utilizing a bitstream format that includes weights for artificial neural networks, allowing for efficient image analysis without the need for additional memory to store weights, and enabling selective processing of enhancement layers based on specific machine analysis tasks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional video coding standards (H.264/AVC, HEVC) are used, then video compression is achieved, but machine-based image analysis capability is insufficient
Solution Approach 1:
The patent extends the video coding system to serve dual purposes: traditional video compression and machine-based image analysis. The encoder generates not only compressed video data but also feature maps and weight data that can be directly used by AI models. The decoder reconstructs both the video and the feature maps needed for analysis, making the system universally applicable to both entertainment and AI applications without requiring separate processing systems.
Solution Approach 2:
The patent introduces feature maps as an intermediary between the compressed video data and the AI analysis process. Instead of requiring raw video data to be processed by AI models, the system generates intermediate feature representations (feature maps) that bridge the gap between traditional video coding and machine learning, enabling efficient analysis while maintaining compression benefits.
2Measurement precision
If high-resolution video data is processed, then image quality is improved, but data processing burden increases
Solution Approach 1:
The patent segments the video data into multiple components: base layer data for compression and enhancement layers for quality improvement. Additionally, it extracts feature maps at different resolution levels, allowing the system to process only the necessary detail levels for AI analysis rather than treating the entire high-resolution video as a single monolithic data structure, thereby reducing processing burden while maintaining quality.
Solution Approach 2:
The patent extracts relevant features from high-resolution video data in the form of feature maps and weight data, separating the essential information needed for AI analysis from the redundant visual data. This extraction allows the system to maintain high-resolution image quality for analysis purposes while reducing the overall data volume that needs to be processed by transferring only the extracted features to the AI model.
3Reliability
If additional memory is allocated to store AI model weights, then machine analysis performance is improved, but device memory requirements increase
Solution Approach 1:
The patent merges the storage of AI model weights with the compressed video data structure. The weight data is integrated into the bitstream alongside the video information, eliminating the need for separate memory allocation. This combining approach allows the system to maintain reliable AI analysis performance by having weight data readily available during decoding while avoiding the additional memory overhead that would otherwise be required.
4Measurement precision
If all enhancement layers are decoded, then image quality is maximized, but processing time increases
Solution Approach 1:
The patent implements dynamic decoding where the system can adaptively select which enhancement layers to decode based on the specific AI analysis task requirements. Rather than statically decoding all available layers, the system dynamically adjusts the decoding depth and layer selection in real-time, allowing flexible optimization between processing time and feature map quality depending on the computational resources and analysis needs at any given moment.
Data Source
AI summary
A neural processing unit (NPU) for decoding video and/or feature map, the NPU may comprise at least one processing element (PE) for an artificial neural network, the at least one PE configured to receive and decode a bitstream. The bitstream may be received in a unit of data frame. One data frame of the bitstream may include a weight for an artificial neural network model, data of a base layer, and data of at least one enhancement layer. The data of the base layer included in the one data frame may include a first feature map, and the data of the at least one enhancement layer included in the one data frame may include a second feature map.


