Picture Decoding With Lightweight Attention Transformation Networks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current composite transformation networks in picture coding and decoding based on deep learning do not consider computation efficiency, leading to low picture coding and decoding performance.
Innovation Solution
Implement a composite transformation network with K types of lightweight attention modules, each with computation complexity less than a preset value, to improve picture coding and decoding efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If a composite transformation network with traditional attention modules is used, then picture processing effect is improved, but computation complexity increases
Solution Approach 1:
The patent segments the traditional attention module into multiple lightweight attention modules, each responsible for specific processing tasks. This segmentation allows the system to achieve comprehensive picture processing effects while keeping each individual module's computation complexity low and controllable.
Solution Approach 2:
The patent applies local quality by designing different lightweight attention modules with specialized functions for different parts of the picture processing pipeline. Each module is optimized for its specific local task, achieving high processing effect in each region while maintaining overall low computation complexity.
2Manufacturing precision
If a composite transformation network with traditional attention modules is used, then picture processing effect is improved, but decoding time increases
Solution Approach 1:
By segmenting the processing into multiple lightweight modules that can be executed in parallel or sequential stages, the patent reduces the overall decoding time while maintaining picture processing quality. Each lightweight module processes faster than traditional monolithic attention modules.
Solution Approach 2:
The patent changes the computational parameters of the attention modules by using lightweight variants with reduced computation complexity. This parameter optimization maintains the essential processing effects while significantly reducing the time required for decoding operations.
3Productivity
If computation complexity is reduced using lightweight modules, then decoding efficiency is improved, but picture processing effect may deteriorate
Solution Approach 1:
The patent merges multiple lightweight attention modules into a composite transformation network, where each module contributes specific processing capabilities. The combination of these lightweight modules achieves comprehensive picture processing effects that match or exceed traditional modules, while maintaining low individual computation complexities.
Solution Approach 2:
The patent creates a composite transformation network by combining different types of lightweight attention modules, each with specialized functions. This composite structure leverages the strengths of each module type to achieve high picture processing effects while maintaining overall decoding efficiency.
Data Source
AI summary
This application provides picture coding and decoding methods performed by a computer device, which may be applied to fields such as picture processing, video coding and decoding, and video livestreaming. The picture decoding method includes: decoding a bitstream of a current picture, to obtain a residual value of the current picture, determining a predicted value of the current picture based on the decoded bitstream, and determining a transformed value of the current picture based on the residual value and the predicted value; and processing the transformed value of the current picture by using a composite transformation network, to obtain a reconstructed picture of the current picture, where the composite transformation network includes K types of lightweight attention modules, and computation complexities of the K types of lightweight attention modules are each less than a preset value.


