A multi-level multi-module collaborative video perception coding optimization method and device
By employing a multi-level, multi-module collaborative video perception coding optimization method, the problems of performance coupling and distortion transmission in video coding are solved, improving video compression efficiency and quality consistency, and achieving more efficient video coding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU DIANZI UNIV
- Filing Date
- 2023-03-02
- Publication Date
- 2026-04-14
AI Technical Summary
Existing video coding technologies suffer from performance coupling and distortion propagation issues in multi-module perceptual coding optimization, which limits the improvement of compression efficiency.
A multi-level, multi-module collaborative video perception coding optimization method is adopted. By decoupling the multi-module coding process through coding distortion prediction, frame-level quantization parameter derivation, residual filtering, perception quantization and quality enhancement network, intra-frame/inter-frame prediction and entropy coding are optimized.
It improves video compression efficiency, reduces encoding distortion, enhances the consistency of subjective and objective quality of compressed video, and alleviates the performance coupling and distortion propagation effects between modules.
Smart Images

Figure CN116193122B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of video compression technology, specifically relating to a multi-level, multi-module collaborative video perception coding optimization method and apparatus. Background Technology
[0002] With the exponential growth of video data, higher demands are being placed on video compression. Continuous and in-depth optimization of lossy video coding technology and improvement of video compression efficiency remain the focus of the industry. The continuous improvement in the performance of video capture and playback equipment has fostered user demand for high-bandwidth, high-definition video services. There is a fierce contradiction between relatively limited storage resources, network bandwidth, and ever-increasing user demand. The main solution is continuous and in-depth optimization of the lossy video codec framework to continuously improve video compression efficiency. Related research indicates that visual perception-based video coding optimization methods will provide a selection for the optimal video compression solution.
[0003] In fact, for a series of compressed videos of varying quality, the human visual system (HVS) can only perceive a finite number of subjective quality levels, each corresponding to a perceptual coding distortion threshold. For video compression, the core objective of perceptual coding optimization is to maximize bitrate savings under constraints of different quality levels. Furthermore, video compression is a multi-module serial coding process; maximizing bitrate savings necessitates collaborative optimization across multiple modules.
[0004] The existing technology has the following technical defects:
[0005] Video coding optimization has always been a hot and challenging research area in multimedia communication. In recent years, standardization organizations have successively released new-generation video coding standards such as AVS3 and H.266 / VVC for lossy video coding. However, these latest standards still employ a block-based hybrid coding framework, optimizing each coding stage to squeeze out remaining information redundancy, without fundamentally improving the coding algorithm. Compared to the exponentially increasing video data, compression efficiency remains insufficient, and further improving coding efficiency remains a focus of industry attention. As the final receiver of images and videos, the human visual system (HVS) has been a hot research topic in video compression, particularly in video coding optimization based on HVS visual perception mechanisms. Existing video perception coding optimization methods can be broadly categorized into two types: encoding preprocessing and encoding / decoding embedding.
[0006] Most preprocessing methods for encoding design directly filter out visual redundancy in images and videos using perceptual filters. These methods do not consider video compression mechanisms, and the preprocessing and encoding processes suffer from performance coupling effects, resulting in limited improvement in encoding gain. Encoding-decoding embedding methods are largely based on optimizing different encoding modules in video coding according to visual perception characteristics, such as residual filtering, perceptual quantization, and perceptual rate-distortion optimization. These methods mostly optimize single modules, leaving room for further improvement in encoding gain. Some techniques have explored multi-module perceptual coding optimization, but they have not fully considered the performance coupling effects and distortion propagation effects between modules. Performance coupling in multi-module perceptual coding optimization is mainly reflected in the following two aspects: 1) Multi-module coding distortion may cause the overall coding distortion to exceed the perceptual coding distortion threshold, leading to a loss in the subjective quality of compressed video; 2) The process of removing perceptual redundancy between modules may affect each other, leading to a loss in the overall coding gain. Distortion propagation in multi-module perceptual coding optimization is mainly reflected in the serial propagation of distortion between modules and the accumulation of distortion propagation in temporal coding.
[0007] Video-aware coding optimization technology is relatively mature, but it still faces performance bottlenecks, such as:
[0008] Performance coupling issues in multi-module perceptual coding: Video coding is a serial multi-module collaborative coding process. Single-module perceptual coding optimization cannot achieve maximum video perceptual compression, while multi-module perceptual coding optimization will cause performance coupling effects, resulting in overall coding distortion exceeding the perceptual coding distortion threshold and reducing the quality of compressed video.
[0009] The distortion propagation problem in multi-module sensing coding is that, on the one hand, the distortion in multi-module sensing coding will be transmitted serially between coding modules, resulting in uncontrollable coding distortion in the current frame; on the other hand, it will be transmitted in the coding of time-series frames, causing coding distortion to accumulate in time, thereby continuously reducing coding quality in time.
[0010] Due to performance coupling and distortion propagation issues between modules, multi-module collaborative sensing coding optimization requires further in-depth research. Exploring decoupled multi-module collaborative sensing coding optimization strategies is expected to further improve video compression efficiency. Summary of the Invention
[0011] To address the shortcomings of existing technologies and improve coding quality, this invention adopts the following technical solution:
[0012] A multi-level, multi-module collaborative video perception coding optimization method includes the following steps:
[0013] Step S1: Perform coding distortion prediction, frame-level coding distortion prediction, and derive frame-level quantization parameters from the original video.
[0014] Step S2: Perform intra-frame / inter-frame prediction on the images of the original video, and calculate the difference between the obtained predicted image and the original image to obtain a residual image. Perform residual filtering on the residual image through the predicted coding distortion. After the filtered residual image is transformed based on the residual block, perceptual quantization is performed according to the predicted frame-level coding distortion and frame-level quantization parameters.
[0015] Step S3: Optimize rate distortion based on perceptual quantization parameters to improve intra-frame / inter-frame prediction;
[0016] Step S4: Perform inverse quantization and inverse transformation on the perceptual quantization result to obtain the reconstructed frame. Construct a perceptual quality enhancement network. Based on the reconstructed frame, perceptual quantization parameters, coupling strength between residual filtering and perceptual quantization, and predicted coding distortion, perform coding distortion prediction to obtain a distortion prediction map. Enhance the quality of the reconstructed frame using the distortion prediction map to obtain a quality enhancement map, which is used to optimize intra / inter-frame prediction.
[0017] Step S5: Based on optimized intra / inter-frame prediction, perform prediction, difference calculation, residual filtering, transformation, and perceptual quantization on the images of the original video, and then perform entropy coding.
[0018] Furthermore, in step S2, residual filtering filters out residuals with amplitude values lower than the coding distortion threshold, based on the coding distortion threshold obtained from the coding distortion prediction.
[0019]
[0020] Where sign(·) represents the sign operation, R0 and J represents the original residual and the filtered residual at the i-th quality level, respectively. i Let (x, y) represent the pixel-level perceptual coding distortion threshold for the i-th quality level, where (x, y) represents the pixel coordinates.
[0021] Further, in step S2, the frame-level quantization distortion is derived based on the frame-level distribution law of the transform coefficients and applied to perceptual quantization; a D-QP model is constructed according to quantization compensation and the frame-level distribution of the transform coefficients, and its coding distortion is calculated under the constraint of the frame-level coding distortion threshold obtained by frame-level coding distortion prediction; the pre-quantization parameters at the transform unit level are derived in reverse from the D-QP model at the transform unit level and the coding distortion threshold at the transform unit level; the final perceptual quantization parameters of the current transform unit are determined according to the frame-level perceptual quantization parameters, the pre-quantization parameters of the current transform unit, and the perceptual quantization parameters of neighboring transform units.
[0022] Specifically, frame-level quantization distortion D f Based on the quantization step size, it can be approximated as:
[0023]
[0024] Where Qstep represents the quantization step size, and f(C) represents the frame-level distribution function of the transform coefficients C. The transform coefficient distributions of different encoders (e.g., AVS3 and H.266 / VVC) are not entirely consistent. Therefore, a D-QP model D adapted to the encoder is constructed. f (f(C),Qstep), where the coding distortion approaches the frame-level perceived coding distortion threshold. Under constraints, solve for the frame-level aware quantization parameters that are just perceptible at the i-th quality level. This process can be represented as:
[0025]
[0026] Where SQ represents the subsequent quantization parameter set, in the actual encoding process, for each transform unit TU, according to the D-QP model D at the transform unit TU level... t (Qstep), the perceptual coding distortion threshold at the transform unit (TU) level. Reverse derive the prequantization parameters of a transform unit TU level. Right now:
[0027]
[0028] To avoid block artifacts during encoding, frame-aware quantization parameters are used. Prequantization parameters of the current transform unit In addition to the perceptual quantization parameters of neighboring transform units, the final perceptual quantization parameters of the current transform unit are determined. The functional expression of this process can be represented as:
[0029]
[0030] in Let Γ(·) represent the set of transform block sensing quantization parameters in the neighborhood, and let Γ(·) represent the functional relationship that needs further investigation.
[0031] Furthermore, in step S2, video compression is a multi-module serial encoding process. The perceptual residual filtering process and the perceptual adaptive quantization process inevitably introduce encoding distortion, and since they implement residual compression in the spatial domain and frequency domain respectively, there is a performance coupling effect. By using the perceptual all-zero block decision unit and the quantization filter, the residual filtering and perceptual quantization are decoupled, and the video perceptual compression limit is approached while eliminating the coupling effect. When the magnitude of the residual coefficients in the residual image block is less than the quantization distortion prediction value, the perceptual all-zero block decision unit forces the current residual image block to all zeros and does not perform subsequent transformation and entropy encoding. Otherwise, the quantization filter introduces quantization distortion to control the residual filtering intensity. When the magnitude of the residual coefficients in the residual image block is less than the product of the encoding distortion threshold and the filter intensity parameter, the quantization filter filters the current residual image block. The quantization distortion prediction value is obtained by quantizing the current residual image block after transformation using the perceptual quantization parameter. The filter intensity parameter is generated based on the coupling degree between residual filtering and perceptual quantization, perceptual encoding distortion, and the quantization distortion prediction value of the current transformed residual block.
[0032] Specifically, the judgment process is represented as follows:
[0033]
[0034] in This indicates that the current residual block is transformed using... When performing quantization, if the magnitude of all residual coefficients in the residual block is less than the predicted quantization distortion value, it means that the residual block will likely be quantized into an all-zero block in subsequent quantization processes. Therefore, in this case, the current residual block will be forced to be all zero, and no further transformation and entropy coding will be performed. The all-zero block perception decision can improve video compression efficiency and speed up video compression.
[0035] If the current residual block is determined to be a non-all-zero block, it means that the current residual block will undergo subsequent transformation and quantization processes, introducing quantization distortion. In this case, the present invention designs a quantization filter to introduce quantization distortion to control the residual filtering strength, weaken the performance coupling effect between modules, and further improve video compression efficiency. The quantization filtering process is as follows:
[0036]
[0037] Where κ represents the filter strength parameter, and C represents the coupling parameter. cr Perceptual coding distortion and the current transform block coding distortion prediction value The function, that is:
[0038]
[0039] Where Ccr This represents the coupling parameter, which can be determined based on video compression experiments. C cr The values 0 and 1 represent the two cases where residual filtering and perceptual quantization are uncoupled and fully coupled, respectively. C cr A value between 0 and 1 represents the degree of coupling between residual filtering and perceptual quantization; based on the coding gain maximization criterion, exploratory research has revealed that the coupling parameter C... cr The optimal value is closely related to the coding prediction method, and the coupling strength in inter-frame coding is significantly higher than that in intra-frame coding.
[0040] Furthermore, the objective distortion criterion based on mean square error (MSE) cannot fully reflect the perceptual characteristics of the Human Visual System (HVS) for compressed video. Therefore, the rate-distortion optimization in step S3 is based on prediction to measure perceptual coding distortion and applies it to the rate-distortion optimization of video coding. Its rate-distortion cost is determined by perceptual coding distortion and the product of the Lagrange multiplier determined by the quantization parameters and the coding bitrate. Perceptual coding distortion is measured by the sum of perceptual squared errors and the sum of perceptual absolute errors.
[0041] Specifically, using the perceived squared error and With perceptual absolute error and To measure encoding distortion, the cost of perceptual rate distortion. Represented as:
[0042]
[0043] in, Let represent the perceived distortion at the i-th quality level, Bit represent the coding rate, and λ(QP) represent the Lagrange multiplier determined by the quantization parameter QP.
[0044] Furthermore, in step S3, the intra / inter-frame prediction is optimized by introducing rate-distortion optimization. The intra / inter-frame prediction mode with the lowest rate-distortion cost is selected from the candidate intra / inter-frame prediction modes to determine the optimal intra / inter-frame prediction mode.
[0045]
[0046] in M represents the optimal intra-frame or inter-frame prediction mode in a perceptual sense. pre Indicates the candidate intra-frame or inter-frame prediction mode. This indicates that the quantization parameters are sensed by the current transform block. Determined Lagrange multipliers.
[0047] Furthermore, in step S4, the transmission of distortion between modules inevitably reduces the subjective and objective quality of the current frame, and the accumulation of distortion in the temporal sequence between frames inevitably affects the overall subjective quality consistency of the compressed video. This invention improves the subjective and objective quality consistency of the compressed video based on a distortion prediction quality enhancement network. Specifically, it optimizes the video perceptual coding process through keyframe quality enhancement technology to improve the subjective and objective quality of the compressed video. Considering the complexity constraints of video compression, this invention adaptively selects keyframes (such as intra-coded frames or some bidirectional prediction frames in a GOP) from the encoded video frames to enhance the subjective and objective quality of the keyframes and weaken the distortion transmission effect in video coding.
[0048] Further, in step S4, the image of an original video frame is divided into multiple non-overlapping coding units. Each coding unit is encoded independently, and encoding processes such as prediction, transformation, quantization, inverse quantization, and inverse transformation are performed. The resulting images are then reconstructed and integrated into a reconstructed frame, which is then subjected to loop filtering and placed in the decoding buffer as a reference for subsequent encoded frames. The perceptual quality enhancement network is applied to the reconstructed image in the loop filtering and / or the decoding buffer. Taking the reconstructed frame, perceptual quantization parameters, coupling strength, and perceptual coding distortion threshold distribution as inputs, the network sequentially performs upsampling, multi-scale feature extraction, and adaptive feature fusion to obtain a distortion prediction map. The distortion prediction map and the reconstructed frame are then input into a compressed video quality enhancement network based on residual dense network (RDN) to obtain a quality enhancement map.
[0049] Specifically, the coding distortion prediction is represented as:
[0050]
[0051] F rec Indicates the reconstructed frame. Let κ represent the set of perceptual quantization parameters for all coding units in the current frame, and let J represent the coupling strength between residual filtering and perceptual quantization. i This represents the perceptual coding distortion threshold for the i-th level. This represents the coding distortion prediction function during the quality enhancement process.
[0052] Furthermore, the adaptive feature fusion is a multi-scale adaptive feature fusion. Due to the diversity of network inputs, features of different scales have different effects on coding distortion prediction. Therefore, this invention introduces an attention mechanism in the adaptive feature fusion network to assign different weights to features of different channels and scales in a learning manner. Features of different scales are sampled to their original size and input into convolutional layers and activation layers to obtain the weight assignment of each element. After the features are multiplied by the weights, average pooling and max pooling are performed to extract features. After learning the channel weights through convolutional layers and activation layers, the coding distortion distribution map, i.e., the distortion prediction map, is finally obtained.
[0053] A multi-level, multi-module collaborative video perception coding optimization device includes a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the multi-level, multi-module collaborative video perception coding optimization method.
[0054] The advantages and beneficial effects of this invention are as follows:
[0055] This invention provides a multi-level, multi-module collaborative video perception coding optimization method and apparatus. Based on single-module perception coding optimization, it avoids performance coupling effects in multi-module coding by decoupling multi-module collaborative perception coding optimization strategies; and improves compressed video quality and alleviates distortion in multi-module perception coding by using intelligent compressed video perception quality enhancement technology.
[0056] Delivery Attached Figure Description
[0057] Figure 1 This is a flowchart of a multi-level, multi-module collaborative perceptual coding optimization method in an embodiment of the present invention.
[0058] Figure 2 This is a schematic diagram of decoupled multi-module collaborative video coding optimization in an embodiment of the present invention.
[0059] Figure 3 This is a schematic diagram of a compressed video quality enhancement network based on perceptual coding distortion prediction in an embodiment of the present invention.
[0060] Figure 4 This is a schematic diagram of the adaptive feature fusion network structure in an embodiment of the present invention.
[0061] Figure 5 This is a schematic diagram of the structure of a multi-level, multi-module collaborative perception coding optimization device in an embodiment of the present invention. Detailed Implementation
[0062] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0063] like Figure 1As shown, a multi-level, multi-module collaborative video perceptual coding optimization method is divided into a preprocessing stage and a perceptual coding stage. In the preprocessing stage, a preprocessing module performs perceptual coding distortion prediction, frame-level coding distortion prediction, and multi-granularity perceptual QP derivation on the input raw video. In the perceptual coding stage, a perceptual coding module optimizes the multi-level single-module video perceptual coding, including perceptual (spatial domain) residual filtering, transform, perceptual adaptive quantization (CU level), inverse quantization / inverse transform, and introduces a perceptual quality enhancement network to the perceptual quality enhancement module (loop filtering, decoding buffer) for intelligent compressed video perceptual quality enhancement, which is applied to the perceptual RDO module (intra-frame / inter-frame prediction). Based on decoupled multi-module collaborative coding optimization, intra-frame / inter-frame prediction based on perceptual rate distortion optimization, decoupled residual filtering, and perceptual adaptive quantization are performed. Specifically, the method includes the following steps:
[0064] Step S1: Perform coding distortion prediction, frame-level coding distortion prediction, and derive frame-level quantization parameters from the original video.
[0065] Step S2: Perform intra-frame / inter-frame prediction on the images of the original video, and calculate the difference between the obtained predicted image and the original image to obtain a residual image. Perform residual filtering on the residual image through the predicted coding distortion. After the filtered residual image is transformed based on the residual block, perceptual quantization is performed according to the predicted frame-level coding distortion and frame-level quantization parameters.
[0066] Perceptual residual filtering: When the magnitude of the predicted residual is lower than the perceptual coding distortion threshold, this part of the residual will not be detected by HVS even without encoding. Therefore, the present invention designs the following filter to filter out residuals with magnitude values lower than the perceptual coding distortion threshold, thereby reducing the amount of video compression data.
[0067] Specifically as follows:
[0068]
[0069] Where sign(·) represents the sign operation, R0 and J represents the original residual and the filtered residual at the i-th quality level, respectively. i Let (x, y) represent the pixel-level perceptual coding distortion threshold for the i-th quality level, where (x, y) represents the pixel coordinates.
[0070] Perceptual adaptive quantization: The coding distortion in traditional video compression mainly comes from the quantization module. Based on the distribution law of the transform coefficients, the frame-level quantization distortion expression is derived and applied to the quantization module of video encoding.
[0071] Specifically, frame-level quantization distortion D f Based on the quantization step size, it can be approximated as:
[0072]
[0073] Where Qstep represents the quantization step size, and f(C) represents the frame-level distribution function of the transform coefficients C. The transform coefficient distributions of different encoders (e.g., AVS3 and H.266 / VVC) are not entirely consistent. Therefore, a D-QP model D adapted to the encoder is constructed. f (f(C),Qstep), where the coding distortion approaches the frame-level perceived coding distortion threshold. Under constraints, solve for the frame-level aware quantization parameters that are just perceptible at the i-th quality level. This process can be represented as:
[0074]
[0075] Where SQ represents the subsequent quantization parameter set, in the actual encoding process, for each transform unit TU, according to the D-QP model D at the transform unit TU level... t (Qstep), the perceptual coding distortion threshold at the transform unit (TU) level. Reverse derive the prequantization parameters of a transform unit TU level. Right now:
[0076]
[0077] To avoid block artifacts during encoding, frame-aware quantization parameters are used. Prequantization parameters of the current transform unit In addition to the perceptual quantization parameters of neighboring transform units, the final perceptual quantization parameters of the current transform unit are determined. The functional expression of this process can be represented as:
[0078]
[0079] in Let Γ(·) represent the set of transform block sensing quantization parameters in the neighborhood, and let Γ(·) represent the functional relationship that needs further investigation.
[0080] Decoupled Residual Filtering and Perceptual Adaptive Quantization: Video compression is a multi-module serial coding process. Perceptual residual filtering and perceptual adaptive quantization inevitably introduce coding distortion, and since they implement residual compression in the spatial and frequency domains respectively, performance coupling effects exist. To address this problem, such as... Figure 2 As shown, this invention designs a perceptual all-zero block decision unit and a quantization filter in perceptual residual filtering, which approaches the limit of video perceptual compression while eliminating coupling effects.
[0081] The original video will undergo multi-level, multi-granularity perceptual coding distortion threshold prediction. Based on the predicted threshold, the rate-distortion cost calculation process in video coding will be optimized and applied to intra-frame or inter-frame prediction. Furthermore, an all-zero block decision will be made for the predicted residual blocks. This decision process is expressed as follows:
[0082]
[0083] in This indicates that the current residual block is transformed using... When performing quantization, if the magnitude of all residual coefficients in the residual block is less than the predicted quantization distortion value, it means that the residual block will likely be quantized into an all-zero block in subsequent quantization processes. Therefore, in this case, the current residual block will be forced to be all zero, and no further transformation and entropy coding will be performed. The all-zero block perception decision can improve video compression efficiency and speed up video compression.
[0084] If the current residual block is determined to be a non-all-zero block, it means that the current residual block will undergo subsequent transformation and quantization processes, introducing quantization distortion. In this case, the present invention designs a quantization filter to introduce quantization distortion to control the residual filtering strength, weaken the performance coupling effect between modules, and further improve video compression efficiency. The quantization filtering process is as follows:
[0085]
[0086] Where κ represents the filter strength parameter, and C represents the coupling parameter. cr Perceptual coding distortion and the current transform block coding distortion prediction value The function, that is:
[0087]
[0088] Where C cr This represents the coupling parameter, which can be determined based on video compression experiments. C cr The values 0 and 1 represent the two cases where residual filtering and perceptual quantization are uncoupled and fully coupled, respectively. C cr A value between 0 and 1 represents the degree of coupling between residual filtering and perceptual quantization; based on the coding gain maximization criterion, exploratory research has revealed that the coupling parameter C... cr The optimal value is closely related to the coding prediction method, and the coupling strength in inter-frame coding is significantly higher than that in intra-frame coding.
[0089] Step S3: Optimize rate distortion based on perceptual quantization parameters to improve intra-frame / inter-frame prediction;
[0090] Perceptual Rate-Distortion Optimization: Objective distortion criteria based on mean squared error (MSE) cannot fully reflect the perceptual characteristics of the Human Visual System (HVS) for compressed video. Therefore, this invention proposes a perceptual coding distortion measurement method based on the predicted perceptual coding distortion threshold and applies it to rate-distortion optimization (RDO) of video coding. In video coding, the sum of squared errors (SSE) and the sum of absolute differences (SAD) are commonly used to measure coding distortion. This invention, for perceptual video coding optimization, uses the perceptual sum of squared errors... With perceptual absolute error and To measure encoding distortion, the cost of perceptual rate distortion. Represented as:
[0091]
[0092] in, Let represent the perceived distortion at the i-th quality level, Bit represent the coding rate, and λ(QP) represent the Lagrange multiplier determined by the quantization parameter QP.
[0093] Intra / Inter-frame Prediction Based on Perceptual Rate Distortion Optimization: In intra-frame and inter-frame prediction, this invention introduces perceptual rate distortion optimization into the predictive coding module to determine the optimal intra / inter-frame prediction mode in a perceptual sense. Considering the subsequent perceptual adaptive quantization module, the optimal intra-frame or inter-frame prediction mode is expressed as:
[0094]
[0095] in M represents the optimal intra-frame or inter-frame prediction mode in a perceptual sense. pre Indicates the candidate intra-frame or inter-frame prediction mode. This indicates that the quantization parameters are sensed by the current transform block. Determined Lagrange multipliers.
[0096] Step S4: Perform inverse quantization and inverse transformation on the perceptual quantization result to obtain the reconstructed frame. Construct a perceptual quality enhancement network. Based on the reconstructed frame, perceptual quantization parameters, coupling strength between residual filtering and perceptual quantization, and predicted coding distortion, perform coding distortion prediction to obtain a distortion prediction map. Enhance the quality of the reconstructed frame using the distortion prediction map to obtain a quality enhancement map, which is used to optimize intra / inter-frame prediction.
[0097] Perceptual quality enhancement networks: Inter-module distortion propagation inevitably reduces the subjective and objective quality of the current frame, and the accumulation of temporal distortion between frames inevitably affects the overall subjective quality consistency of the compressed video, such as... Figure 3As shown, this invention improves the consistency of subjective and objective quality of compressed video based on a quality enhancement network for distortion prediction. Specifically, it optimizes the video perceptual coding process through keyframe quality enhancement technology to improve the subjective and objective quality of compressed video. Considering the complexity constraints of video compression, this invention adaptively selects keyframes (such as intra-coded frames or some bidirectional prediction frames in a GOP) from the encoded video frames to enhance the subjective and objective quality of keyframes and reduce the distortion propagation effect in video coding.
[0098] A single original image frame is divided into multiple non-overlapping coding units (CUs). Each CU is encoded independently, undergoing prediction, transformation, quantization, inverse quantization, and inverse transform processes to obtain a reconstructed CU. Finally, all reconstructed CUs are integrated into a single reconstructed frame, subjected to in-loop filtering, and placed in a buffer as a reference for subsequent encoded frames. The perceptual quality enhancement network can be directly used for in-loop filtering or for quality enhancement of the reconstructed image in the buffer. For example, the perceptual quality enhancement network takes the decoded reconstructed frame, perceptual quantization parameters, coupling strength, and perceptual coding distortion threshold distribution as inputs. It sequentially performs upsampling, multi-scale feature extraction, and adaptive feature fusion to obtain a distortion prediction map. The distortion prediction map and the reconstructed frame are then input into a compressed video quality enhancement network based on a residual dense network (RDN) to obtain a quality enhancement map, which is placed in the decoding buffer for reference in subsequent intra / inter-frame predictions. The coding distortion prediction is represented as:
[0099]
[0100] F rec Indicates the reconstructed frame. Let κ represent the set of perceptual quantization parameters for all coding units in the current frame, and let J represent the coupling strength between residual filtering and perceptual quantization. i This represents the perceptual coding distortion threshold for the i-th level. This represents the coding distortion prediction function during the quality enhancement process.
[0101] Multi-scale adaptive feature fusion: Due to the diversity of network inputs, features at different scales have varying impacts on encoding distortion prediction. This invention introduces an attention mechanism module into the adaptive feature fusion network to learn and assign different weights to features at different channels and scales, such as... Figure 4 As shown, the network first samples features at different scales to the original size and inputs them into convolutional and activation layers to obtain the weight assignment of each element. After multiplying the features with the weights, average pooling and max pooling are performed to extract features. After learning the channel weights through convolutional and activation layers, the encoding distortion distribution map is finally obtained.
[0102] Quality Enhancement Network Coding Applications: This network can be used in the loop filtering module at the encoding end, in conjunction with existing encoder techniques such as deblocking filtering and SAO (sample adaptive offset) to improve the quality of reconstructed frames at the encoding end and reduce the distortion propagation effect in video perceptual coding; at the same time, this network can be used independently in the post-processing stages of the encoding and decoding ends to improve the subjective and objective quality of the reconstructed video.
[0103] Corresponding to the aforementioned embodiment of a multi-level, multi-module collaborative video perception coding optimization method, the present invention also provides an embodiment of a multi-level, multi-module collaborative video perception coding optimization device.
[0104] See Figure 5 The present invention provides a multi-level, multi-module collaborative video perception coding optimization device, including a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, they are used to implement a multi-level, multi-module collaborative video perception coding optimization method in the above embodiment.
[0105] An embodiment of the multi-level, multi-module collaborative video perception coding optimization device of the present invention can be applied to any device with data processing capabilities, such as a computer. The device embodiment can be implemented in software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of any data processing device loading the corresponding computer program instructions from non-volatile memory into memory for execution. From a hardware perspective, such as... Figure 5 The diagram shown is a hardware structure diagram of any device with data processing capabilities, including the multi-level, multi-module collaborative video perception coding optimization device of the present invention. (Except for...) Figure 5 In addition to the processor, memory, network interface, and non-volatile memory shown, any data processing device in the embodiment may also include other hardware depending on the actual function of the data processing device, which will not be described in detail here.
[0106] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0107] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the present invention according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0108] This invention also provides a computer-readable storage medium storing a program thereon, which, when executed by a processor, implements a multi-level, multi-module collaborative video perception coding optimization method as described in the above embodiments.
[0109] The computer-readable storage medium can be an internal storage unit of any data processing device as described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device of any data processing device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of any data processing device. The computer-readable storage medium is used to store the computer program and other programs and data required by the data processing device, and can also be used to temporarily store data that has been output or will be output.
[0110] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-level, multi-module collaborative video perception coding optimization method, characterized in that... Includes the following steps: Step S1: Perform coding distortion prediction, frame-level coding distortion prediction, and derive frame-level quantization parameters from the original video. Step S2: Perform intra-frame / inter-frame prediction on the original video images, and calculate the difference between the predicted images and the original images to obtain residual images. Perform residual filtering on the residual images based on the predicted coding distortion. After the filtered residual images are transformed based on residual blocks, perceptual quantization is performed according to the predicted frame-level coding distortion and frame-level quantization parameters. The residual filtering and perceptual quantization are decoupled through the perceptual all-zero block decision unit and the quantization filter. Step S3: Optimize rate distortion based on perceptual quantization parameters to improve intra-frame / inter-frame prediction; Step S4: Perform inverse quantization and inverse transformation on the perceptual quantization result to obtain the reconstructed frame. Construct a perceptual quality enhancement network. Based on the reconstructed frame, perceptual quantization parameters, coupling strength between residual filtering and perceptual quantization, and predicted coding distortion, perform coding distortion prediction to obtain a distortion prediction map. Enhance the quality of the reconstructed frame using the distortion prediction map to obtain a quality enhancement map, which is used to optimize intra / inter-frame prediction. An image of a raw video frame is divided into multiple non-overlapping coding units. Each coding unit is encoded separately and reconstructed and integrated into a reconstructed frame. Loop filtering is performed and the reconstructed image is placed in the decoding buffer. The perceptual quality enhancement network is applied to the reconstructed image in the loop filter and / or the decoding buffer. The network takes the reconstructed frame, perceptual quantization parameters, coupling strength, and perceptual coding distortion threshold distribution as inputs. It then sequentially performs upsampling, multi-scale feature extraction, and adaptive feature fusion to obtain a distortion prediction map. The distortion prediction map and the reconstructed frame are then input into a compressed video quality enhancement network based on a residual dense network to obtain a quality enhancement map. The adaptive feature fusion is a multi-scale adaptive feature fusion that introduces an attention mechanism. It learns to assign different weights to features of different channels and scales. Features of different scales are sampled to their original size and input into convolutional and activation layers to obtain the weight assignment of each element. After the features are multiplied by the weights, average pooling and max pooling are performed to extract features. After learning the channel weights through convolutional and activation layers, the encoding distortion distribution map, i.e. the distortion prediction map, is finally obtained. Step S5: Based on optimized intra / inter-frame prediction, perform prediction, difference calculation, residual filtering, transformation, and perceptual quantization on the images of the original video, and then perform entropy coding.
2. The multi-level, multi-module collaborative video perception coding optimization method according to claim 1, characterized in that: The residual filtering in step S2 filters out residuals with amplitude values lower than the coding distortion threshold, based on the coding distortion threshold obtained from the coding distortion prediction.
3. The multi-level, multi-module collaborative video perception coding optimization method according to claim 1, characterized in that: In step S2, the frame-level quantization distortion is derived based on the frame-level distribution law of the transform coefficients and applied to perceptual quantization. Based on the frame-level distribution of quantization compensation and transform coefficients, construct... D - QP The model calculates frame-level aware quantization parameters under the constraint of the frame-level coding distortion threshold obtained from frame-level coding distortion prediction; based on the transform unit level... D - QP The encoding distortion thresholds at the model and transform unit levels are used to back-engineer the pre-quantization parameters at the transform unit level. Based on the frame-level perceptual quantization parameters, the pre-quantization parameters of the current transform unit, and the perceptual quantization parameters of neighboring transform units, the final perceptual quantization parameters of the current transform unit are determined.
4. The multi-level, multi-module collaborative video perception coding optimization method according to claim 1, characterized in that: In step S2, when the magnitude of the residual coefficients in the residual image block is less than the quantization distortion prediction value, the perceptual zero block decision device forces the current residual image block to be all zeros; otherwise, the quantization filter introduces quantization distortion to control the residual filtering intensity. When the magnitude of the residual coefficients in the residual image block is less than the product of the encoding distortion threshold and the filtering intensity parameter, the quantization filter filters the current residual image block. The quantization distortion prediction value is obtained by quantizing the current residual image block after transformation using the perceptual quantization parameter. The filtering intensity parameter is generated based on the coupling degree between residual filtering and perceptual quantization, perceptual encoding distortion, and the quantization distortion prediction value of the current transformed residual block.
5. The multi-level, multi-module collaborative video perception coding optimization method according to claim 1, characterized in that: Rate-distortion optimization in step S3 is based on prediction to measure perceptual coding distortion and apply it to the rate-distortion optimization of video coding. Its rate-distortion cost is determined by perceptual coding distortion and the product of the Lagrange multiplier determined by the quantization parameter and the coding bit rate. Perceptual coding distortion is measured by the sum of perceptual squared errors and the sum of perceptual absolute errors.
6. The multi-level, multi-module collaborative video perception coding optimization method according to claim 5, characterized in that: In step S3, intra / inter-frame prediction is performed by introducing rate-distortion optimization to select the intra / inter-frame prediction mode with the lowest rate-distortion cost from the candidate intra / inter-frame prediction modes, thereby determining the optimal intra / inter-frame prediction mode.
7. The multi-level, multi-module collaborative video perception coding optimization method according to claim 1, characterized in that: In step S4, key frames are adaptively selected from the encoded video frames to enhance the subjective and objective quality of the key frames.
8. A multi-level, multi-module collaborative video perception coding optimization device, characterized in that, The device includes a memory and one or more processors, wherein the memory stores executable code, and the one or more processors execute the executable code to implement a multi-level, multi-module collaborative video perception coding optimization method according to any one of claims 1-7.
Citation Information
Patent Citations
Multi-level multi-granularity perceptual coding distortion prediction method
CN116248883A