A multi-level multi-granularity perceptual coding distortion prediction method

By constructing a multi-level, multi-granularity perceptual coding distortion prediction model, combining visual perception and coding mechanisms, and employing deep learning methods, the problem of existing models being unable to accurately predict multi-level, multi-granularity perceptual coding distortion was solved, thereby improving the quality and efficiency of video compression.

CN116248883BActive Publication Date: 2025-11-25HANGZHOU DIANZI UNIV +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310156672.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-17
Publication Date
2025-11-25
Estimated Expiration
2043-02-17

AI Technical Summary

Technical Problem

Existing video coding models fail to effectively predict multi-level, multi-granularity perceptual coding distortion, cannot meet the multi-granularity optimization requirements of video compression, and traditional models fail to accurately reflect the joint constraints of visual perception mechanisms and coding mechanisms.

Method used

A multi-level, multi-granularity perceptual coding distortion prediction model is constructed. Through joint analysis of visual perception and coding mechanisms, statistical analysis and deep learning methods are adopted, combining visual perception features and neural networks to predict the multi-level, multi-granularity JNCD threshold.

Benefits of technology

It achieves accurate prediction of multi-level and multi-granularity video coding distortion, improves video compression quality and efficiency, and meets the needs of multi-granularity perceptual coding optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116248883B_ABST
    Figure CN116248883B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of video perceptual coding optimization, and discloses a multi-level multi-granularity perceptual coding distortion prediction method, comprising the following steps: step 1: visual perception effect and perceptual coding distortion mapping analysis: constructing multi-level just noticeable quantization parameter dataset and just noticeable coding distortion dataset of each source video; step 2: multi-level multi-granularity perceptual coding distortion prediction: based on visual perception mechanism, qualitatively analyzing the mapping relationship between each visual perception characteristic and just noticeable coding distortion by using statistical analysis method. The present application solves the problem that different perceptual effects are not completely consistent in the perceptual effect on compressed video, increases the difficulty of perceptual coding distortion theoretical analysis under the joint constraint of video coding mechanism and visual perception mechanism, and cannot deduce the ideal JNCD threshold model by using traditional theoretical modeling, and meets the demand of multi-granularity perceptual coding optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of video perceptual coding optimization, and in particular relates to a multi-level, multi-granularity perceptual coding distortion prediction method. Background Technology

[0002] The exponential growth of video data places higher demands on video compression, while information theory-based video coding optimization has approached its performance bottleneck. In fact, both visual perception mechanisms and video coding mechanisms jointly determine the subjective perceived quality of compressed video. Therefore, visual perception-based video coding optimization is an effective way to further improve video compression efficiency. The Human Visual System (HVS) can only perceive a limited number of quality levels for a series of compressed videos with different coding configurations. In other words, the maximum coding distortion corresponding to each quality level of compressed video is the Just Noticeable Coding Distortion (JNCD) threshold for that quality level. All compressed videos with coding distortion below this threshold have almost identical subjective quality. Therefore, multi-level perceptual coding distortion of video is of great significance for guiding perceptual video coding. Since video compression is a multi-module collaborative coding process, the image processing unit size of different coding modules is not entirely consistent. Therefore, for perceptual video coding, multi-level, multi-granularity perceptual coding distortion prediction models can help guide the various coding modules of video compression and improve video compression quality. The construction of a multi-level, multi-granularity perceptual coding distortion prediction model faces the following challenges: (1) The magnitude of perceptual coding distortion is closely related to the video content, and different perceptual effects do not have the same perceptual effect on compressed videos; (2) The joint constraints of video coding mechanism and visual perception mechanism increase the difficulty of theoretical analysis of perceptual coding distortion. Traditional theoretical modeling cannot derive an ideal JNCD threshold model, which poses a significant challenge to the accurate prediction of perceptual coding distortion.

[0003] Diverse video signals exhibit various visual perception effects under the influence of visual perception mechanisms, such as brightness adaptive masking, contrast masking, and foveal visual masking. The combined effect of these diverse visual masking effects causes HVS (Hardware Visual System) to fail to detect visual noise below a certain threshold, known as Just Noticeable Distortion (JND). The JND model is commonly used in video perceptual coding optimization design to predict video perceptual coding distortion.

[0004] JND Threshold Prediction: JND threshold prediction models are commonly constructed using two methods: statistical modeling and theoretical modeling. Based on visual experimental data, statistical modeling methods produce JND models with simple expressions and low complexity. However, because visual experiments are influenced by experimental conditions and the coupling relationships between visual perception effects are difficult to analyze, these methods often suffer from overestimation or underestimation of the threshold. Theoretical modeling methods attempt to simulate the visual perception process during JND modeling by introducing more complex visual perception effects and analyzing the correlations between these effects to improve the prediction accuracy of the JND threshold. Both types of JND models typically use random noise for accuracy verification. However, coding distortion in compressed videos does not completely follow a random distribution. The video coding mechanism determines that coding distortion is mainly distributed in areas with complex textures and intense motion, meaning that using the above two models to predict video perceptual coding distortion is inappropriate. Therefore, in recent years, researchers have attempted to explore accurate prediction methods for Just Noticeable Coding Distortion (JNCD) for compressed videos and images.

[0005] JNCD Threshold Prediction: In the exploration of JNCD prediction, researchers were the first to consider the impact of video coding quantization distortion in JNCD threshold prediction, proposing a Just-Aware Quantization Distortion (JNQD) model suitable for 8×8 DCT transform. However, this model is difficult to apply to coding standards such as AVS3 and VVC that support multi-size transforms, as well as multi-granularity video-aware coding optimization. With the rise of machine learning and deep learning, researchers began to explore JNCD prediction models based on support vector machines, predicting just-aware quantization parameters at the first quality level; transforming the image JNCD prediction problem into a binary classification problem, and using deep learning methods to predict the just-aware quantization factor (QF) of images. Other researchers have transformed the JNCD prediction problem into the prediction problem of user satisfaction (SUR) curves, predicting the SUR curves of images or videos based on deep learning.

[0006] In summary, most existing JNCD prediction models rely on learning-based modeling to predict the JNCD threshold of videos or images. However, these models suffer from two main drawbacks: 1) The modeling process is largely data-driven and neglects visual perception mechanisms, making the prediction accuracy of JNCD susceptible to the influence of the dataset; 2) Existing JNCD datasets for images and videos are primarily image-level and video-level, leading to overly granular JNCD prediction models that cannot meet the demands of multi-granularity perceptual coding optimization. To address these shortcomings, this invention proposes a multi-level, multi-granularity perceptual coding distortion prediction model jointly guided by visual perception and video coding mechanisms. Summary of the Invention

[0007] The purpose of this invention is to provide a multi-level, multi-granularity perceptual coding distortion prediction model to address the technical problems that the perceptual effects of different perceptual effects on compressed videos are not entirely consistent, the joint constraints of video coding mechanisms and visual perception mechanisms increase the difficulty of theoretical analysis of perceptual coding distortion, and traditional theoretical modeling cannot derive an ideal JNCD threshold model.

[0008] To address the aforementioned technical problems, the specific technical solution of the multi-level, multi-granularity sensing coding distortion prediction method of the present invention is as follows:

[0009] A multi-level, multi-granularity perceptual coding distortion prediction method includes the following steps: Step 1: Visual perception effect and perceptual coding distortion mapping analysis: Construct a multi-level just-perceptible quantization parameter dataset and just-perceptible coding distortion dataset for each source video;

[0010] Step 2: Multi-level and multi-granularity perceptual coding distortion prediction: Based on the visual perception mechanism, statistical analysis methods are used to qualitatively analyze the mapping relationship between each visual perception feature and just-perceptible coding distortion.

[0011] Further, step 1 includes the following steps:

[0012] Step 1.1: Determine the perceptual coding distortion label;

[0013] Step 1.2: Analyze the mapping relationship between perceptual effects and perceptual coding distortion.

[0014] Further, step 1.1 includes the following specific steps:

[0015] A video sequence consists of a series of video frames. Different video frames encoded with the same quantization parameter (QP) exhibit certain differences in subjective and objective quality. The compression quality of all video frames collectively determines the subjective quality of the compressed video. The publicly available dataset Videosot is used, which is encoded using H.264 / AVC and includes 220 source videos at four resolutions: 1920×1080, 1280×720, 960×540, and 640×360, totaling 220×4=880 source videos. Each source video has 51 compressed versions, corresponding to QPs from 1 to 51, with three quality levels. Each level has a just-perceptible QP. The objective quality of the compressed video corresponding to the just-perceptible QP at each level is used as the multi-level perceptual coding distortion label for the current video.

[0016] Furthermore, step 1.2 includes the following specific steps:

[0017] Visual perception effects are inherent characteristics of the visual system in response to visual signals. Different perception effects do not have the same effect on the perception of compressed video. An exploratory experiment was conducted based on a publicly available compressed video dataset, MCL-JCV. The experiment showed that the luminance adaptive masking effect (LA) has a relatively small impact on the perception of compressed video. The measured JNCD threshold showed a high correlation with the contrast masking effect (CM) and the entropy masking effect (EM), with correlation coefficients (PLCC) of 0.9 and 0.8, respectively. However, the correlation with the luminance adaptive masking effect was only 0.1. A unified approach was used to analyze the mapping relationship between different visual perception effects and perceptual coding distortion.

[0018] Further, step 2 includes the following steps:

[0019] Step 2.1: Adaptive sampling;

[0020] Step 2.2: Perceptual feature extraction;

[0021] Step 2.3: Perceptual feature fusion;

[0022] Step 2.4: Perceptual encoding distortion threshold aggregation.

[0023] Further, step 2.1 includes the following steps:

[0024] The JNCD threshold of the video signal is jointly determined by the coding mechanism and the visual perception mechanism. In the design of the JNCD prediction model, the spatiotemporal perception characteristics of the video are analyzed, and the qualitative relationship between the spatiotemporal perception effect of each video frame and the perception coding distortion is analyzed. In the training of the JNCD model, the compressed video is adaptively sampled to reduce the prediction complexity of the model while ensuring the accuracy of the JNCD prediction model. In the actual coding process, multiple frames with similar time sequences share the same JNCD prediction threshold.

[0025] Further, step 2.2 includes the following steps:

[0026] Based on the analysis of visual perception effects and perceptual coding distortion, visual perception features are manually extracted, and the brightness adaptive masking feature map M of the sampled frame is calculated using existing methods. LA Contrast masking feature map M CM Entropy masking feature map M EM Visual salience map M SA Fixation point distribution map M FD Positive perception effect intensity diagram M P And negative perception effect intensity map M N ;

[0027] The seven handcrafted feature maps extracted and associated with human visual perception effects are concatenated and used as input to the prediction network:

[0028] F input =Concat(M LA M CM M EM M SA M FD M P M N )

[0029] Where Concat represents the connection symbol on the feature dimension.

[0030] Furthermore, step 2.3 includes the following steps:

[0031] Since visual perception is the result of the combined effects of diverse perceptual effects on video signals, CNNs are used to simulate the mechanism of visual perception effects. An attention mechanism is employed to fuse visual perception features, which are then ultimately mapped to pixel-level JNCD thresholds.

[0032] M JNCD =CNN(F input )

[0033] Where CNN(·) is the constructed neural network prediction model, M JNCD This represents the output JNCD threshold map. Further, step 2.4 includes the following steps:

[0034] After perceptual feature fusion, a pixel-level JNCD threshold map is obtained. Based on the pixel-level JNCD threshold map, it is mapped to a block-level JNCD threshold through local average pooling. Under the action of the attention mechanism, operations such as average pooling and max pooling are used to aggregate the JNCD thresholds into frame-level and video-level thresholds.

[0035] J B =A(M JNCD )

[0036] J F =A(J B )

[0037] J V =A(J F )

[0038] Where A(·) represents the average pooling operation, J B J F J V These represent the JNCD prediction thresholds at the block, frame, and video levels, respectively.

[0039] The multi-level, multi-granularity perceptual coding distortion prediction method of this invention has the following advantages: This invention analyzes the mapping relationship between diverse visual perception effects and perceptual coding distortion, and explores the perceptual effects of different visual perception effects on perceptual coding distortion. It extracts various handcrafted features of visual perception to map JNCD thresholds. By combining the attention mechanism of neural networks with handcrafted features, a deep learning-based method is proposed to learn and predict multi-level, multi-granularity JNCD thresholds. This solves the problems that the perceptual effects of different perception effects on compressed videos are not entirely consistent, the joint constraints of video coding mechanisms and visual perception mechanisms increase the difficulty of theoretical analysis of perceptual coding distortion, and traditional theoretical modeling cannot derive an ideal JNCD threshold model, thus meeting the needs of multi-granularity perceptual coding optimization. Attached Figure Description

[0040] Figure 1 This is a flowchart of the multi-level, multi-granularity sensing coding distortion prediction method of the present invention.

[0041] Figure 2 This is a schematic diagram illustrating the mapping relationship between the JNCD threshold and brightness adaptive masking of the present invention, comparing the masking and entropy masking effects.

[0042] Figure 3 This is a schematic diagram of the multi-level, multi-granularity JNCD prediction model of the present invention. Detailed Implementation

[0043] To better understand the purpose, structure, and function of this invention, the following detailed description of a multi-level, multi-granularity sensing coding distortion prediction method is provided in conjunction with the accompanying drawings.

[0044] This invention predicts just-perceptible coding distortion (JCD) of video signals by constructing a perceptual coding distortion model. The overall flowchart for multi-level, multi-module perceptual coding distortion prediction is shown below. Figure 1 As shown, it mainly includes the visual perception effect and perception coding distortion mapping analysis stage, the pixel-level JNCD modeling process, and the multi-particle JNCD aggregation process.

[0045] (1) Visual perception effect and perceptual coding distortion mapping analysis

[0046] Perceptual Coding Distortion Label Determination: A video sequence consists of a series of video frames. Different video frames encoded with the same quantization parameter (QP) exhibit certain differences in subjective and objective quality. The compression quality of all video frames collectively determines the subjective quality of the compressed video. This invention uses the publicly available dataset Videosot, which is encoded using H.264 / AVC and includes 220 source videos at four resolutions: 1920×1080, 1280×720, 960×540, and 640×360, totaling 220×4=880 source videos. Each source video has 51 compressed versions, corresponding to QPs from 1 to 51, with three quality levels. Each level has a just-perceptible QP. The objective quality of the compressed video corresponding to the just-perceptible QP at each level is used as the multi-level perceptual coding distortion label for the current video.

[0047] Analysis of the mapping relationship between perceptual effects and perceptual coding distortion: Visual perceptual effects are inherent characteristics of the visual system in response to visual signals, and different perceptual effects do not have completely consistent effects on the perception of compressed video. To address this issue, this invention conducted exploratory experiments based on a publicly available compressed video dataset, MCL-JCV, such as... Figure 2 As shown, the luminance-adaptive masking effect (LA) has a relatively small impact on the perception of compressed video. The measured JNCD threshold shows a high correlation with contrast masking (CM) and entropy masking (EM), with correlation coefficients (PLCC) of 0.9 and 0.8, respectively, while the correlation with luminance-adaptive masking is only 0.1. This is because coding distortion is mainly concentrated in textured regions, where contrast and entropy masking are relatively higher. Existing research indicates that visual saliency, visual attention, and foveal masking effects play important roles in human visual perception. Therefore, a unified approach can be used to analyze the mapping relationship between different visual perception effects and perceptual coding distortion.

[0048] (2) Multi-level, multi-granularity sensing coding distortion prediction

[0049] Based on the visual perception effect and perceptual coding distortion mapping analysis in (1), this invention designs a neural network model to learn the quantitative mapping relationship between the two and constructs a JNCD prediction model. The preliminary framework of the JNCD prediction model is as follows: Figure 3 As shown, the entire prediction model consists of four steps: adaptive sampling, perceptual feature extraction, perceptual feature fusion, and multi-granularity perceptual coding distortion threshold aggregation.

[0050] Adaptive Sampling: The JNCD threshold for video signals is jointly determined by the encoding mechanism and the visual perception mechanism. To reduce the complexity of the JNCD prediction model, this invention analyzes the spatiotemporal perceptual characteristics of video during the JNCD prediction model design. In the study of the mapping relationship between perceptual effects and perceptual coding distortion, the qualitative relationship between the spatiotemporal perceptual effects and perceptual coding distortion of each video frame is analyzed. Adaptive sampling is performed on the compressed video during JNCD model training to reduce the prediction complexity of the model while ensuring its accuracy. In actual encoding, multiple temporally similar frames can share the same JNCD prediction threshold.

[0051] Perceptual Feature Extraction: Based on visual perception effects and perceptual coding distortion mapping analysis, visual perception features are manually extracted, and the brightness adaptive masking feature map M of the sampled frame is calculated using existing methods. LA Contrast masking feature map M CM Entropy masking feature map M EM Visual salience map M SA Fixation point distribution map M FD Positive perception effect intensity diagram M P And negative perception effect intensity map M N .

[0052] Seven handcrafted feature maps extracted and associated with human visual perception effects are concatenated and used as input to the prediction network.

[0053] F input =Concat(M LA M CM M EM M SA M FD M P M N )

[0054] Where Concat represents the connection symbol on the feature dimension.

[0055] Perceptual Feature Fusion: Since visual perception is the result of the combined effect of diverse perceptual effects on video signals, this invention uses CNN to simulate the mechanism of visual perception effects, employs an attention mechanism to fuse visual perception features, and finally maps them to pixel-level JNCD thresholds.

[0056] M JNCD =CNN(F input )

[0057] Where CNN(·) is the constructed neural network prediction model, M JNCD This represents the output JNCD threshold map.

[0058] Perceptual Coding Distortion Threshold Aggregation: After perceptual feature fusion, we obtain a pixel-level JNCD threshold map. Based on the pixel-level JNCD threshold map, it is mapped to a block-level JNCD threshold through local average pooling. Under the action of the attention mechanism, operations such as average pooling and max pooling are used to aggregate the JNCD thresholds into frame-level and video-level thresholds.

[0059] J B =A(M JNCD )

[0060] J F =A(J B )

[0061] J V =A(J F )

[0062] Where A(·) represents the average pooling operation, J B J F J V These represent the JNCD prediction thresholds at the block, frame, and video levels, respectively.

[0063] It is understood that the present invention has been described through some embodiments, and those skilled in the art will recognize that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the invention. Furthermore, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of the present invention.

Claims

1. A multi-level, multi-granularity sensing coding distortion prediction method, characterized in that, Includes the following steps: Step 1: Visual perception effect and perceptual coding distortion mapping analysis: Construct a multi-level just-perceptible quantization parameter dataset and just-perceptible coding distortion dataset for each source video; Step 2: Multi-level and multi-granularity perceptual coding distortion prediction: Based on the visual perception mechanism, statistical analysis methods are used to qualitatively analyze the mapping relationship between each visual perception feature and just-perceptible coding distortion; Step 2.1: Adaptive sampling; Step 2.2: Perceptual feature extraction; Step 2.3: Perceptual feature fusion; Since visual perception is the result of the combined effects of diverse perceptual effects on video signals, CNNs are used to simulate the mechanism of visual perception effects. An attention mechanism is employed to fuse visual perception features, which are then ultimately mapped to pixel-level JNCD thresholds. Where CNN(·) is the constructed neural network prediction model, This represents the output JNCD threshold map; Step 2.4: Aggregation of perceptual encoding distortion thresholds; After perceptual feature fusion, a pixel-level JNCD threshold map is obtained. Based on the pixel-level JNCD threshold map, it is mapped to a block-level JNCD threshold through local average pooling. Under the action of the attention mechanism, average pooling and max pooling operations are used to aggregate the JNCD thresholds into frame-level and video-level thresholds. in J represents the average pooling operation. B J F J V These represent the JNCD prediction thresholds at the block, frame, and video levels, respectively.

2. The multi-level, multi-granularity sensing coding distortion prediction method according to claim 1, characterized in that, Step 1 includes the following steps: Step 1.1: Determine the perceptual coding distortion label; Step 1.2: Analyze the mapping relationship between perceptual effects and perceptual coding distortion.

3. The multi-level, multi-granularity sensing coding distortion prediction method according to claim 2, characterized in that, Step 1.1 includes the following specific steps: A video sequence consists of a series of video frames. Different video frames encoded with the same quantization parameter (QP) exhibit certain differences in subjective and objective quality. The compression quality of all video frames collectively determines the subjective quality of the compressed video. The publicly available dataset Videosot is used, which is encoded using H.264 / AVC and includes 220 source videos at four resolutions: 1920×1080, 1280×720, 960×540, and 640×360, totaling 220×4=880 source videos. Each source video has 51 compressed versions, corresponding to QPs from 1 to 51, with three quality levels. Each level has a just-perceptible QP. The objective quality of the compressed video corresponding to the just-perceptible QP at each level is used as the multi-level perceptual coding distortion label for the current video.

4. The multi-level, multi-granularity sensing coding distortion prediction method according to claim 2, characterized in that, Step 1.2 includes the following specific steps: Visual perception effects are inherent characteristics of the visual system in response to visual signals. Different perception effects do not have the same effect on the perception of compressed video. An exploratory experiment was conducted based on a publicly available compressed video dataset, MCL-JCV. The experiment showed that the luminance adaptive masking effect (LA) has a relatively small impact on the perception of compressed video. The measured JNCD threshold showed a high correlation with the contrast masking effect (CM) and the entropy masking effect (EM), with correlation coefficients (PLCC) of 0.9 and 0.8, respectively. However, the correlation with the luminance adaptive masking effect was only 0.

1. A unified approach was used to analyze the mapping relationship between different visual perception effects and perceptual coding distortion.

5. The multi-level, multi-granularity sensing coding distortion prediction method according to claim 1, characterized in that, Step 2.1 includes the following steps: The JNCD threshold of the video signal is jointly determined by the coding mechanism and the visual perception mechanism. In the design of the JNCD prediction model, the spatiotemporal perception characteristics of the video are analyzed, and the qualitative relationship between the spatiotemporal perception effect of each video frame and the perception coding distortion is analyzed. In the training of the JNCD model, the compressed video is adaptively sampled to reduce the prediction complexity of the model while ensuring the accuracy of the JNCD prediction model. In the actual coding process, multiple frames with similar time sequences share the same JNCD prediction threshold.

6. The multi-level, multi-granularity sensing coding distortion prediction method according to claim 1, characterized in that, Step 2.2 includes the following steps: Based on the analysis of visual perception effects and perceptual coding distortion, visual perception features were manually extracted, and the brightness adaptive masking feature map of the sampled frame was calculated using existing methods. Contrast masking feature map Entropy masking feature map Visual saliency map fixation point distribution map positive perception effect intensity map and negative perception effect intensity diagram The seven handcrafted feature maps extracted and associated with human visual perception effects are concatenated and used as input to the prediction network: Where Concat represents the connection symbol on the feature dimension.

Citation Information

Patent Citations

  • Multi-level multi-module collaborative video perception coding optimization method and device

    CN116193122A