Method and device for perceptual rate control of screen content video coding based on DCT domain JND

By using a DCT domain JND-based perceptual bitrate control method for screen content video coding, the bitrate allocation and quantization parameters are optimized by utilizing DCT coefficients and human visual characteristics. This solves the problems of low efficiency and high computational complexity in screen content video coding, and achieves more efficient coding quality and rate-distortion performance.

CN120751139BActive Publication Date: 2025-11-11HUAQIAO UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511145934.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-11-11
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

Existing video coding standards have low compression efficiency when processing screen content video, and traditional bitrate control algorithms have high computational complexity, making it difficult to optimize transmission performance while maintaining video quality.

Method used

A perceptual bitrate control method for screen content video coding based on DCT domain JND is adopted. By acquiring the DCT coefficients of the screen content video, calculating texture energy and edge information, constructing a JND threshold in combination with human visual characteristics, optimizing bitrate allocation and quantization parameters, and achieving accurate mapping of perceptual distortion indicators.

Benefits of technology

It improves the quality of video encoding for screen content and rate-distortion performance, enhances the accuracy of perceived bitrate control, reduces computational complexity, and optimizes the efficiency and quality of the encoding process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751139B_ABST
    Figure CN120751139B_ABST
Patent Text Reader

Abstract

A method and apparatus for perceptual bitrate control of screen content video coding based on Discrete Cosine Transform (DCT) domain JND is disclosed, relating to the field of video coding. The method includes: acquiring screen content video; obtaining DCT coefficients of the video through Discrete Cosine Transform; using texture energy weights to guide frame-level and CTU-level target bit allocation; obtaining a JND threshold in the DCT domain based on the product of a spatial contrast sensitivity threshold, a luminance masking modulation factor, and a contrast masking adjustment factor; and deriving a Lagrange multiplier for the JND adjustment factor using the JND threshold to achieve bitrate control. This invention establishes a JND model suitable for screen content video in the DCT domain, guides bit allocation through texture energy weights, and establishes a new rate-distortion model, which can improve coding rate-distortion performance and enhance the bitrate control accuracy of screen content video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video coding, and in particular to a method and apparatus for perceptual bitrate control of screen content video coding based on DCT domain JND (Just Noticeable Difference). Background Technology

[0002] With the continuous development of computer and multimedia technologies, modern society is rapidly entering an era dominated by visual information. Video, as a core carrier of information expression and communication, has been widely used in many fields such as education, communication, and entertainment, becoming an indispensable part. In particular, the booming rise of self-media has profoundly transformed the production and dissemination models of video content, greatly promoting the advancement of video technology.

[0003] In recent years, the continuous evolution of cloud computing and mobile internet technologies has driven the rapid popularization of screen content video (SCV) applications such as screen sharing, wireless projection, and distance education. Unlike natural video captured by camera equipment, screen content video is mainly generated by computers and other devices, and its characteristics include large areas of uniform background, repetitive patterns and characters, limited but highly saturated colors, high image contrast, and sharp edges. These combined characteristics cause a significant decrease in compression efficiency of traditional coding standards designed for natural video when processing screen content video. Therefore, it is necessary to develop dedicated coding strategies that differ from those for natural video to address the characteristics of screen content video.

[0004] The iterative upgrades of video technology are driving a transformation in coding paradigms. The emergence of new-generation video formats, such as high resolution, high dynamic range, strong interactivity, and full-view capabilities, places higher demands on the compression performance and transmission efficiency of traditional video coding standards. To address this challenge, international standards organizations jointly developed the next-generation video coding standard H.266 / VVC (Versatile Video Coding). Compared to its predecessor, HEVC, VVC can achieve approximately 50% bitrate savings while maintaining the same subjective quality. Notably, VVC integrates several tools specifically designed for on-screen content video within its coding architecture, such as Intra Block Copy (IBC), Palette Mode (PLT), Adaptive Color Transform (ACT), and Adaptive Motion Vector Resolution (AMVR), significantly improving the compression efficiency for on-screen content video.

[0005] In the video encoding process, bitrate control is a core component, aiming to efficiently utilize network bandwidth and optimize transmission performance. Its core task is to establish a precise mapping between bitrate and quantization parameters, ensuring the output bitstream meets the target bitrate requirements while striving for the best balance between video quality and compression efficiency. Although existing bitrate control algorithms have achieved a high degree of optimization in removing spatial and temporal redundancy, further compressing redundancy beyond current levels in a hybrid coding framework often means a significant increase in computational complexity.

[0006] Therefore, recent research has increasingly focused on integrating the perceptual characteristics of the human visual system into the bitrate control process. By employing perception-driven bit resource allocation strategies, more bits are allocated to visually sensitive areas (such as areas with complex textures or motion), effectively improving the user's subjective viewing experience with limited or insignificant increases in bitrate. This bitrate control method, which incorporates perceptual characteristics, has become a key development direction for overcoming coding performance bottlenecks and optimizing visual quality. Summary of the Invention

[0007] The main objective of this invention is to combine the characteristics of screen content and the characteristics of human vision to propose a screen content video coding perceptual bitrate control method and device based on DCT domain JND, in order to solve the technical problems mentioned in the background art and improve the screen content video coding quality, rate-distortion performance and perceptual bitrate control accuracy.

[0008] The present invention adopts the following technical solution:

[0009] On the one hand, a screen content video coding-aware bitrate control method based on DCT domain JND includes:

[0010] S101, DCT coefficient acquisition step: acquire screen content video, perform block-based Discrete Cosine Transform (DCT) on the screen content video, and obtain the DCT coefficients of each block.

[0011] S102, Texture Energy and Target Bit Count Acquisition Steps: Based on the DCT coefficients of each block, obtain the sum of the absolute values ​​of the low-frequency, mid-frequency, and high-frequency coefficients in the DCT block respectively; based on the sum of the absolute values ​​of the mid-frequency and high-frequency coefficients, obtain the texture energy of each DCT block; based on the low-frequency, mid-frequency, and high-frequency coefficients, obtain the edge information of each DCT block; based on the sum of the texture energy of all DCT blocks in each coding unit (CTU), obtain the texture energy of each CTU; based on the sum of the texture energy of all CTUs in each frame, obtain the texture energy of each frame; based on the texture energy of each CTU, obtain the CTU-level texture energy weight; based on the CTU-level texture energy weight, perform CTU-level bit allocation to obtain the CTU-level target bit count; based on the texture energy of each frame, obtain the frame-level texture energy weight; based on the frame-level texture energy weight, perform frame-level bit allocation to obtain the frame-level target bit count.

[0012] S103, JND threshold acquisition step: Based on the DCT coefficients of each block, obtain the spatial contrast sensitivity threshold, luminance masking modulation factor, and contrast intensity; Based on the texture energy and edge information of each DCT block, classify each DCT block into flat blocks, edge blocks, or texture blocks; Based on the contrast intensity, classification type, and DCT frequency linear function, obtain the contrast masking adjustment factor; Based on the product of the spatial contrast sensitivity threshold, luminance masking modulation factor, and contrast masking adjustment factor, obtain the JND threshold in the DCT domain.

[0013] S104, the step of obtaining the rate control quantization parameters: based on the JND threshold in the DCT domain, obtain the perceptual distortion index; based on the perceptual distortion index and the traditional Lagrange multiplier, obtain the Lagrange multiplier of the corresponding coding unit (CTU); based on the Lagrange multiplier, obtain the final rate control quantization parameters.

[0014] S105, the encoding control step, guides the actual encoding process and performs encoding control based on the CTU-level target bit count, frame-level target bit count, and bit rate control quantization parameters.

[0015] Preferably, in S101, for DCT block of size, number DCT coefficients The calculation formula is as follows:

[0016] ;

[0017] in, Indicates the first Image pixels, ; and They represent the first The and the first The compensation coefficient is calculated as follows:

[0018] .

[0019] Preferably, in step S102, the texture energy of each DCT block is obtained based on the sum of the absolute values ​​of the intermediate frequency coefficients and the high frequency coefficients, as shown below:

[0020] ;

[0021] in, It represents the sum of the absolute values ​​of the intermediate frequency coefficients, that is, the sum of the absolute values ​​of all DCT coefficients in the intermediate frequency region of the DCT block; It represents the sum of the absolute values ​​of the high-frequency coefficients, that is, the sum of the absolute values ​​of all DCT coefficients in the high-frequency region of the DCT block;

[0022] Based on low-frequency, mid-frequency, and high-frequency coefficients, edge information of each DCT block is obtained, including:

[0023] Calculate the first edge information of each DCT block Second edge information ,as follows:

[0024] ;

[0025] ;

[0026] in, The average value of low-frequency energy is equal to In addition to the number of DCT coefficients in the low-frequency region, This represents the sum of the absolute values ​​of the low-frequency coefficients, that is, the sum of the absolute values ​​of all DCT coefficients in the low-frequency region of the DCT block. The average value of the intermediate frequency energy is equal to... The number of DCT coefficients excluding those in the mid-frequency region; The average value of high-frequency energy is equal to The number of DCT coefficients in the high-frequency region.

[0027] Preferably, in step S102, based on the texture energy of each CTU, a CTU-level texture energy weight is obtained; based on the CTU-level texture energy weight, CTU-level bit allocation is performed to obtain the CTU-level target number of bits; based on the texture energy of each frame, a frame-level texture energy weight is obtained; based on the frame-level texture energy weight, frame-level bit allocation is performed to obtain the frame-level target number of bits, including:

[0028] The texture energy of the corresponding CTU is obtained based on the texture energy of all DCT blocks in each CTU. The texture energy of a frame is obtained based on the texture energy of all CTUs in the frame. ;

[0029] Calculate frame-level texture energy weights separately and CTU-level texture energy weight ,as follows:

[0030] ;

[0031] ;

[0032] in, This represents the total texture energy of all frames in the current GOP, which is obtained by adding the texture energy of each frame in the GOP sequentially. This represents the total texture energy of the encoded frames in the current GOP, which is obtained by adding the texture energy of each frame that has been encoded. This represents the total texture energy of the encoded CTUs within the current frame;

[0033] CTU-level bit allocation is performed based on CTU-level texture energy weights to obtain the target number of CTU-level bits. ,as follows:

[0034] ;

[0035] in, This refers to the total number of bits that the CTUs that have not yet been encoded in the current frame can use. It is obtained by subtracting the total target number of encoded CTUs from the total bit budget of the current frame. This indicates the number of CTUs that have been encoded in the current frame. After each CTU is encoded, Add 1; This represents the total number of target bits for CTUs already encoded in the current frame, obtained by summing the number of target bits previously allocated to these CTUs; This represents the total number of actual bits of the encoded CTU in the current frame, which is obtained by summing the actual number of bits generated by the encoded CTU. This represents a smoothing window used to adjust the target bit allocation of the remaining CTUs, based on... and The deviation is dynamically corrected for subsequent bit allocation;

[0036] Based on the texture energy of each frame, frame-level texture energy weights are obtained. Frame-level bit allocation is then performed based on these weights to obtain the target number of bits at the frame level. ,as follows:

[0037] ;

[0038] in, This represents the total bit budget of the current GOP, calculated by the encoder based on the set bit rate and the number of GOP frames. This represents the total number of bits actually used by the frames already encoded in the current GOP. During encoding, the actual number of bits generated for each frame is summed up after encoding, and this value is updated incrementally. .

[0039] Preferably, in step S103, the spatial contrast sensitivity threshold, luminance masking modulation factor, and contrast intensity are obtained based on the DCT coefficients of each block, specifically including:

[0040] Based on spatial frequency contrast sensitivity, obtain the spatial contrast sensitivity threshold. ,as follows:

[0041] ;

[0042] in, and These are the DCT normalization coefficients; This indicates the tilting effect of the human eye. The value is 0.6; This indicates a collective effect; Indicates the direction angle of the corresponding DCT coefficient;

[0043] The initial spatial contrast sensitivity threshold is represented as follows:

[0044] ;

[0045] in, , and Represents a constant; Indicates the first DCT coefficients The corresponding sub-band frequency;

[0046] Brightness masking modulation factor , means as follows:

[0047] ;

[0048] in, This represents the average intensity value of the nth DCT block; The calculation method is as follows:

[0049] ;

[0050] in, The DC coefficient in the DCT block is used to approximate the average brightness value of the image block; N represents the block size of the DCT block.

[0051] Contrast intensity , means as follows:

[0052] ;

[0053] Where R represents DCT blocks of various sizes Representing a spatial frequency as The DCT coefficient value at that time; K represents the range of pixel intensity values.

[0054] Preferably, in step S103, based on the texture energy and edge information of each DCT block, each DCT block is classified into flat blocks, edge blocks, or texture blocks, specifically including:

[0055] When the texture energy of a certain DCT block satisfies When the time is right, it is directly determined to be a flat block;

[0056] When satisfied When, if the DCT block Or simultaneously satisfy If the condition is met, the DCT block is determined to be an edge block; otherwise, it is a flat block. and These represent the first and second edge information of the DCT block, respectively.

[0057] When satisfied When, if the DCT block Or simultaneously satisfy If it is true, it is classified as an edge block; otherwise, it is classified as a texture block.

[0058] when When, if the DCT block or If it is true, it is determined to be an edge block; otherwise, it is a texture block.

[0059] in, , , , , , .

[0060] Preferably, in step S103, a contrast masking adjustment factor is obtained based on contrast intensity, classification type, and a linear function of DCT frequency. The details are as follows:

[0061] ;

[0062] Where t represents the adjustment factor, which is used to adjust the contrast intensity of flat blocks, texture blocks, and edge blocks; Indicates contrast intensity; This represents the gradient of a linear function of frequency.

[0063] Preferably, in step S104, a perceptual distortion index is obtained based on the JND threshold in the DCT domain; the Lagrange multiplier for the corresponding coding unit (CTU) is obtained based on the perceptual distortion index and the traditional Lagrange multiplier, as shown below:

[0064] ;

[0065] in, Represents the Lagrange multiplier of the m-th coding unit (CTU); The JND threshold of the m-th coding unit is obtained by summing the JND thresholds of the DCT fields within the coding unit. This is a preset constant, set to 0.01; This refers to the traditional Lagrange multipliers.

[0066] Preferably, in step S104, the final rate control quantization parameters are obtained based on the Lagrange multipliers. , means as follows:

[0067] .

[0068] On the other hand, a screen content video encoding-aware bitrate control device based on DCT domain JND includes:

[0069] The DCT coefficient acquisition module is used to acquire screen content video, perform block-based Discrete Cosine Transform (DCT) on the screen content video, and obtain the DCT coefficients of each block.

[0070] The texture energy and target bit count acquisition module is used to obtain the sum of the absolute values ​​of low-frequency, mid-frequency, and high-frequency coefficients in each DCT block based on the DCT coefficients of each block; obtain the texture energy of each DCT block based on the sum of the absolute values ​​of the mid-frequency and high-frequency coefficients; obtain the edge information of each DCT block based on the low-frequency, mid-frequency, and high-frequency coefficients; obtain the texture energy of each CTU based on the sum of the texture energy of all DCT blocks in each coding unit; obtain the texture energy of each frame based on the sum of the texture energy of all CTUs in each frame; obtain the CTU-level texture energy weight based on the texture energy of each CTU; perform CTU-level bit allocation based on the CTU-level texture energy weight to obtain the CTU-level target bit count; and obtain the frame-level texture energy weight based on the texture energy of each frame; perform frame-level bit allocation based on the frame-level texture energy weight to obtain the frame-level target bit count.

[0071] The JND threshold acquisition module is used to obtain the spatial contrast sensitivity threshold, luminance masking modulation factor, and contrast intensity based on the DCT coefficients of each block; classify each DCT block into flat blocks, edge blocks, or texture blocks based on the texture energy and edge information of each DCT block; obtain the contrast masking adjustment factor based on the contrast intensity, classification type, and linear function of DCT frequency; and obtain the JND threshold in the DCT domain based on the product of the spatial contrast sensitivity threshold, luminance masking modulation factor, and contrast masking adjustment factor.

[0072] The rate control quantization parameter acquisition module is used to obtain the perceptual distortion index based on the JND threshold in the DCT domain; obtain the Lagrange multiplier of the corresponding coding unit (CTU) based on the perceptual distortion index and the traditional Lagrange multiplier; and obtain the final rate control quantization parameter based on the Lagrange multiplier.

[0073] The encoding control module is used to guide the actual encoding process and perform encoding control based on the target bit count at the CTU level, the target bit count at the frame level, and the bit rate control quantization parameters.

[0074] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0075] (1) The present invention fully considers the content characteristics of screen content video and the characteristics of human visual system, and uses the texture energy of DCT domain to guide the allocation of target bits. Compared with the allocation method in the standard that does not consider the characteristics of screen content and visual redundancy, the target bit rate allocation scheme of the present invention is more accurate.

[0076] (2) This invention fully considers the problem that the rate-distortion model in the standard does not take into account the characteristics of the screen content video. It uses the JND threshold in the DCT domain to construct a perceptual rate-distortion model of the screen content video, obtains the perceptual distortion index, obtains the Lagrange multiplier of the corresponding coding unit CTU based on the perceptual distortion index and the traditional Lagrange multiplier, and obtains the final bitrate control quantization parameter based on the Lagrange multiplier, thereby improving the quality and rate-distortion performance of the encoded video sequence and improving the perceptual bitrate control accuracy. Attached Figure Description

[0077] Figure 1 A flowchart of a screen content video coding perceptual bitrate control method based on DCT domain JND provided in an embodiment of the present invention;

[0078] Figure 2 A flowchart of the screen content video coding perception bitrate control method based on DCT domain JND provided in an embodiment of the present invention;

[0079] Figure 3A schematic diagram of the DCT block classification process of the screen content video coding perception bitrate control method based on DCT domain JND provided in an embodiment of the present invention;

[0080] Figure 4 A schematic diagram comparing the performance of the screen content video coding perceptual bitrate control method based on DCT domain JND provided in this embodiment of the invention with the default bitrate control algorithm of VTM-20.0 on a Map sequence.

[0081] Figure 5 A schematic diagram comparing the performance of the screen content video coding perceptual bitrate control method based on DCT domain JND provided in this embodiment of the invention with the default bitrate control algorithm of VTM-20.0 on the Web_browsing sequence.

[0082] Figure 6 The diagram shows the structure of a screen content video encoding sensing bitrate control device based on DCT domain JND, as provided in an embodiment of the present invention. Detailed Implementation

[0083] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the teachings of this invention, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined by the appended claims.

[0084] See Figure 1 and Figure 2 As shown in the figure, this embodiment of a screen content video coding perception bitrate control method based on DCT domain JND includes the following steps.

[0085] S101, DCT coefficient acquisition step: acquire screen content video, perform block-based Discrete Cosine Transform (DCT) on the screen content video, and obtain the DCT coefficients of each block.

[0086] Specifically, the input screen content video is subjected to block-based DCT (Discrete Cosine Transform), where the DCT block size is chosen to be 8. For DCT block of size, number The formula for calculating each DCT coefficient is as follows:

[0087] ;

[0088] in, Indicates the first Image pixels, N=8. and The compensation coefficient is calculated as follows:

[0089] .

[0090] S102, Texture Energy and Target Bit Count Acquisition Steps: Based on the DCT coefficients of each block, obtain the sum of the absolute values ​​of the low-frequency, mid-frequency, and high-frequency coefficients in the DCT block respectively; based on the sum of the absolute values ​​of the mid-frequency and high-frequency coefficients, obtain the texture energy of each DCT block; based on the low-frequency, mid-frequency, and high-frequency coefficients, obtain the edge information of each DCT block; based on the sum of the texture energy of all DCT blocks in each coding unit (CTU), obtain the texture energy of each CTU; based on the sum of the texture energy of all CTUs in each frame, obtain the texture energy of each frame; based on the texture energy of each CTU, obtain the CTU-level texture energy weight; based on the CTU-level texture energy weight, perform CTU-level bit allocation to obtain the CTU-level target bit count; based on the texture energy of each frame, obtain the frame-level texture energy weight; based on the frame-level texture energy weight, perform frame-level bit allocation to obtain the frame-level target bit count.

[0091] Specifically, assuming , and Represent The sum of the absolute values ​​of the low-frequency, mid-frequency, and high-frequency coefficients in the block. In a DCT block, the frequency gradually increases from the upper left corner to the lower right corner, dividing it into low-frequency, mid-frequency, and high-frequency regions. It represents the sum of the absolute values ​​of the intermediate frequency coefficients, that is, all DCT coefficients in the intermediate frequency region of the DCT block. The sum of the absolute values ​​of . This represents the sum of the absolute values ​​of the intermediate frequency coefficients, i.e., all DCT coefficients in the intermediate frequency region of the DCT block. The sum of the absolute values; This represents the sum of the absolute values ​​of the high-frequency coefficients, i.e., all DCT coefficients in the high-frequency region of the DCT block. The sum of the absolute values ​​of .

[0092] The texture energy of a DCT block is represented by the sum of the mid-frequency coefficients and the high-frequency coefficients, as calculated below:

[0093] ;

[0094] use and Reflecting edge information, it is defined as follows:

[0095] ;

[0096] ;

[0097] in, The average value of low-frequency energy is equal to In addition to the number of DCT coefficients in the low-frequency region, This represents the sum of the absolute values ​​of the low-frequency coefficients, that is, the sum of the absolute values ​​of all DCT coefficients in the low-frequency region of the DCT block. The average value of the intermediate frequency energy is equal to... The number of DCT coefficients excluding those in the mid-frequency region;

[0098] The average value of high-frequency energy is equal to The number of DCT coefficients in the high-frequency region.

[0099] The texture energy of the corresponding CTU is obtained based on the texture energy of all DCT blocks in each CTU. The texture energy of a frame is obtained based on the texture energy of all CTUs in the frame. ;

[0100] Calculate frame-level texture energy weights separately and CTU-level texture energy weight ,as follows:

[0101] ;

[0102] ;

[0103] in, This represents the total texture energy of all frames in the current GOP, which is obtained by adding the texture energy of each frame in the GOP sequentially. This represents the total texture energy of the encoded frames in the current GOP, which is obtained by adding the texture energy of each frame that has been encoded. This represents the total texture energy of the encoded CTUs within the current frame.

[0104] It should be noted that a GOP (Group of Pictures) is an independent processing unit composed of consecutive frames in video coding, usually consisting of I-frames, P-frames, and B-frames, used to achieve inter-frame compression.

[0105] CTU-level bit allocation is performed based on CTU-level texture energy weights to obtain the target number of CTU-level bits. ,as follows:

[0106] ;

[0107] in, This refers to the total number of bits that the CTUs that have not yet been encoded in the current frame can use. It is obtained by subtracting the total target number of encoded CTUs from the total bit budget of the current frame. This indicates the number of CTUs that have been encoded in the current frame. After each CTU is encoded, Add 1; This represents the total number of target bits for CTUs already encoded in the current frame, obtained by summing the number of target bits previously allocated to these CTUs; This represents the total number of actual bits of the encoded CTU in the current frame, which is obtained by summing the actual number of bits generated by the encoded CTU. This represents a smoothing window used to adjust the target bit allocation of the remaining CTUs, based on... and The deviation is dynamically corrected for subsequent bit allocation;

[0108] Based on the texture energy of each frame, frame-level texture energy weights are obtained. Frame-level bit allocation is then performed based on these weights to obtain the target number of bits at the frame level. ,as follows:

[0109] ;

[0110] in, This represents the total bit budget of the current GOP, calculated by the encoder based on the set bit rate and the number of GOP frames. This represents the total number of bits actually used by the frames already encoded in the current GOP. During encoding, the actual number of bits generated for each frame is summed up after encoding, and this value is updated incrementally. .

[0111] S103, JND threshold acquisition step: Based on the DCT coefficients of each block, obtain the spatial contrast sensitivity threshold, luminance masking modulation factor, and contrast intensity; Based on the texture energy and edge information of each DCT block, classify each DCT block into flat blocks, edge blocks, or texture blocks; Based on the contrast intensity, classification type, and DCT frequency linear function, obtain the contrast masking adjustment factor; Based on the product of the spatial contrast sensitivity threshold, luminance masking modulation factor, and contrast masking adjustment factor, obtain the JND threshold in the DCT domain.

[0112] Specifically, the spatial contrast sensitivity function describes the relationship between contrast sensitivity (commonly measured by the reciprocal of the contrast perception threshold) and spatial frequency under different conditions. Generally, the spatial model of the spatial contrast sensitivity function (CSF) can be expressed as:

[0113] ;

[0114] in, , and These represent constants, with specific values ​​of 1.33, 0.11, and 0.18, respectively. Represents spatial frequency, for a size of DCT block of size, its first DCT coefficients Corresponding sub-band frequency The calculation method is as follows:

[0115] ;

[0116] ;

[0117] in, This indicates the viewpoint of a pixel in the horizontal or vertical direction. This represents the ratio of viewing distance to screen height. This indicates the number of pixels contained in the screen height.

[0118] According to the definition of contrast sensitivity, the spatial contrast sensitivity function and the distortion threshold defined based on the contrast sensitivity function are reciprocals of each other. Therefore, for a given spatial frequency, the initial spatial contrast sensitivity JND threshold can be expressed as:

[0119] ;

[0120] Because the human eye reacts differently to visual stimuli from different directions, it is generally more sensitive to frequency components in the horizontal and vertical directions, but less sensitive to frequency components in the tilt direction. This difference causes the JND threshold to increase accordingly. Therefore, the threshold is adjusted to obtain the spatial contrast sensitivity threshold. ,as follows:

[0121] ;

[0122] in, and It is the DCT normalization coefficient.

[0123] .

[0124] This indicates the tilting effect of the human eye. The value is 0.6. This represents the ensemble effect, with a value of 0.25. The direction angle representing the corresponding DCT coefficient is calculated as follows:

[0125] ;

[0126] The luminance masking effect indicates that as luminance intensity increases, the difference in luminance changes that the human eye can perceive also increases. That is, as the luminance level gradually increases, the JND threshold for luminance changes perceived by the human eye also increases. In 8-bit images, a pixel intensity of 128 is typically used as a baseline to estimate the threshold of the luminance masking modulation factor. The greater the deviation of the pixel value from this baseline, the higher the corresponding luminance masking modulation factor. The actual measured value of the luminance masking modulation factor is set as the ratio of two actual measured JND thresholds, as shown below:

[0127] ;

[0128] in, Indicates the nth block. The luminance intensity value corresponding to each DCT coefficient. The symbol (^) indicates that the value was obtained through measurement. Based on this, Indicates brightness intensity as The measurement of the DCT coefficients and the JND threshold; The measured JND threshold represents the DCT coefficient with a luminance intensity of 128. Analysis reveals that the luminance masking modulation factor exhibits a U-shaped curve; therefore, a piecewise linear function is proposed to describe the relationship between the average background luminance and the visual threshold. This method uses this function to calculate the luminance masking modulation factor, as detailed below:

[0129] ;

[0130] in, This represents the average intensity value of the nth DCT block; The calculation method is as follows:

[0131] ;

[0132] in, This represents the DC coefficient in the DCT block, used to approximate the average brightness value of the image block, i.e., when both i and j are 0. N represents the block size of the DCT block.

[0133] The contrast masking effect of the human eye manifests as follows: when another signal with a different contrast exists simultaneously in the visual field, the visual sensitivity of the human eye to the target signal of interest decreases. Generally, the visual sensitivity of the human eye changes with the complexity of image texture. Therefore, in constructing a contrast modulation factor model, the complexity of image texture must be considered, and it is usually determined based on the calculated contrast intensity. Since the contrast masking modulation factor is positively correlated with the standard deviation and negatively correlated with the peak value of the normalized power spectral density, the contrast intensity... The DCT block is characterized by the ratio of its standard deviation to the peak value of the normalized power spectral density, and is calculated as follows:

[0134] ;

[0135] Among them, parameters and Set them to 1.4 and 0.7 respectively. The standard deviation is expressed as follows:

[0136] ;

[0137] ;

[0138] Where P represents the magnitude of the AC component in the DCT block, and R represents... DCT blocks of various sizes Representing a spatial frequency as DCT coefficient values ​​at time ( It is a two-dimensional form representing position. The DCT coefficients; It is a one-dimensional form, representing a spatial frequency of . The DCT coefficient values ​​at that time. To ensure that the calculation results fall within a reasonable range of 0 to 1, a normalization factor of K / 2 is introduced, where K represents the range of pixel intensity values. For 8-bit image formats, the value of K is 255.

[0139] Normalized power spectral density peak It can be calculated using the following formula:

[0140] ;

[0141] in, It is the k-th moment of the probability distribution (k is usually taken as 1~4), defined as:

[0142] ;

[0143] ;

[0144] In summary, the contrast intensity can be obtained by solving the problem. The specific values ​​are:

[0145] ;

[0146] From the perspective of visual attention mechanisms, edge blocks are more favored by the human eye than texture blocks. Therefore, block classification is performed using DCT domain texture energy, dividing the image into three categories: flat blocks, texture blocks, and edge blocks. Subsequently, a specific adjustment factor t is introduced to adjust the contrast intensity of flat blocks, texture blocks, and edge blocks. The adjusted contrast intensity can be used as... This indicates that, specifically, the value of t is set to 0.5, 1, and 0.5 for edge blocks, texture blocks, and flat blocks, respectively. The specific classification process is as follows: Figure 3 As shown.

[0147] exist Figure 3 In this process, after dividing the video content of the screen into blocks, a DCT transform is applied to each block, and its texture energy is calculated. Next, according to Based on the size of the block and its frequency domain energy characteristics, the block is divided into flat blocks, edge blocks, or texture blocks. The specific determination process is as follows:

[0148] Case 1: When the texture energy of a certain block satisfies ( If the result is "flat", it indicates that the block has little texture information and is directly identified as a flat block.

[0149] Case 2: When satisfied ( When ), if the block ( (or simultaneously satisfy) ( If the condition is met, the block is considered an edge block; otherwise, it is considered a flat block.

[0150] Case 3: When the condition is satisfied ( When this condition is met, the judgment logic is the same as in case 2. If the condition in this block is met... Or simultaneously satisfy If it is true, it is classified as an edge block; otherwise, it is classified as a texture block.

[0151] Case 4: When At that time, if the block or , If it is true, it is determined to be an edge block; otherwise, it is a texture block.

[0152] The contrast masking modulation factor is linearly positively correlated with the background texture contrast. As the contrast increases, its frequency domain response also shows a positive correlation and exhibits bandpass characteristics. Based on this relationship, the contrast masking threshold is modeled as a linear function of contrast intensity and DCT frequency, as follows:

[0153] ;

[0154] in, The gradient of the frequency linear function can be calculated using the following formula:

[0155] ;

[0156] in, express The maximum parameter is set to 4.12 cpd (cycles per degree). , and The calculations are based on the least squares solution, and the specific values ​​are shown below:

[0157] ;

[0158] ;

[0159] Solve , and Then, the final JND threshold of the DCT domain is calculated using the following formula:

[0160] ;

[0161] in, This represents the JND threshold (spatial contrast sensitivity threshold) in the DCT domain, where n is the block index and i and j are the indices of the DCT coefficients. Indicates the spatial CSF threshold. Indicates the brightness masking modulation factor. This represents the contrast masking adjustment factor.

[0162] S104, the step of obtaining the rate control quantization parameters, is as follows: based on the JND threshold in the DCT domain, the perceptual distortion index is obtained; based on the perceptual distortion index and the traditional Lagrange multiplier, the Lagrange multiplier of the corresponding coding unit (CTU) is obtained; based on the Lagrange multiplier, the final rate control quantization parameters are obtained.

[0163] Specifically, the JND threshold for each coding unit is determined using the DCT domain JND model. (The JND threshold DJND of the coding unit is obtained by summing the JND thresholds of the DCT domain within the coding unit.) This threshold is then used as a perceptual sensitivity factor to measure objective distortion, correcting the objective distortion in the rate distortion formula, as shown below:

[0164] ;

[0165] Where c is a constant, set to 0.01; It is the sum of squared errors, representing the sum of squared pixel differences between the reconstructed image and the original image.

[0166] For each coding unit, the proposed perceptual distortion metric replaces the traditional distortion metric, and a new rate-distortion equation is constructed accordingly, as follows:

[0167]

[0168] in, Represents the perception of Lagrange multipliers; This indicates the actual number of encoded bits.

[0169] By further solving for the adjustment factor of the Lagrange multiplier, when perceptual distortion is used as an indicator of distortion, the following formula can be derived:

[0170] ;

[0171] Where m represents the sequence number of the coding unit, the m-th coding unit; Represents the m-th coding unit ; The m-th coding unit consumes the number of bits; c is a smoothing constant to prevent the denominator from being zero; Nblock represents the total number of coding units.

[0172] By differentiating the above equation, we can obtain the optimal solution that minimizes the rate-distortion cost for each coding unit, as shown below:

[0173] ;

[0174] Furthermore, the formula for calculating the Lagrange multipliers of the m-th coding unit can be derived as follows:

[0175] ;

[0176] in, Represents the Lagrange multiplier of the m-th coding unit (CTU); The JND threshold of the m-th coding unit is obtained by summing the JND thresholds of the DCT fields within the coding unit. This is a preset constant, set to 0.01; This refers to the traditional Lagrange multipliers.

[0177] From the above formula, we can see that the adaptive Lagrange multiplier adjustment factor of the m-th coding unit based on the DCT domain JND is... :

[0178] ;

[0179] The final quantization parameters for the rate control model are:

[0180] .

[0181] S105, the encoding control step, guides the actual encoding process and performs encoding control based on the CTU-level target bit count, frame-level target bit count, and bit rate control quantization parameters.

[0182] Based on the derivation of S104, the quantization parameters applicable to the current coding unit are obtained. , Used to control quantization strength and guide the actual coding process.

[0183] according to Calculate the quantization step size of the current CTU, and then perform encoding quantization based on the quantization step size. To maintain the continuity of encoding quality between adjacent CTUs, the current CTU's... As the initial quantization parameters for the next CTU.

[0184] After completing the actual encoding for each CTU, the corresponding actual encoded output data, such as the actual number of bits and perceived distortion, will be recorded. and the quantization parameters used Then, based on the encoded output data, the parameters of the perceptual rate distortion model are updated to optimize the accuracy of model predictions and the performance of parameter estimation in subsequent encoding processes.

[0185] This closed-loop control process, encompassing "perceptual quantization parameter generation—encoding execution—actual encoding data acquisition—model update," is performed sequentially according to the CTU's scanning order until the current frame is fully encoded. This achieves adaptive estimation of perceptual quantization parameters and continuous updating of the rate-distortion model, thereby improving the overall optimization capability of screen content video encoding between bitrate control accuracy and subjective visual quality.

[0186] In summary, this application provides a screen content video coding-aware bitrate control method based on DCT domain JND, including:

[0187] Constructing a DCT-domain JND model: A JND model in the DCT domain is constructed from three aspects: spatial contrast sensitivity function, luminance masking effect, and contrast masking effect. This model describes the human eye's perceptual threshold for distortion under different frequencies, luminance, and texture complexity. Finally, a JND threshold is formed for each DCT coefficient. ;

[0188] Perceptual Distortion Modeling: In traditional rate-distortion optimization models, the distortion term is typically represented as the mean square error in the pixel domain. This invention introduces the JND model into the distortion modeling process, proposing a perceptual distortion term. This reflects the difference in sensitivity of the human eye to different types of distortion, ensuring that imperceptible distortion does not excessively affect bit allocation.

[0189] Constructing a perceptual Lagrange multiplier model Perceptual distortion is introduced into the rate-distortion cost function. Then, by differentiating and reasoning about the rate-distortion cost function, the perceptual Lagrange multipliers are obtained. , For adaptive Lagrange multiplier adjustment factor, It is a traditional Lagrange multiplier.

[0190] Mapped to perceptual quantization parameters The final By fitting function Mapped to perceptually optimized quantization parameters It is used to guide the bit allocation and quantization control of each CTU in the actual encoding process.

[0191] Rate-distortion performance is a crucial metric for evaluating the quality of rate control algorithms, reflecting the balance between compression efficiency and reconstruction quality. Commonly used evaluation parameters include Bjøntegaard Delta Bit Rate (BDBR) and Bjøntegaard Delta PSNR (BDPSNR). A lower BDBR indicates greater bitrate savings for the same quality, while a higher BDPSNR indicates better image quality at the same bitrate. These two metrics provide a comprehensive assessment of an algorithm's rate-distortion performance.

[0192] Table 1 below shows a comparison of the rate-distortion performance of the proposed DJND-RC algorithm and the default rate control algorithm of VTM-20.0. Based on the default algorithm of the VTM-20.0 platform, the standard test set video sequences show an average bitrate saving of 2.78% and an average BDPSNR improvement of 0.34 dB. Experimental results demonstrate that the proposed rate control method has significant advantages and effectively improves bandwidth utilization efficiency.

[0193] Figure 4 and Figure 5 The figure shows a comparison of rate-distortion performance curves obtained using the DJND-RC method of this invention and the default bitrate control algorithm of the VTM-20.0 platform for different screen content video test sequences. Under the same bitrate conditions, a higher PSNR value indicates better encoding quality. In the figure, the red curve represents the performance of the DJND-RC method, and the blue curve represents the performance of the default bitrate control algorithm of VTM-20.0. As can be seen from the figure, the red curve is significantly higher than the blue curve in all the listed test sequences, which fully demonstrates that the encoding performance of this invention is superior to the platform's default algorithm, further verifying that DJND-RC has good encoding performance.

[0194] Table 1. Comparison of rate-distortion performance of DJND-RC and VTM-20.0 of the present invention;

[0195]

[0196] See Figure 6 As shown, the present invention also discloses a screen content video coding-aware bitrate control device based on DCT domain JND, comprising:

[0197] DCT coefficient acquisition module 601 is used to acquire screen content video, perform block-based discrete cosine transform (DCT) on the screen content video, and obtain the DCT coefficients of each block.

[0198] The texture energy and target bit count acquisition module 602 is used to: obtain the sum of the absolute values ​​of low-frequency, mid-frequency, and high-frequency coefficients in each DCT block based on the DCT coefficients of each block; obtain the texture energy of each DCT block based on the sum of the absolute values ​​of the mid-frequency and high-frequency coefficients; obtain the edge information of each DCT block based on the low-frequency, mid-frequency, and high-frequency coefficients; obtain the texture energy of each CTU based on the sum of the texture energy of all DCT blocks in each coding unit; obtain the texture energy of each frame based on the sum of the texture energy of all CTUs in each frame; obtain the CTU-level texture energy weight based on the texture energy of each CTU; perform CTU-level bit allocation based on the CTU-level texture energy weight to obtain the CTU-level target bit count; and obtain the frame-level texture energy weight based on the texture energy of each frame; perform frame-level bit allocation based on the frame-level texture energy weight to obtain the frame-level target bit count.

[0199] The JND threshold acquisition module 603 is used to obtain the spatial contrast sensitivity threshold, luminance masking modulation factor, and contrast intensity based on the DCT coefficients of each block; classify each DCT block into flat blocks, edge blocks, or texture blocks based on the texture energy and edge information of each DCT block; obtain the contrast masking adjustment factor based on the contrast intensity, classification type, and DCT frequency linear function; and obtain the JND threshold in the DCT domain based on the product of the spatial contrast sensitivity threshold, luminance masking modulation factor, and contrast masking adjustment factor.

[0200] The rate control quantization parameter acquisition module 604 is used to obtain the perceptual distortion index based on the JND threshold in the DCT domain; obtain the Lagrange multiplier of the corresponding coding unit (CTU) based on the perceptual distortion index and the traditional Lagrange multiplier; and obtain the final rate control quantization parameter based on the Lagrange multiplier.

[0201] The encoding control module 605 is used to guide the actual encoding process and perform encoding control based on the target bit count at the CTU level, the target bit count at the frame level, and the bit rate control quantization parameters.

[0202] The specific implementation of each module of the screen content video coding perception bitrate control device based on DCT domain JND is the same as the screen content video coding perception bitrate control method based on DCT domain JND, and will not be described again in this embodiment.

[0203] Furthermore, this embodiment of the invention also provides an electronic device. The electronic device of this embodiment includes: a processor and a memory; wherein the memory is used to store computer execution instructions; and the processor is used to execute the computer execution instructions stored in the memory to implement the various steps performed by the electronic device in the above embodiment. For details, please refer to the relevant descriptions in the foregoing method embodiments.

[0204] Alternatively, the memory can be either standalone or integrated with the processor.

[0205] When the memory is configured independently, the electronic device also includes a bus for connecting the memory and the processor.

[0206] This invention also provides a computer storage medium storing computer execution instructions, which, when executed by a processor, implement the method described above.

[0207] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0208] In the embodiments provided by this invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.

[0209] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to implement the solution of this embodiment according to actual needs.

[0210] Furthermore, the functional modules in the various embodiments of this invention can be integrated into one processing unit, or each module can exist physically separately, or two or more modules can be integrated into one unit. The unit formed by the above modules can be implemented in hardware or in the form of hardware plus software functional units.

[0211] The integrated modules described above, implemented as software functional modules, can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application.

[0212] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly manifested as execution by the hardware processor, or execution by a combination of hardware and software modules within the processor.

[0213] The memory may include high-speed RAM, and may also include non-volatile storage (NVM), such as at least one disk storage device, and may also be a USB flash drive, external hard drive, read-only memory, disk or optical disc, etc.

[0214] The bus can be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0215] The aforementioned storage medium can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium can be any available medium accessible to general-purpose or special-purpose computers.

[0216] An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium can be an integral part of the processor. Both the processor and the storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and storage medium can exist as discrete components in an electronic device or host device.

[0217] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0218] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A screen content video coding-aware bitrate control method based on DCT domain JND, characterized in that, include: S101, DCT coefficient acquisition step: acquire screen content video, perform block-based Discrete Cosine Transform (DCT) on the screen content video, and obtain the DCT coefficients of each block. S102, the step of obtaining texture energy and target bit count, based on the DCT coefficients of each block, obtains the sum of the absolute values ​​of the low-frequency, mid-frequency and high-frequency coefficients in the DCT block respectively; The texture energy of each DCT block is obtained by summing the absolute values ​​of the intermediate frequency coefficients and the high frequency coefficients; the edge information of each DCT block is obtained by summing the low frequency, intermediate frequency, and high frequency coefficients; the texture energy of each CTU is obtained by summing the texture energy of all DCT blocks in each coding unit; the texture energy of each frame is obtained by summing the texture energy of all CTUs in each frame; the CTU-level texture energy weight is obtained based on the texture energy of each CTU; and the CTU-level bit allocation is performed based on the CTU-level texture energy weight to obtain the target number of bits at the CTU level. Based on the texture energy of each frame, obtain the frame-level texture energy weight, perform frame-level bit allocation based on the frame-level texture energy weight, and obtain the frame-level target number of bits. S103, JND threshold acquisition step: Based on the DCT coefficients of each block, obtain the spatial contrast sensitivity threshold, luminance masking modulation factor and contrast intensity; Based on the texture energy and edge information of each DCT block, classify each DCT block into flat blocks, edge blocks or texture blocks; Based on the contrast intensity, classification type and DCT frequency linear function, obtain the contrast masking adjustment factor. The JND threshold in the DCT domain is obtained based on the product of the spatial contrast sensitivity threshold, the luminance masking modulation factor, and the contrast masking adjustment factor. S104, the step of obtaining the rate control quantization parameters: based on the JND threshold in the DCT domain, obtain the perceptual distortion index; based on the perceptual distortion index and the traditional Lagrange multiplier, obtain the Lagrange multiplier of the corresponding coding unit (CTU); based on the Lagrange multiplier, obtain the final rate control quantization parameters. S105, encoding control step, guides the actual encoding process and performs encoding control based on CTU-level target bit count, frame-level target bit count and bit rate control quantization parameters; In step S103, based on the texture energy and edge information of each DCT block, the DCT blocks are classified into flat blocks, edge blocks, or texture blocks, specifically including: When the texture energy of a certain DCT block satisfies E tex When μ1 is less than or equal to 1, it is directly determined to be a flat block; When μ1 is satisfied <E tex When ≤μ2, if E1≥θ2 or max(E1,E2)≥α1&min(E1,E2)≥α2 simultaneously, the DCT block is determined to be an edge block; otherwise, it is a flat block. E1 and E2 represent the first edge information and the second edge information of the DCT block, respectively. When μ2 is satisfied <E tex When E ≤ μ3, if the DCT block's E1 ≥ θ2 or simultaneously satisfies max(E1, E2) ≥ α1 & min(E1, E2) ≥ α2, it is classified as an edge block; otherwise, it is classified as a texture block. tex When the value is greater than μ3, if E1 ≥ θ2 or max(E1,E2) ≥ θ1α1 & min(E1,E2) ≥ θ1α2, the DCT block is determined to be an edge block; otherwise, it is a texture block. Where μ1 = 14, μ2 = 45, θ2 = 22, α1 = 7, α2 = 5, μ3 = 79, and θ1 = 0.

05.

2. The screen content video coding-aware bitrate control method based on DCT domain JND according to claim 1, in step S101, for an N×N DCT block, the (i,j)th DCT coefficient X i,j The calculation formula is as follows: in, I(k,l) represents the (k,l)th image pixel, where k,l = 0,1,...,N-1; λ i and λ j Let i and j represent the compensation coefficients respectively, calculated as follows:

3. The screen content video coding perceptual bitrate control method based on DCT domain JND according to claim 1, in step S102, the texture energy of each DCT block is obtained based on the sum of the absolute values ​​of the intermediate frequency coefficient and the high frequency coefficient, as shown below: AND Tex =And M +E H ; in, E M This represents the sum of the absolute values ​​of the intermediate frequency coefficients, specifically the sum of the absolute values ​​of all DCT coefficients within the intermediate frequency region of the DCT block; E H It represents the sum of the absolute values ​​of the high-frequency coefficients, that is, the sum of the absolute values ​​of all DCT coefficients in the high-frequency region of the DCT block; Based on low-frequency, mid-frequency, and high-frequency coefficients, edge information of each DCT block is obtained, including: The first edge information E1 and the second edge information E2 of each DCT block are calculated as follows: in, This represents the average value of low-frequency energy, equal to E. L Except for the number of DCT coefficients in the low-frequency region, E L This represents the sum of the absolute values ​​of the low-frequency coefficients, that is, the sum of the absolute values ​​of all DCT coefficients in the low-frequency region of the DCT block. This represents the average value of the intermediate frequency energy, equal to E. M The number of DCT coefficients excluding those in the mid-frequency region; The average value of high-frequency energy is E. H The number of DCT coefficients in the high-frequency region.

4. The screen content video coding perception bitrate control method based on DCT domain JND according to claim 1, in step S102, based on the texture energy of each CTU, a CTU-level texture energy weight is obtained, and based on the CTU-level texture energy weight, CTU-level bit allocation is performed to obtain the CTU-level target number of bits. Based on the texture energy of each frame, frame-level texture energy weights are obtained. Frame-level bit allocation is then performed based on these weights to obtain the target number of bits at the frame level, including: The texture energy E of the corresponding CTU is obtained based on the texture energy of all DCT blocks in each CTU. TexCTU The texture energy E of a frame is obtained based on the texture energy of all CTUs in the frame. Texframe ; Calculate the frame-level texture energy weights ω respectively CurPic and CTU-level texture energy weight ω CurCTU ,as follows: Among them, E TexGOP E represents the total texture energy of all frames in the current GOP, obtained by adding the texture energy of each frame in the GOP sequentially; Texcodedframe E represents the total texture energy of the encoded frames in the current GOP, obtained by adding the texture energy of each frame that has been encoded; TexcodedCTU This represents the total texture energy of the encoded CTUs within the current frame; CTU-level bit allocation is performed based on CTU-level texture energy weights to obtain the target number of CTU-level bits R. CurCTU ,as follows: Among them, R curleft The total number of bits available for CTUs that have not yet been encoded in the current frame is calculated by subtracting the total target number of encoded CTUs from the total bit budget of the current frame; CodedCTUs represents the number of CTUs that have been encoded in the current frame, incremented by 1 for each CTU encoded; T codedCTU This represents the total number of target bits encoded in the current frame for each CTU, obtained by summing the target bits previously allocated to these CTUs; A codedCTU T represents the total number of actual bits of the encoded CTUs in the current frame, obtained by summing the actual number of bits generated by the encoded CTUs; SW represents the smoothing window, used to adjust the target bit allocation of the remaining CTUs, based on T. codedCTU and A codedCTU The deviation is dynamically corrected for subsequent bit allocation; based on the texture energy of each frame, a frame-level texture energy weight is obtained, and frame-level bit allocation is performed based on the frame-level texture energy weight to obtain the frame-level target number of bits R. CurPic ,as follows: R CurPic =(R GOP -Coded GOP )·ω CurPic ; Among them, R GOP This represents the total bit budget of the current GOP, calculated by the encoder based on the set bitrate and the number of GOP frames; Coded GOP This represents the total number of bits actually used by the frames already encoded in the current GOP. During encoding, after each frame is encoded, the actual number of bits generated in that frame is summed up, and the resulting Coded value is updated incrementally. GOP .

5. The screen content video coding perceptual bitrate control method based on DCT domain JND according to claim 1, wherein in step S103, the spatial contrast sensitivity threshold, luminance masking modulation factor, and contrast intensity are obtained based on the DCT coefficients of each block, specifically including: Based on spatial frequency contrast sensitivity, the spatial contrast sensitivity threshold J is obtained. CSF (n,i,j), as follows: Where, φ i and φ j These are the DCT normalization coefficients; The value of r represents the tilt effect of the human eye, with a value of 0.6; s represents the aggregation effect. Indicates the direction angle of the corresponding DCT coefficient; J(ω i,j The initial spatial contrast sensitivity threshold is represented as follows: Where a1, a2, and a3 represent constants; ω i,j X represents the (i,j)th DCT coefficient. i,j Corresponding sub-band frequency; brightness masking modulation factor M lum , means as follows: in, This represents the average intensity value of the nth DCT block; The calculation method is as follows: Among them, X 0,0 The DC coefficient in the DCT block is used to approximate the average brightness value of the image block; N represents the block size of the DCT block. The contrast intensity τ(n) is expressed as follows: Where R represents an N×N DCT block, X(ω) represents the DCT coefficient value at a certain spatial frequency ω, and K represents the range of pixel intensity values.

6. The screen content video coding perceptual bitrate control method based on DCT domain JND according to claim 1, in step S103, the contrast masking adjustment factor M is obtained based on contrast intensity, classification type, and a linear function of DCT frequency. con The details are as follows: M con =t·τ(n)·f(ω)+1; in, t represents the adjustment factor used to adjust the contrast intensity of flat blocks, texture blocks, and edge blocks; τ(n) represents the contrast intensity; f(ω) represents the gradient of the frequency linear function.

7. The screen content video coding perceptual bitrate control method based on DCT domain JND according to claim 1, in step S104, a perceptual distortion index is obtained based on the JND threshold of the DCT domain; the Lagrange multiplier of the corresponding coding unit (CTU) is obtained based on the perceptual distortion index and the traditional Lagrange multiplier, as follows: in, λ mp DJND represents the Lagrange multiplier of the m-th coding unit (CTU); m λ represents the JND threshold of the m-th coding unit, obtained by summing the JND thresholds of the DCT fields within the coding unit; c is a preset constant, set to 0.01; mSSE This refers to the traditional Lagrange multipliers.

8. The screen content video coding-aware bitrate control method based on DCT domain JND according to claim 7, in step S104, the final bitrate control quantization parameter QP is obtained based on Lagrange multipliers. mp , means as follows: QP mp =4.2005·λ mp +13.7122+0.5。 9. A screen content video encoding sensing bitrate control device based on DCT domain JND, characterized in that, The method based on any one of claims 1 to 8 includes: The DCT coefficient acquisition module is used to acquire screen content video, perform block-based Discrete Cosine Transform (DCT) on the screen content video, and obtain the DCT coefficients of each block. The texture energy and target bit count acquisition module is used to obtain the sum of the absolute values ​​of low-frequency, mid-frequency, and high-frequency coefficients in each DCT block based on the DCT coefficients of each block; obtain the texture energy of each DCT block based on the sum of the absolute values ​​of the mid-frequency and high-frequency coefficients; obtain the edge information of each DCT block based on the low-frequency, mid-frequency, and high-frequency coefficients; obtain the texture energy of each CTU based on the sum of the texture energy of all DCT blocks in each coding unit; obtain the texture energy of each frame based on the sum of the texture energy of all CTUs in each frame; obtain the CTU-level texture energy weight based on the texture energy of each CTU; perform CTU-level bit allocation based on the CTU-level texture energy weight to obtain the CTU-level target bit count; and obtain the frame-level texture energy weight based on the texture energy of each frame; perform frame-level bit allocation based on the frame-level texture energy weight to obtain the frame-level target bit count. The JND threshold acquisition module is used to obtain the spatial contrast sensitivity threshold, luminance masking modulation factor, and contrast intensity based on the DCT coefficients of each block; classify each DCT block into flat blocks, edge blocks, or texture blocks based on the texture energy and edge information of each DCT block; obtain the contrast masking adjustment factor based on the contrast intensity, classification type, and DCT frequency linear function; and obtain the JND threshold in the DCT domain based on the product of the spatial contrast sensitivity threshold, luminance masking modulation factor, and contrast masking adjustment factor. The rate control quantization parameter acquisition module is used to obtain the perceptual distortion index based on the JND threshold in the DCT domain; obtain the Lagrange multiplier for the corresponding coding unit (CTU) based on the perceptual distortion index and the traditional Lagrange multiplier; and obtain the final rate control quantization parameters based on the Lagrange multiplier. The encoding control module is used to guide the actual encoding process and perform encoding control based on the target bit count at the CTU level, the target bit count at the frame level, and the bit rate control quantization parameters.

Citation Information

Patent Citations

  • Method for guiding multi-view video coding quantization process by visual perception characteristics

    CN103124347A

  • Visual perceptual coding method based on multi-domain JND (Just Noticeable Difference) model

    CN107241607A