Construction Method of Just-Noticeable Distortion Model in DCT Domain Based on Entropy Masking
By constructing the Bayesian prediction model and entropy masking model of the DCT domain, the problem that the existing DCT domain JND model fails to fully consider the entropy masking effect is solved, and the accuracy of JND threshold estimation is improved, especially in the processing of image edges and texture areas, and better noise hiding effect and practicality are achieved.
Patent Information
- Application Number
- CN202210662131.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-13
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-06-13
AI Technical Summary
The existing DCT domain JND model fails to fully consider the entropy masking effect when processing image blocks, resulting in inaccurate JND threshold estimation, especially in the processing of image edge areas and texture areas.
By constructing a Bayesian prediction model of the DCT domain, the similarity is calculated using the texture energy difference between the central block and the surrounding block, and a DCT domain autoregressive prediction model is established. Then, the disorder function is constructed based on the residual block, and the regulator of the entropy masking effect is constructed based on the disorder, which is fused into the DCT domain JND model.
The accuracy of the DCT domain JND model is improved, especially in the processing of image edge areas and texture areas, which can better hide the noise of ordered texture areas, and at the same time, the calculation process does not have cross-domain operations, which is more practical.
Smart Images

Figure CN115086682B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of perceptual video coding, and particularly relates to a method for constructing a just-noticeable distortion model in the DCT domain based on entropy masking. Background Art
[0002] Due to its potential physiological and psychological mechanisms, the human visual system (HVS) cannot perceive image distortions below a certain threshold, which is called the just-noticeable distortion (JND) threshold. The JND model directly utilizes the characteristics of the human visual system and can thus be widely applied to perceptual image and video processing applications, such as image and video compression / coding, quality evaluation, and watermarking.
[0003] Generally speaking, existing DCT-domain JND models are composed of three main influencing factors, namely the contrast sensitivity function (CSF), the luminance adaptation (LA) effect, and the contrast masking (CM) effect. Existing DCT-domain JND models mainly focus on the CM effect. The CM model in the DCT domain usually divides image blocks into flat blocks, edge blocks, and texture blocks, and sets different weights for different types of blocks to highlight the texture. However, in fact, the HVS can adaptively predict ordered texture content, and the JND thresholds in these regions are often overestimated. Some literature has explored incorporating the entropy masking (EM) effect, which can characterize this property, into the estimation of the DCT-domain JND threshold, but there are more or less some defects. The EM effect is considered in the CM model of Bae et al., but due to the relatively larger standard deviation in the edge region than in the texture region during the calculation process of the CM model, the quality of the edge region of the distorted image is poor; Wan et al. classified image blocks into five types by fusing the CM and EM effects based on the texture energy (TE) classification method, but due to the limited number of contrast intensity levels, the accuracy of the JND model is still limited; Wang et al. proposed a distance-based disorder evaluation method and constructed an EM effect adjustment factor, but their EM model must operate in a cross-domain according to the calculation process. Since most image / video resources are effectively stored and transmitted in a compressed form, it becomes very important and urgent to directly operate on the compressed data rather than on its decompressed version. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for constructing a just-noticeable distortion model in the DCT domain based on entropy masking to solve the above technical problems.
[0005] To solve the above technical problems, the specific technical solution of the method for constructing a just-noticeable distortion model in the DCT domain based on entropy masking of the present invention is as follows:
[0006] A method for constructing a just-noticeable distortion model in the DCT domain based on entropy masking includes the following steps:
[0007] Step 1: Construct a Bayesian prediction model in the DCT domain;
[0008] Step 1.1: Calculate the similarity between the central block and the surrounding blocks by using the texture energy difference between them;
[0009] Step 1.2: Construct a Bayesian prediction model in the DCT domain by using the surrounding blocks and their similarity with the central block;
[0010] Step 1.3: Classify the image into edge blocks and non-edge blocks by using the TE block classification method in the DCT domain for prediction;
[0011] Step 2: Construct an entropy masking model;
[0012] Step 2.1: Obtain the prediction residual through the current block and its predicted block;
[0013] Step 2.2: Measure the disorder degree of the current block by using the prediction residual and the DCT coefficients of the current block;
[0014] Step 2.3: Construct a regulator for the EM effect according to the disorder degree;
[0015] Step 3: Construct an EM-JND model.
[0016] Furthermore, the specific steps of the said Step 1.1 are as follows:
[0017] Since the N×N DCT transform breaks the correlation between image blocks, the DCT coefficients of the central block are predicted by using the corresponding DCT coefficients of the surrounding 8 blocks χ={B 1 ,…,B 8};
[0018] For an N = 8 DCT block, L, M, and H respectively represent the sum of the absolute DCT coefficient values in the low-frequency, medium-frequency, and high-frequency groups, and the texture energy of the DCT block is calculated as:
[0019] E tex = M + H (1)
[0020] The texture energy difference between the central block and its surrounding blocks is then calculated as:
[0021] ΔE tex =|E ctex -E stex | (2)
[0022] In the formula, E ctex and E stex are respectively the texture energies of the central block and the surrounding blocks;
[0023] If the central block is very similar to the surrounding blocks with little uncertainty, the current block can be accurately inferred from the surrounding blocks, i.e., ΔE tex is close to or equal to 0; conversely, the lower the similarity between the central block and the surrounding blocks, the tex larger ΔE tex is. Therefore, it can be inferred that ΔE tex is an effective measure of the order degree of the image content and is inversely proportional to it. Therefore, an inverse proportional function is constructed using ΔTexE to measure the similarity between DCT blocks. Due to the existence of the situation where ΔE tex = 0, and the exponential function can more effectively assign similarity according to the
[0024]
[0025] Furthermore, step 1.2 includes the following specific steps:
[0026] By imitating the reasoning mechanism of IGM, a prediction model is created in the DCT domain. This model predicts the DCT coefficients of the current block based on the DCT coefficients of the surrounding DCT blocks and their similarity. The higher the similarity S(B,χ) between the central block and its surrounding blocks, the lower the degree of disorder of the image content; and the k larger the S(B c ,B k ) of the surrounding block B c , the greater its role in the prediction process of IGM. Therefore, the similarity S(B k ,B k ) between the central block and its surrounding blocks is used as the autoregressive coefficient to establish a Bayesian prediction model in the DCT domain:
[0027]
[0028] where P(X) is the predicted DCT block, B k (X) is the input DCT block, is the normalized similarity, and ε is white noise.
[0029] Furthermore, step 1.3 includes the following specific steps:
[0030] For the edge region, HVS should be able to make predictions adaptively. When the central block is on a straight line in the flat region, only the surrounding blocks in the edge region can be used for prediction. Therefore, the TE block classification method in the DCT domain is used to divide the image into edge blocks and non-edge blocks for prediction. When the central block is an edge block, only the surrounding blocks that are also edge blocks are used for similarity weighted prediction; when the central block is a non-edge block, only the surrounding blocks that are also non-edge blocks among the surrounding blocks are used for similarity weighted prediction. Except for the above two situations, the central block is predicted by the 8 surrounding blocks together with similarity weighting.
[0031] Furthermore, the specific steps of step 2.1 are as follows:
[0032] The residual between the central block B c (X) and the predicted block P(X) is regarded as the residual block R(X), and the calculation is as follows:
[0033] R(X) = |B c (X) - P(X)|. (5)
[0034] Furthermore, the specific steps of step 2.2 are as follows:
[0035] Dividing the residual block R(X) by the central block B c (X) and normalizing the division result to obtain the disorder block. The disorder of the nth DCT block is obtained by averaging its disorder block, and the calculation is as follows:
[0036]
[0037] In the formula, is the normalization function.
[0038] Furthermore, the specific steps of step 2.3 are as follows:
[0039] The JND model in the DCT domain is expressed as a product of multiple adjustment factors. Therefore, the entropy masking adjustment factor J EM for the ordered region without entropy masking is expressed as 1, while the J EM for the disordered region with entropy masking should be greater than 1. The disorder ζ(n) represents the uncertainty of the input image content. The larger ζ(n) is, the more uncertain information the input image contains, that is, the more noise can be hidden. Since the exponential function image is concave upward, more noise can be distributed to the image blocks with larger ζ(n). Therefore, J EM is constructed as:
[0040] J EM = 1 + α · e ζ(n) (7)
[0041] In the formula, α is the proportionality factor, which is set to 0.15 according to subjective experiments here.
[0042] Furthermore, the specific steps of step 3 are as follows:
[0043] The DCT-domain JND model is expressed as a product of multiple adjustment factors. The JND model considers four masking effects, and the calculation is as follows:
[0044] J = J base × J LA × J CM × JEM (8)
[0045] Wherein, J base That is, CSF-JND, which is mainly estimated based on the spatial contrast sensitivity function CSF and does not include other masking effects. J LA Corresponds to the luminance adaptation effect, which reveals the sensitivity of the HVS to the background luminance. J CM Corresponds to the contrast masking effect, that is, the visibility attenuation effect of the current visual signal in the presence of other visual signals.
[0046] The method for constructing a just-noticeable distortion model in the DCT domain based on entropy masking of the present invention has the following advantages:
[0047] 1. The DCT-domain Bayesian prediction model proposed by the present invention fully considers the characteristics of the human visual system actively predicting the input scene and the correlation between DCT blocks;
[0048] 2. The DCT-domain EM-JND model proposed by the present invention is more in line with the human visual characteristics and can better hide the noise in the ordered texture area while ensuring the image quality;
[0049] 3. The DCT-domain EM-JND model proposed by the present invention does not have cross-domain operations throughout the calculation process and is more practical in practice. Brief Description of the Drawings
[0050] Figure 1 is the flow chart for constructing the entropy masking model of the present invention;
[0051] Figure 2(a) is the schematic principle diagram of the DCT-domain Bayesian prediction example of the present invention;
[0052] Figure 2(b) is the schematic principle diagram of the DCT-domain edge region prediction example of the present invention. Detailed Embodiments
[0053] In order to better understand the purpose, structure and function of the present invention, the following further describes in detail a method for constructing a just-noticeable distortion model in the DCT domain based on entropy masking of the present invention with reference to the drawings.
[0054] The existing DCT-domain JND model is mainly affected by the contrast sensitivity function (CSF), luminance adaptation (LA) effect and contrast masking (CM) effect. However, in fact, the internal generation mechanism (IGM) of the HVS can adaptively predict the input scene under the guidance of the free energy principle, that is, the HVS processes as much structural information as possible and avoids uncertain information. This characteristic reveals the perceptual limitation of the HVS, and the JND threshold of the disordered texture area containing a large amount of uncertain information is often relatively high. Therefore, it is necessary to introduce the entropy masking (EM) effect to improve the accuracy of the DCT-domain JND model.
[0055] Since Bayesian inference is a powerful tool for information prediction, the present invention uses the Bayesian brain theory to simulate the IGM in the human brain for orderly information prediction and establish an autoregressive prediction model in the DCT domain. Then, a disorder function is constructed based on residual blocks, and the EM effect is constructed as a regulator of disorder according to subjective experiments. The construction process of the entropy masking model is as Figure 1 shown. Finally, the DCTune model that can characterize the masking effect in the DCT domain is used to fuse the EM effect with the spatial contrast sensitivity function, brightness adaptation effect, and contrast masking effect to simulate the estimated threshold of EM-JND.
[0056] Specifically, the method for constructing a just-noticeable distortion model in the DCT domain based on entropy masking of the present invention includes the following steps:
[0057] Step 1: Construct a Bayesian prediction model in the DCT domain
[0058] 1. Calculate the similarity between the central block and the surrounding blocks using the texture energy difference between them.
[0059] Since the N×N DCT transform breaks the correlation between image blocks, the present invention uses the co-located DCT coefficients of the surrounding 8 blocks (χ = {B 1 , …, B 8}) to predict the DCT coefficients of the central block. For example, the DCT coefficient at the (0, 0) position of the central block is jointly predicted by the DCT coefficients at the (0, 0) positions of the surrounding 8 blocks, as shown in Fig. 2(a).
[0060] For an N = 8 DCT block, L, M, and H can respectively represent the sum of the absolute DCT coefficient values in the low-frequency (LF), mid-frequency (MF), and high-frequency (HF) groups. The texture energy of the DCT block can be approximately calculated as:
[0061] E tex = M + H (9)
[0062] The texture energy difference between the central block and its surrounding blocks can then be calculated as:
[0063] ΔE tex = |E ctex - E stex | (10)
[0064] In the formula, E ctex and E stex are the texture energies of the central block and the surrounding blocks, respectively.
[0065] If the central block and the surrounding blocks are very similar and there is almost no uncertainty, the current block can be accurately inferred using the surrounding blocks, i.e., ΔE texis close to or equal to 0. On the contrary, if the similarity between the central block and the surrounding blocks is lower, then ΔE tex is larger. Therefore, it can be inferred that ΔE tex is an effective measure of the order degree of the image content and is inversely proportional to it. Therefore, the present invention uses ΔTexE to construct an inverse proportional function to measure the similarity between DCT blocks. Since there is a situation where ΔE tex = 0, and the exponential function can more effectively allocate similarity according to the size of ΔE tex , the final similarity calculation is as follows:
[0066]
[0067] 2. Use the surrounding blocks and their similarity to the central block to construct a Bayesian prediction model in the DCT domain.
[0068] By imitating the inference mechanism of IGM, the present invention attempts to create a prediction model in the DCT domain, which predicts the DCT coefficients of the current block based on the DCT coefficients of the surrounding DCT blocks and their similarity. The higher the similarity S(B,χ) between the central block and its surrounding blocks, the lower the degree of disorder of the image content; and the larger the S(B k ) of the surrounding block B c ,B k ), the greater the role in the prediction process of IGM. Therefore, the present invention uses the similarity S(B c ,B k ) between the central block and its surrounding blocks as the autoregressive coefficient to establish a Bayesian prediction model in the DCT domain:
[0069]
[0070] In the formula, P(X) is the predicted DCT block, B k (X) is the input DCT block, is the normalized similarity, and ε is white noise.
[0071] 3. Use the TE block classification method in the DCT domain to divide the image into edge blocks and non-edge blocks for prediction.
[0072] For the edge region, HVS should be able to perform adaptive prediction. As shown in Fig. 2(b), when the central block is on a straight line in the flat region (i.e., the edge region), only the surrounding blocks in the same edge region (the two diagonal blocks on the straight line) should be used for prediction. Therefore, the present invention uses the TE block classification method in the DCT domain to divide the image into edge blocks and non-edge blocks for prediction. When the central block is an edge block, only the surrounding blocks that are also edge blocks are used for similarity weighted prediction; when the central block is a non-edge block, only the surrounding blocks that are also non-edge blocks among the surrounding blocks are used for similarity weighted prediction. Except for the above two cases, the central block is jointly predicted by 8 surrounding blocks with similarity weighting.
[0073] Step 2: Construct an entropy masking model
[0074] 1. Obtain the prediction residual from the current block and its predicted block.
[0075] According to the IGM theory, the brain works as an active inference system, which can accurately predict ordered content and avoid disordered information. Therefore, in the present invention, the residual between the central block B c (X) and the predicted block P(X) is regarded as the residual block R(X), and the calculation is as follows:
[0076] R(X) = |B c (X) - P(X)| (13)
[0077] 2. Measure the disorder degree of the current block by using the prediction residual and the DCT coefficients of the current block.
[0078] Since most of the energy is concentrated in the low-frequency coefficients in the upper left corner after the DCT transformation, there is a large difference in the amplitude between the DCT coefficients. And when the amplitude of the DCT coefficient is larger, its prediction residual is usually larger. Therefore, in the present invention, the residual block R(X) is divided by the central block B c (X), and the division result is normalized to obtain the disorder degree block. The disorder degree of the nth DCT block is obtained by averaging its disorder degree block, and the calculation is as follows:
[0079]
[0080] In the formula, is the normalization function.
[0081] 3. Construct the adjustment factor of the EM effect according to the disorder degree.
[0082] The JND model based on the DCT domain is generally expressed in the form of a product of multiple adjustment factors. Therefore, the entropy masking adjustment factor (J EM ) of the ordered region without entropy masking can be expressed as 1, while the J EM in the disordered region with entropy masking should be greater than 1. The disorder degree ζ(n) represents the uncertainty of the content of the input image. The larger ζ(n) is, the more uncertain information the input image contains, that is, the more noise can be hidden. Since the exponential function image is concave upward, more noise can be distributed to the image blocks with larger ζ(n). Therefore, in the present invention, J EM is constructed as:
[0083] J EM = 1 + α·e ζ(n) (15)
[0084] In the formula, α is the proportionality factor, which is set to 0.15 according to the subjective experiment of the present invention.
[0085] Step 3: Construct the EM-JND model
[0086] Generally, the DCT-domain JND model is expressed as the product of multiple adjustment factors, which is derived from the DCTune model proposed by Watson et al. The JND model proposed by the present invention mainly considers four masking effects and is calculated as follows:
[0087] J = J base ×J LA ×J CM ×J EM (16)
[0088] In the formula, J base , that is, CSF-JND, is mainly estimated based on the spatial contrast sensitivity function (CSF) and does not include other masking effects. J LA corresponds to the luminance adaptation effect, which reveals the sensitivity of the HVS to the background luminance. J CM corresponds to the contrast masking effect, that is, the visibility attenuation effect of the current visual signal in the presence of other visual signals.
[0089] It can be understood that the present invention is described through some embodiments. Those skilled in the art know that, without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. In addition, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.
Claims
1. A method for constructing a just noticeable distortion model in the DCT domain based on entropy masking, characterized in that, it includes the following steps: Step 1: Construct a Bayesian prediction model in the DCT domain; Step 1.1: Calculate the similarity between the central block and the surrounding blocks using the texture energy difference between them; since the N×N DCT transform breaks the correlation between image blocks, the DCT coefficients of the central block are predicted using the co-located DCT coefficients of the surrounding 8 blocks χ = {B 1 , …, B 8}. For an N = 8 DCT block, L, M, and H represent the sums of the absolute DCT coefficient values in the low-frequency, mid-frequency, and high-frequency groups respectively, and the texture energy of the DCT block is calculated as: E tex = M + H (1) The texture energy difference between the central block and its surrounding blocks is calculated as: ΔE tex = |E ctex - E stex | (2) where E ctex and E stex are the texture energies of the central block and the surrounding blocks, respectively; If the central block is very similar to the surrounding blocks and there is little uncertainty, the current block can be accurately inferred from the surrounding blocks, i.e., ΔE tex is close to or equal to 0; conversely, if the similarity between the central block and the surrounding blocks is lower, then ΔE tex is larger. Therefore, it can be inferred that ΔE tex is an effective measure of the order degree of the image content and is inversely proportional to it. Therefore, use ΔE tex to construct an inverse proportional function to measure the similarity between DCT blocks. Since there is a situation where ΔE tex = 0, and the exponential function can more effectively allocate similarity according to the size of ΔE tex , the final similarity calculation is as follows: Step 1.2: Use the surrounding blocks and their similarity to the central block to construct a Bayesian prediction model in the DCT domain; By mimicking the inference mechanism of IGM, a prediction model is created in the DCT domain. This model predicts the DCT coefficients of the current block based on the DCT coefficients of the surrounding DCT blocks and their similarity. The higher the similarity S(B,χ) between the central block and its surrounding blocks, the lower the degree of disorder of the image content; and the surrounding block B k with a larger S(B c ,B k ) plays a greater role in the prediction process of IGM. Therefore, the similarity S(B c ,B k ) between the central block and its surrounding blocks is used as the autoregressive coefficient to establish a Bayesian prediction model in the DCT domain: Wherein, P(X) is the predicted DCT block, and B k (X) is the input DCT block, is the normalized similarity, and ε is white noise; Step 1.3: Use the TE block classification method in the DCT domain to divide the image into edge blocks and non-edge blocks for prediction; For the edge region, the HVS should be able to make predictions adaptively. When the central block is on a straight line in the flat region, only the surrounding blocks in the same edge region are used for prediction. Therefore, the TE block classification method in the DCT domain is used to divide the image into edge blocks and non-edge blocks for prediction. When the central block is an edge block, only the surrounding blocks that are also edge blocks are used for similarity weighted prediction; when the central block is a non-edge block, only the surrounding blocks that are also non-edge blocks among the surrounding blocks are used for similarity weighted prediction. Except for the above two cases, the central block is jointly predicted by 8 surrounding blocks with similarity weighting; Step 2: Construct an entropy masking model; Step 2.1: Obtain the prediction residual through the current block and its prediction block; Take the center block B c (X) The residual between the predicted block P(X) is regarded as the residual block R(X), which is calculated as follows: R(X) = |B c (X) - P(X)|; (5) Step 2.2: Use the prediction residual and the DCT coefficients of the current block to measure the disorder degree of the current block; Dividing the residual block R(X) by the central block B c (X), and normalizing the division result to obtain the disorder block. The disorder of the nth DCT block is obtained by averaging its disorder block, and the calculation is as follows: In the formula, is a normalization function; Step 2.3: Construct an adjustment factor for the EM effect according to the disorder degree; The JND model based on the DCT domain is expressed as a product of multiple adjustment factors. Therefore, there is no entropy masking adjustment factor J for the ordered region of entropy masking EM which is expressed as 1, while J for the disordered region with entropy masking EM should be greater than 1. The degree of disorder ζ(n) represents the uncertainty of the input image content. The larger ζ(n) is, the more uncertain information the input image contains, that is, the more noise can be hidden. Since the exponential function image is concave upward, more noise can be distributed to the image blocks with larger ζ(n). Therefore, J EM is constructed as: J EM = 1 + α·e ζ(n) (7) In the formula, α is a scaling factor, which is set to 0.15 according to subjective experiments here; Step 3: Construct an EM-JND model; The JND model in the DCT domain is expressed as the product of multiple adjustment factors. The JND model considers four masking effects and is calculated as follows: J = J base × J LA × J CM × J EM (8) where J base i.e., CSF-JND, is mainly estimated based on the contrast sensitivity function (CSF) in the spatial domain and does not include other masking effects. J LA corresponds to the luminance adaptation effect, which reveals the sensitivity of the HVS to the background luminance. J CM corresponds to the contrast masking effect, i.e., the visibility attenuation effect of the current visual signal in the presence of other visual signals.