A top-down approach to estimating just-perceptible distortion threshold for natural images
Through top-down design ideas and KLT transformation technology, the appropriate perceived distortion threshold map of natural images is estimated, which solves the problem that existing models are difficult to accurately reflect the characteristics of human visual systems, and achieves more efficient visual signal processing and image compression.
Patent Information
- Application Number
- CN202210025549.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-11
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-01-11
AI Technical Summary
The existing accurate perceptible distortion (JND) threshold estimation model is difficult to accurately reflect the visual masking characteristics of human visual systems and the visual perception redundancy of natural images, resulting in poor results.
Using the top-down design idea and using KLT transformation technology, the KLT coefficient energy and normalized energy of the image are calculated, and the perceived distortion-free critical point is estimated, thereby obtaining the appropriate perceived distortion threshold map of the natural image.
This method can more accurately reflect the visual masking characteristics of the human visual system and the visual perception redundancy of natural images, improve the accuracy of the appropriately perceived distortion threshold estimate, and hide more noise and save bit rate while maintaining consistency of visual quality.
Smart Images

Figure CN114519668B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a natural image just noticeable distortion (JND) threshold estimation technology, and in particular to a top-down natural image just noticeable distortion threshold estimation method, which is based on a top-down design idea and utilizes the KLT (Karhunen-Loéve Transform) transformation technology to achieve just noticeable distortion threshold estimation of natural images. Background Art
[0002] Just Noticeable Distortion (JND) refers to the maximum change amplitude of the visual signal that cannot be perceived by the human visual system (HVS). It reflects the sensitivity of the human visual system (HVS) to changes in visual information and the potential perceptual redundancy in visual signals. This makes it widely used in many image / video perception processing tasks, including image / video compression, image / video enhancement, information hiding, and image / video evaluation. Precisely because of its wide application, the JND threshold estimation of natural images has received extensive attention and research.
[0003] Existing JND threshold estimation models can be divided into two categories: JND threshold estimation models based on pixel domain and JND threshold estimation models based on transform domain. The JND threshold estimation model based on pixel domain mainly considers factors such as luminance adaptation (LA), contrast masking (CM) and pattern complexity (PC). The JND threshold estimation model based on transform domain converts the image to a specific transform domain and estimates the JND threshold for each subband, which mainly considers factors such as contrast sensitivity function (CSF). From the design perspective, the existing pixel-domain based JND threshold estimation model is roughly the same as the transform-domain based JND threshold estimation model. Specifically, the visual masking effects (such as LA, CM, PC, CSF) with different effects are first modeled, and then the different visual masking effect models are fused to obtain the final JND threshold estimation model. This design idea can be regarded as a bottom-up strategy, that is, starting from the bottom of the contributing factors, the final JND threshold estimation model is derived. However, this design idea has some inherent limitations: first, due to the lack of a deep and comprehensive understanding of the characteristics of the human visual system (HVS), it is difficult to take all the potentially relevant influencing factors into account; second, the influencing factors that are taken into account are often difficult to accurately describe through simple mathematical models; third, the relationship between different influencing factors is also difficult to model. Therefore, the existing JND threshold estimation model often fails to achieve satisfactory results. Although people can explore more influencing factors and the corresponding visual masking effects through experiments and model them with more accurate mathematical models, such work is endless. Therefore, it is of great significance to overcome the above shortcomings and design a more advanced JND threshold estimation model. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a top-down natural image just perceptible distortion threshold estimation method, which can well reflect the visual masking characteristics of the human visual system and can well characterize the visual perception redundancy of natural images, and thus can provide effective guidance for various visual signal perception processing tasks.
[0005] The technical solution adopted by the present invention to solve the above technical problem is: a top-down natural image just perceptible distortion threshold estimation method, characterized by comprising the following steps:
[0006] Step 1: Take a natural image to be processed as the source image; then convert the source image into a grayscale image, denoted as I Y ; Among them, the source image is an RGB color image, the source image and I Y The width is W and the height is H;
[0007] Step 2: I Y Divide into Num non-overlapping sizes Then, I Y Each image block in is vectorized to obtain I Y The column vector corresponding to each image block in I Y The column vector corresponding to the nth image block in is denoted as x n ; Then I Y The column vectors corresponding to all image blocks in are concatenated to form a vectorized matrix, denoted as X, X = [x 1 ,x 2 ,...,x n ,...,x Num ];in, Set W and H to be Divisible, the value of K is 4 2 or 5 2 or 6 2 or 7 2 or 8 2 or 9 2 or 10 2 , 1≤n≤Num,x 1 I Y The column vector corresponding to the first image block in 2 I Y The column vector corresponding to the second image block in Num I Y The column vector corresponding to the Numth image block in x 1 、x 2 、x n 、x Num The dimensions of are all K×1, the dimension of X is K×Num, and the symbol “[]” represents a vector or matrix;
[0008] Step 3: Calculate the covariance matrix of X, denoted as C; then use the eigenvalue decomposition technique to process C to obtain K eigenvalues and corresponding K eigenvectors of C; then sort the K eigenvectors of C in descending order according to the corresponding K eigenvalues, and use the matrix composed of the K eigenvectors of C according to the sorting results as the prior information extracted from X; where the dimension of C is K×K, the eigenvector is a column vector, the dimension of the eigenvector is K×1, each column of the prior information extracted from X is an eigenvector of C, and the dimension of the prior information extracted from X is K×K;
[0009] Step 4: Extract the prior information from X as I Y The KLT kernel is denoted as P; then based on P and X, calculate I Y The KLT coefficient matrix is denoted as Q, Q = (P) T X; then Q is expressed as Q = [q 1 ,q 2 ,...,q k ,…,q K ] T ; Among them, the dimension of P is K×K, the dimension of Q is K×Num, 1≤k≤K, q 1 represents the first-dimensional KLT spectral component in Q, q 2 represents the second-dimensional KLT spectral component in Q, q k represents the k-th dimension KLT spectral component in Q, q K represents the K-th dimension KLT spectral component in Q, q 1 ,q 2 ,q k ,q K The dimensions of are all Num×1;
[0010] Step 5: Calculate the KLT coefficient energy of each dimension of the KLT spectral component in Q, and convert q k The KLT coefficient energy is denoted as E k , Then calculate the normalized KLT coefficient energy of each KLT spectral component in Q, and convert q k The normalized KLT coefficient energy is recorded as Then calculate the cumulative normalized KLT coefficient energy of each dimension of the KLT spectrum component in Q, and convert q k The cumulative normalized KLT coefficient energy is recorded as Finally, the accumulated normalized KLT coefficient energies of all KLT spectral components in Q are combined into the accumulated normalized KLT coefficient energy vector, denoted as E cum , Among them, q k (n) represents q kThe value of the nth element in E, 1≤ζ≤K, ζ represents the ζ-th KLT spectral component q in Q ζ The KLT coefficient energy, Indicates q 1 The normalized KLT coefficient energy of Indicates q 2 The normalized KLT coefficient energy, E cum The dimension of is 1×K, Indicates q 1 The cumulative normalized KLT coefficient energy of Indicates q 2 The cumulative normalized KLT coefficient energy of Indicates q K The cumulative normalized KLT coefficient energy of;
[0011] Step 6: E cum Substituted as input into the perceptual distortion-free critical point calculation model, the calculation results are I Y The perceptual distortion-free critical point is recorded as L; then the perceptual distortion-free coefficient reconstruction matrix is constructed based on L, recorded as Then adopt The reconstructed perceptually distortion-free coefficient matrix is recorded as Then Expressed as Where L is a positive integer, 1≤L≤K, The dimension is K×Num, The "=" in q is an assignment symbol. L represents the L-th dimension KLT spectral component in Q, q L The dimension is Num×1, to All are zero vectors, The dimensions of are all Num×1, The dimension is K×Num, express The first dimension of the perceptually distortion-free coefficient vector in, express The second dimension perceptually distortion-free coefficient vector in, express The n-th dimension perceptually distortion-free coefficient vector in, express The Num-th dimension perceptually distortion-free coefficient vector in, The dimensions of are all K×1;
[0012] Step 7: Perform the inverse operation of the vectorization process in step 2, Each dimension of the perceptually distortion-free coefficient vector in is converted into a vector of size The image block will The converted image block is taken as the nth image block; then The image blocks converted from all perceptually distortion-free coefficient vectors in are spliced into an image as the perceptually distortion-free critical image, denoted as I L ; Then according to I Y and I L , calculate the just perceptible distortion threshold map, denoted as M, and denote the pixel value of the pixel point with coordinate position (a, b) in M as M(a, b), M(a, b) = |I Y (a,b)-I L (a,b)|; Among them, 1≤a≤W, 1≤b≤H, I Y (a,b) represents I Y The pixel value of the pixel point with coordinate position (a, b) in the middle, I L (a,b) represents I L The pixel value of the pixel point with coordinate position (a, b). The symbol “||” is the absolute value symbol.
[0013] In step 2, x n The acquisition process is: I Y The pixel values of all pixels in the nth image block are arranged in a column to form x n .
[0014] In step 3, the calculation formula of C is: The superscript "T" indicates the transpose of a vector or matrix. represents the mean vector obtained by taking the mean of X by row, The dimension of is K×1.
[0015] In step 4, P = [p 1 ,p 2 ,…,p K ], where p 1 represents the first eigenvector after sorting the K eigenvectors of C in descending order according to the corresponding K eigenvalues, p 2 It represents the second eigenvector after sorting the K eigenvectors of C in descending order according to the corresponding K eigenvalues, p K represents the Kth eigenvector of C after sorting the K eigenvectors of C in descending order according to the corresponding K eigenvalues, p 1 、p 2 、p K The dimensions of are all K×1.
[0016] In step 6, the process of obtaining the perceptual distortion-free critical point calculation model is as follows:
[0017] Step 6_1: Select S high-definition images and convert each high-definition image into a grayscale image; then follow the process from step 2 to step 4 to obtain the KLT coefficient matrix of the grayscale image of each high-definition image in the same way, and record the KLT coefficient matrix of the grayscale image of the i-th high-definition image as Q' i , Q' i Expressed as Q' i =[q' i,1 ,q' i,2 ,…,q' i,k ,…,q' i,K ] T ; Where S ≥ 100, the HD image is an RGB color image, the width of the HD image is W' and the height is H', and both W' and H' can be Divisible by, the size of the image block of the grayscale image segmentation of the high-definition image is 1≤i≤S,Q' i The dimension is K×Num', where Num' represents the total number of image blocks of the grayscale image segmentation of the high-definition image. q' i,1 Indicates Q' i The first dimension KLT spectral component, q' i,2 Indicates Q' i The second-dimensional KLT spectral component in , q' i,k Indicates Q' i The k-th KLT spectral component in , q' i,K Indicates Q' i The K-th dimension KLT spectral component, q' i,1 ,q' i,2 ,q' i,k ,q' i,K The dimensions of are all Num'×1;
[0018] Step 6_2: Construct K perceptual coefficient reconstruction matrices corresponding to the grayscale image of each high-definition image, and record the kth perceptual coefficient reconstruction matrix corresponding to the grayscale image of the i-th high-definition image as Then, the K perception coefficient matrices corresponding to the grayscale image of each high-definition image are reconstructed, and the kth perception coefficient matrix corresponding to the grayscale image of the i-th high-definition image is recorded as Will Expressed as in, The dimension is K×Num', The "=" in is the assignment symbol. to All are zero vectors, The dimensions of are all Num'×1, The dimension is K×Num', P' i represents the KLT kernel of the grayscale image of the i-th HD image, P' i The dimension is K×K, 1≤n'≤Num', n' is a positive integer, express The first dimension perceptual coefficient vector in, express The second dimension perceptual coefficient vector in, express The n'th dimension perceptual coefficient vector in, express The Num'th dimension perceptual coefficient vector in, The dimensions of are all K×1;
[0019] Step 6_3: According to the inverse operation of the vectorization process in step 2, each dimension of the perceptual coefficient vector in each perceptual coefficient matrix corresponding to the grayscale image of each high-definition image is converted into a perceptual coefficient vector of size Then, all the image blocks converted from the perceptual coefficient vectors in each perceptual coefficient matrix corresponding to the grayscale image of each high-definition image are stitched into one image, and the kth perceptual coefficient matrix corresponding to the grayscale image of the i-th high-definition image is The image formed by stitching together the image blocks converted from all the perceptual coefficient vectors in is used as the kth reconstructed image corresponding to the grayscale image of the i-th high-definition image, denoted as in, The width is W' and the height is H';
[0020] Step 6_4: Call D volunteers, each volunteer compares the grayscale image of each high-definition image with its corresponding reconstructed images in turn by naked eye observation, each volunteer determines a reconstructed image from the K reconstructed images corresponding to the grayscale image of each high-definition image as the perceptual distortion-free critical image corresponding to the grayscale image, and uses the serial number of the determined reconstructed image as the perceptual distortion-free critical point corresponding to the grayscale image; for the d-th volunteer and the grayscale image of the ith high-definition image, the d-th volunteer compares the grayscale image of the ith high-definition image with its corresponding 1st reconstructed image, 2nd reconstructed image, ..., Kth reconstructed image in turn by naked eye observation, and stops the comparison process once the d-th volunteer cannot distinguish the grayscale image of the ith high-definition image from one of its corresponding reconstructed images, and assumes that the reconstructed image is the kth reconstructed image corresponding to the grayscale image of the ith high-definition image. Then As the perceptual distortion-free critical image corresponding to the grayscale image of the i-th HD image observed by the d-th volunteer, the value k is taken as the perceptual distortion-free critical point corresponding to the grayscale image of the i-th HD image observed by the d-th volunteer, denoted as Then the vector consisting of the perceptual distortion-free critical points corresponding to the grayscale images of the i-th high-definition image observed by all volunteers is recorded as J i , Where D>1, 1≤d≤D, The "=" in J is an assignment symbol. i The dimension of is 1×D, represents the perceptual distortion-free critical point corresponding to the grayscale image of the i-th HD image observed by the first volunteer, represents the perceptual distortion-free critical point corresponding to the grayscale image of the i-th high-definition image observed by the second volunteer, represents the perceptual distortion-free critical point corresponding to the grayscale image of the i-th high-definition image observed by the D-th volunteer;
[0021] Step 6_5: Calculate the mean and standard deviation of all perceptual distortion-free critical points in the vector composed of the perceptual distortion-free critical points corresponding to the grayscale images of each high-definition image observed by all volunteers, and convert J i The mean and standard deviation of all perceptually distortion-free critical points in are recorded as and Then, outliers are removed from the vector of perceptual distortion-free critical points corresponding to the grayscale images of each HD image observed by all volunteers. i ,like Dissatisfied
[0022] Then judge is an outlier, From J i The vector obtained after removing outliers is recorded as
[0023] Step 6_6: Calculate the vector of perceptual distortion-free critical points corresponding to the grayscale image of each high-definition image observed by all volunteers, and the mean of all perceptual distortion-free critical points after removing outliers. The mean of all perceptually distortion-free critical points in is denoted as Then, the perceptual distortion-free critical point corresponding to the grayscale image of each high-definition image is obtained, and the perceptual distortion-free critical point corresponding to the grayscale image of the i-th high-definition image is recorded as J i , Among them, the symbol is the symbol for the rounding up operation;
[0024] Step 6_7: Calculate the KLT coefficient energy of each dimension of the KLT spectral component in the KLT coefficient matrix of the grayscale image of each high-definition image, and convert Q' i q' i,k The KLT coefficient energy is denoted as U i,k , Then calculate the normalized KLT coefficient energy of each dimension of the KLT spectral component in the KLT coefficient matrix of the grayscale image of each HD image, and convert Q' i q' i,k The normalized KLT coefficient energy is recorded as Then calculate the cumulative normalized KLT coefficient energy at the perceptual distortion-free critical point corresponding to the grayscale image of each HD image, and convert the perceptual distortion-free critical point J corresponding to the grayscale image of the i-th HD image into i The cumulative normalized KLT coefficient energy at Finally, the accumulated normalized KLT coefficient energy at the perceptual distortion-free critical point corresponding to the grayscale images of all high-definition images is formed into a vector, denoted as U cum , Among them, q' i,k (n') means q' i,k The value of the n'th element in U, 1≤ζ≤K, i,ζ Indicates Q' i The ζ-th KLT spectral component q' in i,ζ The KLT coefficient energy, Indicates Q' i q' i,1 The normalized KLT coefficient energy of Indicates Q' i q' i,2 The normalized KLT coefficient energy of Indicates Q' i In The normalized KLT coefficient energy of Indicates Q' i The J i dimensional KLT spectral component, U cum The dimension of is 1×S, The perceptual distortion-free critical point J corresponding to the grayscale image of the first high-definition image is represented by 1 The accumulated normalized KLT coefficient energy at , The perceptual distortion-free critical point J corresponding to the grayscale image of the second high-definition image is represented by 2 The accumulated normalized KLT coefficient energy at , The perceptual distortion-free critical point J corresponding to the grayscale image of the Sth high-definition image is represented by S The accumulated normalized KLT coefficient energy at ;
[0025] Step 6_8: Calculate U cum The mean and standard deviation of all the cumulative normalized KLT coefficient energies in are recorded as and Then according to and The calculation model of the perceived distortion-free critical point is obtained, which is described as: Where L represents I Y The perceptual distortion-free critical point.
[0026] In the step 6_1, Q' i The acquisition process is:
[0027] Step 6_1a: The grayscale image of the i-th high-definition image is recorded as I' Y,i ; Then I' Y,i Divide into Num' non-overlapping sizes image block; then I' Y,i Each image block in is vectorized to obtain I' Y,i The column vector corresponding to each image block in I' Y,i The column vector corresponding to the n'th image block in is denoted as x' i,n' ; Then I' Y,i The column vectors corresponding to all image blocks in are concatenated to form a vectorized matrix, denoted as X' i , X' i =[x' i,1 ,x' i,2 ,…,x' i,n' ,…,x' i,Num' ];in, The value of K is 4 2 or 5 2 or 6 2 or 7 2 or 8 2 or 9 2 or 10 2 , 1≤n'≤Num', x' i,1 Indicates I' Y,i The column vector corresponding to the first image block in, x' i,2 Indicates I' Y,i The column vector corresponding to the second image block in, x' i,Num' Indicates I' Y,i The column vector corresponding to the Num'th image block in x' i,1 、x' i,2 、x' i,n' 、x' i,Num' The dimensions of X' are K×1. iThe dimension is K×Num';
[0028] Step 6_1b: Calculate X' i The covariance matrix of i ; Then use the eigenvalue decomposition technique to calculate C' i Process and get C' i The K eigenvalues and corresponding K eigenvectors of C' i Sort the K eigenvectors of C' in descending order according to the corresponding K eigenvalues. i The matrix composed of the K eigenvectors of X' is constructed by sorting them. i The prior information extracted from i The dimension of is K×K, the eigenvector is a column vector, and the dimension of the eigenvector is K×1. i Each column of the prior information extracted is C' i 1 eigenvector of X' i The dimension of the prior information extracted is K×K;
[0029] Step 6_1c: Will be from X' i The prior information extracted from Y,i The KLT kernel is denoted as P' i ; Then according to P' i and X' i , calculate I' Y,i The KLT coefficient matrix Q' i , Q' i =(P' i ) T X' i ; Among them, P' i The dimension is K×K, Q' i The dimension is K×Num'.
[0030] Compared with the prior art, the advantages of the present invention are:
[0031] 1) Starting from the definition of just perceptible distortion, the method of the present invention converts the estimation problem of just perceptible distortion threshold into the estimation problem of perceptual distortion-free critical image. Perceptual distortion-free critical image refers to a distorted image that the human visual system just cannot perceive distortion. Compared with the source image, the perceptual distortion-free critical image is still distorted, but from the perspective of perception, its distortion cannot be perceived by the human eye. Therefore, the source image can be subtracted from the perceptual distortion-free critical image to obtain the just perceptible distortion threshold map corresponding to the source image. This method can effectively avoid the errors caused by the complex and inaccurate HVS visual masking factor modeling and its fusion in the traditional just perceptible distortion threshold estimation model, so as to more accurately reflect the visual masking characteristics of the human visual system and the visual perception redundancy of natural images, and further provide effective guidance for various visual signal perception processing tasks. The just perceptible distortion threshold map estimated by the method of the present invention is used to guide image denoising and JPEG compression. Under the premise of keeping the subjective perception quality almost completely consistent, the method of the present invention can hide more noise and save more bit rate.
[0032] 2) The method of the present invention uses KLT transformation technology to estimate the perceptual distortion-free critical image of the source image, which is independent of specific image processing tasks and has better universality. First, the source image is subjected to KLT transformation, and the perceptual distortion-free critical point of the source image is found according to the convergence characteristics of the accumulated normalized KLT coefficient energy. Under the condition of the perceptual distortion-free critical point, the perceptual distortion-free critical image of the source image can be reconstructed by performing KLT inverse transformation. The perceptual distortion-free critical image obtained in this way is universal and can be well applied to different visual information perception processing tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is the overall flow chart of the method of the present invention;
[0034] Figure 2a It is the first source image img1;
[0035] Figure 2b It is the second source image img2;
[0036] Figure 2c It is the third source image img3;
[0037] Figure 2d It is the fourth source image img4;
[0038] Figure 2e for Figure 2a , Figure 2b , Figure 2c , Figure 2d The corresponding cumulative normalized KLT coefficient energy curve;
[0039] Figure 3a is the source image I03;
[0040] Figure 3b In order to use the method of the present invention Figure 3a The just perceptible distortion threshold map obtained by processing the source image shown;
[0041] Figure 3c For use Figure 3b Guided generated noisy image under PSNR=26dB condition;
[0042] Figure 3d for Figure 3c A magnified view of the portion within the middle frame;
[0043] Figure 4a The source image I01 is directly compressed using JPEG.
[0044] Figure 4b The JPEG compression result guided by the just perceptible distortion threshold map obtained by the method of the present invention for the source image I01;
[0045] Figure 4c The JPEG compression result guided by the just perceptible distortion threshold map obtained by Wu2017 method for the source image I01;
[0046] Figure 5 This is the curve showing the change of the average value of gain Gain of the method of the present invention and the method of Wu2017 with the quality factor QP. DETAILED DESCRIPTION
[0047] The present invention is further described in detail below with reference to the accompanying drawings.
[0048] The present invention proposes a top-down natural image just perceptible distortion threshold estimation method, the overall flow chart of which is as follows: Figure 1 As shown, it includes the following steps:
[0049] Step 1: Take a natural image to be processed as the source image; then convert the source image into a grayscale image, denoted as I Y ; Among them, the source image is an RGB color image, the source image and I Y The width is W and the height is H, I Y The pixel value I of the pixel point with coordinate position (a, b) Y The calculation formula of (a,b) is: Y (a,b)=0.299I R (a,b)+0.587I G (a,b)+0.114I B (a,b),1≤a≤W,1≤b≤H,I R(a,b) represents the pixel value of the pixel with coordinates (a,b) in the red channel of the source image. G (a,b) represents the pixel value of the pixel with coordinates (a,b) in the green channel of the source image. B (a,b) represents the pixel value of the pixel with coordinates (a,b) in the blue channel of the source image.
[0050] Step 2: I Y Divide into Num non-overlapping sizes Then, I Y Each image block in is vectorized to obtain I Y The column vector corresponding to each image block in I Y The column vector corresponding to the nth image block in is denoted as x n ; Then I Y The column vectors corresponding to all image blocks in are concatenated to form a vectorized matrix, denoted as X, X = [x 1 ,x 2 ,…,x n ,…,x Num ];in, Set W and H to be Divisible, the value of K is 4 2 or 5 2 or 6 2 or 7 2 or 8 2 or 9 2 or 10 2 , usually take 8 2 , 1≤n≤Num,x 1 Indicates I Y The column vector corresponding to the first image block in 2 Indicates I Y The column vector corresponding to the second image block in Num Indicates I Y The column vector corresponding to the Num-th image block in x 1 、x 2 、x n 、x Num The dimensions of are K×1, the dimension of X is K×Num, and the symbol “[]” represents a vector or matrix.
[0051] In this embodiment, in step 2, x n The acquisition process is: I Y The pixel values of all pixels in the nth image block are arranged in a column to form x n .
[0052] In the field of image processing, it is a conventional technical means to perform vectorization processing on image blocks, that is, to arrange the pixel values of all pixels in the image block in a certain order (such as in the order of row scanning, first scanning the first row, then scanning the second row, and so on, that is, a Z-shaped scanning method) to form a column vector; when multiple column vectors are spliced into a vectorized matrix, they can be spliced according to the order of the image blocks, such as the first column of the vectorized matrix is the column vector corresponding to the first image block, and the last column of the vectorized matrix is the column vector corresponding to the Num-th image block.
[0053] Step 3: Calculate the covariance matrix of X, denoted as C; then use the existing eigenvalue decomposition technology to process C to obtain K eigenvalues and corresponding K eigenvectors of C; then sort the K eigenvectors of C in descending order according to the corresponding K eigenvalues from large to small, and use the matrix composed of the K eigenvectors of C according to the sorting results as the prior information extracted from X; wherein the dimension of C is K×K, the eigenvector is a column vector, the dimension of the eigenvector is K×1, each column in the prior information extracted from X is an eigenvector of C, and the dimension of the prior information extracted from X is K×K.
[0054] In this embodiment, in step 3, the calculation formula of C is: The superscript "T" indicates the transpose of a vector or matrix. represents the mean vector obtained by taking the mean of X by row, The dimension of is K×1.
[0055] Step 4: Extract the prior information from X as I Y The KLT (Karhunen-Loéve Transform) kernel is denoted as P; then based on P and X, I is calculated Y The KLT coefficient matrix is denoted as Q, Q = (P) T X; then Q is expressed as Q = [q 1 ,q 2 ,…,q k ,…,q K ] T ; Among them, the dimension of P is K×K, the dimension of Q is K×Num, 1≤k≤K, q 1 represents the first-dimensional KLT spectral component in Q, q 2 represents the second-dimensional KLT spectral component in Q, q k represents the k-th dimension KLT spectral component in Q, q K represents the K-th dimension KLT spectral component in Q, q 1 ,q 2 ,q k ,q KThe dimension of is Num×1.
[0056] In this embodiment, in step 4, P = [p 1 ,p 2 ,…,p K ], where p 1 represents the first eigenvector after sorting the K eigenvectors of C in descending order according to the corresponding K eigenvalues, p 2 It represents the second eigenvector after sorting the K eigenvectors of C in descending order according to the corresponding K eigenvalues, p K represents the Kth eigenvector of C after sorting the K eigenvectors of C in descending order according to the corresponding K eigenvalues, p 1 、p 2 、p K The dimensions of are all K×1.
[0057] Step 5: Calculate the KLT coefficient energy of each dimension of the KLT spectral component in Q, and convert q k The KLT coefficient energy is denoted as E k , Then calculate the normalized KLT coefficient energy of each KLT spectral component in Q, and convert q k The normalized KLT coefficient energy is recorded as Then calculate the cumulative normalized KLT coefficient energy of each dimension of the KLT spectrum component in Q, and convert q k The cumulative normalized KLT coefficient energy is recorded as Finally, the accumulated normalized KLT coefficient energies of all KLT spectral components in Q are combined into the accumulated normalized KLT coefficient energy vector, denoted as E cum , Among them, q k (n) represents q k The value of the nth element in E, 1≤ζ≤K, ζ represents the ζ-th KLT spectral component q in Q ζ The KLT coefficient energy, Indicates q 1 The normalized KLT coefficient energy of Indicates q 2 The normalized KLT coefficient energy, E cum The dimension of is 1×K, Indicates q 1 The cumulative normalized KLT coefficient energy of Indicates q 2 The cumulative normalized KLT coefficient energy of Indicates q K The cumulative normalized KLT coefficient energy of; Figure 2aGiven the first source image img1, Figure 2b Given the second source image img2, Figure 2c Given the third source image img3, Figure 2d Given the 4th source image img4, Figure 2e Given Figure 2a , Figure 2b , Figure 2c , Figure 2d The corresponding cumulative normalized KLT coefficient energy curve. Figure 2e It can be seen that the cumulative normalized KLT coefficient energy of the first dimension KLT spectral component is the largest, and the cumulative normalized KLT coefficient energy increases with the increase of the KLT spectral component index. As the KLT spectral component index increases, the cumulative normalized KLT coefficient energy gradually tends to saturation and eventually becomes 1. Although the cumulative normalized KLT coefficient energy curves corresponding to different images have the same trend, they are not completely consistent.
[0058] Step 6: E cum Substituted as input into the perceptual distortion-free critical point calculation model, the calculation results are I Y The perceptual distortion-free critical point is recorded as L; then the perceptual distortion-free coefficient reconstruction matrix is constructed based on L, recorded as Then adopt The reconstructed perceptually distortion-free coefficient matrix is recorded as Then Expressed as Where L is a positive integer, 1≤L≤K, The dimension is K×Num, The "=" in q is an assignment symbol. L represents the L-th dimension KLT spectral component in Q, q L The dimension is Num×1, to All are zero vectors, The dimensions of are all Num×1, The dimension is K×Num, express The first dimension of the perceptually distortion-free coefficient vector in, express The second dimension perceptually distortion-free coefficient vector in, express The n-th dimension perceptually distortion-free coefficient vector in, express The Num-th dimension perceptually distortion-free coefficient vector in, The dimensions of are all K×1.
[0059] In this embodiment, in step 6, the process of acquiring the perceptual distortion-free critical point calculation model is as follows:
[0060] Step 6_1: Select S high-definition images and convert each high-definition image into a grayscale image; then follow the process from step 2 to step 4 to obtain the KLT coefficient matrix of the grayscale image of each high-definition image in the same way, and record the KLT coefficient matrix of the grayscale image of the i-th high-definition image as Q' i , Q' i Expressed as Q' i =[q' i,1 ,q' i,2 ,…,q' i,k ,…,q' i,K ] T ; Where S ≥ 100, in the experiment, S = 500 can be taken, the high-definition image is an RGB color image, the width of the high-definition image is W' and the height is H', W' and H' can be Divisible by, the size of the image block of the grayscale image segmentation of the high-definition image is 1≤i≤S,Q' i The dimension is K×Num', where Num' represents the total number of image blocks of the grayscale image segmentation of the high-definition image. q' i,1 Indicates Q' i The first dimension KLT spectral component, q' i,2 Indicates Q' i The second-dimensional KLT spectral component in , q' i,k Indicates Q' i The k-th KLT spectral component in , q' i,K Indicates Q' i The K-th dimension KLT spectral component, q' i,1 ,q' i,2 ,q' i,k ,q' i,K The dimension is Num'×1.
[0061] In this embodiment, in step 6_1, Q' i The acquisition process is:
[0062] Step 6_1a: The grayscale image of the i-th high-definition image is recorded as I' Y,i ; Then I' Y,i Divide into Num' non-overlapping sizes image block; then I' Y,i Each image block in is vectorized to obtain I' Y,i The column vector corresponding to each image block in I' Y,i The column vector corresponding to the n'th image block in is denoted as x'i,n' ; Then I' Y,i The column vectors corresponding to all image blocks in are concatenated to form a vectorized matrix, denoted as X' i , X' i =[x' i,1 ,x' i,2 ,…,x' i,n' ,…,x' i,Num' ];in, The value of K is 4 2 or 5 2 or 6 2 or 7 2 or 8 2 or 9 2 or 10 2 , usually take 8 2 , 1≤n'≤Num', x' i,1 Indicates I' Y,i The column vector corresponding to the first image block in, x' i,2 Indicates I' Y,i The column vector corresponding to the second image block in, x' i,Num' Indicates I' Y,i The column vector corresponding to the Num'th image block in x' i,1 、x' i,2 、x' i,n' 、x' i,Num' The dimensions of X' are K×1. i The dimension is K×Num'.
[0063] Step 6_1b: Calculate X' i The covariance matrix of i ; Then use the existing eigenvalue decomposition technology to C' i Process and obtain C' i The K eigenvalues and corresponding K eigenvectors of C' i Sort the K eigenvectors of C' in descending order according to the corresponding K eigenvalues. i The matrix composed of the K eigenvectors of X' is constructed by sorting them. i The prior information extracted from i The dimension of is K×K, the eigenvector is a column vector, and the dimension of the eigenvector is K×1. i Each column of the prior information extracted is C' i 1 eigenvector of X' i The dimension of the prior information extracted is K×K.
[0064] Step 6_1c: Will be from X' iThe prior information extracted from Y,i The KLT kernel is denoted as P' i ; Then according to P' i and X' i , calculate I' Y,i The KLT coefficient matrix Q' i , Q' i =(P' i ) T X' i ; Among them, P' i The dimension is K×K, Q' i The dimension is K×Num'.
[0065] Step 6_2: Construct K perceptual coefficient reconstruction matrices corresponding to the grayscale image of each high-definition image, and record the kth perceptual coefficient reconstruction matrix corresponding to the grayscale image of the i-th high-definition image as Then, the K perception coefficient matrices corresponding to the grayscale image of each high-definition image are reconstructed, and the kth perception coefficient matrix corresponding to the grayscale image of the i-th high-definition image is recorded as Will Expressed as in, The dimension is K×Num', The "=" in is the assignment symbol. to All are zero vectors, The dimensions of are all Num'×1, The dimension is K×Num', P'i represents the KLT kernel of the grayscale image of the i-th high-definition image, P' i The dimension is K×K, 1≤n'≤Num', n' is a positive integer, express The first dimension perceptual coefficient vector in, express The second dimension perceptual coefficient vector in, express The n'th dimension perceptual coefficient vector in, express The Num'th dimension perceptual coefficient vector in, The dimensions of are all K×1.
[0066] Step 6_3: According to the inverse operation of the vectorization process in step 2, each dimension of the perceptual coefficient vector in each perceptual coefficient matrix corresponding to the grayscale image of each high-definition image is converted into a perceptual coefficient vector of size Then, according to the order of image block segmentation in step 2, all the image blocks converted from the perceptual coefficient vectors in each perceptual coefficient matrix corresponding to the grayscale image of each high-definition image are spliced into one image, and the kth perceptual coefficient matrix corresponding to the grayscale image of the i-th high-definition image is The image formed by stitching together the image blocks converted from all the perceptual coefficient vectors in is used as the k-th reconstructed image corresponding to the grayscale image of the i-th high-definition image, denoted as in, The width is W' and the height is H'.
[0067] Step 6_4: Call D volunteers, each volunteer compares the grayscale image of each high-definition image with its corresponding reconstructed images in turn by naked eye observation, each volunteer determines a reconstructed image from the K reconstructed images corresponding to the grayscale image of each high-definition image as the perceptual distortion-free critical image corresponding to the grayscale image, and uses the serial number of the determined reconstructed image as the perceptual distortion-free critical point corresponding to the grayscale image; for the d-th volunteer and the grayscale image of the ith high-definition image, the d-th volunteer compares the grayscale image of the ith high-definition image with its corresponding 1st reconstructed image, 2nd reconstructed image, ..., Kth reconstructed image in turn by naked eye observation, and stops the comparison process once the d-th volunteer cannot distinguish the grayscale image of the ith high-definition image from one of its corresponding reconstructed images, and assumes that the reconstructed image is the kth reconstructed image corresponding to the grayscale image of the ith high-definition image. Then As the perceptual distortion-free critical image corresponding to the grayscale image of the i-th HD image observed by the d-th volunteer, the value k is taken as the perceptual distortion-free critical point corresponding to the grayscale image of the i-th HD image observed by the d-th volunteer, denoted as Then the vector consisting of the perceptual distortion-free critical points corresponding to the grayscale images of the i-th high-definition image observed by all volunteers is recorded as J i , Among them, D>1, in the experiment, D=30, 1≤d≤D, The "=" in J is an assignment symbol. i The dimension of is 1×D, represents the perceptual distortion-free critical point corresponding to the grayscale image of the i-th HD image observed by the first volunteer, represents the perceptual distortion-free critical point corresponding to the grayscale image of the i-th high-definition image observed by the second volunteer, It represents the perceptual distortion-free critical point corresponding to the grayscale image of the i-th high-definition image observed by the D-th volunteer.
[0068] Step 6_5: Calculate the mean and standard deviation of all perceptual distortion-free critical points in the vector composed of the perceptual distortion-free critical points corresponding to the grayscale images of each high-definition image observed by all volunteers, and convert J i The mean and standard deviation of all perceptually distortion-free critical points in are recorded as and Then, outliers are removed from the vector of perceptual distortion-free critical points corresponding to the grayscale images of each HD image observed by all volunteers. i ,like Dissatisfied Then judge is an outlier, From J i The vector obtained after removing outliers is recorded as
[0069] Step 6_6: Calculate the vector of perceptual distortion-free critical points corresponding to the grayscale image of each high-definition image observed by all volunteers, and the mean of all perceptual distortion-free critical points after removing outliers. The mean of all perceptually distortion-free critical points in is denoted as Then, the perceptual distortion-free critical point corresponding to the grayscale image of each high-definition image is obtained, and the perceptual distortion-free critical point corresponding to the grayscale image of the i-th high-definition image is recorded as J i , Among them, the symbol The symbol for the rounding up operation.
[0070] Step 6_7: Calculate the KLT coefficient energy of each dimension of the KLT spectral component in the KLT coefficient matrix of the grayscale image of each high-definition image, and convert Q' i q' i,k The KLT coefficient energy is denoted as U i,k , Then calculate the normalized KLT coefficient energy of each dimension of the KLT spectral component in the KLT coefficient matrix of the grayscale image of each HD image, and convert Q' i q' i,k The normalized KLT coefficient energy is recorded as Then calculate the cumulative normalized KLT coefficient energy at the perceptual distortion-free critical point corresponding to the grayscale image of each HD image, and convert the perceptual distortion-free critical point J corresponding to the grayscale image of the i-th HD image into i The cumulative normalized KLT coefficient energy at Finally, the accumulated normalized KLT coefficient energy at the perceptual distortion-free critical point corresponding to the grayscale images of all high-definition images is formed into a vector, denoted as U cum , Among them, q' i,k (n') means q'i,k The value of the n'th element in U, 1≤ζ≤K, i,ζ Indicates Q' i The ζ-th KLT spectral component q' in i,ζ The KLT coefficient energy, Indicates Q' i q' i,1 The normalized KLT coefficient energy of Indicates Q' i q' i,2 The normalized KLT coefficient energy of Indicates Q' i In The normalized KLT coefficient energy of Indicates Q' i The J i dimensional KLT spectral component, U cum The dimension of is 1×S, The perceptual distortion-free critical point J corresponding to the grayscale image of the first high-definition image is represented by 1 The accumulated normalized KLT coefficient energy at , The perceptual distortion-free critical point J corresponding to the grayscale image of the second high-definition image is represented by 2 The accumulated normalized KLT coefficient energy at , The perceptual distortion-free critical point J corresponding to the grayscale image of the Sth high-definition image is represented by S The accumulated normalized KLT coefficient energy at ;
[0071] Step 6_8: Calculate U cum The mean and standard deviation of all the cumulative normalized KLT coefficient energies in are recorded as and Then according to and The calculation model of the perceived distortion-free critical point is obtained, which is described as: Where L represents I Y The perceptual distortion-free critical point.
[0072] Step 7: Perform the inverse operation of the vectorization process in step 2, Each dimension of the perceptually distortion-free coefficient vector in is converted into a vector of size The image block will The converted image block is taken as the nth image block; then the image blocks are segmented in the order of the image blocks in step 2. The image blocks converted from all perceptually distortion-free coefficient vectors in are spliced into an image as the perceptually distortion-free critical image, denoted as I L ; Then according to I Y and I L, calculate the just perceptible distortion threshold map, denoted as M, and denote the pixel value of the pixel point with coordinate position (a, b) in M as M(a, b), M(a, b) = |I Y (a,b)-I L (a,b)|; Among them, 1≤a≤W, 1≤b≤H, I Y (a,b) represents I Y The pixel value of the pixel point with coordinate position (a, b) in the middle, I L (a,b) represents I L The pixel value of the pixel point with coordinate position (a, b). The symbol “||” is the absolute value symbol.
[0073] In order to further illustrate the feasibility and effectiveness of the method of the present invention, experiments were conducted on the method of the present invention.
[0074] The first part of the experiment is the noisy image generation experiment guided by the just perceptible distortion threshold estimation model. A source image is input, and the corresponding just perceptible distortion threshold map is obtained using the method of the present invention. The red channel, green channel, and blue channel of the source image are denoised using the just perceptible distortion threshold map. The noisy image corresponding to the red channel of the source image is recorded as Will The pixel value of the pixel point with coordinate position (a, b) is recorded as The noisy image corresponding to the green channel of the source image is recorded as Will The pixel value of the pixel point with coordinate position (a, b) is recorded as The noisy image corresponding to the blue channel of the source image is recorded as Will The pixel value of the pixel point with coordinate position (a, b) is recorded as Finally, the noisy image of the source image can be obtained; where 1≤a≤W, 1≤b≤H, I R (a,b) represents the pixel value of the pixel with coordinate position (a,b) in the red channel of the source image, M(a,b) represents the pixel value of the pixel with coordinate position (a,b) in the just perceptible distortion threshold map, N(a,b) represents the value of the element with subscript (a,b) in a random binary matrix with dimension W×H, N(a,b) is +1 or -1, θ is the adjustment parameter of the noise size, and the amount of noise injected can be controlled by changing the value of θ, I G (a,b) represents the pixel value of the pixel with coordinates (a,b) in the green channel of the source image. B (a,b) represents the pixel value of the pixel with coordinates (a,b) in the blue channel of the source image.
[0075] Here, S=500, D=60, K=64 are set, and the PSNR of the noised image of the source image is controlled to be around 26dB. PSNR, or peak signal-to-noise ratio, is an existing image quality evaluation index. 20 source images are selected for the experiment. A comparative study was conducted using six existing JND threshold map generation methods, which are: the first method, Yang 2005 (X. Yang, W. Lin, Z. Lu, E. Ong, and S. Yao, "Motion-compensated residue pre-processing in video coding based on just-noticeable-distortion profile," IEEE Transactions on Circuits and Systems for Video Technology, vol. 15, no. 6, pp. 742-752, 2005. (Motion-compensated residue pre-processing in video coding based on just-noticeable-distortion profile)); the second method, Zhang 2005 (X. Zhang, W. Lin, and P. Xue, "Improved estimation for just-noticeable visual distortion," Signal Processing, vol. 15, no. 6, pp. 742-752, 2005.) Processing, vol.85, no.4, pp.795-808, 2005. (Improved estimation of just noticeable visual distortion)); the third one, Wu 2013 (J.Wu, G.Shi, W.Lin, A.Liu, and F.Qi, "Just noticeable difference estimation for images with free-energy principle," IEEE Transactions on Multimedia, vol.15, no.7, pp.1705-1710, 2013. (Just noticeable distortion estimation for images based on the free energy principle)); the fourth one, Wu 2017 (J.Wu, L.Li, W.Dong, G.Shi, W.Lin, and C.-CJKuo, "Enhanced just noticeable difference model for images with pattern complexity," IEEE Transactions on Image Processing, vol.26, no.6, pp.2682-2693, 2017. (Just perceptible distortion model for image enhancement based on pattern complexity)); the fifth, Jakhetiya 2018 (V.Jakhetiya, W. Lin, S. Jaiswal, K. Gu, and SCG Guntuku, "Just noticeable difference for natural images using rms contrast and feed-back mechanism," Neurocomputing, vol. 275, pp. 366-376, 2018. (Just noticeable distortion of natural images based on rms contrast and feedback mechanism)); the sixth type, Chen 2020 (Z. Chen and W. Wu, "Asymmetric foveated just-noticeable-difference model for images with visual field inhomogeneities," IEEE Transactions on Circuits and Systems for Video Technology, vol. 30, no. 11, pp. 4064-4074, 2020. (Asymmetric foveated just-noticeable-difference model for images with visual field inhomogeneities)). For each source image, the just noticeable distortion threshold map is obtained using the method of the present invention and the existing six JND threshold map generation methods; then the noisy image of each source image is obtained, that is, for each source image, seven noisy images with approximately the same peak signal-to-noise ratio can be obtained.
[0076] Table 1 Experimental comparison of noisy images obtained by using the method of the present invention and six existing JND threshold map generation methods
[0077]
[0078]
[0079] 30 volunteers were recruited to conduct a subjective experiment. Each volunteer was asked to compare each source image with its corresponding 7 noisy images, and to rate the noisy images between 0 and -1. If the score is 0, it means that the visual quality of the noisy image is very close to that of the source image; if the score is -1, it means that the visual quality of the noisy image is very poor compared to the source image. Each noisy image received 30 scores, and the subjective score of the noisy image was obtained by taking the average value after removing outliers, which was recorded as the MOS value. The higher the MOS value, the better the visual quality of the noisy image. In addition, the existing image quality objective evaluation algorithm MS-SSIM (Z. Wang, EPSimoncelli, and ACBovik, "Multiscale structural similarity for image quality assessment," in The Thrity-Seventh Asilomar Conference on Signals, Systems Computers, 2003, vol. 2, 2003, pp. 1398-1402 Vol. 2. (Multiscale structural similarity for image quality evaluation)) is used to evaluate the quality of the noisy image. The higher the MS-SSIM value, the better the quality of the noisy image. The MOS value and MS-SSIM value of each noisy image are listed in Table 1.
[0080] In Table 1, I01 represents the first source image. There are 20 source images involved in the experiment, so they are marked as I01 to I20 in sequence. Figure 3a Given a source image I03, Figure 3b The method of the present invention is given Figure 3a The just perceptible distortion threshold map is obtained by processing the source image shown in Figure 3c Given the use Figure 3b Guided generated noisy image under PSNR=26dB condition, Figure 3d Given Figure 3c The enlarged view of the part in the middle frame, from Figure 3dIt can be seen that the noise is almost invisible, and the visual quality is good. As shown in Table 1, the noisy image generated by the just perceptible distortion threshold map obtained by the method of the present invention exceeds the noisy image generated by the just perceptible distortion threshold map obtained by other methods in terms of both the MS-SSIM value of a single noisy image and the MS-SSIM average value. On the noisy image corresponding to the source image I12, the MOS value of the noisy image generated by the just perceptible distortion threshold map obtained by the method of the present invention is only lower than that of the noisy image generated by the just perceptible distortion threshold map obtained by the method of Chen2020, but higher than that of the noisy image generated by the just perceptible distortion threshold map obtained by other methods. And on the noisy images corresponding to other source images, the MOS value and MOS average value of the noisy image generated by the just perceptible distortion threshold map obtained by the method of the present invention are higher than those of other methods. This experiment shows that, under the premise of keeping the same amount of noise, the noisy image generated by the just perceptible distortion threshold map obtained by the method of the present invention has a better visual effect and can inject noise guidance into places that are not easily perceived by vision. Therefore, the method of the present invention can more accurately reflect the visual masking characteristics of HVS and characterize its visual perception redundancy.
[0081] The second part of the experiment is the JPEG compression experiment guided by the just noticeable distortion threshold map.
[0082] According to the requirements of the JPEG compression model, the red channel, green channel, blue channel of the source image and the corresponding just perceptible distortion threshold map obtained by the method of the present invention are divided into image blocks in the same order, and the red channel, green channel, blue channel of the source image and the corresponding just perceptible distortion threshold map obtained by the method of the present invention are divided into non-overlapping image blocks of size 8×8.
[0083] Taking image blocks as units, smoothing is performed on the red channel, green channel and blue channel of the source image respectively, guided by the just noticeable distortion threshold map.
[0084] Taking the image block as the unit, the calculation formula for smoothing the j-th image block in the red channel of the source image guided by the j-th image block in the just perceptible distortion threshold map is: Taking the image block as the unit, the calculation formula for smoothing the j-th image block in the green channel of the source image guided by the j-th image block in the just perceptible distortion threshold map is: Taking the image block as the unit, the calculation formula for smoothing the j-th image block in the blue channel of the source image guided by the j-th image block in the just perceptible distortion threshold map is: in, Represents the pixel value of the pixel with coordinate position (a', b') in the image block obtained after smoothing of the j-th image block in the red channel of the source image, 1≤a'≤8,1≤b'≤8, I R,j (a', b') represents the pixel value of the pixel with coordinates (a', b') in the jth image block in the red channel of the source image, M j (a',b') represents the pixel value of the pixel with coordinate position (a',b') in the jth image block in the just perceptible distortion threshold map, Represents the average pixel value of all pixels in the jth image block in the red channel of the source image, represents the pixel value of the pixel with coordinate position (a', b') in the image block obtained after smoothing of the jth image block in the green channel of the source image, I G,j (a',b') represents the pixel value of the pixel with coordinate position (a',b') in the jth image block in the green channel of the source image. Represents the average pixel value of all pixels in the jth image block in the green channel of the source image, represents the pixel value of the pixel with coordinate position (a', b') in the image block obtained after smoothing of the j-th image block in the blue channel of the source image, I B,j (a',b') represents the pixel value of the pixel with coordinates (a',b') in the jth image block in the blue channel of the source image. Represents the average pixel value of all pixels in the jth image block in the blue channel of the source image. After the smoothing process is completed, the image blocks are reassembled in the order of image block segmentation to obtain the smoothed red channel, green channel, and blue channel, and finally a smoothed image is obtained.
[0085] Set the quality factor (encoding quantization parameter) QP = 1 and perform JPEG compression. Perform JPEG compression on the input source image to obtain the JPEG compressed image of the source image. Calculate the compression bit rate and PSNR value of the JPEG compressed image of the source image, which is recorded as Bitrate 1 and PSNR 1 ; Perform JPEG compression on the smoothed image of the source image to obtain a JPEG compressed image of the smoothed image of the source image, and calculate the compression bit rate and PSNR value of the JPEG compressed image of the smoothed image of the source image, which are recorded as Bitrate 2 and PSNR 2. Calculate the bit rate saving percentage and PSNR loss percentage, which are recorded as ΔBitrate and ΔPSNR respectively. Then calculate the gain, denoted as Gain, The larger the ΔBitrate is, the more bitrate is saved, and the smaller the ΔPSNR is, the less PSNR is lost. If the gain is larger, it means the least possible PSNR loss in exchange for the greatest possible compression bitrate savings, which means the image compression performance guided by the just perceptible distortion threshold map is better.
[0086] Table 2 Experimental comparison of JPEG compression guided by the just perceptible distortion threshold map obtained by the method of the present invention and JPEG compression guided by the just perceptible distortion threshold map obtained by the method of Wu2017
[0087]
[0088]
[0089] Here, we set S = 500, D = 60, and K = 64. We selected the 20 source images used in the first part of the experiment for the experiment, and selected the Wu2017 method as the comparison method. Figure 4a The result of directly using JPEG compression for the source image I01 is given. Figure 4b The JPEG compression result guided by the just perceptible distortion threshold map obtained by the method of the present invention for the source image I01 is given. Figure 4c The JPEG compression results guided by the just perceptible distortion threshold map of the source image I01 obtained by the Wu2017 method are given. Figure 4a , Figure 4b , Figure 4c The right frame in each example is a partial enlargement of the left frame. Figure 4a , Figure 4b , Figure 4cIn the right box, it can be seen that the visual effect of the JPEG compression result guided by the just perceptible distortion threshold map obtained by the method of the present invention for the source image I01 is similar to the result of the source image I01 directly compressed by JPEG, and its visual quality is almost not reduced. However, the JPEG compression result guided by the just perceptible distortion threshold map obtained by the method of Wu2017 for the source image I01 is very obviously blurred, and its visual quality is much different from the result of the source image I01 directly compressed by JPEG. The experimental data for each source image are shown in Table 2. It can be seen that although the JPEG compression guided by the just perceptible distortion threshold map obtained by the method of the present invention is lower than that of the Wu2017 method in terms of bit rate savings, it is less than that of the Wu2017 method in terms of PSNR loss. In addition, except for the I11 source image, the JPEG compression gain Gain guided by the just perceptible distortion threshold map obtained by the method of the present invention is slightly weaker than that of the Wu2017 method, and the gain Gain and the average value of the gain Gain on all other source images are stronger than that of the Wu2017 method. Furthermore, the quality factor QP value is changed, and the average gain of the method of the present invention and the method of Wu2017 on 20 source images is tested, and a line graph is drawn, as shown in FIG. Figure 5 As shown in the figure, it can be seen that at each QP value, the average gain of the method of the present invention is higher than that of the Wu2017 method. This experiment shows that the JPEG compression guided by the just perceptible distortion threshold map obtained by the method of the present invention can ensure the visual quality of the compressed image as much as possible while reducing the compression bit rate as much as possible. Therefore, the method of the present invention can more accurately reflect the visual masking characteristics of HVS and characterize its visual perception redundancy.
Claims
1. A top-down approach to estimating the threshold of just perceptible distortion in natural images. Features The following steps are involved: Step 1: Take a natural image to be processed as a source image; Then the source image is converted into a grayscale image, denoted as I Y ; Among them, the source image is an RGB color image, the source image and I Y The width is W and the height is H; Step 2: I Y Divide into Num non-overlapping sizes Then, I Y Each image block in is vectorized to obtain I Y The column vector corresponding to each image block in I Y The column vector corresponding to the nth image block in is denoted as x n ; Then I Y The column vectors corresponding to all image blocks in are concatenated to form a vectorized matrix, denoted as X, X = [x 1 ,x 2 ,…,x n ,…,x Num ];in, Set W and H to be Divisible, the value of K is 4 2 or 5 2 or 6 2 or 7 2 or 8 2 or 9 2 or 10 2 , 1≤n≤Num,x 1 I Y The column vector corresponding to the first image block in 2 I Y The column vector corresponding to the second image block in Num I Y The column vector corresponding to the Numth image block in x 1 、x 2 、x n 、x Num The dimensions of are all K×1, the dimensions of X are K×Num, and the symbol "[]" represents a vector or matrix. Step 3: Calculate the covariance matrix of X, denoted as C; then use the eigenvalue decomposition technique to process C to obtain K eigenvalues and corresponding K eigenvectors of C; then sort the K eigenvectors of C in descending order according to the corresponding K eigenvalues, and use the matrix composed of the K eigenvectors of C according to the sorting results as the prior information extracted from X; where the dimension of C is K×K, the eigenvector is a column vector, the dimension of the eigenvector is K×1, each column of the prior information extracted from X is an eigenvector of C, and the dimension of the prior information extracted from X is K×K; Step 4: Extract the prior information from X as I Y The KLT kernel is denoted as P; then based on P and X, calculate I Y The KLT coefficient matrix is denoted as Q, Q = (P) T X; then Q is expressed as Q = [q 1 ,q 2 ,…,q k ,…,q K ] T ; Among them, the dimension of P is K×K, the dimension of Q is K×Num, 1≤k≤K, q 1 represents the first-dimensional KLT spectral component in Q, q 2 represents the second-dimensional KLT spectral component in Q, q k represents the k-th dimension KLT spectral component in Q, q K represents the K-th dimension KLT spectral component in Q, q 1 ,q 2 ,q k ,q K The dimensions of are all Num×1; Step 5: Calculate the KLT coefficient energy of each dimension of the KLT spectrum components in Q, and denote the KLT coefficient energy of q k as E k , Then calculate the normalized KLT coefficient energy of each dimension of the KLT spectrum components in Q, and denote the normalized KLT coefficient energy of q k as Next, calculate the cumulative normalized KLT coefficient energy of each dimension of the KLT spectrum components in Q, and denote the cumulative normalized KLT coefficient energy of q k as Finally, form the cumulative normalized KLT coefficient energy vector from the cumulative normalized KLT coefficient energies of all the KLT spectrum components in Q, and denote it as E cum , where q k (n) represents the value of the nth element in q k , 1 ≤ ζ ≤ K, E ζ represents the KLT coefficient energy of the ζth dimension KLT spectrum component q ζ in Q, represents the normalized KLT coefficient energy of q 1 , represents the normalized KLT coefficient energy of q 2 , E cum has a dimension of 1 × K, represents the cumulative normalized KLT coefficient energy of q 1 , represents the cumulative normalized KLT coefficient energy of q 2 , represents the cumulative normalized KLT coefficient energy of q K ; Step 6: E cum Substituted as input into the perceptual distortion-free critical point calculation model, the calculation results are I Y The perceptual distortion-free critical point is recorded as L; then the perceptual distortion-free coefficient reconstruction matrix is constructed based on L, recorded as Then adopt The reconstructed perceptually distortion-free coefficient matrix is recorded as Then Expressed as Where L is a positive integer, 1≤L≤K, The dimension is K×Num, The "=" in q is the assignment symbol. L represents the L-th dimension KLT spectral component in Q, q L The dimension is Num×1, to All are zero vectors, The dimensions of are all Num×1, The dimension is K×Num, express The first dimension of the perceptually distortion-free coefficient vector in, express The second dimension perceptually distortion-free coefficient vector in, express The n-th dimension perceptually distortion-free coefficient vector in, express The Num-th dimension perceptually distortion-free coefficient vector in, The dimensions of are all K×1; Step 7: Perform the inverse operation of the vectorization process in step 2, Each dimension of the perceptually distortion-free coefficient vector in is converted into a vector of size The image block will The converted image block is taken as the nth image block; then The image blocks converted from all perceptually distortion-free coefficient vectors in are spliced into an image as the perceptually distortion-free critical image, denoted as I L ; Then according to I Y and I L , calculate the just perceptible distortion threshold map, denoted as M, and denote the pixel value of the pixel point with coordinate position (a, b) in M as M(a, b), M(a, b) = |I Y (a,b)-I L (a,b)|; Among them, 1≤a≤W, 1≤b≤H, I Y (a,b) represents I Y The pixel value of the pixel point with coordinate position (a, b) in the middle, I L (a,b) represents I L The pixel value of the pixel point with coordinate position (a, b). The symbol "| |" is the absolute value symbol.
2. The top-down natural image just perceptible distortion threshold estimation method according to claim 1, Features In step 2, x n The acquisition process is: I Y The pixel values of all pixels in the nth image block are arranged in a column to form x n .
3. A top-down natural image just perceptible distortion threshold estimation method according to claim 1 or 2, Features In step 3, the calculation formula of C is: The superscript "T" indicates the transpose of a vector or matrix. represents the mean vector obtained by taking the mean of X by row, The dimension of is K×1.
4. The top-down natural image just perceptible distortion threshold estimation method according to claim 3, Features In step 4, P = [p 1 ,p 2 ,…,p K ], where p 1 represents the first eigenvector after sorting the K eigenvectors of C in descending order according to the corresponding K eigenvalues, p 2 It represents the second eigenvector after sorting the K eigenvectors of C in descending order according to the corresponding K eigenvalues, p K represents the Kth eigenvector of C after sorting the K eigenvectors of C in descending order according to the corresponding K eigenvalues, p 1 、p 2 、p K The dimensions of are all K×1.
5. The top-down natural image just perceptible distortion threshold estimation method according to claim 4, Features In step 6, the process of obtaining the perceptual distortion-free critical point calculation model is as follows: Step 6_1: Select S high-definition images and convert each high-definition image into a grayscale image; then follow the process from step 2 to step 4 to obtain the KLT coefficient matrix of the grayscale image of each high-definition image in the same way, and record the KLT coefficient matrix of the grayscale image of the i-th high-definition image as Q' i , Q' i Expressed as Q' i =[q' i,1 ,q' i,2 ,…,q' i,k ,…,q' i,K ] T ; Where S ≥ 100, the HD image is an RGB color image, the width of the HD image is W' and the height is H', W' and H' can be Divisible by, the size of the image block of the grayscale image segmentation of the high-definition image is 1≤i≤S,Q' i The dimension is K×Num', where Num' represents the total number of image blocks of the grayscale image segmentation of the high-definition image. q' i,1 Indicates Q' i The first dimension KLT spectral component, q' i,2 Indicates Q' i The second-dimensional KLT spectral component in , q' i,k Indicates Q' i The k-th KLT spectral component in , q' i,K Indicates Q' i The K-th dimension KLT spectral component, q' i,1 ,q' i,2 ,q' i,k ,q' i,K The dimensions of are all Num'×1; Step 6_2: Construct K perceptual coefficient reconstruction matrices corresponding to the grayscale image of each high-definition image, and record the kth perceptual coefficient reconstruction matrix corresponding to the grayscale image of the i-th high-definition image as Then, the K perception coefficient matrices corresponding to the grayscale image of each high-definition image are reconstructed, and the kth perception coefficient matrix corresponding to the grayscale image of the i-th high-definition image is recorded as Will Expressed as in, The dimension is K×Num', The "=" in is the assignment symbol. to All are zero vectors, The dimensions of are all Num'×1, The dimension is K×Num', P' i represents the KLT kernel of the grayscale image of the i-th HD image, P' i The dimension is K×K, 1≤n'≤Num', n' is a positive integer, express The first dimension perceptual coefficient vector in, express The second dimension perceptual coefficient vector in, express The n'th dimension perceptual coefficient vector in, express The Num'th dimension perceptual coefficient vector in, The dimensions of are all K×1; Step 6_3: According to the inverse operation of the vectorization process in step 2, each dimension of the perceptual coefficient vector in each perceptual coefficient matrix corresponding to the grayscale image of each high-definition image is converted into a perceptual coefficient vector of size Then, all the image blocks converted from the perceptual coefficient vectors in each perceptual coefficient matrix corresponding to the grayscale image of each high-definition image are stitched into one image, and the kth perceptual coefficient matrix corresponding to the grayscale image of the i-th high-definition image is The image formed by stitching together the image blocks converted from all the perceptual coefficient vectors in is used as the kth reconstructed image corresponding to the grayscale image of the i-th high-definition image, denoted as in, The width is W' and the height is H'; Step 6_4: Call D volunteers, each volunteer compares the grayscale image of each high-definition image with its corresponding reconstructed images in turn by naked eye observation, each volunteer determines a reconstructed image from the K reconstructed images corresponding to the grayscale image of each high-definition image as the perceptual distortion-free critical image corresponding to the grayscale image, and uses the serial number of the determined reconstructed image as the perceptual distortion-free critical point corresponding to the grayscale image; for the d-th volunteer and the grayscale image of the ith high-definition image, the d-th volunteer compares the grayscale image of the ith high-definition image with its corresponding 1st reconstructed image, 2nd reconstructed image, ..., Kth reconstructed image in turn by naked eye observation, and stops the comparison process once the d-th volunteer cannot distinguish the grayscale image of the ith high-definition image from one of its corresponding reconstructed images, and assumes that the reconstructed image is the kth reconstructed image corresponding to the grayscale image of the ith high-definition image. Then As the perceptual distortion-free critical image corresponding to the grayscale image of the i-th HD image observed by the d-th volunteer, the value k is taken as the perceptual distortion-free critical point corresponding to the grayscale image of the i-th HD image observed by the d-th volunteer, denoted as Then the vector consisting of the perceptual distortion-free critical points corresponding to the grayscale images of the i-th high-definition image observed by all volunteers is recorded as J i , Where D>1, 1≤d≤D, The "=" in J is an assignment symbol. i The dimension of is 1×D, represents the perceptual distortion-free critical point corresponding to the grayscale image of the i-th HD image observed by the first volunteer, represents the perceptual distortion-free critical point corresponding to the grayscale image of the i-th high-definition image observed by the second volunteer, represents the perceptual distortion-free critical point corresponding to the grayscale image of the i-th high-definition image observed by the D-th volunteer; Step 6_5: Calculate the mean and standard deviation of all perceptual distortion-free critical points in the vector composed of the perceptual distortion-free critical points corresponding to the grayscale images of each high-definition image observed by all volunteers, and convert J i The mean and standard deviation of all perceptually distortion-free critical points in are recorded as and Then, outliers are removed from the vector of perceptual distortion-free critical points corresponding to the grayscale images of each HD image observed by all volunteers. i ,like Dissatisfied Then judge is an outlier, From J i The vector obtained after removing outliers is recorded as Step 6_6: Calculate the vector of perceptual distortion-free critical points corresponding to the grayscale image of each high-definition image observed by all volunteers, and the mean of all perceptual distortion-free critical points after removing outliers. The mean of all perceptually distortion-free critical points in is denoted as Then, the perceptual distortion-free critical point corresponding to the grayscale image of each high-definition image is obtained, and the perceptual distortion-free critical point corresponding to the grayscale image of the i-th high-definition image is recorded as J i , Among them, the symbol is the symbol for the rounding up operation; Step 6_7: Calculate the KLT coefficient energy of each dimension of the KLT spectral component in the KLT coefficient matrix of the grayscale image of each high-definition image, and convert Q' i q' i,k The KLT coefficient energy is denoted as U i,k , Then calculate the normalized KLT coefficient energy of each dimension of the KLT spectral component in the KLT coefficient matrix of the grayscale image of each HD image, and convert Q' i q' i,k The normalized KLT coefficient energy is recorded as Then calculate the cumulative normalized KLT coefficient energy at the perceptual distortion-free critical point corresponding to the grayscale image of each HD image, and convert the perceptual distortion-free critical point J corresponding to the grayscale image of the i-th HD image into i The cumulative normalized KLT coefficient energy at Finally, the accumulated normalized KLT coefficient energy at the perceptual distortion-free critical point corresponding to the grayscale images of all high-definition images is formed into a vector, denoted as U cum , Among them, q' i,k (n') means q' i,k The value of the n'th element in U, 1≤ζ≤K, i,ζ Indicates Q' i The ζ-th KLT spectral component q' in i,ζ The KLT coefficient energy, Indicates Q' i q' i,1 The normalized KLT coefficient energy of Indicates Q' i q' i,2 The normalized KLT coefficient energy of Indicates Q' i In The normalized KLT coefficient energy of Indicates Q' i The J i dimensional KLT spectral component, U cum The dimension of is 1×S, The perceptual distortion-free critical point J corresponding to the grayscale image of the first high-definition image is represented by 1 The accumulated normalized KLT coefficient energy at , The perceptual distortion-free critical point J corresponding to the grayscale image of the second high-definition image is represented by 2 The accumulated normalized KLT coefficient energy at , The perceptual distortion-free critical point J corresponding to the grayscale image of the Sth high-definition image is represented by S The accumulated normalized KLT coefficient energy at ; Step 6_8: Calculate U cum The mean and standard deviation of all the cumulative normalized KLT coefficient energies in are recorded as and Then according to and The calculation model of the perceived distortion-free critical point is obtained, which is described as: Where L represents I Y The perceptual distortion-free critical point.
6. The top-down natural image just perceptible distortion threshold estimation method according to claim 5, Features In the step 6_1, Q' i The acquisition process is: Step 6_1a: The grayscale image of the i-th high-definition image is recorded as I' Y,i ; Then I' Y,i Divide into Num' non-overlapping sizes image block; then I' Y,i Each image block in is vectorized to obtain I' Y,i The column vector corresponding to each image block in I' Y,i The column vector corresponding to the n'th image block in is denoted as x' i,n' ; Then I' Y,i The column vectors corresponding to all image blocks in are concatenated to form a vectorized matrix, denoted as X' i , X' i =[x' i,1 ,x' i,2 ,…,x' i,n' ,…,x' i,Num' ];in, The value of K is 4 2 or 5 2 or 6 2 or 7 2 or 8 2 or 9 2 or 10 2 , 1≤n'≤Num', x' i,1 Indicates I' Y,i The column vector corresponding to the first image block in, x' i,2 Indicates I' Y,i The column vector corresponding to the second image block in, x' i,Num' Indicates I' Y,i The column vector corresponding to the Num'th image block in x' i,1 、x' i,2 、x' i,n' 、x' i,Num' The dimensions of X' are K×1. i The dimension is K×Num'; Step 6_1b: Calculate X' i The covariance matrix of i ; Then use the eigenvalue decomposition technique to calculate C' i Process and obtain C' i The K eigenvalues and corresponding K eigenvectors of C' i Sort the K eigenvectors of C' in descending order according to the corresponding K eigenvalues. i The matrix composed of the K eigenvectors of X' is constructed by sorting them. i The prior information extracted from i The dimension of is K×K, the eigenvector is a column vector, and the dimension of the eigenvector is K×1. i Each column of the prior information extracted is C' i 1 eigenvector of X' i The dimension of the prior information extracted is K×K; Step 6_1c: Will be from X' i The prior information extracted from Y,i The KLT kernel is denoted as P' i ; Then according to P' i and X' i , calculate I' Y,i The KLT coefficient matrix Q' i , Q' i =(P' i ) T X' i ; Among them, P' i The dimension is K×K, Q' i The dimension is K×Num'.
Citation Information
Patent Citations
Natural image perceptible distortion threshold estimation method based on sparse representation
CN109872302A
KR20190062284A