No-Reference Screen Content Image Quality Assessment Method Based on Global and High-Impact Region Analysis
Through the overall and high-impact area analysis methods, combined with phase consistency, local binary mode and gradient characteristics, the problem of inaccurate screen content image quality evaluation in the prior art is solved, and a more refined image quality evaluation is achieved.
Patent Information
- Application Number
- CN202211471737.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-23
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-11-23
AI Technical Summary
The existing image quality evaluation method of screen content without reference cannot accurately reflect the overall quality of the image, and ignores direction information and regional differences, resulting in inaccurate evaluation results.
Using a method based on overall and high-impact area analysis, the text and image areas are divided by local image activity metric algorithm, and the information entropy is used to select high-impact areas, combining phase consistency, local binary mode and gradient features to extract structural features, calculate color features based on opposite color spaces, and using AdaBoosting BP neural network to train the regression model to obtain the scores of the overall and high-impact areas, and finally adjust the overall score through a weighted strategy.
It improves the accuracy of image quality evaluation, takes into account the sensitivity of human visual system to structural information, and can more accurately capture the structural degradation of screen content images, providing a more refined quality evaluation.
Smart Images

Figure CN115937108B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image quality analysis, and specifically to a no-reference screen content image quality evaluation method based on holistic and high-impact region analysis. Background Art
[0002] In recent years, with the improvement of network transmission performance, various multimedia applications and broadcast services have been widely developed. In many scenarios such as distance education, telecommuting, screen sharing, and online advertising, users need to interact with the local display interface, so end-users expect higher-quality screen content images.
[0003] However, due to factors such as device performance, information loss during storage or transmission, the captured screen content images may suffer from various types of distortions, namely contrast changes, JPEG compression, etc. Undoubtedly, the degradation of image quality will have a direct impact on the user experience of consumers.
[0004] In the past few decades, research on image quality evaluation has mainly focused on natural images. However, it has been found that the average de-contrast normalization coefficient extracted from natural images follows the generalized Gaussian distribution well, while the distribution curve of the average de-contrast normalization coefficient extracted from screen content images fluctuates greatly, and some distortion types do not affect the statistical distribution. Therefore, the no-reference quality evaluation methods designed for natural images cannot be directly applied to screen content images.
[0005] Screen content images are mainly composed of text regions and image regions, and the human eye has different perceptions of different regions of the image. Therefore, it is crucial to consider the visual impact of different regions on screen content images. However, most no-reference screen content image quality evaluation methods focus only on feature extraction of the entire image. In addition, in segmentation-based methods, the quality scores of text and image regions are usually predicted separately, and then the scores are fused using an adaptive weighting strategy to obtain the final quality score. But this fusion method may cause the object to lose integrity, making the final result inaccurate. Secondly, when measuring the image structure in existing no-reference screen content image quality evaluation methods, only spatial intensity and distribution are considered, ignoring the direction information. Therefore, the measurement of screen content image structure information can still be further optimized.
[0006] In summary, how to design an effective screen content image quality evaluation method to improve the quality of experience and optimize the multimedia processing system has become an urgent technical problem to be solved. Summary of the Invention
[0007] The objective of the present invention is to solve the defect that the fractional fusion method and the extracted statistical features in the prior art cannot comprehensively reflect the image quality, and to provide a no-reference screen content image quality evaluation method based on global and high-impact region analysis to solve the above problems.
[0008] To achieve the above objective, the technical solution of the present invention is as follows:
[0009] A no-reference screen content image quality evaluation method based on global and high-impact region analysis, comprising the following steps:
[0010] Partitioning of the screen content image: First, obtain the rough text layer of the screen content image through the local image activity metric algorithm; further segment the pure text region from the rough text layer through the refinement process based on text connection components, and the remaining regions are identified as image regions;
[0011] Partitioning of the high-impact region: Select the high-impact region with a higher impact on the overall quality between the pure text and image regions through information entropy;
[0012] Extraction of structural features: Combine phase congruency, local binary pattern, and multiple gradient features to obtain a gradient-weighted histogram, and extract the same structural features from the screen content image and the high-impact region;
[0013] Extraction of color features: Calculate saturation and color entropy based on the opponent color space, and extract the same color features from the screen content image and the high-impact region;
[0014] Obtaining the global image score and the high-impact region score: Use the AdaBoosting BP neural network to train a regression model to obtain the global image score and the high-impact region score;
[0015] Obtaining the final visual quality score: Locally adjust the global score with the score of the high-impact region through a weighting strategy to obtain the final visual quality score.
[0016] The partitioning of the high-impact region includes the following steps:
[0017] Calculate the information entropy of the pure text region and the image region respectively, and the information entropy formula is expressed as:
[0018]
[0019] where p(h) is the probability of pixel intensity h,
[0020] Obtain the value p of the information entropy of the pure text region tex and the value p of the information entropy of the image region pic ;
[0021] Compare the information entropy values of the pure text region and the image region:
[0022] When p tex ≥ p pic , the text region is identified as a high-impact region; when p tex < p pic , the image region is identified as a high-impact region.
[0023] The extraction of the structural features includes the following steps:
[0024] For the screen content image and the high-impact region, input the grayscale image I gray (i, j), and perform wavelet transform on it using a log-Gabor filter.
[0025] The log-Gabor filter uses a Gaussian function as the expansion function and extends the filter to two dimensions. The two-dimensional log-Gabor filter extended by the Gaussian function has the following transfer function:
[0026]
[0027] where θ o = oπ / 0, o = {0, 1,..., O - 1} is the direction angle of the filter, O is the number of directions, and σ θ determines the angular bandwidth of the filter, ω0 represents the center frequency of the filter, and σ r is for controlling the bandwidth of the filter;
[0028] Calculate the amplitude A no (i, j) and the phase angle
[0029] The amplitude A no (i, j) and the phase angle at a given wavelet scale n and direction o are defined by the following formulas:
[0030]
[0031]
[0032] where, and represent the even-symmetric cosine and odd-symmetric sine wavelets of two-dimensional Log-Gabor respectively;
[0033] Calculate the phase congruency PC(i, j).
[0034] The calculation method of PC(i, j) at various scales and directions is as follows:
[0035]
[0036] Among them, ε is a constant introduced to avoid a zero denominator;
[0037] Calculate the local binary pattern of each pixel in the phase congruency map to obtain the spatial distribution PLBP of the structure in the phase congruency domain, and the calculation is as follows:
[0038]
[0039] where PC t is the value of the surrounding pixels, PC c is the value of the central pixel, n is the number of equally spaced adjacent pixels around the central pixel, set to 8, and r is the radius of the neighborhood, set to 1;
[0040] The definitions of the formulas W(·) and θ(·) are as follows:
[0041]
[0042]
[0043] Since n is set to 8, the value range of the LBP mapping is [0, 9], which means there are 10 distribution patterns;
[0044] Calculate the cumulative value of the gradient magnitude, relative gradient magnitude, and gradient direction of the pixels with the same local binary pattern in the phase congruency domain, that is, the PLBP histogram with gradient weighting, and the calculation process is as follows:
[0045]
[0046] where f(·) is the Kronecker function, M and N respectively represent the height and width of the image, and m is the possible pattern of PLBP, with a range from 0 to 9;
[0047] G(i, j) ∈ {GM(i, j), RGM(i, j), GO(i, j)} is the weight of the PLBP of each pixel, and the calculation process of G(i, j) is as follows:
[0048]
[0049]
[0050]
[0051] where and represent the horizontal gradient and vertical gradient, dx′ and dy′ are the local average gradients of dx and dy after being filtered by an averaging filter, and the size of the filter is 3x3;
[0052] 30 structural features and perceptual feature vectors are respectively obtained from the overall image and the high-impact regions of a distorted screen content image m ∈ [0, 9];
[0053] Finally, 120 structural features can be extracted respectively at four scales for each overall image and high-impact region
[0054] The extraction of the color features includes the following steps:
[0055] For the screen content image and the high-impact region, input the color image I RGB (i, j), and extract the color features of the image based on the red-green and blue-yellow channels, where the red-green and blue-yellow channels are represented by O1 and O2 respectively, and the calculation process is as follows:
[0056]
[0057]
[0058] The saturation feature IS is obtained by calculating the global mean of the image saturation:
[0059]
[0060] where is the saturation value at the position (i, j);
[0061] The color richness of the image is measured by calculating the color entropy in the O1 and O2 color channels:
[0062] The mutual influence of the image neighborhood information in the color space is excluded respectively on the O1 and O2 color channels by using the de-mean normalization coefficient:
[0063]
[0064] where, I(i, j) and represent the pixel value and the normalized value at the position (i, j) on the color channel, and u(i, j) and δ(i, j) represent the local region mean and the standard variance;
[0065] Use the mean filter to calculate and the mean of the neighborhood pixels of each pixel point to capture the structural information of the image in different color channels, and the calculation is as follows:
[0066]
[0067] where,
[0068] Denote the mean coefficient of MSCN, and {S(b, v)|b = -B, ..., B; v = -V, ..., V} represents the mean filter, where B and V represent the height and width of the mean filter, which are set to 3;
[0069] Calculate the entropy in different color channels using the information entropy formula and to obtain the color entropy in the O1 channel and the color entropy in the O2 channel
[0070] Obtain 5 color features and perceptual feature vectors from the overall image and high-impact regions of a distorted screen content image respectively Finally, 20 color features can be extracted at four scales for each overall image and high-impact region respectively.
[0071] The obtaining of the overall image score and high-impact region score includes the following steps:
[0072] Set the training set: Connect the structural feature FS and the color feature FC to generate as the feature vector of the final input; Randomly select m groups of data from the sample space where is the input feature vector, Y x is the corresponding subjective score, and the value range of x is from 1 to m;
[0073] Select the BP neural network as the weak classifier, and the final strong classifier is obtained by iteratively training several weak classifiers. Here, the number of weak classifiers is set to 32;
[0074] Estimate its evaluation error E i by integrating the difference between the original quality evaluation score and the predicted quality evaluation score of the i-th weak classifier, and the corresponding distribution;
[0075] Use the convex function to transform the weight of a single weak classifier into the corresponding weight, which is beneficial to ensuring that the lower the weak classifier, the greater the weight, and the lower the weak classifier, the smaller the weight;
[0076] Finally, combine the weight with the predicted output to calculate the overall image score Q ent and the high-impact region score Q hig .
[0077] In the step of obtaining the final visual quality score,
[0078] The specific score adjustment formula is as follows:
[0079] Q = Q ent - c·sign(Q ent-Q hig )·δ
[0080] Wherein, Q is the final quality score of the screen content image, Q ent and Q hig represent the quality scores of the overall and high-impact regions, sign(·) represents the sign function, and δ is the standard variance of Q ent and Q hig . c is a constant used to avoid excessive fluctuations in the scores of the overall image, and is set to 0.5.
[0081] Beneficial effects
[0082] The no-reference screen content image quality evaluation method based on overall and high-impact region analysis of the present invention adopts a score fusion strategy that takes the overall image score as the dominant and uses the scores of high-impact regions to locally adjust the overall score. It not only fully considers the differences between the two regions, but also ensures the integrity of the image. Compared with the existing text and image two-region score fusion strategies, it can more accurately reflect the perceptual quality of the image.
[0083] Considering the fact that the human visual system is very sensitive to structural information, the present invention combines phase consistency, local binary pattern and multiple gradient features that can effectively represent underlying features (structures), and finally obtains a gradient-weighted histogram, fully considering multiple important structural factors such as spatial intensity, orientation, and spatial distribution, and can quantify the fine structure of the image from multiple aspects, so as to more accurately capture the structural degradation of the screen content image. Description of the drawings
[0084] Figure 1 is the method sequence diagram of the present invention;
[0085] Figure 2 is the logical framework diagram of the present invention;
[0086] Figure 3 is the performance comparison diagram of the present invention using different feature combinations in the database SIQAD. Specific implementation manners
[0087] To further understand and recognize the structural features and achieved effects of the present invention, the following is a detailed description with reference to preferred embodiments and drawings:
[0088] The present invention divides the screen content image into text and image regions, selects the regions that have a higher impact on the overall quality through information entropy, and then performs the same feature extraction on the entire image and the high-impact regions. For structural features, we effectively interpret the low-level features of the image through phase congruency, and then introduce local binary pattern and various gradient features to combine with phase congruency, complementing each other to further capture the structural degradation of the image. For color features, we calculate the saturation and color entropy based on the opponent color space that is consistent with the human visual characteristics. Then, we use the AdaBoosting BP neural network to predict the scores of the entire image and the high-impact regions. Finally, through a weighting strategy, we locally adjust the score of the entire image using the quality score of the high-impact regions to obtain the final score.
[0089] As Figure 1 and Figure 2 shown, the no-reference screen content image quality evaluation method based on the analysis of the overall and high-impact regions described in the present invention includes the following steps:
[0090] The first step, the division of the screen content image: First, obtain the rough text layer (including text regions and image regions with high activity) of the screen content image through the local image activity metric algorithm; further segment the pure text regions from the rough text layer through the refinement process based on text connected components, and the remaining regions are identified as image regions.
[0091] The second step, the division of the high-impact regions: Select the high-impact regions that have a higher impact on the overall quality between the pure text and image regions through information entropy.
[0092] The screen content image is mainly composed of text and image regions. Considering the perceptual differences of the human eye for different regions, it is first necessary to find the regions that have a higher impact on the visual quality of the screen content image. Information entropy can reflect the amount of information provided by the image. From the perspective of physical measurement, the greater the information entropy of the image, the more information is transmitted, and the higher the impact on the visual perception of the screen content image. Therefore, we calculate the information entropy of the text and image regions to determine which one is the high-impact region.
[0093] The division of the high-impact regions includes the following steps:
[0094] (1) Calculate the information entropy of the pure text region and the image region respectively. The information entropy formula is expressed as:
[0095]
[0096] where p(h) is the probability of the pixel intensity h,
[0097] obtain the value p of the information entropy of the pure text region tex and the value p of the information entropy of the image regionpic .
[0098] (2) Compare the information entropy values of the pure text region and the image region:
[0099] When p tex ≥ p pic , the text region is identified as a high-impact region; when p tex < p pic , the image region is identified as a high-impact region.
[0100] Step 3: Extraction of structural features: Combine phase congruency, local binary pattern, and multiple gradient features to obtain a gradient-weighted histogram, and extract the same structural features from the screen content image and the high-impact region.
[0101] The screen content image contains a large amount of edge thin lines and texture information, and the human visual system is highly sensitive to such structural information for visual perception and understanding. Based on this fact, first extract the structural features of the screen content image. Phase congruency can reliably detect the structural information of the image. Therefore, we first use phase congruency to capture the low-level structural features of the image and calculate the local binary pattern in the phase congruency domain to obtain the distribution of the underlying features. Considering the contrast invariance of phase congruency, we introduce the gradient magnitude and relative gradient magnitude as supplementary features to further capture the intensity changes of the structure. In addition, the gradient direction is introduced to measure the direction information of the structure, and finally a gradient-weighted histogram is obtained to quantify the fine structure of the screen content image.
[0102] (1) For the screen content image and the high-impact region, input the grayscale image I gray (i, j), and perform wavelet transform on it using a log-Gabor filter.
[0103] The log-Gabor filter uses the Gaussian function as the expansion function and extends the filter to two dimensions. The two-dimensional log-Gabor filter extended by the Gaussian function has the following transfer function:
[0104]
[0105] where θ o = oπ / O, o = {0, 1,..., O - 1} is the direction angle of the filter, O is the number of directions, and σ θ determines the angular bandwidth of the filter, ω0 represents the center frequency of the filter, and σ r is for controlling the bandwidth of the filter.
[0106] (2) Calculate the amplitude A no (i, j) and the phase angle
[0107] The amplitude A at a given wavelet scale n and direction o no (i, j) and the phase angle are defined by the following formula:
[0108]
[0109]
[0110] where and respectively represent the even-symmetric cosine and odd-symmetric sine wavelets of two-dimensional Log-Gabor.
[0111] (3) Calculate the phase congruency PC(i, j),
[0112] The calculation method of PC(i, j) at various scales and directions is as follows:
[0113]
[0114] where ε is a constant introduced to avoid a zero denominator.
[0115] (4) Calculate the local binary pattern of each pixel in the phase congruency map to obtain the spatial distribution PLBP of the structure in the phase congruency domain, and the calculation is as follows:
[0116]
[0117] where PC t is the value of the surrounding pixels, PC c is the value of the central pixel, n is the number of equally spaced adjacent pixels around the central pixel, set to 8, and r is the radius of the neighborhood, set to 1;
[0118] The definitions of the formulas W(·) and θ(·) are as follows:
[0119]
[0120]
[0121] Since n is set to 8, the value range of the LBP mapping is [0, 9], which means there are 10 distribution patterns.
[0122] (5) Calculate the cumulative values of the gradient magnitude, relative gradient magnitude, and gradient direction of the pixels with the same local binary pattern in the phase congruency domain, that is, the PLBP histogram with gradient weighting, and the calculation process is as follows:
[0123]
[0124] Among them, f(·) is the Kronecker function, M and N respectively represent the height and width of the image, and m is the possible pattern of PLBP, with a range from 0 to 9;
[0125] G(i, j) ∈ {GM(i, j), RGM(i, j), GO(i, j)} is the weight of the PLBP of each pixel, and the calculation process of G(i, j) is as follows:
[0126]
[0127]
[0128]
[0129] Among them, and represent the horizontal gradient and the vertical gradient, dx′ and dy′ are the local average gradients of dx and dy after being filtered by an average filter, and the size of the filter is 3x3.
[0130] (6) Obtain 30 structural features and perceptual feature vectors from the overall image and the high-impact regions of a distorted screen content image respectively m ∈ [0, 9];
[0131] Finally, 120 structural features can be extracted at four scales for each overall image and high-impact region respectively.
[0132] Step 4: Extraction of color features: Calculate the saturation and color entropy based on the opponent color space, and extract the same color features for the screen content image and the high-impact regions.
[0133] The screen content image includes discontinuous tone content (text, graphics) and continuous tone content (natural scene images), and it has both large tiled color patches and rich color details. Therefore, the measurement of the image color information is also particularly crucial. Saturation is an attribute of color purity, and it can reflect the color relative to its own brightness. In addition, color entropy can be used as a measure of the color diversity of an image, and the larger the value of color entropy, the more types of colors there are in the image. Therefore, we introduce the opponent color space and calculate the color entropy and chromaticity based on the red-green and blue-yellow channels that are closely related to the image chromaticity, so as to effectively capture the changes in the image color information.
[0134] (1) For the screen content image and the high-impact regions, input the color image I RGB (i, j), extract the color features of the image based on the red-green and blue-yellow channels, and the red-green and blue-yellow channels are represented by O1 and O2 respectively, and the calculation process is as follows:
[0135]
[0136]
[0137] The saturation feature IS is obtained by calculating the global mean of the image saturation:
[0138]
[0139] where is the saturation value at the position (i, j).
[0140] (2) The color richness of the image is measured by calculating the color entropy in the O1 and O2 color channels:
[0141] A1) The mutual influence of the image neighborhood information in the color space is excluded by using the mean-removed normalization coefficient on the O1 and O2 color channels respectively:
[0142]
[0143] where, I(i, j) and represent the pixel value and the normalized value at the position (i, j) on the color channel, and u(i, j) and δ(i, j) represent the local region mean and the standard variance;
[0144] A2) The mean filter is used to calculate the mean of the neighboring pixels of each pixel point in, so as to capture the structural information of the image in different color channels, and the calculation is as follows:
[0145]
[0146] where,
[0147] represents the mean coefficient of MSCN, {S(b, v)|b = -B,..., B; v = -V,..., V} represents the mean filter, and B and V represent the height and width of the mean filter, which are set to 3;
[0148] A3) The entropy formula of information is used to calculate the entropy of and in different color channels, and the color entropy in the O1 channel the color entropy in the O2 channel
[0149] Five color features are obtained from the overall image and the high-impact regions of a distorted screen content image respectively, and the perceptual feature vector Finally, 20 color features can be extracted at four scales for each overall image and high-impact region respectively.
[0150] Step 5. Obtaining the overall image score and the high-impact region score: The overall image score and the high-impact region score are obtained by training a regression model using the AdaBoosting BP neural network.
[0151] (1) Set the training set: Connect the structural feature FS and the color feature FC to generate the feature vector as the final input;
[0152] Randomly select m groups of data from the sample space where is the input feature vector, and Y x is the corresponding subjective score, and the value range of x is from 1 to m.
[0153] (2) Using traditional methods, select the BP neural network as the weak classifier, and the final strong classifier is obtained by iteratively training several weak classifiers. Here, the number of weak classifiers is set to 32.
[0154] (3) Estimate its evaluation error E i and the corresponding distribution by integrating the difference between the original quality evaluation score and the predicted quality evaluation score of the i-th weak classifier.
[0155] (4) Use a convex function to transform the weight of a single weak classifier into the corresponding weight, which is beneficial to ensuring that the lower the weak classifier, the greater the weight, and the lower the weak classifier, the smaller the weight.
[0156] (5) Finally, combine the weight with the predicted output to calculate the overall image score Q ent and the high-impact region score Q hig .
[0157] Step 6. Obtaining the final visual quality score: Through a weighting strategy, locally adjust the overall score with the score of the high-impact region to obtain the final visual quality score. The specific score adjustment formula is as follows:
[0158] Q = Q ent - c·sign(Q ent - Q hig )·δ
[0159] where Q is the final quality score of the screen content image, Q ent and Q hig represent the quality scores of the overall and high-impact regions, sign(·) represents the sign function, δ is the standard variance of Q ent and Q hig , and c is a constant used to avoid excessive fluctuations in the score of the overall image, set to 0.5.
[0160] As shown in Table 1, it shows all the extracted features (structural features, color features) and the corresponding IDs and feature descriptions. As Figure 3 shown, it shows the performance comparison of different feature combinations in the database SIQAD. From Figure 3 it can be seen that most of the methods using combined features can obtain better performance than the methods using single features, and the performance is the best when all features are combined. Therefore, all these features are complementary and can effectively predict the visual quality of screen content images.
[0161] As shown in Table 2, it shows the result comparison between the proposed score fusion method and other score fusion methods in the database SIQAD. It can be seen from the table that the proposed score fusion method achieves the best performance. This method finely adjusts the score of the entire image through the scoring of high-impact regions, taking into account both the visual impact of local regions and ensuring the integrity of the image, thus improving the accuracy.
[0162] As shown in Table 3, it shows the performance comparison between the proposed method and other different no-reference screen content image quality assessment method models in the databases SIQAD and SCID. It can be seen from the table that the proposed method has the best performance on SIQAD and SCID. In particular, the PLCC of the two databases both exceed 0.93, and the RMSE is 0.3274 and 0.2344 lower than that of MtDl respectively, showing obvious performance advantages.
[0163] Table 1 List of Extracted Features
[0164]
[0165] Table 2 Result Comparison Table between the Score Fusion Method Described in the Present Invention and Other Score Fusion Methods in the Database SIQAD
[0166]
[0167] Table 3 Performance Comparison Table between the Method Described in the Present Invention and Other Different No-reference Screen Content Image Quality Assessment Method Models in the Databases SIQAD and SCID
[0168]
[0169]
[0170] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed. The scope of protection required by the present invention is defined by the appended claims and their equivalents.
Claims
1. A no-reference screen content image quality evaluation method based on global and high-impact region analysis, characterized in that Including the following steps: 11) Division of the screen content image: First, obtain the rough text layer of the screen content image through the local image activity measurement algorithm; further segment the pure text area from the rough text layer through the refinement process based on text connected components, and the remaining areas are identified as image areas; 12) Division of high-impact areas: Select high-impact areas that have a higher impact on the overall quality between the pure text and image areas through information entropy; The division of the high-impact areas includes the following steps: 121) Calculate the information entropy of the pure text area and the image area respectively, and the information entropy formula is expressed as: where p(h) is the probability of pixel intensity h, Obtain the value p of the information entropy of the pure text region tex and the value p of the information entropy of the image region pic ; 122) Compare the information entropy values of the pure text area and the image area: When p tex ≥ p pic , the text area is recognized as a high-impact area; when p tex < p pic , the image area is recognized as a high-impact area; 13) Extraction of structural features: Combine phase congruency, local binary pattern, and various gradient features to obtain a gradient-weighted histogram, and extract the same structural features from the screen content image and the high-impact areas; 14) Extraction of color features: Calculate saturation and color entropy based on the opponent color space, and extract the same color features from the screen content image and the high-impact areas; 15) Obtaining the overall image score and the high-impact area score: Use the AdaBoosting BP neural network to train a regression model to obtain the overall image score and the high-impact area score; 16) Obtaining the final visual quality score: Locally adjust the overall score with the score of the high-impact area through a weighting strategy to obtain the final visual quality score.
2. The no-reference screen content image quality evaluation method based on global and high-impact region analysis according to claim 1, wherein The extraction of the structural features includes the following steps: 21) For the screen content image and the high-impact area, input the grayscale image I gray (i, j), and perform wavelet transform on it using a log-Gabor filter The log-Gabor filter uses the Gaussian function as the expansion function and extends the filter to two dimensions. The two-dimensional log-Gabor filter extended by the Gaussian function has the following transfer function: where, θ o = oπ / O, o = {0, 1, ..., O - 1} is the orientation angle of the filter, O is the number of orientations, and σ θ determines the angular bandwidth of the filter, ω0 represents the center frequency of the filter, and σ r is for controlling the bandwidth of the filter; 22) Calculate the amplitude A no (i, j) and the phase angle Amplitude A at a given wavelet scale n and direction o no (i, j) and phase angle are defined by the following formula: where, and respectively represent the even-symmetric cosine and odd-symmetric sine wavelets of two-dimensional Log-Gabor; 23) Calculate the phase congruency PC(i,j), The calculation method of PC(i,j) at various scales and directions is as follows: where ε is a constant introduced to avoid the denominator being zero; 24) Calculate the local binary pattern of each pixel in the phase congruency map to obtain the spatial distribution PLBP of the structure in the phase congruency domain, and the calculation is as follows: Among them, PC t is the value of the surrounding pixels, and PC c is the value of the central pixel. n is the number of equally spaced adjacent pixels around the central pixel, set to 8, and r is the radius of the neighborhood, set to 1; The definitions of the formulas W(·) and θ(·) are as follows: Since n is set to 8, the value range of the LBP mapping is [0, 9], which means there are 10 distribution patterns; 25) Calculate the pixel cumulative values of the gradient magnitude, relative gradient magnitude, and gradient direction with the same local binary pattern in the phase congruency domain, that is, the PLBP histogram with gradient weighting, and the calculation process is as follows: where f(·) is the Kronecker function, M and N respectively represent the height and width of the image, and m is the mode of PLBP, with a range of 0 to 9; G(i,j) ∈ {GM(i,j), RGM(i,j), GO(i,j)} is the weight of the PLBP of each pixel, and the calculation process of G(i,j) is as follows: Among them, and represent the horizontal gradient and the vertical gradient. dx′ and dy′ are the local average gradients of dx and dy after being filtered by an averaging filter, and the size of the filter is 3x3; 26) Obtain 30 structural features and perceptual feature vectors from the overall image and high-impact regions of a distorted screen content image respectively Finally, 120 structural features can be extracted at four scales for each overall image and high-impact area respectively.
3. The no-reference screen content image quality evaluation method based on global and high-impact region analysis according to claim 1, characterized in that The extraction of the color features includes the following steps: 31) For the screen content image and the high-impact area, input the color image I RGB (i, j), extract the color features of the image based on the red-green and blue-yellow channels, where the red-green and blue-yellow channels are represented by O1 and O2 respectively. The calculation process is as follows: Obtain the saturation feature IS by calculating the global mean of the image saturation: wherein is the saturation value at the (i, j) position; 32) Measure the color richness of the image by calculating the color entropy in the O1 and O2 color channels: 321) Use the de-mean normalization coefficient on the O1 and O2 color channels respectively to exclude the mutual influence of the image neighborhood information in the color space: where I(i,j) and represent the pixel value and the normalized value at the position (i,j) on the color channel, and u(i,j) and δ(i,j) represent the local region mean and the standard variance; 322) Calculate using the mean filter the mean of the neighboring pixels of each pixel in, so as to capture the structural information of the image in different color channels, and the calculation is as follows: Among them, Indicates the mean coefficient of MSCN, {S(b,v)|b = -B,...,B; v = -V,...,V} represents the mean filter, and B and V represent the height and width of the mean filter, which are set to 3; 323) Calculate the entropy in different color channels using the information entropy formula and to obtain the color entropy in the O1 channel the color entropy in the O2 channel Obtain 5 color features and the perceptual feature vector respectively from the overall image and the high-impact region of a distorted screen content image Finally, 20 color features can be extracted at four scales for each overall image and high-impact region respectively.
4. The no-reference screen content image quality evaluation method based on global and high-impact region analysis according to claim 1, characterized in that The obtaining of the overall image score and the high-impact area score includes the following steps: 41) Set up the training set: Connect the structural feature FS and the color feature FC to generate the feature vector as the final input; Randomly select m groups of data from the sample space where is the input feature vector, and Y x is the corresponding subjective score, and the value range of x is from 1 to m; 42) Select the BP neural network as the weak classifier, and the final strong classifier is obtained by iteratively training several weak classifiers. Here, the number of weak classifiers is set to 32; 43) Estimate its evaluation error E by integrating the difference between the original quality evaluation score and the predicted quality evaluation score of the i-th weak classifier i and the corresponding distribution; 44) Use the convex function to transform the weights of individual weak classifiers into corresponding weights, which is beneficial to ensuring that the weights of weak classifiers with low errors are larger and the weights of weak classifiers with high errors are smaller; 45) Finally, the weights are combined with the predicted output to calculate the overall image score Q ent and the high-impact region score Q hig .
5. The no-reference screen content image quality evaluation method based on global and high-impact region analysis according to claim 1, wherein In the step of obtaining the final visual quality score, The specific score adjustment formula is as follows: Q = Q ent -c·sign(Q ent -Q hig )·δ Among them, Q is the final quality score of the screen content image, Q ent and Q hig represent the quality scores of the overall and high-impact regions, sign(·) represents the sign function, and δ is the standard variance of Q ent and Q hig The standard deviation of, c is a constant used to avoid excessive fluctuations in the scores of the overall image, and is set to 0.5.
Citation Information
Patent Citations
Non-reference screen image quality evaluation method based on global information.
CN107274388A
Full-reference image quality evaluation method based on visual attention characteristics
CN109859157A