Video sharpness evaluation method, apparatus, device, storage medium, and product
By introducing saliency detection into video sharpness assessment and utilizing discrete wavelet transform of grayscale images and saliency calculation, the accuracy of video sharpness assessment is improved, meeting the needs of human visual perception.
Patent Information
- Application Number
- CN202310713304.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-15
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2043-06-15
AI Technical Summary
Existing video sharpness assessment methods ignore the human eye's perception of local image features, resulting in low accuracy in sharpness estimation.
This paper introduces saliency detection into video sharpness evaluation. By converting video image frames into grayscale images, performing discrete wavelet transform to obtain high-frequency sub-band images, calculating saliency maps and converting them into weight matrices, and combining the high-frequency sub-band images to calculate the total high-frequency energy value to evaluate sharpness.
It improves the accuracy of sharpness estimation, making the assessment more in line with human visual perception and providing better feedback on the clarity of the video.
Smart Images

Figure CN116739931B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of video, and in particular to a video sharpness evaluation method, device, equipment, storage medium and product. BACKGROUND
[0002] As an index reflecting the sharpness of image edges and details, sharpness is an important part of definition evaluation. Video sharpness evaluation, as one of the important indicators of video quality evaluation, is of great significance to video encoding, transmission, storage and other links. Especially for live video, sharpness evaluation can directly judge the definition of the audience end, and then improve the definition of the video with insufficient definition through certain picture quality enhancement technology to improve the user experience. At the same time, sharpness estimation can also evaluate the effect of picture quality enhancement technology, and guide the algorithm adjustment of technical personnel. Further, sharpness estimation can also score the audience end of the live video, and give the detected low-definition video to the recommendation side to assist the low-definition suppression strategy of the recommendation end. Therefore, video sharpness evaluation is of great significance to the revenue of live or short videos.
[0003] In related technologies, sharpness evaluation methods are roughly divided into three categories. One is a spatial domain method, which uses mathematical parameters to represent the transition zone width of the image edge. The idea is that the more intense the normal gray scale change of the edge is, the narrower the edge width is, and the higher the sharpness evaluation value of the image is, and the clearer the image is; one is a frequency domain method, which judges the sharpness based on the peak value of the transform coefficient of discrete cosine transform, discrete wavelet transform, etc.; one is the fusion of edge feature-based and transform-based methods, such as using the statistical information of image gradient histogram and high-frequency detail information based on wavelet transform to evaluate sharpness. The hybrid method based on spatial and frequency domains has been proven to usually perform better than the method based on edges or the method based on transforms only, but often at the cost of increased computational complexity. The foregoing sharpness evaluation methods only consider the global information of the image, and the accuracy of sharpness estimation is low. SUMMARY
[0004] Embodiments of the present application provide a video sharpness evaluation method, device, equipment, storage medium and product, which solve the problem that the existing video sharpness evaluation ignores the perception of the human eye to the local features of the image, introduce saliency detection into video sharpness evaluation, make the video sharpness evaluation pay more attention to the part of the image that attracts the attention of the human eye, and be more consistent with the human eye perception, improve the accuracy of sharpness estimation, and thus better feedback the definition of the video.
[0005] In a first aspect, embodiments of the present application provide a video sharpness evaluation method, which comprises:
[0006] Converting a video image frame to be evaluated into a gray-scale image, and performing discrete wavelet transform on the gray-scale image to obtain a high-frequency sub-band image;
[0007] performing saliency calculation on the gray image to obtain a saliency map, and performing scaling on the saliency map to convert the saliency map into a weight matrix with the same size as the high-frequency subband map;
[0008] based on the weight matrix and the high-frequency subband map, calculating a total high-frequency energy value corresponding to the video image frame to represent the sharpness evaluation result of the video image frame.
[0009] In a second aspect, an embodiment of the present application further provides a video sharpness evaluation device, comprising:
[0010] a high-frequency subband determination module configured to convert a video image frame to be evaluated into a gray image, and perform discrete wavelet transform on the gray image to obtain a high-frequency subband map;
[0011] a saliency determination module configured to perform saliency calculation on the gray image to obtain a saliency map, and perform scaling on the saliency map to convert the saliency map into a weight matrix with the same size as the high-frequency subband map;
[0012] a high-frequency energy calculation module configured to calculate a total high-frequency energy value corresponding to the video image frame based on the weight matrix and the high-frequency subband map, to represent the sharpness evaluation result of the video image frame.
[0013] In a third aspect, an embodiment of the present application further provides a video sharpness evaluation device, comprising:
[0014] one or more processors;
[0015] a storage device configured to store one or more programs,
[0016] when the one or more programs are executed by the one or more processors, the one or more processors implement the video sharpness evaluation method provided by the embodiments of the present application.
[0017] In a fourth aspect, an embodiment of the present application further provides a nonvolatile storage medium storing computer executable instructions, which, when executed by a computer processor, are configured to perform the video sharpness evaluation method provided by the embodiments of the present application.
[0018] In a fifth aspect, an embodiment of the present application further provides a computer program product, which comprises a computer program stored in a computer readable storage medium, and at least one processor of a device reads and executes the computer program from the computer readable storage medium, so that the device performs the video sharpness evaluation method provided by the embodiments of the present application.
[0019] In the embodiment of the present application, the video image frame to be evaluated is converted into a gray scale image, and a high frequency sub-band image is obtained by performing discrete wavelet transform on the gray scale image. Then, a saliency map is obtained by performing saliency calculation on the gray scale image, and the saliency map is scaled to convert it into a weight matrix with the same size as the high frequency sub-band image. Finally, the high frequency energy total value corresponding to the video image frame is calculated based on the weight matrix and the high frequency sub-band image, so as to represent the sharpness evaluation result of the video image frame. By introducing saliency detection into the video sharpness evaluation, the video sharpness evaluation pays more attention to the part of the image that attracts the attention of the human eye, and is more consistent with the human eye perception, thereby improving the accuracy of the sharpness estimation, and better feedback the definition of the video. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 A flowchart of a video sharpness evaluation method provided by the embodiment of the present application;
[0021] Figure 2 A schematic diagram of an exemplary three-level Haar wavelet transform;
[0022] Figure 3 A flowchart of a method for calculating a high frequency energy total value provided by the embodiment of the present application;
[0023] Figure 4 A structural block diagram of a video sharpness evaluation device provided by the embodiment of the present application;
[0024] Figure 5 A structural schematic diagram of a video sharpness evaluation apparatus provided by the embodiment of the present application. DETAILED DESCRIPTION
[0025] The embodiment of the present application will be further described in detail below in conjunction with the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the embodiment of the present application, but not to limit the embodiment of the present application. In addition, it should be noted that, in order to facilitate the description, only the parts related to the embodiment of the present application are shown in the drawings, but not all the structures.
[0026] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiment of the present application can be implemented in an order other than that illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually a class, and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / ", generally represents a "or" relationship between the front and rear associated objects.
[0027] The video sharpness evaluation method provided in the embodiments of the present application is used for evaluating the definition of a video or an image, and can be applied to multiple links such as video coding, transmission and storage. The specific application scenarios can include video conferencing, video call, indoor live broadcast, outdoor live broadcast and short video, etc. Taking indoor live broadcast as an example, the method can be used to judge the real-time definition of the live broadcast video watched by the audience end, and can also be used to evaluate the actual effect of the live broadcast picture after quality enhancement processing. The foregoing several application scenarios are only exemplary and explanatory, and in actual application, the video sharpness evaluation method can also be used in video sharpness evaluation in other scenarios, and the embodiments of the present application do not limit this. The present application aims to provide a video sharpness evaluation method, which introduces saliency detection into video sharpness evaluation, so that the video sharpness evaluation pays more attention to the part of the image that attracts the attention of the human eye, is more consistent with the human eye perception, improves the accuracy of sharpness estimation, and thus better feedbacks the definition of the video.
[0028] In addition, the video sharpness evaluation method provided in the embodiments of the present application aims to take the total high-frequency energy value of the video image frame as an evaluation index of the video sharpness, and by setting a test set, the accuracy of the evaluation index corresponding to the embodiments of the present application can be verified. For example, the objective index and the quality correlation evaluation index in image quality evaluation, such as SROCC (Spearman rank order correlation coefficient, Spearman rank order correlation coefficient) and PLCC (Pearson linear correlation coefficient, Pearson linear correlation coefficient), can be used to evaluate the monotonicity and linear correlation degree of the total high-frequency energy value and the subjective MOS (Mean Opinion Score, Mean Opinion Score) score, and the calculation results are shown in the following table:
[0029] SROCC PLCC Ave 0.99 0.96
[0030] Therefore, the video sharpness evaluation method provided in the embodiments of the present application takes the total high-frequency energy value as an evaluation index of the video sharpness, which has high consistency with the subjective MOS score, and the evaluation index has good accuracy and reliability.
[0031] Further, in order to verify the stability of the evaluation index in the time domain of the video, the live video can also be used as test data, and the total high-frequency energy value of each image frame in the live video is calculated by using the video sharpness evaluation method provided in the embodiments of the present application. By integrating and statistically analyzing the calculation results, the following conclusions can be obtained: when the video picture changes little, the difference of the evaluation index values between frames is small, which indicates that the evaluation index has good stability in the time domain; when the video anchor is in a state of intense exercise, the video anchor appears in the picture and then leaves, and the difference of the evaluation index values between frames is large, which indicates that the evaluation index is consistent with the volatility between video frames; when the video background is in a state of change, the image of the video anchor in the picture is basically unchanged, and the change of the evaluation index values between frames is small, which indicates that the video sharpness evaluation method provided in the embodiments of the present application introduces saliency detection into video sharpness evaluation, can better focus on the part of the image that attracts the attention of the human eye, and is more consistent with the human eye perception, and meets the design expectation.
[0032] The execution subject of each step of the video sharpness evaluation method provided in the embodiments of the present application can be a computer device, which refers to any electronic device with data calculation, processing and storage capabilities, such as mobile phones, PC (Personal Computer), tablet computers and other terminal devices, and can also be a server or other device, which is not limited in the embodiments of the present application.
[0033] Figure 1 The flowchart of the video sharpness evaluation method provided in the embodiments of the present application specifically includes the following steps:
[0034] Step S101: converting the video image frame to be evaluated into a gray-scale image, and performing discrete wavelet transform on the gray-scale image to obtain a high-frequency sub-band image.
[0035] In one embodiment, since sharpness mainly describes the definition of edges and details of an image, which is represented as high frequency information of the image, it is necessary to transform the video image signal to the frequency domain, and the embodiment of the present application adopts discrete wavelet transform to perform frequency domain transform on the video image frame. The video image frame to be evaluated is usually a color image, since each pixel point in the color image corresponds to three color channels of R, G and B. Directly performing discrete wavelet transform on the color image requires separate processing of the three color channels and selecting a suitable color space for processing, which has high calculation cost and low processing efficiency. Therefore, the embodiment of the present application first converts the video image frame to be evaluated into a grayscale image, and when performing discrete wavelet transform, each pixel value can be directly processed. The specific discrete wavelet transform can be Haar wavelet, Daubechies wavelet, etc., which is not limited in the present application. In addition, since sharpness feedbacks high frequency information of the image, the high frequency subband image corresponding to the grayscale image can be obtained by discrete wavelet transform, and the high frequency subband image reflects the fluctuation value of the local signal and stores the detail information of the picture.
[0036] In one embodiment, in order to refer to the multi-scale high frequency information and improve the accuracy of sharpness estimation, multi-level discrete wavelet transform can be performed on the grayscale image to obtain high frequency subband images of different scales. Optionally, the discrete wavelet transform of the grayscale image to obtain the high frequency subband image includes: performing multi-level discrete wavelet transform on the grayscale image to obtain multi-layer high frequency subband images, and each layer of high frequency subband image includes a horizontal high frequency subband image, a vertical high frequency subband image and a diagonal high frequency subband image. Taking a three-level Haar wavelet transform as an example, Figure 2 An exemplary schematic diagram of a three-level Haar wavelet transform is shown in FIG. 1. Figure 2As shown, the first-level Haar wavelet transform of the gray-scale image can obtain the first-layer high-frequency sub-band image: the horizontal high-frequency sub-band image 1011, the vertical high-frequency sub-band image 1012, and the diagonal high-frequency sub-band image 1013, and the low-frequency sub-band image. On the basis of the low-frequency sub-band image of the first layer, the second-level Haar wavelet transform can obtain the second-layer high-frequency sub-band image: the horizontal high-frequency sub-band image 1021, the vertical high-frequency sub-band image 1022, and the diagonal high-frequency sub-band image 1023, and the low-frequency sub-band image. On the basis of the low-frequency sub-band image of the second layer, the third-level Haar wavelet transform can obtain the third-layer high-frequency sub-band image: the horizontal high-frequency sub-band image 1031, the vertical high-frequency sub-band image 1032, and the diagonal high-frequency sub-band image 1033, and the low-frequency sub-band image. Thus, the multi-scale three-layer high-frequency sub-band image can be obtained, and each layer of the high-frequency sub-band image can include the horizontal high-frequency sub-band image, the vertical high-frequency sub-band image, and the diagonal high-frequency sub-band image. The horizontal high-frequency sub-band image can be a wavelet coefficient image generated by using the low-pass wavelet filter to convolve in the row direction and then using the high-pass wavelet filter to convolve in the column direction, and represents the horizontal direction singular characteristics of the image. The vertical high-frequency sub-band image can be a wavelet coefficient image generated by using the high-pass wavelet filter to convolve in the row direction and then using the low-pass wavelet filter to convolve in the column direction, and represents the vertical direction singular characteristics of the image. The diagonal high-frequency sub-band image can be a wavelet coefficient image generated by using the high-pass wavelet filter to convolve in both the row and column directions, and represents the diagonal edge characteristics of the image. Therefore, the multi-level discrete wavelet transform can decompose the gray-scale image into high-frequency information of different scales, so that the high-frequency information is more obvious, more detailed information is retained, and the accuracy of the sharpness estimation is improved.
[0037] In step S102, the saliency of the gray-scale image is calculated to obtain a saliency map, and the saliency map is scaled to convert it into a weight matrix with the same size as the high-frequency sub-band image.
[0038] In an embodiment, the saliency can be the degree of attracting the human eye in visual perception, and can reflect the visual attraction and visual quality of the image. The saliency map obtained by the saliency calculation of the gray-scale image can quickly locate the part of the gray-scale image that attracts the human eye. The saliency calculation can use a method based on frequency domain analysis, a method based on region segmentation, or a method based on global contrast, and the present application is not limited herein. The saliency map is converted into a weight matrix with the same size as the high-frequency sub-band image through scaling, and the fusion calculation of the high-frequency sub-band image and the saliency map is realized through the weight matrix. The scaling method can be nearest neighbor interpolation, bilinear interpolation, and bidirectional linear interpolation, and the present application is not limited herein.
[0039] In one embodiment, since the video sharpness evaluation method provided by the embodiment of the present application uses the frequency domain information of the video image frame, the spatial domain information of the video image frame is easily ignored, and therefore a method based on global contrast can be used to calculate the saliency of the gray image, and the specific implementation process includes:
[0040] Each pixel of the gray image is traversed, and the sum of the color distance values of the current pixel and other pixels is taken as the saliency value during the traversal process, so as to calculate the saliency map, and the calculation formula is as follows:
[0041]
[0042] wherein I k is the current pixel, I i is the other pixel in the gray image, and ||*|| represents the Euclidean distance. It can be understood that the luminance difference degree between different regions in the gray image reflects the spatial domain information of the image. Therefore, by calculating the luminance value difference degree of all pixels in the gray image, i.e., the global contrast, the spatial domain information between different regions in the image can be obtained. Based on the saliency of the reference spatial domain information, the saliency can better fit the habit of human eye attention, and the accuracy of the sharpness estimation can be improved.
[0043] Further, since there can be the same pixel value in the gray image, and for the same pixel value, the calculated Euclidean distance result is the same, in order to improve the calculation efficiency, the specific implementation process of the saliency calculation of the gray image includes:
[0044] The color value corresponding to the pixel in the gray image is obtained, the occurrence frequency corresponding to each color value is counted, the saliency value corresponding to each color value is calculated based on the occurrence frequency corresponding to each color value and the color distance between each color value and other color values, and the saliency value corresponding to each color value is assigned to the corresponding pixel to obtain the saliency map. The calculation formula is as follows:
[0045]
[0046] wherein I k is the current pixel, C l is the color value corresponding to the current pixel, C j is the color value corresponding to the other pixel in the gray image, f j is the occurrence frequency of the color value C j , and ||*|| represents the Euclidean distance. It can be understood that the saliency value is calculated by the color value and the corresponding occurrence frequency, without the need to traverse each pixel in the gray image for calculation, thereby reducing the calculation complexity, improving the efficiency of the saliency calculation, and accelerating the processing process of the video image frame to be evaluated.
[0047] In one embodiment, in the case of multi-level discrete wavelet transform on the gray scale image, in one embodiment, the saliency map is scaled and converted into a weight matrix with the same size as the high frequency sub-band map, including:
[0048] The saliency map is scaled and converted into a weight matrix with the same size as each high frequency sub-band map in each layer, respectively.
[0049] It can be understood that multi-level discrete wavelet transform on the gray scale image can obtain multi-scale high frequency information, and accordingly, the high frequency information at each scale can set a corresponding weight matrix, so the saliency map needs to be converted into a weight matrix with the same size as each high frequency sub-band map in each layer. Thus, the embodiments of the present application provide multi-scale saliency information corresponding to the multi-scale high frequency information, and through multi-scale fusion, the sharpness estimation is more accurate.
[0050] In step S103, based on the weight matrix and the high frequency sub-band map, the total high frequency energy value corresponding to the video image frame is calculated to represent the sharpness evaluation result of the video image frame.
[0051] The total high frequency energy value can be the energy proportion of the high frequency signal in the video image frame, representing the energy distribution of the detail part in the video image frame. It can be understood that the higher the total high frequency energy value of the video image frame, the clearer the image details and the richer the texture information. Thus, the total high frequency energy value can be used as the sharpness evaluation result of the video image frame.
[0052] Figure 3 The flow chart of the method for calculating the total high frequency energy value provided by the embodiments of the present application is shown in FIG. 1, in the case of multi-level discrete wavelet transform on the gray scale image, the specific steps for calculating the total high frequency energy value include: Figure 3
[0053] In step S1031, based on each high frequency sub-band map in each layer and the weight matrix with the same size as the high frequency sub-band map, the high frequency energy value corresponding to each high frequency sub-band map in each layer is calculated.
[0054] The multi-level discrete wavelet transform corresponds to multi-layer high frequency sub-band maps, and the number of high frequency sub-band maps in each layer can be multiple, and the high frequency energy value of the high frequency sub-band map can be calculated according to the weight matrix corresponding to each high frequency sub-band map and the high frequency information of the high frequency sub-band map itself.
[0055] For example, the calculation formula of the high frequency energy value corresponding to each high frequency sub-band map in each layer is as follows:
[0056]
[0057] Wherein, En represents the high frequency energy value of each high frequency sub-band map in the nth layer, This represents the high-frequency sub-band diagrams in the nth layer. This represents the pixel value of the pixel located in the i-th row and j-th column of the currently calculated high-frequency subband image. N represents the value of the element in the i-th row and j-th column of the weight matrix corresponding to the currently calculated high-frequency subband map. n This represents the total number of pixels in the currently calculated high-frequency sub-band image. It's understandable that each pixel value in each high-frequency sub-band image corresponds to the element value at the same position in the weight matrix as its weight. For the saliency map, the saliency values of the high-frequency components are more prominent, corresponding to larger weights in the weight matrix. Therefore, the high-frequency energy value calculated by combining the high-frequency sub-band image with the weight matrix focuses more on the parts of the image that attract human attention, making it more consistent with human perception.
[0058] Step S1032: Calculate the high-frequency energy values corresponding to each high-frequency sub-band diagram in each layer by weighting the calculation.
[0059] Each high-frequency sub-band map can include a horizontal, vertical, and diagonal high-frequency sub-band map. The horizontal high-frequency sub-band map reflects the high-frequency portion of the grayscale image in the horizontal direction, and its high-frequency energy is usually proportional to the number and intensity of horizontal edges appearing in the grayscale image. The vertical high-frequency sub-band map reflects the high-frequency portion of the grayscale image in the vertical direction, and its high-frequency energy is usually proportional to the number and intensity of vertical edges appearing in the grayscale image. The diagonal high-frequency sub-band map reflects the high-frequency portion of the grayscale image in the diagonal direction, and its high-frequency energy is usually proportional to the number and intensity of diagonal edges appearing in the grayscale image. Therefore, by combining the edge distribution information of the video image frames to be evaluated in the actual application scenario, different weights can be set for each high-frequency sub-band map, thereby obtaining the high-frequency energy value of each high-frequency sub-band map.
[0060] For example, the formula for calculating the high-frequency energy value corresponding to each high-frequency sub-band diagram is as follows:
[0061]
[0062] Among them, E n This represents the high-frequency energy value of the nth high-frequency subband diagram. This represents the high-frequency energy value of the nth horizontal high-frequency sub-band diagram. This represents the high-frequency energy value of the nth vertical high-frequency sub-band diagram. represents the high frequency energy value of the nth layer diagonal high frequency sub-band diagram, and a represents the energy weight coefficient. It can be understood that the value range of a is between 0 and 1. For the portrait or object in the video image frame to be evaluated, when the portrait or object is horizontally or vertically distributed in the picture, the number and intensity of the horizontal and vertical edges of the video image frame will be more prominent than the diagonal edges, and therefore a can be valued between 0 and 0.5, and the high frequency energy values of the horizontal and vertical high frequency sub-band diagrams are more referenced. For the portrait or object in the video image frame to be evaluated, when the portrait or object is more diagonally distributed in the picture, the number and intensity of the diagonal edges of the video image frame will be more prominent than the horizontal and vertical edges, and therefore a can be valued between 0.5 and 1, and the high frequency energy value of the diagonal high frequency sub-band diagram is more referenced. Therefore, according to the characteristics of the video content in the actual application scene, including the distribution of the portrait and object, the weight values of different high frequency sub-band diagrams are better allocated, the human eye attention habit is fitted, and the accuracy of calculating the high frequency energy value is improved.
[0063] In step S1033, the high frequency energy values corresponding to each layer high frequency sub-band diagram are weighted and calculated to obtain a high frequency energy total value.
[0064] In one embodiment, taking a three-level discrete wavelet transform as an example, the high frequency energy value of the first layer high frequency sub-band diagram represents the coarsest detail information in the video image frame, i.e. the high frequency part on a large scale, which contains the basic contour information of the video image frame; the high frequency energy value of the second layer high frequency sub-band diagram is obtained by performing a discrete wavelet transform on the detail information at the scale of the first layer high frequency sub-band diagram, and therefore represents more detailed information; the high frequency energy value of the third layer high frequency sub-band diagram is obtained by performing a discrete wavelet transform on the detail information at the scale of the second layer high frequency sub-band diagram, and therefore represents the most detailed information, i.e. the high frequency part on the finest scale, which contains the fine texture and noise information of the image. Therefore, different weight values can be set for high frequency energy values of different scales, wherein the higher the level, the more detailed the high frequency part concerned, and the influence degree is relatively reduced, and therefore a layer-by-layer decreasing weight setting method can be adopted to better fit the part of human eye attention and improve the accuracy of sharpness evaluation.
[0065] For example, the calculation formula of the high frequency energy total value is as follows:
[0066]
[0067] wherein Index represents the high frequency energy total value, E nThe high-frequency energy value of the nth layer high-frequency sub-band image, and N represents that the gray image adopts N-level discrete wavelet transform. In an embodiment, taking a three-level discrete wavelet transform as an example, the calculation formula of the total high-frequency energy value is as follows:
[0068]
[0069] It can be understood that, according to the formula, the weight value corresponding to the high-frequency energy value of the first layer high-frequency sub-band image is 4, the weight value corresponding to the high-frequency energy value of the second layer high-frequency sub-band image is 2, and the weight value corresponding to the high-frequency energy value of the first layer high-frequency sub-band image is 1. On the basis that the high-frequency energy value corresponding to the large-scale detail part is the main component of the total high-frequency energy value, the high-frequency energy value corresponding to the finer-scale detail part is increased, the effect of the high-frequency energy value of each layer high-frequency sub-band image is reasonably ensured, and the sharpness estimation is more accurate.
[0070] The above, by converting the video image frame to be evaluated into a gray image, the high-frequency sub-band image is obtained by performing discrete wavelet transform on the gray image; then the saliency map is obtained by performing saliency calculation on the gray image, and the saliency map is scaled to convert the saliency map into a weight matrix with the same size as the high-frequency sub-band image; finally, based on the weight matrix and the high-frequency sub-band image, the total high-frequency energy value corresponding to the video image frame is calculated to represent the sharpness evaluation result of the video image frame. By introducing the saliency detection into the video sharpness evaluation, the video sharpness evaluation pays more attention to the part of the image that attracts the attention of the human eye, is more in line with the human eye perception, improves the accuracy of the sharpness estimation, and thus better feedbacks the definition of the video.
[0071] Figure 4 A structural block diagram of a video sharpness evaluation device provided by an embodiment of the present application is provided, which is configured to execute the video sharpness evaluation method provided by the above-mentioned embodiment, and has the corresponding function modules and beneficial effects of the execution method. As shown in the figure, the device specifically includes: Figure 4
[0072] The high-frequency sub-band determination module 201 is configured to convert the video image frame to be evaluated into a gray image, and obtain a high-frequency sub-band image by performing discrete wavelet transform on the gray image;
[0073] The saliency determination module 202 is configured to obtain a saliency map by performing saliency calculation on the gray image, and scale the saliency map to convert the saliency map into a weight matrix with the same size as the high-frequency sub-band image;
[0074] The high-frequency energy calculation module 203 is configured to calculate the total high-frequency energy value corresponding to the video image frame based on the weight matrix and the high-frequency sub-band image, so as to represent the sharpness evaluation result of the video image frame.
[0075] From the above scheme, by converting the video image frame to be evaluated into a gray image, a high-frequency sub-band image is obtained by performing discrete wavelet transform on the gray image; then a saliency map is obtained by performing saliency calculation on the gray image, and the saliency map is scaled to convert it into a weight matrix with the same size as the high-frequency sub-band image; finally, based on the weight matrix and the high-frequency sub-band image, the total high-frequency energy value corresponding to the video image frame is calculated to represent the sharpness evaluation result of the video image frame. By introducing saliency detection into video sharpness evaluation, the video sharpness evaluation pays more attention to the part of the image that attracts the attention of the human eye, which is more consistent with the human eye perception, improves the accuracy of sharpness estimation, and thus better feedbacks the definition of the video.
[0076] In one possible embodiment, the high-frequency sub-band determination module 201 is specifically configured to:
[0077] performing multi-level discrete wavelet transform on the gray image to obtain a plurality of high-frequency sub-band images, each layer of the high-frequency sub-band image including a horizontal high-frequency sub-band image, a vertical high-frequency sub-band image and a diagonal high-frequency sub-band image.
[0078] Correspondingly, the saliency determination module 202 is specifically configured to:
[0079] scaling the saliency map to convert it into a weight matrix with the same size as each high-frequency sub-band image in each layer.
[0080] In one possible embodiment, the high-frequency energy calculation module 203 is specifically configured to:
[0081] calculating a high-frequency energy value corresponding to each high-frequency sub-band image in each layer based on each high-frequency sub-band image in each layer and the weight matrix with the same size as the high-frequency sub-band image;
[0082] performing weighted calculation on the high-frequency energy values corresponding to each high-frequency sub-band image in each layer to obtain a high-frequency energy value corresponding to each layer of high-frequency sub-band image;
[0083] performing weighted calculation on the high-frequency energy values corresponding to each layer of high-frequency sub-band image to obtain the total high-frequency energy value.
[0084] In one possible embodiment, the calculation formula of the high-frequency energy value corresponding to each high-frequency sub-band image in each layer is as follows:
[0085]
[0086] wherein, represents the high-frequency energy value of each high-frequency sub-band image in the nth layer, represents each high-frequency sub-band image in the nth layer, represents the pixel value of the pixel located at the i-th row and the j-th column in the current calculated high-frequency sub-band image, represents the element value of the i-th row and the j-th column of the weight matrix corresponding to the current calculated high-frequency sub-band image, N represents the total number of pixels of the current calculated high-frequency sub-band image. n represents the element value of the i-th row and the j-th column of the weight matrix corresponding to the current calculated high-frequency sub-band image, N represents the total number of pixels of the current calculated high-frequency sub-band image.
[0087] In one possible embodiment, the calculation formula of the high-frequency energy value corresponding to each layer of high-frequency sub-band image is as follows:
[0088]
[0089] wherein, E n represents the high-frequency energy value of the n-th layer of high-frequency sub-band image, represents the high-frequency energy value of the n-th layer of horizontal high-frequency sub-band image, represents the high-frequency energy value of the n-th layer of vertical high-frequency sub-band image, represents the high-frequency energy value of the n-th layer of diagonal high-frequency sub-band image, and a represents an energy weight coefficient.
[0090] In one possible embodiment, the calculation formula of the total high-frequency energy value is as follows:
[0091]
[0092] wherein, Index represents the total high-frequency energy value, E n represents the high-frequency energy value of the n-th layer of high-frequency sub-band image, and N represents the number of levels of discrete wavelet transform adopted by the gray image.
[0093] In one possible embodiment, the saliency determination module 202 is specifically configured as:
[0094] The sum of the color distance values of the current pixel and other pixels in the gray image is taken as the saliency value to calculate the saliency map.
[0095] In one possible embodiment, the saliency determination module 202 is specifically configured as:
[0096] The color values corresponding to the pixels in the gray image are obtained, and the occurrence frequency corresponding to each color value is counted;
[0097] The saliency value corresponding to each color value is calculated based on the occurrence frequency corresponding to each color value and the color distance between each color value and other color values;
[0098] The saliency value corresponding to each color value is assigned to the corresponding pixel to obtain the saliency map.
[0099] Figure 5 A structural schematic diagram of a video sharpness evaluation device provided by the embodiments of the present application is shown in FIG. 1. Figure 5As shown, the device includes a processor 301, a memory 302, an input device 303 and an output device 304; the number of processors 301 in the device can be one or more, Figure 5 The processor 301 in the device is taken as an example in the embodiment. The processor 301, the memory 302, the input device 303 and the output device 304 in the device can be connected through a bus or other means, Figure 5 The memory 302 is taken as an example in the embodiment. The memory 302 can be configured to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the video sharpness evaluation method in the embodiment. The processor 301 can execute various function applications and data processing of the device by running the software programs, instructions and modules stored in the memory 302, that is, the video sharpness evaluation method described above is implemented. The input device 303 can be configured to receive input digital or character information, and generate key signal input related to user settings and function control of the device. The output device 304 can include a display device such as a display screen.
[0100] The embodiment of the application also provides a non-volatile storage medium containing computer executable instructions, which are configured to execute a video sharpness evaluation method described in the above embodiment when executed by a computer processor, wherein the method comprises:
[0101] The video image frame to be evaluated is converted into a gray scale image, and a high frequency sub-band image is obtained by performing discrete wavelet transform on the gray scale image;
[0102] The saliency of the gray scale image is calculated to obtain a saliency map, and the saliency map is scaled to convert it into a weight matrix with the same size as the high frequency sub-band image;
[0103] Based on the weight matrix and the high frequency sub-band image, the total high frequency energy value corresponding to the video image frame is calculated to represent the sharpness evaluation result of the video image frame.
[0104] It is worth noting that the embodiments of the video sharpness evaluation device described above include various units and modules only according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy distinction, and do not configure to limit the protection scope of the embodiment.
[0105] In some possible implementation, each of the aspects of the method provided by the present application can also be implemented as a program product in the form of a computer program product, which includes program codes configured to cause a computer device to perform the steps of the method according to various exemplary embodiments of the present application described above in the specification when the program product is run on the computer device, for example, the computer device can perform the video sharpness evaluation method described in the embodiments of the present application. The program product can be implemented by any combination of one or more computer readable media.
Claims
1. A method of video sharpness evaluation, characterized by, The method comprises the following steps: transforming a video image frame to be evaluated into a gray image, performing multi-level discrete wavelet transform on the gray image to obtain a plurality of layers of high-frequency sub-band images, each layer of the high-frequency sub-band images comprising a horizontal high-frequency sub-band image, a vertical high-frequency sub-band image and a diagonal high-frequency sub-band image; performing saliency calculation on the gray image to obtain a saliency map, and scaling the saliency map to be respectively converted into a weight matrix with the same size as each high-frequency sub-band image in each layer; based on each high-frequency sub-band image in each layer and the weight matrix with the same size as each high-frequency sub-band image, calculating a high-frequency energy value corresponding to each high-frequency sub-band image in each layer, performing weighted calculation on the high-frequency energy values corresponding to each high-frequency sub-band image in each layer to obtain a high-frequency energy value corresponding to each layer of high-frequency sub-band images, and performing weighted calculation on the high-frequency energy values corresponding to each layer of high-frequency sub-band images to obtain a total high-frequency energy value.
2. The video sharpness evaluation method of claim 1, wherein, The calculation formula of the high-frequency energy value corresponding to each high-frequency sub-band image in each layer is as follows: wherein, represents the layer of the high-frequency sub-band image, represents the layer of the high-frequency sub-band image, represents the pixel value of the pixel located at the row and the column of the currently calculated high-frequency sub-band image, represents the element value of the element located at the row and the column of the weight matrix corresponding to the currently calculated high-frequency sub-band image, represents the total number of pixels of the currently calculated high-frequency sub-band image.
3. The video sharpness evaluation method of claim 1, wherein, The calculation formula of the high-frequency energy value corresponding to each layer of high-frequency sub-band images is as follows: wherein, represents the high frequency energy value of the high frequency subband map of the layer represents the high frequency energy value of the high frequency subband map of the layer represents the high frequency energy value of the high frequency subband map of the layer represents the high frequency energy value of the high frequency subband map of the layer represents the high frequency energy value of the high frequency subband map of the layer represents the high frequency energy value of the high frequency subband map of the layer represents the high frequency energy value of the high frequency subband map of the layer represents the high frequency energy value of the high frequency subband map of the layer represents the energy weight coefficient.
4. The video sharpness evaluation method of claim 1, wherein, The calculation formula of the total high-frequency energy value is as follows: wherein represents the total value of high frequency energy, represents the first high frequency subband map, represents that the gray scale map employs a level discrete wavelet transform.
5. The video sharpness evaluation method according to any of claims 1-4, characterized by, The saliency calculation on the gray image to obtain a saliency map comprises: traversing each pixel of the gray image, and taking the sum of color distance values of the current pixel and other pixels as a saliency value to calculate the saliency map.
6. The video sharpness evaluation method of any of claims 1-4, wherein, The saliency calculation on the gray image to obtain a saliency map comprises: obtaining color values corresponding to pixels in the gray image, and counting the occurrence frequency of each color value; based on the occurrence frequency of each color value and the color distance between each color value and other color values, calculating a saliency value corresponding to each color value; assigning the saliency value corresponding to each color value to the corresponding pixel to obtain a saliency map.
7. Apparatus for video sharpness evaluation, characterized in that The method comprises the following steps: a high-frequency sub-band determination module configured to transform a video image frame to be evaluated into a gray image, perform multi-level discrete wavelet transform on the gray image to obtain a plurality of layers of high-frequency sub-band images, and each layer of the high-frequency sub-band images comprising a horizontal high-frequency sub-band image, a vertical high-frequency sub-band image and a diagonal high-frequency sub-band image; a saliency determination module configured to perform saliency calculation on the gray image to obtain a saliency map, and scale the saliency map to be respectively converted into a weight matrix with the same size as each high-frequency sub-band image in each layer; a high-frequency energy calculation module configured to, based on each high-frequency sub-band image in each layer and the weight matrix with the same size as each high-frequency sub-band image, calculate a high-frequency energy value corresponding to each high-frequency sub-band image in each layer, perform weighted calculation on the high-frequency energy values corresponding to each high-frequency sub-band image in each layer to obtain a high-frequency energy value corresponding to each layer of high-frequency sub-band images, and perform weighted calculation on the high-frequency energy values corresponding to each layer of high-frequency sub-band images to obtain a total high-frequency energy value.
8. A video sharpness assessment device, the device comprising: one or more processors; a storage device configured to store one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the video sharpness evaluation method in any one of claims 1-6.
9. A non-transitory storage medium storing computer-executable instructions that, when executed by a computer processor, are configured to perform the video sharpness evaluation method of any one of claims 1-6.
10. A computer program product comprising a computer program, characterized in that, The computer program, when executed by the processor, implements the video sharpness evaluation method of any one of claims 1-6.
Citation Information
Patent Citations
Application method of vision multichannel model in stereoscopic video quality objective evaluation
CN107071423A
Image processing method and device, equipment and storage medium
CN114612336A
Method for evaluating non-reference multi-layer perception quality of panoramic image
CN115588001A