Method for automatic video quality assessment
The temporal clustering method, integrating the INRF model with advanced temporal pooling, addresses the challenge of assessing HFR, HDR, and UHD 8K video quality by achieving a high correlation with human perception, surpassing existing metrics in performance.
Patent Information
- Application Number
- PCT/ES2023/070676
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-13
- Publication Date
- 2025-05-22
AI Technical Summary
Existing video quality metrics struggle to effectively assess the quality of high frame rate (HFR), high dynamic range (HDR), and ultra high definition 8K (UHD 8K) videos, as they do not correlate well with human observer perception and perform poorly on videos with emerging formats.
A temporal clustering method that combines the Intrinsically Nonlinear Receptive Field (INRF) model with a sophisticated temporal pooling strategy, calculating mean, standard deviation, skew, kurtosis, and maximum differences for each frame to create a video quality assessment metric that outperforms state-of-the-art metrics.
The method achieves a high correlation with human observer scores, approximately 0.9, significantly outperforming existing metrics like INRF-VQA and VMAF, demonstrating its ability to predict visual quality perception accurately.
Smart Images

Figure ES2023070676_22052025_PF_FP_ABST
Abstract
Description
[0001]METHOD FOR AUTOMATIC EVALUATION OF VIDEO QUALITY OF HIGH FRAME RATE (HFR), HIGH DYNAMIC RANGE (HDR) AND ULTRA HIGH DEFINITION 8K (UHD 8K) VIDEOS DESCRIPTION FIELD OF THE INVENTION The object of the invention is a temporal clustering method for the evaluation of video quality, particularly high frame rate (HFR), high dynamic range (HDR) and ultra high definition 8K (UHD 8K) videos, with which we can create a video metric from any arbitrary image metric, surpassing the video quality metrics of the state of the art.BACKGROUND OF THE INVENTION The development of automatic methods for image and video quality assessment that correlate well with the perception of human observers is a very challenging open problem in vision sciences, with numerous practical applications in disciplines such as image processing and computer vision, as well as in the media industry. Image quality assessment methods can be divided into three categories: full-reference methods, which compare an original image with a distorted version of it; reduced-reference methods, which compare some features of the distorted and reference images, since the full reference image is not available; and reference-free methods (also called blind models), which operate only on the distorted image.Full-reference methods constitute the vast majority of image quality approaches, and this is the category to which the present invention belongs. Over the past two decades, the focus of image quality research has been on improving classical metrics by developing models that emulate some aspects of the visual system. While considerable progress has been made, state-of-the-art quality assessment methods are still in use, which share several shortcomings, such as their performance degrading significantly when tested on a database very different from the one used to train them, or their significant limitations in predicting observer scores for videos in emerging formats. This is the case with videos with high frame rates, high dynamic range, or ultra-high definition in 8K.For this type of video content, existing video quality metrics do not perform well. The intrinsically nonlinear receptive field (INRF) model, on which the present invention is based, is a physiologically plausible single-neuron summation model that incorporates the efficient representation principle and can remain constant in situations where linear models must change with input. The model is very powerful in representing nonlinear functions and is consistent with more recent studies on dendritic computations. The INRF equation for the response of a single neuron at location x is: where ^^ and ^^ ^ denote pixel locations, ^^ and ^^ are 2D kernels, ^^ is a scalar, ^^ ^is 2D kernel ^^^^^, ^^^^, ^^ represents a nonlinearity and ^^^^^ denotes pixel values. The model is based on knowledge about neuronal dendrites: some dendritic branches act as nonlinear units, a single nonlinearity is not sufficient to model dendritic computations and there is feedback from the neuronal soma to the dendrites. This feedback is reflected in the changing nonlinearity term^^, expressed by the term ^^ ∗ ^^^^^^, which has the effect of making the nonlinearity change for each contributing location ^^ ^depending on the values at location ^^ and its neighbors. Using a single fixed INRF module where the kernels ^^, ^^ and ^^ are Gaussian in shape and the nonlinearity is a sigmoid power law, and applying it to grayscale images, the model's response emulates the perceived image and can explain several visual perception phenomena that other models cannot, such as: the "crunchy effect", the White illusion under noise or light / dark asymmetry. Since applying the INRF model to a grayscale image produces a result that closely resembles how the image is perceived, a very simple image quality metric has been proposed: given an image ^^, and its distorted version ^^ ^ , the INRF transformation is applied to both images, obtaining ^^ and ^^ ^, and then preferably the root mean square error (RMSE) between the processed images is calculated. This method, known as INRF Image Quality Assessment (INRF-IQA), when tested on a database of natural images has been shown to perform very similarly to the state-of-the-art deep learning perceptual metric LPIPS (Learned Perceptual Image Patch Similarity). Based on this result, a full-reference INRF-based IQA method has been proposed as follows: i) Given a grayscale image I, the INRF transformation applied to it produces an image ^^ whose value at each pixel location ^^ is calculated as: where the kernels ^^, ^^ and ^^ are 2D Gaussians of standard deviations ^^ ^ , ^^ ௪ and ^^ ^, respectively, ^^ is a scalar, ^^ is a sigmoid which has the form of a function ^^^^^^^^, and the symbol ∗ denotes 2D convolution. ii) The comparison image of INRF-IQA value ^^ and its distorted version ^^ ^ It is preferably calculated as: where the INRF transform of an image is calculated as in step i) and RMSE represents the square root of the mean square error. iii) Preferably, the values of the four parameters of the metric, specifically ^^ ^ , ^^ ௪ , ^^ ^and ^^, are chosen to maximize the correlation between the INRF-IQA perceptual distance from Eq. 3 and the mean opinion scores (MOS) of human observers on a large-scale natural image database TID2008. To apply this model to video, it was combined with a very simple temporal clustering strategy: image metrics are applied frame by frame, and then the results are averaged, in what is known as INRF video quality assessment (INRF-VQA). That is, given a reference video. and a distorted version ^^ ^ , the result of the INRF-VQA metric is: INRF െ VQA^^^ோ, ^^^^ ൌ mean^INRF െ IQAௌ൫^^^, ^^^ ^ ൯^ where the average is calculated over the entire duration of the video, ^^ denotes the number of frames,^^ ^ and ^^ ^ ^ are, respectively, the i-th frame and ^^ ^and the IQA INRF metric െ IQAௌ has the same form as in Eq. 3 and the same value for the parameter^^, but the spatial deviations of the kernel size^^ ^ , ^^ ௪ and ^^ ^ are scaled by a factor ^^ which denotes the relationship between the size of the video frames and the size of the images in TID2008 (which were used to optimize ^^^, ^^௪, ^^^). For example, if ^^ is a 2K video with a resolution of 1920x1080 and since the images in TID2008 are 512x384 in size, the scale factor is ^^ ൌ^ଽଶ^ ൌ 3,75.However, due to the simplicity of the temporal pooling strategy (just an average), this method lacks a certain degree of sophistication and could therefore be further improved to achieve better results. VMAF (Video Multi-Method Assessment Fusion...), a state-of-the-art video quality metric, also averages VMAF scores per frame. DESCRIPTION OF THE INVENTION The present invention relates to a temporal pooling method for video quality assessment, particularly High Frame Rate (HFR), High Dynamic Range (HDR), and Ultra High Definition 8K (UHD 8K) videos, with which we can create a video metric from any arbitrary image metric. The invention enables a video quality assessment that correlates remarkably well with the perception of human observers.There are two main features of the present invention that differentiate it from other prior-art methods, especially when combined with the INFR-VQA (Intrinsically Nonlinear Received Field Video Quality Assessment) video quality metric. First, it replaces the time-averaging stage of the INRF-VQA algorithm, described above. Second, it is able to outperform by a very wide margin all prior-art video quality metrics, including INRF-VQA and the VMAF (Video Multi-Method Assessment Fusion) metric, proposed and used by some major streaming services to automatically select their streaming parameters. As for the possible uses of the invention, since image quality assessment is of crucial importance in the media industry, it has numerous practical applications there.For example, streaming services use video quality metrics to optimize resources and achieve savings in energy consumption and internet bandwidth usage, because video quality metrics allow them to maximize compression while ensuring that the video's visual quality is above the required level. Visual quality assessment is also critical in applied disciplines such as image processing and computer vision, where it plays an important role in algorithm development, optimization, and testing.In particular, the present invention replaces the very simple existing temporal averaging methods with a temporal binning technique which now involves computing over all frames not only the mean but also the standard deviation (std), skewness (skew), kurtosis (curt) and maximum differences (from one frame to the next) of arbitrary values of IQA and {IQA^^ ൌ ^IQA^ , IQAଶ ,…, IQAே^, where IQA^ denotes an IQA value of the ith frame of a distorted video ^^. ^ . Therefore, a value of the VQA video quality metric is obtained by the formula: ^^ ^ MaxDiff^IQA^^ ^^^ ^ MaxMinusDiff^IQA^^ ^^^ ^ skew^IQA^^ ^^^ ^ curt^IQA^^where MaxDiff is the maximum value of Diff (increase from one frame to the next), MaxMinusDiff is the maximum value of LeastDiff (decrease (in absolute value) from one frame to the next), and ^^^, ^^, ^^, ^^, ^^, ^^^ are real-valued parameters.The mean, std, skew, curt, Diff, and LeastDiff statistics for^ IQA ^^ are calculated using the following expressions: Diff^IQA^^ ൌ ^IQAଶ െ IQA^, IQAଷ െ IQAଶ,…, IQAே െ IQAேି^^LessDiff^IQA^^ ൌ ^IQA^ െ IQAଶ, IQAଶ െ IQAଷ,…, IQAேି^ െ IQAே^ where y ^^ is the number of frames. Preferably, the value of the parameters ^^, ^^, ^^, ^^, ^^, ^^ could be calculated by following the steps of: - preparing a dataset containing a reference video and its distorted version ^^ ^ and a subjective experiment result comprising the mean opinion score (MOS) and 95% confidence interval (95CI) of MOS for ^^ ^ , - calculate the mean statistics, estd, bias, curt, MaxDif and MaxMinusDif of {IQA^^ ൌ ^IQA^ , IQAଶ ,…, IQAே^ , where IQA^ ൌ IQA ( ^^^, ^^^ ^ ) for every ith frame of ^^ ^; - splitting the dataset into a test set and an auxiliary set - randomly separating the auxiliary set into a training set and a validation set; - obtaining the parameters ^^ to ^^ from the expression that minimizes the sum of the insensitive squared errors epsilon (∑^ ^^ଶ୰୰୭୰ ^^^^ ) where ^^ ୰୰୭୰^^^^ ൌ max ^0, ห൫^^^^^^ െiMOS^^^^൯ห െ 95CI^^^^^ for all ^^^ videos ^^ in the training set,where iMOS ൌ 6 െ MOS; - calculating the Pearson linear correlation coefficient (PLCC) between iMOS and ^^ for the validation set, in order to validate the performance of ^^; - repeating a predefined number of times the steps of random separation of the auxiliary group, obtaining parameters ^^ to ^^ that minimize the sum of the insensitive squared error epsilon and calculating the Pearson linear correlation coefficient (PLCC), - setting the parameters of ^^ to ^^ to the mean of the results obtained from the training set having a PLCC>0.8 in the step of calculating the Pearson linear correlation coefficient (PLCC), and - evaluating the performance of ^^ using PLCC and Spearman's rank order correlation coefficient (SRCC) for the test set. In preferred embodiments, IQA could be INRF (intrinsically nonlinear receptive field) - IQA, and an INRF-IQA value for images ^^ and ^^^ can be obtained by: INRF െ IQA൫^^^, ^^ ^ ^^ ൯ ൌ RMSE൫INRF^^^ ^, INRF^^^^ ^^൯ where the INRF value൫^^ ^ ൯ and INRF൫^^ ^ ^ ൯ at each pixel location ^^ is: where ^^, ^^ and ^^ are kernels with 2D Gaussians of standard deviations ^^ ^ , ^^ ௪ and ^^ ^, respectively, ^^ is a scalar, ^^ is a sigmoid having the form of a function ^^^^^^^^, and the symbol ∗ is a 2D convolution, The invention also relates to a computer program for video quality evaluation configured to carry out the defined method and to a storage unit configured to store said computer program. BRIEF DESCRIPTION OF THE DRAWINGS In order to complement the description made and to help towards a better understanding of the characteristics of the invention, according to a preferred example of practical embodiment thereof, a set of drawings are attached as an integral part of said description in which, with illustrative and non-limiting character, the following are presented: Figure 1.- Shows a table showing the Pearson Linear Correlation Coefficient as well as the Spearman Rank Order Correlation Coefficient values obtained from the VMAF, VMAF in conjunction with the invention, and INRF-IQA tests in conjunction with the invention. Figure 2.- Shows the correspondence between the MOS (Mean Opinion Score) values (vertical axis) of human observers and the metric values obtained by the same three aforementioned methods (horizontal axis). Figure 3.- Shows a schematic view of a preferred embodiment of the invention. PREFERRED EMBODIMENT OF THE INVENTION A preferred embodiment of the temporal clustering method for assessing video quality, particularly high frame rate (HFR), high dynamic range (HDR), and ultra high definition 8K (UHD 8K) videos, object of the present invention, is described below with the help of Figures 1 and 2.This preferred embodiment of the invention primarily focuses on combining the novel temporal clustering method of the present invention with image quality (IQA) metrics known in the art. For example, the INRF (Intrinsically Nonlinear Receptive Field) method and the VMAF (Video Multi-Method Assessment Fusion...) method. In a first particular embodiment, the method comprises a first step of providing a reference video. and its distorted version ^^ ^ (for example, a compressed video), where ^^ ^ and ^^ ^ ^ are, respectively, the i-th frame of ^^ ோ and ^^ ^ A second stage consists of applying ^^ and ^^ to the images ^ an INRF (intrinsically nonlinear receptive field) transformation and generate two images INRF൫^^^൯ and INRF൫^^ ^ ^ ൯ whose value at each pixel location ^^ is: where ^^, ^^ and ^^ are kernels with 2D Gaussians of standard deviations ^^ ^ , ^^ ௪ and ^^ ^ , respectively, ^^ is a scalar, ^^ is a sigmoid which has the form of a function^^^^^^^^, and the symbol ∗ is a 2D convolution. Here, the scale factor ^^ ^ for ^^^, ^^௪ and^^ ^ is set to 4 for 8K videos viewed at 0.75H (times the screen height). Then, an INRF-IQA value of the ^^ and ^^^ images is preferably obtained by applying the following formula: Finally, a video quality assessment (VQA) is obtained by applying the following formula: Where estd denotes the standard deviation, skew denotes the skewness, curt denotes the kurtosis, MaxDiff is the maximum value of Diff (the increase from one frame to the next), MaxMinusDiff is the maximum value of LeastDiff (the decrease in absolute value from one frame to the next), and ^^^, ^^, ^^, ^^, ^^, ^^^ are real-valued parameters. To obtain the values of the parameters ^^, ^^, ^^, ^^, ^^, ^^, the method may comprise additional steps. First, a dataset consisting of several reference videos may be prepared. (e.g. Sequence1, Sequence2,…), each of which has several distorted versions ^^ ^(e.g. for Sequence 1 there are distorted versions encoded with different bit rates, where higher bit rates imply smaller distortions). For example, the results of the subjective experiment might comprise the Mean Opinion Score (MOS) from 1 to 5 and the 95% Confidence Interval (95CI) of MOS, as shown in the following table: MOS is a measure used in the field of quality of experience and telecommunications engineering, which represents the overall quality of a stimulus or system. It is the arithmetic mean of all individual values on a predefined scale that a subject assigns to their opinion of a system's quality performance. These ratings are usually collected in a subjective quality assessment test, but can also be estimated algorithmically. Then, for each frame of the compressed videos, an image quality metric is calculated, e.g., INRF-IQA or VMAF. If the metric is a full-reference metric, the metric value for the i-th IQA frame ୧ It can be calculated from the i-th frame of the original and compressed videos. For each compressed video, the six types of statistics for ^IQA^^ ൌ ^IQA^, IQAଶ,…,IQA ே ^ are preferably calculated as: Diff^IQA^^ ൌ ^IQAଶ െ IQA^, IQAଷ െ IQAଶ,…, IQAே െ IQAேି^^LessDiff^IQA^^ ൌ ^IQA^ െ IQAଶ, IQAଶ െ IQAଷ,…, IQAேି^ െ IQAே^ , where ^^^^^^^^^ ^ ^ ^= ^∑ ^^^^^^ ^ ^ ^^^ଶே^ ^ െ ^^^^^^^^^^ ^^^^^^where ^^ is the number of frames. The datasets are then preferably split into "test" (approximately 20% of all data) and "other" which are used for training and validation, while avoiding the same sequence from being separated into "test" and "other". The "other" category is then randomly split into "training" and "validation" (approximately 20% of "other"), while avoiding the same sequence from being separated into "training" and "validation". The parameters ^^ to ^^ are then preferably obtained from: which minimize the sum of the insensitive squared errors epsilon ∑ ଶ^ ^^^^^^^ ^^^^ where for all ^^^ videos ^^ in "training". Subsequently, the performance of ^^ is preferably validated by calculating the Pearson linear correlation coefficient (PLCC) between iMOS and ^^. This cross-validation is repeated for multiple trials, e.g., 500 times. Finally, the parameters ^^ to ^^ are preferably set to the mean of the trained parameters that show PLCC>0.8, and the performance of ^^ is evaluated using the PLCC and Spearman's rank-order correlation coefficient (SRCC) for "testing". The explained parameters have been optimized to maximize the method's performance on a database of HFR, HDR, and 8K videos. An outstanding correlation with human observer scores, approximately 0.9, has been obtained, while INRF-VQA and other commonly used methods usually obtain approximately 0.6.This is highly significant because a correlation of 1 means that the two phenomena considered behave identically (i.e., understanding one determines the other), and thus the excellent correlation of the present method with observer scores means that the invention is able to predict very well how observers perceive the visual quality of videos. As an example, Figure 1 shows an outstanding correlation for both Pearson's Linear Correlation Coefficient and Spearman's Rank-Order Correlation Coefficient, with scores of 0.926 and 0.917, respectively, easily outperforming prior-art alternatives. Specifically, the combination of the method with the VMAF metric proves to perform exceptionally well.Regarding Figure 2, it shows the correspondence between the MOS (Mean Opinion Score) values (vertical axis) from human observers and the metric values obtained by the same three aforementioned methods (horizontal axis). In order to easily compare them (i.e., all graphs are aligned from bottom left to top right), the vertical axis of the present invention was set to the inverse MOS (iMOS), where iMOS ൌ 6 െ MOS. Finally, Figure 3 shows a schematic view of a preferred embodiment of the invention. Here, a dataset (1) containing a reference video is shown. and its distorted version ^^ ^ and a subjective experiment result comprising the MOS and 95CI of MOS for ^^ ^. From this dataset (1), the training and validation data are transferred to a training / validation splitting medium (2). After splitting, the training data are transferred to training media (3) and the validation data are transferred to validation media (5), which also make use of the ^^ a ^^ values, which are the output of the training media (3). The results of this validation (PLCC(MOS, V) and ^^ a ^^ ) are stored in a validation results storage (6). The last stage of the training and validation process for testing parameters is to choose the most suitable values by making use of parameter decision media (7). The testing data from the dataset (1) and the chosen parameter values are transferred to temporal clustering media (8).Finally, the result ^^ is evaluated with the evaluation means (9), using the MOS and the 95CI of MOS for ^^. ^ , producing the evaluation results (10) (PLCC (MOS, V) and RMSE (MOS, V)).
Claims
CLAIMS n method of evaluating video quality comprising the steps of: i) providing a reference video ^^ ோ and its distorted version ^^ ^ , where ^^ ^ and ^^ ^ ^ are, respectively, the i-th frame of ^^ ோ and ^^ ^ , ii) apply to ^^ ^ and ^^ ^ ^ an IQA image quality metric that measures the level of degradation of ^^ ^ ^ relative to ^^ ^ and that generates an IQA metric value ^ ൌ I QA^^^^, ^^^ ^ ^ for each i-th frame iii) obtain a value of the VQA video quality metric by applying the temporal grouping defined by the formula: d onde MáxDif es el valor máximo de Dif , incremento de un fotograma con with respect to the next one, MaxMinusDif is the maximum value of MinusDif, decrease e n valor absoluto de un fotograma con respecto al siguiente, y ^^^, ^^, ^^, ^^, ^^, ^^^ son real-valued parameters; and en donde las estadísticas de ^IQA^^ ൌ ^IQA^, IQAଶ,…, IQAே^ se calculan como: ∑ I ^ ^ ^ QA media^IQA ^ ൌ ^^ M enosDif^IQA^^ ൌ ^IQA^ െ IQAଶ, IQAଶ െ IQAଷ,…, IQAேି^ െ IQAே^ where y ^^ es el número de fotogramas. 2.- The method of claim 1, further comprising the step of obtaining los valores de parámetros ^^, ^^, ^^, ^^, ^^, ^^, que comprende las fases de: - prepare a dataset containing a reference video and its distorted version ^^ ^ and a subjective experiment result comprising the mean opinion score (MOS) and 95% confidence interval (95CI) of MOS for ^^ ^ , - calculate the statistics of mean, std, skew, curt, MaxDif and MaxMinusDif of { IQA^^ ൌ ^IQA^, IQAଶ,…, IQAே^; - split the data set into a test group and an auxiliary group - randomly separate the auxiliary group into a training group and a validation group, - obtain the parameters ^^ a ^^ from the expression that minimizes the sum of the error c uadrático insensible épsilon (∑ ଶ ^ ^^ ୰୰୭୰ ^^^^ ) en donde ^^ ୰୰୭୰^^^^ ൌ máx ^0, ห൫^^^^^^ െ iMOS^^^^൯ห െ 95CI^^^^^ y iMOS ൌ 6 െ MOS para todos ^^^ los vídeos ^^ en el grupo training, V QA^^^ ^ ^ ^ ோ, ^^^ ൌ ^^ ^ media^IQA ^ ^ ^^ ^ estd^IQA ^ ^ ^^ ^ máx൫Dif^IQA^^൯ ^ ^^ ^ máx൫iDif^IQA^^൯ ^ ^^ ^ ^^^^^^^^^^^IQA^^ ^ ^^ ^ curt^IQA^^ - calculate the Pearson linear correlation coefficient (PLCC) between iMOS and ^^ para el grupo de validación, con el fin de validar el desempeño de VQA^^^, ^^^^ - repeating the steps of random separation of the auxiliary group a predefined number of times, obtaining the parameters ^^ to ^^ that minimize the sum of the insensitive squared errors epsilon and calculating the Pearson linear correlation coefficient (PLCC); - setting the parameters ^^ to ^^ to the average of the results obtained from the training group having a PLCC>0.8 in the step of calculating the Pearson linear correlation coefficient (PLCC), and - evaluar el desempeño de ^^ usando PLCC y el coeficiente de correlación de Spearman rank order (SRCC) for the test group. 3.- The method of claims 1 and 2, wherein IQA is INRF (intrinsically non-linear received field) - IQA, and an INRF-IQA value for the images ^^ and ^^ ^ can be obtained by I NRF െ IQA൫^^^, ^^ ^ ^ ൯ ൌ RMSE൫INRF^^^^^, INRF^^^^ ^ ^൯ where the INRF value൫^^ ^ ൯ and INRF൫^^ ^ ^ ൯ at each pixel location ^^ is: where ^^, ^^ and ^^ are kernels with 2D Gaussians of standard deviations ^^ ^ , ^^ ௪ and ^^ ^ , respectively, ^^ is a scalar, ^^ is a sigmoid having the form of a function ^^^^^^^^, and the symbol ∗ is a 2D convolution.
4. A computer program for video quality assessment configured to perform the steps of the method according to claims 1 to 3.
5. A storage unit configured to store the computer program according to claim 4.