Techniques for predicting video quality across different viewing parameters

Through the trained perceptual quality model, combined with the characteristics in the visual quality metric, approximate the relationship between PPD value and human perception, the problem of inaccurate perceptual quality scores in the prior art is solved, and accurate prediction of different viewing parameters and better streaming experience are achieved.

CN119968843APending Publication Date: 2025-05-09NETFLIX INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380069685.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-30
Filing Date
2023-09-27
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art cannot accurately handle changes in different viewing parameters when predicting the perceived video quality of reconstructed videos, resulting in inaccuracy of perceived quality scores, affecting the quality/bit rate trade-off and streaming experience during the encoding process.

Method used

Using a trained perceptual quality model, which combines features in visual quality metrics, approximates the relationship between unit angle pixels (PPD) values ​​and human perception, and is able to generate different perceptual quality scores for different combinations of display resolutions and standardized viewing distances.

Benefits of technology

By using trained perceptual quality models, the perceived visual quality score of the reconstructed video can be accurately predicted over a wider range of viewing parameters, improving the quality/bit rate trade-off during the encoding process and the viewer's streaming experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119968843A_ABST
    Figure CN119968843A_ABST
Patent Text Reader

Abstract

In various embodiments, a quality reasoning application estimates a perceived video quality of a reconstructed video. The quality reasoning application computes a set of feature values corresponding to a set of visual quality metrics based on the reconstructed frame, the source frame, the display resolution, and the standardized viewing distance. The quality reasoning application executes a trained perceptual quality model on a set of feature values to generate a perceptual quality score indicating a perceptual visual quality level of the reconstructed frame. The quality reasoning application performs one or more operations associated with the encoding process based on the perceived quality score.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of U.S. patent application serial number 17 / 937,033, filed on September 30, 2022, which is incorporated herein by reference. Technical Field

[0003] Various embodiments relate generally to computer science and video coding techniques, and more particularly to techniques for predicting video quality across different viewing parameters. Background Art

[0004] Efficiently and accurately encoding video content is an important aspect of real-time streaming of high-quality video. Typically, when an encoded version of a video is streamed to a playback device, the encoded video content is decoded to generate a reconstructed video that is played back on the playback device. In order to increase the degree of compression and correspondingly reduce the size of the encoded video, the encoder typically performs a lossy data compression algorithm to eliminate certain selected information. As a general matter, eliminating information during the encoding process may cause visual impairment or "distortion", which reduces the overall visual quality of the reconstructed video derived from the encoded video.

[0005] Since the number and type of distortion introduced during the encoding process can vary, quality control is usually implemented to ensure that the visual quality of the reconstructed video perceived by the actual audience of the reconstructed video is at an acceptable level. This type of visual quality is often referred to as "perceived video quality" or "perceived video quality". However, manually verifying the perceived video quality of the reconstructed video may lead to inaccurate evaluations, and may also be very time-consuming and expensive. Therefore, sometimes some form of automatic perceptual video quality assessment is integrated into the video encoding and transmission process. For example, when a given video is encoded to generate an encoded version of the video (the encoded version of the video attempts to achieve the highest predicted level of overall visual quality during the playback of the video for the corresponding total bit rate), automatic perceptual video quality assessment can be used. In another example, when evaluating different encoders / decoders, automatic perceptual video quality assessment can be used to estimate the number of bits used by different data compression algorithms to achieve certain perceptual video quality levels.

[0006] One method of automatically assessing perceived video quality involves computing feature values ​​for different features that are input into a conventional perceptual quality model based on the reconstructed video and the associated source video. The feature values ​​are then mapped to a perceptual quality score by the conventional perceptual quality model that estimates the human perceived video quality of the reconstructed video. The conventional perceptual quality model is trained based on assessments of the perceived visual quality provided by humans viewing the reconstructed training videos according to a set of "training" viewing parameters. Two examples of training viewing parameters are a digital resolution height of 1080 pixels and a normalized viewing distance (NVD) that is three times the physical height of the display screen.

[0007] One disadvantage of the above method is that the trained conventional perceptual quality model cannot accurately predict the perceptual video quality of the reconstructed video content when the viewing parameters associated with the actual viewing experience are different from the viewing parameters used during training. For example, as the digital resolution height and / or NVD associated with the actual viewing experience increases, the number of pixels or "pixels per angle" (PPD) within each angle of the viewer's viewing angle also increases. Within the range of PPD values ​​typically associated with the actual viewing experience, as the PPD value increases, the distortion in the reconstructed video content becomes less noticeable to the viewer, which should increase the perceptual video quality of the reconstructed video content. However, for a given actual viewing experience, if the feature values ​​of the relevant conventional perceptual quality model are not calculated based on the PPD value associated with the actual viewing experience, the perceptual quality score ultimately calculated by the trained conventional perceptual quality model is only consistent with the PPD value used during training, rather than the PPD value associated with the actual viewing experience. In this case, the accuracy of the perceptual quality score calculated by the trained conventional perceptual video quality model may be significantly reduced. Inaccurate perceptual quality scores may result in erroneous visual quality / bitrate tradeoffs during the encoding process, which in turn may result in encoded videos having unnecessarily low overall visual quality levels being streamed to actual viewers for playback. Given the relatively large number of combinations of digital resolution heights and normalized viewing distances that can exist within the range of PPD values ​​typically associated with actual viewing experiences, it is currently impossible to properly train conventional perceptual quality models over these ranges of PPD, which exacerbates the negative impacts on the encoding process and subsequent streaming experience described above.

[0008] As previously stated, what is needed in the art are more efficient techniques for predicting the perceptual quality of reconstructed video. Summary of the invention

[0009] One embodiment describes a computer-implemented method for estimating the perceived video quality of a reconstructed video. The method includes calculating a first set of feature values ​​corresponding to a set of visual quality metrics based on a first reconstructed frame, a first source frame, a first display resolution, and a first normalized viewing distance; executing a trained perceptual quality model on the first set of feature values ​​to generate a first perceptual quality score indicating a first perceptual visual quality level of the first reconstructed frame; and performing one or more operations associated with an encoding process based on the first perceptual quality score.

[0010] At least one technical advantage of the disclosed technology over the prior art is that, with the disclosed technology, a trained perceptual quality model can be used to predict more accurate perceptual visual quality scores for reconstructed video over a wider range of viewing parameters than can be achieved using prior art methods. Specifically, unlike features of conventional perceptual quality models, the disclosed technology incorporates features in a trained perceptual quality model that approximate a fundamental relationship between PPD values ​​and human perception of the quality of reconstructed video content. Thus, the trained perceptual quality model can be used to generate a different perceptual quality score for each different combination of display resolution and normalized viewing distance over a range of PPD values ​​typically associated with an actual viewing experience. Thus, implementing the trained perceptual quality model during the encoding process can improve the quality / bitrate tradeoff typically made when encoding video content, and thereby improve the overall streaming experience for the viewer. These technical advantages provide one or more technical advances over prior art methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order that the above-mentioned features of various embodiments may be understood in detail, a more particular description of the inventive concept briefly summarized above may be made by reference to various embodiments, some of which are illustrated in the accompanying drawings. However, it should be noted that the drawings illustrate only typical embodiments of the inventive concept and are therefore not to be considered as limiting the scope in any way and that there are other equally effective embodiments.

[0012] Figure 1 is a conceptual illustration of a system configured to implement one or more aspects of various embodiments;

[0013] Figure 2 According to various embodiments Figure 1 A more detailed illustration of one of the feature engines of ;

[0014] Figure 3 According to various embodiments Figure 2 A more detailed illustration of one of our visual information fidelity (VIF) extractors;

[0015] Figure 4 According to various embodiments Figure 2A more detailed illustration of one of the PPD-perceptual detail loss metric (DLM) extractors of ;

[0016] Figure 5 is a flow chart of method steps for generating a trained perceptual quality model accounting for NVD and display resolution according to various embodiments; and

[0017] Figure 6 is a flow chart of method steps for estimating the perceptual video quality of a reconstructed video using a trained perceptual quality model according to various embodiments. DETAILED DESCRIPTION

[0018] In the following description, numerous specific details are set forth to provide a more thorough understanding of various embodiments. However, it will be apparent to one skilled in the art that the inventive concept may be practiced without one or more of these specific details. For purposes of explanation, multiple instances of similar objects are represented by reference numerals identifying the object and alphanumeric characters in parentheses identifying the instance when necessary.

[0019] A typical video streaming service provides access to a library of videos that can be viewed on a range of different playback devices, each of which is typically connected to the video streaming service under different connections and network conditions. In order to efficiently transmit the video to the playback device, the video streaming service provider encodes the video and then streams the resulting encoded video to the playback device. Each playback device decodes the encoded video data stream and displays the resulting reconstructed video to the viewer. In order to reduce the size of the encoded video, the encoder typically utilizes lossy data compression techniques that eliminate selected information. Typically, eliminating information during encoding may result in visual quality impairment or "distortion," which may reduce the visual quality of the reconstructed video derived from the encoded video.

[0020] Because the amount and type of distortion introduced when encoding a video varies, video streaming services typically implement quality control to ensure that the visual quality of the reconstructed video as perceived by actual viewers ("perceived video quality") is acceptable. In practice, because manually evaluating the perceptual video quality of a reconstructed video can be very time consuming, some video streaming services integrate conventional perceptual quality models that estimate the perceptual video quality of the reconstructed video into the video encoding and transmission process. For example, some video streaming services use conventional perceptual quality models to set the degree(s) of compression when encoding a video to ensure a target perceptual video quality level during playback of the associated reconstructed video content.

[0021] One drawback of using conventional perceptual quality models to estimate the perceptual video quality of reconstructed videos is that conventional perceptual quality models cannot accurately predict the perceptual video quality of reconstructed video content when viewing parameters associated with actual viewing experience differ from the viewing parameters used during training. In this regard, each conventional perceptual quality model is trained based on evaluations of the perceived visual quality provided by people viewing the reconstructed training videos according to different sets of "training" viewing parameters. For example, one perceptual quality model can be trained based on standard laptop-like full high-definition (FHD) viewing conditions, where the NVD is 3 (e.g., the viewing distance from the display is three times the physical height of the display) and the resolution of the display is 1920 pixels by 1080 pixels.

[0022] It is not feasible to obtain the large number of human-observed visual quality scores that are required to properly train different conventional perceptual quality models for the combinations of DRH and NVD that are typically associated with actual viewing experiences. As a result, typical video encoding and transmission processes use a very small number of conventional perceptual quality models to generate perceptual quality scores that accurately predict the perceptual video quality for the corresponding set of trained viewing parameters, but inaccurately predict the perceptual video quality for many actual viewers. Inaccurate perceptual quality scores may lead to incorrect visual quality / bitrate tradeoffs during the encoding process, which in turn may result in encoded videos with unnecessarily low overall visual quality levels being streamed to actual viewers for playback.

[0023] However, using the disclosed techniques, a training application generates a single trained machine learning model that takes into account different viewing parameters when predicting perceptual quality scores. In one embodiment, the training application computes feature vectors corresponding to subjective quality scores in two subjective quality score groups associated with different combinations of DRH and NVD. In the subjective quality score groups, each subjective quality score is associated with a different reconstructed video sequence.

[0024] For each subjective quality score, the training application configures the feature engine to compute a feature vector based on the associated reconstructed video sequence, the corresponding source video sequence, the associated DRH, and the associated NVD. It is worth noting that each feature vector is a set of feature values ​​corresponding to a set of features including the VIF index at four NVD-specific spatial scales, the DLM modified to consider PPD, and the conventional temporal information metric.

[0025] The VIF index is an image quality metric that can be used to estimate the visual quality of a "reconstructed" frame of a reconstructed video sequence at an associated spatial scale based on an approximation of the amount of information shared between the reconstructed frame and the corresponding source frame as perceived by the human visual system (HVS). In order to approximate the multi-scale effect of NVD on visual quality quantified by the VIF index, the feature engine calculates the value of the VIF index at four spatial scales, each of which varies in the same direction as the NVD. In other words, the feature engine varies the four spatial scales based on the NVD.

[0026] DLM is an image quality metric that can be used to estimate the visual quality of a reconstructed frame based on an approximation of the loss of useful visual information or "detail loss" in the reconstructed frame relative to the corresponding source frame. In order to model the impact of NVD and DRH quantified by DLM on visual quality, the feature engine implements a modified version of the conventional DLM (referred to herein as "PPD-aware DLM", which takes PPD into account when approximating detail loss). More specifically, the PPD-aware DLM approximates the impact of PPD on the visibility of video distortion. For example, as PPD increases, distortion visibility tends to decrease, and the overall perceived quality increases.

[0027] The training engine performs a machine learning algorithm on the feature vectors and the corresponding subjective quality scores to generate a trained perceptual quality model that maps feature vectors corresponding to a set of feature values ​​to perceptual quality scores via a learning function. Subsequently, the quality reasoning application uses the trained perceptual quality model to calculate any number of overall perceptual quality scores, where each overall perceptual quality score corresponds to a different combination of reconstructed video sequence, DRH, and NVD.

[0028] In some embodiments, to calculate the overall perceptual quality score, the quality reasoning application configures the feature engine to calculate different feature vectors for each reconstructed frame of the reconstructed video sequence based on the reconstructed frame, the corresponding source frame, the DRH, and the NVD. For each reconstructed frame, the quality reasoning application inputs the associated feature vector into the trained perceptual quality model. In response, the trained perceptual quality model outputs a perceptual quality score for the reconstructed frame. The quality reasoning application calculates the overall perceptual quality score for the reconstructed video sequence based on the perceptual quality scores for the reconstructed frames.

[0029] At least one technical advantage of the disclosed techniques over the prior art is that, with the disclosed techniques, a single trained perceptual quality model can be used to accurately predict different perceived visual quality levels of a reconstructed video sequence when viewed according to different combinations of DRH and NVD. Thus, implementing the trained perceptual quality model during encoding can improve the quality / bitrate tradeoff typically made when encoding video content, and thereby improve the overall streaming experience for viewers. These technical advantages provide one or more technical advances over prior art approaches.

[0030] System Overview

[0031] Figure 1 is a conceptual illustration of a system 100 configured to implement one or more aspects of various embodiments. As shown, in some embodiments, system 100 includes, but is not limited to, computing instance 110(1), computing instance 1102(2), trained perceptual quality model 158, viewing parameter group 120(1), reconstructed video sequence group 128(1), source video sequence group 126(1), subjective quality score group 102(1), viewing parameter group 120(2), reconstructed video sequence group 128(2), source video sequence group 126(2), subjective quality score group 102(2), viewing parameter group 120(3), reconstructed video sequence 170, source video sequence 160, overall perceptual quality score 190(1), viewing parameter group 120(4), and overall perceptual quality score 190(2).

[0032] In some embodiments, system 100 may include, but is not limited to, any number of other computation instances, any number of other viewing parameter groups, any number of other reconstructed video sequence groups, any number of other source video sequence groups, any number of subjective quality score groups, any number of other reconstructed videos, any number of other source videos, and any number of overall perceptual quality scores. In some other embodiments, system 100 may omit computation instance 110(1) or computation instance 110. In the same or other embodiments, system 100 may omit viewing parameter group 120(1), reconstructed video sequence group 128(1), source video sequence group 126(1), subjective quality score group 102(1) and / or viewing parameter group 120(2), reconstructed video sequence group 128(2), source video sequence group 126(2), and subjective quality score group 102(2). In some embodiments, system 100 may omit viewing parameter set 120(3), reconstructed video sequence 170, source video sequence 160, and overall perceptual quality score 190(1) and / or viewing parameter set 120(4) and overall perceptual quality score 190(2).

[0033] Any number of components of system 100 may be distributed across multiple geographic locations or implemented in any combination in one or more cloud computing environments (e.g., packaged shared resources, software, and data). In some embodiments, each of computing instance 110(1), computing instance 110(2), and zero or more other computing instances may be implemented in a cloud computing environment, or as part of any other distributed computing environment, or in a standalone manner.

[0034] As shown, computing instance 110(1) includes, but is not limited to, processor 112(1) and memory 116(1), and computing instance 110(2) includes, but is not limited to, processor 112(2) and memory 116(2). For purposes of explanation, computing instance 110(1) and computing instance 110(2) are also referred to herein individually as "computing instance 110" or collectively as "computing instance 110". Processor 112(1) and processor 112(2) are also referred to herein individually as "processor 112" or collectively as "processor 112". Memory 116(1) and memory 116(2) are also referred to herein individually as "memory 116" or collectively as "memory 116".

[0035] Processor 112 may be any instruction execution system, device, or apparatus capable of executing instructions. For example, processor 112 may be a central processing unit, a graphics processing unit, a controller, a microcontroller, a state machine, or any combination thereof. Memory 116 of computing instance 110 stores content such as software applications and data for use by processor 112 of computing instance 110. Memory 116 may be one or more of readily available memories, such as random access memory, read-only memory, floppy disk, hard disk, or any other form of local or remote digital storage.

[0036] In some other embodiments, each computing instance 110 may include any combination of any number of processors 112 and any number of memories 116. In particular, any number of computing instances 110 (including one) and / or any number of other computing instances may provide a multi-processing environment in any technically feasible manner.

[0037] In some embodiments, storage devices (not shown) may supplement or replace memory 116 of computing instance 110. The storage devices may include, but are not limited to, any number and type of external memory accessible to processor 112 of computing instance 110. For example, and without limitation, the memory may include a secure digital card, an external flash memory, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0038] In some embodiments, each computing instance 110 can be integrated into a user device along with any number and / or type of other devices (e.g., one or more other computing instances and / or I / O devices). Some examples of user devices include, but are not limited to, desktop computers, laptops, smartphones, smart TVs, game consoles, and tablet computers.

[0039] Typically, each computing instance 110 is configured to implement one or more software applications. For purposes of explanation only, each software application is described as residing in memory 116 of a single computing instance (e.g., computing instance 110(1) or computing instance 110(2)) and executing on processor 112 of a single computing instance. In some embodiments, any number of instances of any number of software applications may reside in memory 116 and in any number of other memories associated with any number of other computing instances, and execute on processor 112 of computing instance 110 and any number of other processors associated with any number of other computing instances in any combination. In the same or other embodiments, the functionality of any number of software applications may be distributed across any number of other software applications residing in memory 116 and in any number of other memories associated with any number of other computing instances, and execute on processor 112 and any number of other processors associated with any number of other computing instances in any combination. In addition, subsets of the functionality of multiple software applications may be combined into a single software application.

[0040] Specifically, in some embodiments, computing instance 110(2) is configured to estimate a level of perceived video quality of a reconstructed video sequence derived from a source video sequence. As used herein, an estimate of the "perceived video quality" of a reconstructed video sequence refers to an estimate of the visual quality of the reconstructed video sequence as perceived by a typical viewer of the reconstructed video sequence during playback of the reconstructed video sequence.

[0041] Source video sequences include, but are not limited to, sequences of one or more still images, often referred to as "frames". Some examples of source video sequences include, but are not limited to, a single frame, a shot, any portion (including all) of a feature-length movie, any portion (including all) of an episode of a television program, and any portion of a music video. As used herein, frames in a "shot" typically have similar spatiotemporal characteristics and run continuously for a period of time. Source video sequences are also referred to herein as "source video".

[0042] Each source video sequence can be in any number and / or type of video format and / or conform to any number and / or type of video standard. Two different source video sequences can be in the same or different video formats and / or conform to the same or different video standards. For example, two source video sequences can be in the same high dynamic range (HDR) format, a third source video sequence can be in a different HDR format, and a fourth source video sequence can be in a different standard dynamic range (SDR) format. Each frame in each source video sequence can be in any number and / or type of video format and / or conform to any number and / or type of video standard.

[0043] The reconstructed video sequence is an approximate reconstruction of the corresponding source video sequence, and includes, but is not limited to, a different "reconstructed" frame for each "source" frame in the source video sequence. The reconstructed video sequence can be derived from the corresponding source video sequence in any technically feasible manner. In some embodiments, the source video sequence is encoded to generate an encoded video sequence (not shown), and the encoded video sequence is then decoded to generate the corresponding reconstructed video sequence. In this way, the reconstructed video sequence approximates the corresponding source video sequence transmitted to the viewer via the encoding and streaming infrastructure and the playback device. The reconstructed video sequence is also referred to herein as a "reconstructed video sequence".

[0044] For explanation purposes, the resolution of the reconstructed video is the same as the resolution of the corresponding source video. In some embodiments, the "original" source video can be downsampled and / or upsampled to generate one or more other source videos with different resolutions. Therefore, multiple reconstructed videos can be derived from different source videos, which in turn are derived from a single original source video.

[0045] As previously described, in some conventional systems, a conventional perceptual quality model trained according to a single "training" set of viewing parameters is used to estimate the perceptual video quality of a reconstructed video sequence. Some examples of viewing parameters include, but are not limited to, NVD and digital resolution height (DRH). In some embodiments, NVD is equal to the viewing distance (expressed in units of length) divided by the physical height of the display (expressed in the same units of length). The viewing distance can be expressed as a multiple of "H", where H represents the physical height of the display. For example, a viewing distance of "3H" specifies that the viewing distance from the display is three times the physical height of the display, so the NVD is 3. As used herein, "display" refers to a portion of any type of display device that provides visual output.

[0046] To estimate the perceptual video quality of the reconstructed video using a conventional perceptual quality model, feature values ​​of different features of the conventional perceptual quality model are calculated based on the reconstructed video sequence and the corresponding source video sequence. The feature values ​​are then input into the conventional perceptual quality model. In response, the conventional perceptual quality model calculates and outputs a conventional perceptual quality score that estimates the perceptual video quality of the reconstructed video sequence.

[0047] However, none of the feature values ​​of conventional perceptual quality models are calculated based on viewing parameters associated with actual viewing experiences. Therefore, if the PPD of a given actual viewing experience is different from the PPD associated with the training parameters, the perceptual quality score ultimately calculated by the trained conventional perceptual quality model is consistent only with the PPD used during training, rather than with the PSD associated with the actual viewing experience. In this case, the accuracy of the perceptual quality score calculated by the trained conventional perceptual video quality model may be significantly reduced. Inaccurate perceptual quality scores may lead to incorrect visual quality / bitrate tradeoffs during the encoding process, which in turn may result in encoded videos with unnecessary low overall visual quality levels being streamed to actual viewers for playback. Given that a relatively large number of combinations of DRH and NVD can exist within the range of PPD values ​​typically associated with actual viewing experiences, it is currently impossible to correctly train conventional perceptual quality models within these ranges of PPD values, which exacerbates the negative impact on the above-mentioned encoding process and subsequent streaming experience.

[0048] Considering viewing parameters when estimating perceived video quality

[0049] To address the above issues, the system 100 includes, but is not limited to, a training application 130 that trains an untrained machine learning model (not shown) to consider the effects of different viewing parameters associated with different viewing experiences when estimating the perceptual video quality of a reconstructed video sequence. The resulting trained version of the machine learning model is also referred to herein as a trained perceptual quality model 158 and a unified perceptual quality model. Subsequently, in some embodiments, one or more instances of a quality reasoning application (not explicitly shown) can use the trained perceptual quality model 158 to consider different viewing parameters when estimating the perceptual video quality of a reconstructed video sequence.

[0050] As shown, in some embodiments, the training application 130 resides in the memory 116(1) of the computing instance 110(1) and executes on the processor 112(1) of the computing instance 110(1). As shown, in some embodiments, the training application 130 generates a trained perceptual quality model 158 based on the following items: the viewing parameter group 120(1), the reconstructed video sequence group 128(1), the source video sequence group 126(1), the subjective quality score group 102(1), the viewing parameter group 120(2), the reconstructed video sequence group 128(2), the source video sequence group 126(2), and the subjective quality score group 102(2).

[0051] As shown, in some embodiments, viewing parameter group 120(1) includes, but is not limited to, DRH 122(1) and NVD 124(1). DRH is also referred to herein as "display resolution". In some other embodiments, DRH 122(1) and / or any number of other DRHs may be replaced with any number and / or type of display resolutions, and the techniques described herein may be modified accordingly.

[0052] Source video sequence group 126(1) includes, but is not limited to, a plurality of source video sequences having a vertical resolution equal to DRH 122(1). Reconstructed video sequence group 128(1) includes, but is not limited to, C different reconstructed video sequences (not shown) having a vertical resolution equal to DRH 122(1), where C may be any positive integer. The reconstructed video sequences included in reconstructed video sequence group 128(1) are derived from the source video sequences included in source video sequence group 126(1).

[0053] The reconstructed video sequences included in the reconstructed video sequence set 128(1) may be derived from the source video sequences included in the source video sequence set 126(1) in any technically feasible manner. For example, in some embodiments, each source video sequence included in the source video sequence set 126(1) is encoded based on one or more sets of encoding parameters (not shown) to generate one or more encoded video sequences. The encoded video sequences are then decoded to generate the reconstructed video sequences included in the reconstructed video sequence set 128(1).

[0054] As shown, in some embodiments, the subjective quality score group 102(1) includes, but is not limited to, subjective quality scores 104(1)-104(C), each of which is associated with a viewing parameter group 120(1) and a different reconstructed video sequence included in the reconstructed video sequence group 128(1). The subjective quality score group 102(1) is generated based on human-assigned individual quality scores (not shown) that specify the visual quality level of the reconstructed video sequences included in the reconstructed video sequence group 128(1) when viewed according to the viewing parameter group 120(1). The individual quality scores and the subjective quality scores 104(1)-104(C) may be determined in any technically feasible manner.

[0055] In some embodiments, individual quality scores are assigned by human participants in a subjective quality experiment. During the subjective quality experiment, the participants view reconstructed video sequences included in the reconstructed video sequence set 128(1) from a distance of the NVD 124(1) via a display having a DRH 122(1). The participants assign individual quality scores that rate the visual quality of the reconstructed video sequences. The participants may evaluate and rate the visual quality of the reconstructed video sequences based on any type of rating system.

[0056] The subjective quality scores 104(1)-104(C) may be generated based on the individual quality scores in any technically feasible manner. In some embodiments, each of the subjective quality scores 104(1)-104(C) is set equal to an average or "mean opinion score" of the individual quality scores of the associated reconstructed video sequence. In some other embodiments, the subjective quality scores 104(1)-104(C) are generated based on any type of subjective data model that takes into account the individual quality scores of the associated reconstructed video sequences.

[0057] As shown, in some embodiments, viewing parameter set 120(2) includes, but is not limited to, DRH 122(2) and NVD 124(2). In some embodiments, viewing parameter set 120(2) is different from viewing parameter set 120. As shown in italics, in some embodiments, DRH 122(2) and NVD 124(2) are 2160 pixels and 1.5, respectively, while DRH 122(1) and NVD 124(1) are 1080 pixels and 3.0, respectively.

[0058] Source video sequence group 126(2) includes, but is not limited to, a plurality of source video sequences having a vertical resolution equal to DRH 122(2). Reconstructed video sequence group 128(2) includes, but is not limited to, (TC) different reconstructed video sequences (not shown) having a vertical resolution equal to DRH 122(2), where T may be any positive integer greater than C. The reconstructed video sequences included in reconstructed video sequence group 128(2) are derived from the source video sequences included in source video sequence group 126(2).

[0059] As shown, in some embodiments, the subjective quality score group 102(2) includes, but is not limited to, subjective quality scores 104(C+1)-subjective quality scores 104(T), each of which is associated with a viewing parameter group 120(2) and a different reconstructed video sequence included in the reconstructed video sequence group 128(2). The subjective quality score group 102(2) is generated based on human-assigned individual quality scores (not shown) that specify the visual quality level of the reconstructed video sequences included in the reconstructed video sequence group 128(2) when viewed according to the viewing parameter group 120(2). The individual quality scores and the subjective quality scores 104(C+1)-subjective quality scores 104(T) may be determined in any technically feasible manner.

[0060] For purposes of explanation, subjective quality score 104(1)-subjective quality score 104(C) and subjective quality score 104(C+1)-subjective quality score 104(T) are also referred to herein individually as "subjective quality score 104", or collectively as "subjective quality score 104" and "subjective quality scores 104(1)-104(T)".

[0061] In some embodiments, each source video sequence in the set of source video sequences 126(1) and each reconstructed video sequence in the set of reconstructed video sequences 128(1) has a resolution of 1920 pixels x 1080 pixels and is in SDR format. The subjective quality score set 102(1) is derived from individual quality scores collected during an "SDR FHD laptop" experiment. Each individual quality score assigned during the SDR FHD laptop experiment is an assessment of the quality of the reconstructed video sequence in the set of reconstructed video sequences 128(1) when viewed on a laptop display having a resolution of 1920 pixels x 1080 pixels at a viewing distance of 3 times the physical height of the laptop display. In contrast, each source video sequence in the set of source video sequences 126(2) and each reconstructed video sequence in the set of reconstructed video sequences 128(2) has a resolution of 3840 pixels x 2160 pixels and is in HDR format. The subjective quality score group 102(2) is derived from individual quality scores collected during an "HDR 4K Home Cinema" experiment. Each individual quality score assigned during the HDR 4K Home Cinema experiment is an assessment of the quality of a reconstructed video sequence 128(2) in the group of reconstructed video sequences when viewed on a television display having a resolution of 3840 pixels x 2160 pixels at a viewing distance of 1.5 times the physical height of the television display.

[0062] As shown, in some embodiments, the training application 130 includes, but is not limited to, a feature engine 140(1), a feature engine 140(2), a feature pooling engine 134, and a training engine 150. The feature engine 140(1) and the feature engine 140(2) are two instances of a single software application, referred to herein as the feature engine 140.

[0063] In some embodiments, each instance of feature engine 140 calculates a different set of feature vectors for each of one or more reconstructed video sequences based on (one or more) reconstructed video sequences, (one or more) corresponding source video sequences, and a viewing parameter set. Each feature vector set includes, but is not limited to, a different feature vector for each reconstructed frame of the associated reconstructed video sequence. For example, if the reconstructed video sequence will include 8640 frames, the feature vector set associated with the reconstructed video sequence will include, but is not limited to, 8640 feature vectors.

[0064] In some other embodiments, each feature engine 140 can calculate (one or more) feature vectors at any granularity level, and the techniques described herein are modified accordingly. For example, in some embodiments, the feature engine 140 calculates a single feature vector for each reconstructed video sequence, regardless of the total number of frames included in the reconstructed video sequence.

[0065] Each feature vector includes but is not limited to the different values ​​of each feature contained in the feature set ( Figure 1 ). The value of a feature is also referred to herein as a "feature value". Each feature is a quantifiable metric that can be used to evaluate at least one aspect of the visual quality associated with the reconstructed video content. In some embodiments, the feature set includes, but is not limited to, any number and / or type of features associated with any quantitative aspect of video quality, wherein at least one feature is calculated based on at least one of NVD, DRH, or any other display resolution. As used herein, "video quality" refers to the level of visual quality associated with a reconstructed video sequence or any other video sequence.

[0066] The features included in the feature vectors may be determined in any technically feasible manner based on any number and / or type of criteria. In some embodiments, the features included in the feature vectors are selected empirically to provide valuable insights into the visual quality of the reconstructed video sequences included in the reconstructed video sequence group 128(1) and the reconstructed video sequence group 128(2). In the same or other embodiments, the features included in the feature set are selected empirically to provide insights into the impact of any number and / or type of artifacts on the perceived video quality. For example, the selected features may provide insights into blocking, stair noise, color bleeding, and flickering on the perceived visual quality, but are not limited to such.

[0067] In some embodiments, the feature set includes, but is not limited to, temporal features and any number of objective image quality metrics, each of which exhibits strengths and weaknesses. To exploit strengths and mitigate weaknesses, the feature set includes, but is not limited to, multiple objective image quality metrics with complementary strengths.

[0068] As below combined Figure 2 Described in more detail, in some embodiments, the feature set includes, but is not limited to, six features represented herein as NVD_VIF1, NVD_VIF2, NVD_VIF3, NVD_VIF4, PPD_DLM, and TI. For purposes of explanation, the values ​​corresponding to NVD_VIF1, NVD_VIF2, NVD_VIF3, NVD_VIF4, PPD_DLM, and TI are also referred to herein as NVD_VIF1 value, NVD_VIF2 value, NVD_VIF3 value, NVD_VIF4 value, PPD_DLM value, and TI value, respectively.

[0069] The features NVD_VIF1, NVD_VIF2, NVD_VIF3, and NVD_VIF4 are VIF indices at four spatial scales (not explicitly shown), where the four spatial scales are determined based on NVD. In some embodiments, the feature engine 140 calculates NVD_VIF1 values, NVD_VIF2 values, NVD_VIF3 values, and NVD_VIF4 values ​​based on the reconstructed frame, the corresponding source frame, and the NVD.

[0070] The feature denoted PPD_DLM is a PPD-aware DLM. Unlike the value of a conventional version of the DLM associated with some conventional perceptual quality models, the feature engine 140 calculates the PPD_DLM value based on the reconstructed frame, the corresponding source frame, the PPD value and optionally the NVD.

[0071] The feature representing TI is a temporal information metric that captures temporal distortions associated with and / or causing motion, quantified by differences between adjacent pairs of reconstructed frames included in the reconstructed video sequence.

[0072] As shown, in some embodiments, the feature engine 140 (1) calculates feature vector groups 132 (1)-132 (C) based on the reconstructed video sequence group 128 (1), the source video sequence group 126 (1) and the viewing parameter group 120 (1). Each of the feature vector groups 132 (1)-132 (C) includes, but is not limited to, different feature vectors associated with the subjective quality scores 104 (1)-104 (C) for each reconstructed frame included in the reconstructed video sequence. The feature groups associated with the reconstructed frames include feature values, which are also individually referred to as "frame feature values" herein. Multiple feature values ​​associated with the same reconstructed frame, different reconstructed frames, or any combination thereof are also collectively referred to as "frame feature values" herein.

[0073] In the same or other embodiments, feature engine 140(2) calculates feature vector group 132(C+1)-feature vector group 132(T) based on reconstructed video sequence group 128(2), source video sequence group 126(2), and viewing parameter group 120(2). Each of feature vector group 132(C+1)-feature vector group 132(T) includes, but is not limited to, a different feature vector associated with subjective quality score 104(C+1)-subjective quality score 104(T) for each reconstructed frame included in the reconstructed video sequence.

[0074] As shown, in some embodiments, the training application 130 inputs the feature vector group 132 (1) - the feature vector group 133 (C) and the feature vector group 134 (C+1) - the feature vector group 132 into the feature pooling engine 134. In response, the feature pooling engine 134 generates and outputs the overall feature vector 136 (1) - the overall feature vector 136 (T), respectively. For the purpose of explanation, the overall feature vector 136 (1) - the overall feature vector 138 (T) are also individually referred to as "overall feature vectors 136" or collectively referred to as "overall feature vectors 136" and "overall feature vectors 136 (1) - 136 (T)" herein.

[0075] Each of the overall feature vectors 136 is a set of different overall feature values ​​for a feature vector. In some embodiments, the overall feature vectors 136(1)-136(T) are sets of overall feature values ​​associated with the reconstructed video sequences corresponding to the subjective quality scores 104(1)-104(T), respectively. As used herein, feature vectors associated with a reconstructed video sequence are also referred to herein as "overall feature values."

[0076] The feature pooling engine 134 may calculate the overall feature vector 136(j) based on the feature vector group 132(j) in any technically feasible manner, where j is an integer from 1 to T. In some embodiments, the feature pooling engine 134 sets each overall feature value in the overall feature vector 136(j) to be equal to the arithmetic mean of the associated feature values ​​in the feature vectors included in the feature vector group 132(j). In this manner, each overall feature value in the overall feature vector 136(j) represents the average feature value of the associated feature over the frames included in the reconstructed video sequence associated with the subjective quality score 104(j).

[0077] As shown, in some embodiments, the training engine 150 generates a trained perceptual quality model 158 based on the overall feature vectors 136(1)-136(T) and the subjective quality scores 104(1)-104(T). For purposes of explanation, for an integer j from 1 to T, the jth "training sample" refers to a pair of the overall feature vector 136(j) and the subjective quality score 104(j). Thus, the 1st to Tth training samples include the overall feature vectors 136(1)-136(T) and the subjective quality scores 104(1)-104(T), respectively.

[0078] The training application 130 can execute any number and / or type of supervised machine learning algorithms on the training samples to train the untrained machine learning model to generate a trained machine learning model. As described above, the trained perceptual quality model 158 constitutes a trained machine learning model.

[0079] The trained perceptual quality model 158 maps a feature vector corresponding to a feature set, a portion of a reconstructed video, and a viewing parameter set to a perceptual quality score of the portion of the reconstructed video when viewed according to the viewing parameter set. Both the untrained machine learning model and the trained perceptual quality model 158 are associated with the same feature set.

[0080] Some examples of supervised machine learning algorithms that can be implemented with various embodiments include, but are not limited to, support vector regression (SVR), artificial neural network algorithms (e.g., multilayer perceptron regression), tree-based regression algorithms, and tree-based ensemble methods (e.g., random forest algorithms or gradient boosting algorithms). The training application can use SVR, artificial neural network algorithms, or tree-based regression algorithms to generate a trained SVR model, a trained artificial neural network, or a trained regression tree, respectively. The SVR model is also referred to herein as a trained support vector machine (SVM).

[0081] In one embodiment, the training application 130 uses SVR to train or "fit" an SVR model to a hyperplane based on an optimization objective of maximizing the number of training samples within a threshold value (usually denoted as ε) of the hyperplane. The hyperplane defines a "learning" function based on a subset of training samples, learning weights, learning biases, and an optional kernel function that calculates the inner product between two feature vectors in a suitable feature space. In some embodiments, the SVR model implements a radian basis function (RBF) kernel, which can be represented using the following equation:

[0082] K(x, x′ )= exp( -γ∥ x - x′ ∥2 ) (1)

[0083] In equation (1), x and x′ represent two feature vectors, and γ is a hyperparameter of the SVR model that defines the extent of the influence of a single training sample.

[0084] Before the training application 130 begins executing the SVR algorithm on the training samples, the SVR model is referred to as an "untrained" SVR model. After the training application 130 completes executing the SVR algorithm on the training samples, the SVR model is referred to as a "trained" SVR model. The trained SVR model constitutes the trained perceptual quality model 158. The training application 130 can configure any type of SVR to train any type of SVR model based on the overall feature vectors 136(1)-136(T) and the subjective quality scores 104(1)-104(T) in any technically feasible manner.

[0085] Specifically, continuing with the same example, the training application 130 executes a Scikit-Learn implementation of SVR to train a Scikit-Learn implementation of SVM using an RBF kernel (commonly referred to as an RBF kernel SVM). The following article describes several Scikit-Learn implementations of SVR, SVM, and kernel functions in more detail:

[0086] https: / / scikit-learn.org / stable / modules / svm.html#regression

[0087] After completing the above steps, the trained perceptual quality model 158 implements a learning function denoted herein as score(x), which calculates a perceptual quality score based on the feature vector x and can be represented using the following equation:

[0088]

[0089] In equation (2), T is the total number of training samples, x1-x T Respectively represent the overall feature vectors 136(1)-136(T), y1-y T They represent the subjective quality scores 104(1)-104(T) respectively. T is the learning weight, b is the learning bias, K(x,x j ) is a kernel function (e.g., the RBF kernel defined by Equation (1)). For integer j from 1 to T, if α j is equal to zero, the jth training sample (i.e., the pair of x(j) and y(j)) is not used to calculate the perceptual quality score. Otherwise, the jth training sample is also referred to as a "support vector" and is used to calculate the perceptual quality score. As will be appreciated by those skilled in the art, equation (2) defines the perceptual quality score corresponding to the reconstructed video sequence and the viewing parameter group as a weighted combination of the feature vectors corresponding to the reconstructed video sequence, the viewing parameter group, and the feature group associated with the trained perceptual quality model 158. Therefore, equation (2) defines the perceptual quality score as the weighted sum of the feature values ​​included in the corresponding feature vectors.

[0090] Advantageously, the trained perceptual quality model 158 is able to estimate the impact of different DRH and NVD (and therefore different PPD values) on visual quality as perceived by actual human viewers of the reconstructed video. In this regard, since the feature vectors are calculated based on DRH and NVD, the trained perceptual quality model 158 can calculate significantly different perceptual quality scores corresponding to the same reconstructed video sequence and different DRH and / or NVD. Therefore, unlike prior art perceptual quality models, the trained perceptual quality model 158 can accurately estimate the perceptual video quality level of the reconstructed video for a variety of viewing parameters.

[0091] As shown, in some embodiments, the training application 130 transmits the trained perceptual quality model 158 to the quality reasoning application 180(1) and the quality reasoning application 180(2). The quality reasoning application 180(1) and the quality reasoning application 180(2) are two instances of a single software application, which is also referred to herein as the "quality reasoning application 180" (not explicitly shown). In some embodiments, instead of or in addition to transmitting the trained perceptual quality model 158 to the quality reasoning application 180(1) and / or the quality reasoning application 180(2), the training application 130 may transmit the trained perceptual quality model 158 to any number of other instances of the quality reasoning application. In the same or other embodiments, the training application 130 may transmit the trained perceptual quality model 158 to any number and / or type of software applications. In some embodiments, the training application 130 may store the trained perceptual quality model 158 in any number and / or type of available memory.

[0092] As shown, in some embodiments, quality reasoning application 180(1) and quality reasoning application 180(2) reside in memory 116(2) of computing instance 110(2) and execute on processor 112(2) of computing instance 110(2). For purposes of explanation, the functionality of each instance of quality reasoning application 180 is described herein in the context of quality reasoning application 180(1).

[0093] As shown, in some embodiments, quality reasoning application 180(1) calculates an overall perceptual quality score 190(1) based on viewing parameter set 120(3), source video sequence 160, and reconstructed video sequence 170. Overall perceptual quality score 190(1) predicts the perceived video quality of reconstructed video sequence 170 when viewed according to viewing parameter set 120(3).

[0094] Viewing parameter group 120(3) includes, but is not limited to, DRH 122(3) and NVD 124(3). In some embodiments, as shown in italics, DRH 122(3) is 1080 pixels and NVD 124(3) is 4.5. It is worth noting that viewing parameter group 120(3) is different from both viewing parameter group 120(1) and viewing parameter group 120(2). Source video sequence 160 includes, but is not limited to, source frame 162(1)-source frame 162(F), where F can be any positive integer. Reconstructed video sequence 170 includes, but is not limited to, reconstructed frame 172(1)-reconstructed frame 172(F), where F can be any positive integer.

[0095] For purposes of explanation, source frames 162(1)-source frames 162(F) are also referred to herein individually as “source frames 162”, or collectively as “source frames 162” and “source frames 162(1)-162(F)”. For purposes of explanation, reconstructed frames 172(1)-reconstructed frames 172(F) are also referred to herein individually as “reconstructed frames 172”, or collectively as “reconstructed frames 172” and “reconstructed frames 172(1)-172(F)”.

[0096] Reconstructed frames 172(1)-172(F) are approximations of source frames 162(1)-162(F), respectively, derived in any technically feasible manner. For example, in some embodiments, source frames 162(1)-162(F) are encoded to generate source coded frames, and the source coded frames are decoded to generate reconstructed frames 172(1)-172(F), respectively.

[0097] As shown, in one embodiment, the quality reasoning application 180(1) includes, but is not limited to, a feature engine 140(0), a trained perceptual quality model 158, and a score pooling engine 186. The feature engine 140(0) is an instance of the feature engine 140 described previously herein in conjunction with the training application 130. In some embodiments, the feature engine 140(0) computes feature vectors 182(1)-182(T) based on source frames 162(1)-162(F), respectively, reconstructed frames 172(1)-172(F), respectively, and a viewing parameter set 120(3).

[0098] For purposes of explanation, feature vectors 182(1)-182(T) are also referred to herein individually as “feature vectors 182”, or collectively as “feature vectors 182” and “feature vectors 182(1)-182(F)”. Each feature vector 82 includes, but is not limited to, different feature values ​​for each feature in the feature set associated with the trained perceptual quality model 158.

[0099] As shown, the quality reasoning application 180(1) inputs the feature vectors 182(1)-182(F) into the trained perceptual quality model 158. In response, the trained perceptual quality model 158 calculates and outputs perceptual quality scores 184(1)-perceptual quality scores 184(F). For purposes of explanation, the perceptual quality scores 184(1)-perceptual quality scores 184(F) are also referred to herein individually as "perceptual quality scores 184", or collectively as "perceptual quality scores 184" and "perceptual quality scores 184(l)-184(F)".

[0100] As shown, the quality reasoning application 180(1) includes multiple instances of the trained perceptual quality model 158, and the trained perceptual quality model 158 inputs the feature vectors 182(1)-182(F) into any number of instances of the trained perceptual quality model 158 sequentially, simultaneously, or any combination thereof. For example, in some embodiments, the quality reasoning application 180(1) simultaneously inputs the feature vectors 182(1)-182(F) into F instances of the trained perceptual quality model 158. In response, the F instances of the trained perceptual quality model 158 simultaneously output perceptual quality scores 184(1)-184(F).

[0101] As shown, quality reasoning application 180(1) inputs perceptual quality scores 184(1)-184(F) into score pooling engine 186. In response, score pooling engine 186 generates and outputs an overall perceptual quality score 190(1). Score pooling engine 186 may calculate overall perceptual quality score 190(1) based on perceptual quality scores 184(1)-184(F) in any technically feasible manner.

[0102] In some embodiments, the score pooling engine 186 performs any number and / or type of temporal pooling operations on the perceptual quality scores 184(1)-184(F) to calculate the overall perceptual quality score 190(1). For example, in some embodiments, the score pooling engine 186 sets the overall perceptual quality score 190(1) equal to the arithmetic mean of the perceptual quality scores 184(1)-184(F). Thus, the overall perceptual quality score 190(1) represents the average perceptual video quality over the frames included in the reconstructed video sequence when viewed according to the viewing parameter set 120(3).

[0103] In some embodiments, the score pooling engine 186 performs any number and / or type of hysteresis pooling operations that simulate a relatively smooth variance of human opinion scores in response to changes in video quality. For example, in some embodiments, the score pooling engine 186 can perform a linear low-pass operation and a non-linear (order) weighting operation on the perceptual quality scores 184(1)-184(F) to calculate the overall perceptual quality score 190(1).

[0104] As shown, quality reasoning application 180(1) sends overall perceptual quality score 190(1) and / or any number of perceptual quality scores 184(1)-184(F) to any number and / or type of software applications. In the same or other embodiments, quality reasoning application 180(1) stores overall perceptual quality score 190(1) and / or any number of perceptual quality scores 184(1)-184(F) in any number and / or type of available memory.

[0105] In some embodiments, quality reasoning application 180(1) may aggregate any number of perceptual quality scores corresponding to any number of reconstructed frames and viewing parameter sets 120(1) in any technically feasible manner to generate any number and / or type of aggregated perceptual quality scores. For example, in some embodiments, quality reasoning application 180(1) generates S per-shot perceptual quality scores for S different shots included in reconstructed video sequence 170.

[0106] In some other embodiments, the quality reasoning application 180(1) may input any number of feature vectors associated with any level of granularity into the trained perceptual quality model 158, and the techniques described herein are modified accordingly. For example, in some embodiments, the quality reasoning application 180 inputs the feature vectors 182(1)-184(F) in combination with the training application 130 into an instance of the feature pooling engine 134 previously described herein. In response, the feature pooling engine 134 outputs an overall feature vector (not shown). The quality reasoning application 180(1) then inputs the overall feature vector into the trained perceptual quality model 158. In response, the trained perceptual quality model 158 outputs an overall perceptual quality score.

[0107] As shown, in some embodiments, quality reasoning application 180(2) calculates an overall perceptual quality score 190(2) based on viewing parameter set 120(4), source video sequence 160, and reconstructed video sequence 170. Overall perceptual quality score 190(2) predicts the perceived video quality of reconstructed video sequence 170 when viewed according to viewing parameter set 120(4).

[0108] Viewing parameter group 120(4) includes, but is not limited to, DRH 122(4) and NVD 124(4). In some embodiments, as shown in italics, DRH 122(4) is 1080 pixels and NVD 124(3) is 2.0. It is worth noting that viewing parameter group 120(4) is different from viewing parameter group 120(1), viewing parameter group 120(2), and viewing parameter group 120(3).

[0109] As shown in italics, in some embodiments, the overall perceptual quality score 190(1) and the overall perceptual quality score 190(2) corresponding to the reconstructed video sequence 170 when viewed according to the viewing parameter set 120(3) and the viewing parameter set 120(4) are 93 and 77, respectively.

[0110] Advantageously, unlike some prior art perceptual quality models, because the quality reasoning application 180 can use the trained perceptual quality model 158 to calculate different perceptual quality scores based on different viewing parameters, better decisions can be made based on the perceptual quality scores, which can increase the overall quality of experience when streaming video.

[0111] In this regard, any number of perceptual quality scores (e.g., perceptual quality scores 184(1)-184(F), overall perceptual quality score 190(1), and overall perceptual quality score 190(2)) can be used in various encoding process operations. For example, when encoding a source video sequence, any number of perceptual quality scores can be used to calculate a tradeoff between quality and bitrate. In another example, any number of perceptual quality scores can be used to evaluate and / or fine-tune an encoder, a decoder, an adaptive bitrate streaming algorithm, or any combination thereof.

[0112] In yet another example, the quality reasoning application 180 may be configured to calculate an overall perceptual quality score corresponding to a plurality of reconstructed video sequences derived from different encoded versions of the source video sequence 160 when viewed according to different viewing parameter groups. For each viewing parameter group, a different bitrate ladder may be generated based on the overall perceptual quality score corresponding to the viewing parameter group. Each bitrate ladder specifies a subset of encoded versions of the source video sequence 160 corresponding to a different viewing parameter group, an associated bitrate, and an associated overall perceptual quality score. Subsequently, a client application (e.g., executing an adaptive bitrate streaming algorithm) may use the bitrate ladder corresponding to the source video sequence 160 and the viewing parameter group that best matches the actual viewing parameters to select one of the encoded versions of the source video sequence 60 and / or switch between them for streaming and playback based on bitrate.

[0113] Although not shown, in some embodiments, the quality reasoning application 180(1) offsets and / or clips the perceptual quality scores 184(1)-184(F) or the overall perceptual quality score 190(1) based on a set of PPD-specific offsets that reflect subjective quality scores corresponding to different combinations of reconstructed video clips, DRHs, and NVDs. In some embodiments, each PPD-specific offset in the set of PPD-specific offsets is a negative or positive integer corresponding to a different PPD value.

[0114] Although not shown, in one embodiment, quality reasoning application 180(1) calculates the actual PPD value based on DRH 122(3) and NVD 124(3). More specifically, quality reasoning application 180(1) uses the following in combination with Figure 2 The actual PPD value is calculated using equation (4) described in more detail. In other embodiments, other technically feasible techniques for calculating the actual PPD value may be implemented. The quality reasoning application 180(1) selects a PPD specific offset from the set of PPD specific offsets based on the actual PPD value. The quality reasoning application 180(1) adds the PPD specific offset to the overall perceptual quality score 190(1) to generate an "absolute" perceptual visual quality score aligned on the PPD value.

[0115] In some embodiments, PPD values ​​are calculated for subjective quality scores corresponding to different combinations of reconstructed video clips, DRH, and NVD. Based on a subset of subjective quality scores corresponding to the highest quality video sequence for the PPD value of 120 and the highest quality image sequence for the lower PPD values, a quality degradation percentage is calculated between the PPD of 120 and each lower PPD value. Based on the quality degradation percentage and zero or more constraints, a curve of the highest perceptual quality score compared to the PPD value is generated. For example, two constraints can be maximum perceptual quality scores of 100 and 95 for PPD values ​​of 120 and 60, respectively. For each PPD value, a PPD specific offset is set equal to the highest perceptual quality score of the PPD value minus 100. In other embodiments, other technically feasible techniques for calculating PPD specific offsets can be implemented.

[0116] In some embodiments, both the source video sequence 160 and the reconstructed video sequence 170 are in HDR format. In some other embodiments, both the source video sequence 160 and the reconstructed video sequence 170 are in SDR format. In general, the quality reasoning application 180 or any other software application can use the trained perceptual quality model 158 to calculate an overall perceptual quality score for any reconstructed video sequence, having any vertical resolution and any horizontal resolution, in any number and / or types of formats, and associated with any number and / or types of video standards.

[0117] Note that the techniques described herein are illustrative rather than restrictive and may be modified without departing from the broader spirit and scope of the invention. Many modifications and variations to the functionality provided by the training application 130, feature engine 140, feature pooling engine 134, training engine 150, trained perceptual quality model 158, quality reasoning application 180, and score pooling engine 186 will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.

[0118] It should be understood that the system 100 shown herein is illustrative and that variations and modifications are possible. For example, in some embodiments, the training application 130 can apply any number and / or type of machine learning algorithms to any number and / or type of training samples representing any number of features at any level of granularity to generate any type of trained machine learning model implementing any type of video quality metric. In addition, Figure 1 The connection topology between the various components in the training application 130 can be modified as needed. For example, in some embodiments, feature engine 140 (1) and feature engine 140 (2) reside and execute in a cloud environment instead of residing in training application 130.

[0119] Feature vector calculation based on visual parameters

[0120] Figure 2 According to various embodiments Figure 1 A more detailed illustration of one of the feature engines 140 is provided for explanation purposes. Figure 1 The feature engine 140(0) is combined in the context of Figure 2 The functionality of the feature engine 140 is described, which computes one or more feature vectors corresponding to a feature group, one or more reconstructed frames, DRH, and NVD.

[0121] In some embodiments, feature engine 140(0) is an instance of feature engine 140 that computes feature vectors 182(1)-182(F) based on reconstructed frames 172(1)-172(F), source frames 162(1)-162(F), DRH 122(3), and NVD 124(3). In the same or other embodiments, feature vectors 182(1)-182(F) correspond to feature group 202, reconstructed frames 172(1)-172(F), DRH 122(3), and NVD 124(3). As shown in the figure, in some embodiments, the feature engine 140 (0) includes but is not limited to the feature vector 182, the feature group 202, the VIF engine 240 (1)-VIF engine 240 (4), the VIF kernel size adjustment engine 230, the PPD value 208, the PPD-aware DLM extractor 260 (1)-PPD-aware DLM extractor 260 (F), and the TI extractor 270.

[0122] In some embodiments, each feature vector 182 includes, but is not limited to, different feature values ​​for each feature in the feature set 202. As shown, in some embodiments, the feature set 202 includes, but is not limited to, six features represented herein as NVD_VIF1, NVD_VIF2, NVD_VIF3, NVD_VIF4, PPD_DLM, and TI. In some other implementations, the feature set 202 may include any number and / or type of features related to any number of video quality aspects, wherein at least one of the features is calculated based on at least one of NVD, DRH, or any other display resolution.

[0123] The features denoted as NVD_VIF1, NVD_VIF2, NVD_VIF3, NVD_VIF4 are visual information fidelity (VIF) indices at four NVD-specific spatial scales (not explicitly shown) corresponding to scale indices (not shown) of 1-4, respectively. In this article, the VIF index at the NVD-specific spatial scale is also referred to as the "NVD-perceived VIF index" at the spatial scale. Unlike the VIF index associated with some conventional perceptual quality models, the NVD-perceived VIF index indirectly considers NVD via the associated spatial scale.

[0124] As will be appreciated by those skilled in the art, the VIF index is an image quality metric that estimates the visual quality of a distorted image at an associated spatial scale based on an approximation of the amount of information shared between a reference image and the distorted image as perceived by the HVS. As used herein, an "image quality metric" is a visual quality metric that is used to estimate the visual quality of a distorted image (e.g., a reconstructed frame).

[0125] In some embodiments, VIF extractor 250 (not explicitly described) implements a pixel domain version of the VIF index. Figure 3 Described in more detail, in some embodiments, the VIF extractor 250 calculates the VIF value based on the reconstructed frame, the corresponding source frame, and the VIF kernel size that inherently defines the NVD specific spatial scale. More specifically, in some embodiments, the VIF extractor 250 generates a Gaussian 2D kernel with a size and standard deviation proportional to the VIF kernel size. The VIF extractor 250 convolves the reconstructed frame with the Gaussian 2D kernel to generate a filtered reconstructed image, and convolves the source frame with the Gaussian 2D kernel to generate a filtered source image. The VIF extractor 250 then estimates the portion of the information in the filtered source image that the HVS can extract from the filtered reconstructed image to calculate the VIF value of the reconstructed frame at the NVD specific spatial scale.

[0126] As shown, in some embodiments, the VIF kernel resizing engine 230 calculates VIF kernel sizes 238(1)-VIF kernel sizes 238(4) based on the NVD 124(3). For purposes of explanation, the VIF kernel sizes 238(1)-VIF kernel sizes 238(4) are also referred to herein individually as "VIF kernel sizes 238", or collectively as "VIF kernel sizes 238" and "VIF kernel sizes 238(1)-238(4)". In some embodiments, the VIF kernel sizes 238(1)-230(4) correspond to scaling indices 1-4, respectively. In the same or other embodiments, the VIF kernel sizes 238(1)-230(4) inherently define NVD-specific spatial scales corresponding to NVD_VIF1, NVD_VIF2, NVD_VIF3, NVD_VIF4, respectively, for the NVD 124(3).

[0127] The VIF kernel size engine 230 may determine the VIF kernel size 238(1)-238(4) consistent with the VIF index in any technically feasible manner. In some embodiments, the VIF kernel size adjustment engine 230 implements the following equation to calculate the VIF kernel size (denoted as K) corresponding to the NVD and the VIF scale index (denoted as idx):

[0128] K = roundUpToOdd( (2^(5-idx) + 1) * NVD / 3.0 ) (3)

[0129] As shown, in some embodiments, VIF kernel resizing engine 230 calculates VIF kernel sizes 238(1)-238(4) based on scaling exponents of 1-4 and NVD 124(3), respectively. For example, in some embodiments, VIF kernel resizing engine 230 calculates VIF kernel sizes 238(1)-238(4) of 27, 15, 9, and 5, respectively, based on the example value 4.5 (in italics) of NVD 124(3).

[0130] It is noteworthy that, according to equation (3), as NVD changes, each VIF kernel size 238 changes in the same direction. And those skilled in the art will recognize that as each VIF kernel size 238 changes, the smoothness associated with both the filtered reconstructed image and the filtered source image changes in the same direction. Therefore, NVD_VIF1, NVD_VIF2, NVD_VIF3, and NVD_VIF4 accurately reflect that, within the range of actual NVDs typically associated with actual viewing experience, as the actual NVD increases, the distortion in the reconstructed frame becomes less noticeable to the viewer.

[0131] In some embodiments, VIF engine 240(1)-VIF engine 240(4) are different instances of VIF engine 140 (not explicitly shown) that configure one or more instances of VIF extractor 250 to calculate NVD_VIF1 value, NVD_VIF2 value, NVD_VIF3 value, and NVD_VIF4 value, respectively. As shown, in some embodiments, VIF engine 240(1) calculates NVD_VIF1 value 252(1)-NVD_VIF1 value 252(F) based on VIF kernel size 238(1), reconstructed frames 172(1)-172(F), and source frames 162(1)-162(F). In the same or other embodiments, VIF engine 240(2) calculates NVD_VIF2 values ​​254(1)-NVD_VIF2 values ​​254(F) based on VIF kernel size 238(2), reconstructed frames 172(1)-172(F), and source frames 162(1)-162(F). In some embodiments, VIF engine 240(3) calculates NVD_VIF3 values ​​256(1)-NVD_VIF3 values ​​256(F) based on VIF kernel size 238(3), reconstructed frames 172(1)-172(F), and source frames 162(1)-162(F). In the same or other embodiments, VIF engine 240(4) calculates NVD_VIF4 values ​​258(1)-NVD_VIF4 values ​​258(F) based on VIF kernel size 238(4), reconstructed frames 172(1)-172(F), and source frames 162(1)-162(F).

[0132] For purposes of explanation, the VIF engine 240 is described herein in the context of the VIF engine 240(1). As shown, in some embodiments, the VIF engine 240(1) includes, but is not limited to, VIF extractors 250(1)-VIF extractors 250(F). The VIF extractors 250(1)-VIF extractors 250(F) are different instances of the VIF extractor 250. As described above, the VIF extractor 250 calculates VIF values ​​based on the reconstructed frame, the corresponding source frame, and the VIF kernel size that inherently defines a specific spatial scale of the NVD. Figure 3 The VIF extractor 250 is described in more detail.

[0133] As shown, in some embodiments, VIF engine 240(1) simultaneously calculates NVD_VIF1 values ​​252(1)-NVD_VIF1 values ​​252(F) based on VIF kernel size 238(1), reconstructed frames 172(1)-172(F), and source frames 162(1)-162(F), respectively. More generally, in some embodiments, VIF engine 240(1) may configure any number of instances of VIF extractor 250 to simultaneously, sequentially, or any combination thereof, calculate NVD_VIF1 values ​​252(1)-NVD_VIF1 values ​​252(F).

[0134] The feature denoted as PPD_DLM is the PPD perceived detail loss metric (DLM). As will be appreciated by those skilled in the art, DLM is part of an image quality metric known as the "Additive Distortion Metric" ("ADM"). ADM is based on the premise that humans respond differently to detail loss in a distorted image and to additional impairments in a distorted image. As used herein, "detail loss" refers to the loss of useful visual information in a distorted image relative to a corresponding reference image. Detail loss may have a negative impact on content visibility, thereby affecting perceived visual quality. As used herein, "additional impairment" refers to redundant visual information that is present in a distorted image but not in a corresponding reference image. Additional impairments may distract the viewer.

[0135] As below combined Figure 4 As described in more detail, according to the DLM, additional impairments are removed from the distorted image to generate a restored image. Based on the estimated response of the HVS to the restored image and the corresponding reference image, a DLM value is calculated that estimates the percentage of perceived visual information loss associated with the distorted image. The PPD-perceptual DLM is a version of the DLM that is modified to account for different PPD values ​​when estimating the response of the HVS to the restored image and the corresponding reference image. As described above, the "PPD" value refers to the number of pixels within each angle of the viewer's viewing angle.

[0136] In some embodiments, feature engine 140(0) calculates PPD value 208 based on DRH 122(3) and NVD 124(3) in any technically feasible manner. As shown, in some embodiments, feature engine 140(0) implements the following equation to calculate PPD value 208, denoted as PPD, based on DRH 122(3) and NVD 124(3):

[0137] PPD = DRH / ( arctan( 1 / (NVD * 2) ) * 360 / π (4)

[0138] It is noteworthy that as the DRH 122(3) or NVD 124(3) changes, the PPD value 208 changes in the same direction. And within the range of actual PPD values ​​typically associated with actual viewing experience, as the actual PPD value increases, the loss of detail in the reconstructed frame becomes less noticeable to the viewer, which increases the perceived video quality of the reconstructed frame. Unlike the values ​​of the conventional version of the DLM associated with some conventional perceptual quality models, the PPD_DLM values ​​are calculated based on the corresponding PPD values ​​and the corresponding NVD.

[0139] In some embodiments, PPD-aware DLM extractors 260(1)-PPD-aware DLM extractors 260(F) are different instances of PPD-aware DLM extractors 260 (not explicitly shown) that calculate PPD_DLM values ​​268(1)-PPD_DLM values ​​268(F), respectively. Figure 4 The PPD-aware DLM extractor 260(1) is described in more detail.

[0140] For purposes of explanation, the PPD-aware DLM extractors 260(1)-PPD-aware DLM extractors 260(F) are collectively referred to herein as "PPD-aware DLM extractors 260(1)-260(F)". The PPD_DLM values ​​268(1)-PPD_DLM values ​​268(F) are collectively referred to herein as "PPD_DLM values" and "PPD_DLM values ​​268(1)-268(F)".

[0141] As shown, in some embodiments, feature engine 140(0) configures PPD-aware DLM extractors 260(1)-260(F) to simultaneously calculate PPD_DLM values ​​268(1)-268(F) based on PPD value 208, NVD 124(3), reconstructed frames 172(1)-172(F), respectively, and source frames 162(1)-162(F), respectively. More generally, in some embodiments, feature engine 140(0) may configure any number of instances of PPD-aware DLM extractor 260 to simultaneously, sequentially, or any combination thereof.

[0142] The feature denoted TI is a temporal information measure that captures the amount of motion present in a scene and is quantified by the difference between adjacent pairs of source frames 162 included in source video sequence 160. The value of TI is also referred to herein as a "TI value."

[0143] As shown, in some embodiments, TI extractor 270 calculates TI values ​​278(1)-TI values ​​278(F) based on source video sequence 160. TI extractor 270 can quantify the differences between adjacent pairs of source frames 162 in any technically feasible manner. In some embodiments, TI extractor 270 performs any amount and / or type of filtering on source frames 162 to generate filtered source frames (not shown). TI extractor 270 then sets each TI value equal to the average absolute pixel difference of the luminance component between the adjacent pairs of the corresponding filtered source frames.

[0144] In some embodiments, the feature engine 140(0) generates a feature vector 182 based on the ordering of features within the feature group 202. As shown, in some embodiments, the feature engine 140(0) generates a feature vector 182(1), which sequentially includes, but is not limited to, NVD_VIF1 value 252(1), NVD_VIF2 value 254(1), NVD_VIF3 value 256(1), NVD_VIF4 value 258(1), PPD_DLM value 268(1), and TI value 278(1). Similarly, the feature engine 140(0) generates a feature vector 182(F), which sequentially includes, but is not limited to, NVD_VIF1 value 252(F), NVD_VIF2 value 254(F), NVD_VIF3 value 256(F), NVD_VIF4 value 258(F), PPD_DLM value 268(F), and TI value 278(F). Although not shown, the feature engine 140 (0) generates a feature vector 182 (i), where i is an integer from 2 to (F-1), and the feature vector 182 (i) includes, in order, but is not limited to, the NVD_VIF1 value 252 (i), the NVD_VIF2 value 254 (i), the NVD_VIF3 value 256 (i), the NVD_VIF4 value 258 (i), the PPD_DLM value 268 (i), and the TI value 278 (i). The feature engine 140 (0) then outputs the feature vectors 182 (1)-182 (F).

[0145] Figure 3 According to various embodiments Figure 2 A more detailed diagram of one of the VIF extractors 250 of FIG. 2 is provided for explanation purposes. Figure 2 The VIF extractor 250(1) is combined in the context of Figure 3 The functionality of VIF extractor 250 that implements NVD_VIF in some embodiments is described. VIF extractor 250(1) is an example of VIF extractor 250 that calculates NVD_VIF1 value 252(1) based on VIF kernel size 238(1), source frame 162(1), and reconstructed frame 172(1).

[0146] As previously combined Figure 2 As described, VIF kernel size 238(1) corresponds to both a VIF scaling index of 1 and NVD 124(3). More generally, in some embodiments, any number of instances of VIF extractor 250 may be configured to compute NVD-aware VIF values ​​for any number of reconstructed frames based on any number of VIF kernel sizes corresponding to any number of combinations of scaling indexes and NVDs.

[0147] In some embodiments, the VIF extractor 250(1) implements a pixel domain version of the VIF index to calculate an NVD_VIF1 value 252(1) corresponding to the reconstructed frame 172(1) based on the reconstructed frame 172(1), the source frame 162(1), and the VIF kernel size 238(1). More specifically, in the same or other embodiments, the VIF extractor 250(1) calculates the NVD_VIF1 value 252(1) based on an estimate of the amount of human perceptible information shared between the source frame 162(1) and the reconstructed frame 172(1) at a spatial scale defined by the VIF kernel size 238(1). As shown, in some embodiments, the VIF extractor 250(1) includes, but is not limited to, a Gaussian 2D kernel 310, a filtered source image 320, a filtered reconstructed image 330, gains 350(1)-gains 350(M), additive noise variances 370(1)-additive noise variances 370(M), and VIF values ​​390.

[0148] In some embodiments, VIF extractor 250(1) generates a Gaussian 2D kernel 310 with rotational symmetry based on VIF kernel size 238(1). Figure 3 In the context of VIF, VIF kernel size 238(1) is denoted as K. VIF extractor 250(1) may generate Gaussian 2D kernel 310 in any technically feasible manner. In some embodiments, VIF extractor 250(1) generates Gaussian 2D kernel 310 based on a Gaussian 1D kernel of size K and standard deviation (K / 5.0).

[0149] In some embodiments, the VIF extractor 250(1) computes a 2D convolution of the source frame 162(1) with the Gaussian 2D kernel 310 to generate a filtered source image 320. Pixels of the filtered source image 320 are denoted herein as p1-pM, where M represents the number of pixels in the filtered source image 320. As also shown, in some embodiments, the VIF extractor 250(1) computes a 2D convolution of the reconstructed frame 172(1) using the Gaussian 2D kernel 310 to generate a filtered reconstructed image 330. Pixels of the reconstructed image 330 are denoted herein as q1-qM.

[0150] Advantageously, as will be appreciated by those skilled in the art, the smoothness associated with both the filtered source image 320 and the filtered reconstructed image 330 corresponds to the VIF kernel size 238(1). And, returning to reference Figure 2 In some embodiments, VIF kernel size 238(1) is calculated based on NVD 124(3). Thus, for a VIF scaling index of 1, VIF kernel size 238(1) enables VIF extractor 250(1) to approximate the human-perceived visibility of distortion associated with the reconstructed image when viewed at NVD 124(3).

[0151] In some embodiments, the VIF extractor 250(1) models the distortion associated with the filtered reconstructed image 330 as a signal gain and an additive Gaussian noise with zero mean and an additive noise variance. Gain 350(1)-Gain 350(M) is a local estimate of the signal gain associated with the filtered reconstructed image 330. Additive noise variance 370(1)-Additive noise variance 370(M) is a local estimate of the additive noise variance associated with the filtered reconstructed image 330.

[0152] For purposes of explanation, gains 350(1)-gains 350(M) are denoted herein as g1-gM, respectively. In some embodiments, VIF extractor 250(1) estimates gains 350(1)-gains 350(M) based on filtered source image 320 and filtered reconstructed image 330. VIF extractor 250(1) may estimate gains 350(1)-gains 350(M) in any technically feasible manner. As shown, in some embodiments, VIF extractor 250(1) implements the following equations to estimate gains 350(1)-gains 350(M) corresponding to values ​​1-M, respectively, for integer variable i.

[0153]

[0154] For the purpose of explanation, the additive noise variance 370(1)-additive noise variance 370(M) are respectively represented as In some embodiments, the VIF extractor 250 ( 1 ) estimates additive noise variances 370 ( 1 )-370 (M) based on the gains 350 ( 1 )-350 (M), based on the filtered source image 320 and the filtered reconstructed image 330 , respectively.

[0155] The VIF extractor 250(1) may estimate the additive noise variance 370(1)-additive noise variance 370(M) in any technically feasible manner. As shown, in some embodiments, the VIF extractor 250(1) implements the following equation to estimate the additive noise variance 370(1)-additive noise variance 370(M) corresponding to values ​​1-M, respectively, for integer variable i.

[0156]

[0157] In some embodiments, to calculate the NVD_VIF1 value 252(1), the VIF extractor 250(1) estimates a portion of the information in the filtered source image 320 that the HVS can extract from the filtered reconstructed image 330. In the same or other embodiments, the VIF extractor 250(1) calculates the NVD_VIF1 value 252(1) based on the gain 350(1)-gain 350(M), the additive noise variance 370(1)-additive noise variance 370(M), and an additive Gaussian white noise that models the uncertainty in the perception of the visual signal associated with the HVS. In some embodiments, the additive Gaussian white noise has a variance of a constant value (e.g., 2.0) and is denoted herein as σ N 2 As shown, in the same or other embodiments, VIF extractor 250(1) implements the following equation to calculate NVD_VIF1 value 252(1), which is the VIF value 390 corresponding to reconstructed frame 172(1), a VIF scale index of 1, and NVD 124(3):

[0158]

[0159] Figure 4 According to various embodiments Figure 2 A more detailed illustration of one of the PPD-aware DLM extractors 260 of FIG. 2 is provided for explanation purposes. Figure 2 In the context of the PPD-aware DLM extractor 260(1), Figure 4 The functionality of the PPD-aware DLM extractor 260 that implements the PPD_DLM in some embodiments is described.

[0160] The PPD-aware DLM extractor 260(1) is an instance of the PPD-aware DLM extractor 260 that calculates the PPD_DLM value 268(1) based on the PPD value 208, the DRH 122(3), the reconstructed frame 172(1), and the source frame 162(1). More generally, in some embodiments, any number of instances of the PPD-aware DLM extractor 260 may be configured to calculate the PPD_DLM values ​​for any number of reconstructed frames based on any combination of PPD values ​​and NVD.

[0161] In some embodiments, the PPD-aware DLM extractor 260(1) implements a version of the DLM that operates in the wavelet domain and is modified to account for different PPD values. As shown, in the same or other embodiments, the PPD-aware DLM extractor 260(1) includes, but is not limited to, source discrete wavelet transform (DWT) coefficients 412, reconstructed DWT coefficients 414, restored DWT coefficients 416, contrast sensitivity (CS) weighting engine 420, contrast sensitivity function (CSF) graph 430, noise coefficient 470, noise term 480, and DLM value 490.

[0162] In some embodiments, the PPD-aware DLM extractor 260(1) uses a 4-scale Daubechies db2 wavelet, where db2 indicates the number of vanishing moments is 2, to calculate the source DWT coefficients 412 and the reconstructed DWT coefficients 414 based on the source frame 162(1) and the reconstructed frame 172(1), respectively. In some other embodiments, the PPD-aware DLM extractor 260(1) may use any type of technically feasible wavelet (e.g., biorthogonal wavelets). For purposes of explanation only, the source DWT coefficients 412 are also referred to herein individually as "source DWT coefficients 412". And the reconstructed DWT coefficients 414 are also referred to herein individually as "reconstructed DWT coefficients 414".

[0163] In some embodiments, the PPD-aware DLM extractor 260(1) decomposes the reconstructed frame 172(1) into a restored image (not shown) and an additive image (not shown) based on the source DWT coefficients 412 and the reconstructed DWT coefficients 414. The reconstructed frame 172(1) is equal to the sum of the restored image and the additive image. The restored image includes the same detail loss as the reconstructed frame 172(1), but does not include any additional impairments. To determine the restored image, the PPD-aware DLM extractor 260(1) may calculate restored DWT coefficients 416 for the restored image in any technically feasible manner. The restored DWT coefficients 416 are also referred to herein individually as "restored DWT coefficients 416".

[0164] For explanation purposes, S(λ,θ,i,j), T(λ,θ,i,j) and R(λ,θ,i,j) represent source DWT coefficients 412, reconstructed DWT coefficients 414 and recovered DWT coefficients 416, respectively, for a two-dimensional position (i,j) in a subband with a subband index of θ for a DWT scale with a DWT scale index of λ. In some embodiments, the DWT scale index λ is a value of 1-4, wherein each of the DWT scale indices 1-4 represents a different scale in the four scales associated with the Daubechies db2 wavelet.

[0165] In some embodiments, the subband index θ is a value from 1 to 4, where subband index 1 represents an approximate subband, subband index 2 represents a vertical subband, subband index 3 represents a diagonal subband, and subband index 4 represents a horizontal subband. It is worth noting that the PPD-aware DLM extractor 260(1) does not use the restored DWT coefficients 416 of the approximate subband (represented by a subband index θ of 1) to calculate the PPD_DLM value 268(1). Therefore, in some embodiments, the PPD-aware DLM extractor 260(1) does not calculate the restored DWT coefficients 416R(λ,1,i,j).

[0166] As shown, in some embodiments, the PPD-aware DLM extractor 260(1) calculates the recovered DWT coefficients 416R(λ,θ,i,j) for each 2-D position in the vertical, diagonal, and horizontal subbands with subband indices 2-4 at each of the four scales based on the following equations.

[0167] R(λ,θ,i,j) = clip [0,1] (T(λ,θ,i,j) / S(λ,θ,i,j))·S(λ,θ,i,j) (8)

[0168] As shown, in some embodiments, the CS weighting engine 420 generates a source signal term 440 and a restored signal term 450 based on the source DWT coefficients 412 and the restored DWT coefficients 416, respectively, based on the PPD value 208, the vertical / horizontal CSF 436, and the diagonal CSF 438. The source signal term 440 and the restored signal term 450 are also denoted as u and v, respectively, herein.

[0169] For purposes of explanation, as used herein, "contrast level" refers to the relative difference in brightness, "contrast threshold" is the minimum contrast level of an object required for a typical person to detect the object, and "contrast sensitivity" is the inverse of the contrast threshold. As used herein, "spatial frequency" refers to the number of cycles or line pairs (e.g., black and white lines in a grid of alternating black and white lines) that fall within a range of viewing angles. In some embodiments, spatial frequency is measured in cycles / degree. Spatial frequency is also denoted herein as "SF".

[0170] As used herein, a "CSF" describes the relationship between spatial frequency and contrast sensitivity of the HVS. In some embodiments, the vertical / horizontal CSF 436 describes the relationship between spatial frequency and contrast sensitivity of the HVS to vertical and horizontal structures. Thus, the vertical / horizontal CSF 436 is associated with both the vertical and horizontal directions. In the same or other embodiments, the diagonal CSF 438 describes the relationship between spatial frequency and contrast sensitivity of the HVS to diagonal structures. Thus, the diagonal CSF 438 is associated with the diagonal direction.

[0171] For purposes of explanation, CSF chart 430 depicts vertical / horizontal CSF 436 and diagonal CSF 438 implemented by PPD-aware DLM extractor 260(1) in some embodiments relative to spatial frequency (SF) axis 432 and contrast sensitivity (CS) axis 434. As shown by vertical / horizontal CSF 436 and diagonal CSF 438, the contrast sensitivity of the HVS to vertical and horizontal structures at a given spatial frequency is higher than the contrast sensitivity of the HVS to diagonal structures at the spatial frequency. In some embodiments, diagonal CSF 438 is a scaled version of vertical / horizontal CSF 436. As described herein, in a "scaled version" of vertical / horizontal CSF 436, the contrast sensitivity of vertical / horizontal CSF 436 is reduced by a global scale factor, while the functional shape of vertical / horizontal CSF 436 remains unchanged.

[0172] In some embodiments, vertical / horizontal CSF 436 and diagonal CSF 438 are improved versions of conventional CSFs (e.g., BartenCSF) that have been tuned to better correlate with empirical results. Advantageously, in some embodiments, vertical / horizontal CSF 436 and diagonal CSF 438 accurately represent contrast sensitivity at spatial frequencies associated with most viewing experiences.

[0173] In some embodiments, the CS weighting engine 420 calculates a different nominal spatial frequency (not shown) for each of the four DWT scales based at least in part on the PPD value 208. The CS weighting engine 420 can calculate the nominal spatial frequency in any technically feasible manner. For example, in some embodiments, the CS weighting engine 420 can calculate the nominal spatial frequency based on the DWT scale index, the PPD value 208, and the number of cycles per picture height or the number of line pairs per picture height. As used herein, "picture height" refers to the physical height of a display screen.

[0174] In some embodiments, for each DWT scale, the CS weighting engine 420 maps the nominal spatial resolution associated with the DWT scale to the vertical / horizontal contrast sensitivity associated with the DWT scale via the vertical / horizontal CSF 436. In the same or other embodiments, for each DWT scale, the CS weighting engine 420 maps the nominal spatial resolution associated with the DWT scale to the diagonal contrast sensitivity associated with the DWT scale via the diagonal CSF 438. For purposes of explanation, the exemplary point P depicted along the vertical / horizontal CSF 436 is λ=1 -P λ=4 A mapping of nominal spatial resolution to vertical / horizontal contrast sensitivity for DWT scales with DWT scale indices of 1-4 in some embodiments is shown.

[0175] In some embodiments, the CS weighting engine 420 discards the source DWT coefficients 412S (λ, 1, i, j) in the approximate frequency band and multiplies each remaining source DWT coefficient by the corresponding contrast sensitivity to generate the CS weighted source coefficients 422. For each source DWT coefficient 412S (λ, 2, i, j) in the vertical subband and each source DWT coefficient 412S (λ, 4, i, j) in the horizontal subband, the corresponding contrast sensitivity is the vertical / horizontal contrast sensitivity associated with the DWT scale with a DWT scale index of λ. For each source DWT coefficient 412S (λ, 3, i, j) in the diagonal subband, the corresponding contrast sensitivity is the diagonal contrast sensitivity associated with the DWT scale with a DWT scale index of λ.

[0176] In some embodiments, the CS weighting engine 420 discards the restored DWT coefficients 416R(λ,1,i,j) in the approximate frequency band and multiplies each of the remaining restored DWT coefficients by the corresponding contrast sensitivity to generate a CS weighted restoration coefficient (not shown). For each restored DWT coefficient 416R(λ,2,i,j) in the vertical subband and each restored DWT coefficient 416R(λ,4,i,j) in the horizontal subband, the corresponding contrast sensitivity is the vertical / horizontal contrast sensitivity associated with the DWT scale with a DWT scale index of λ. For each restored DWT coefficient 416R(λ,3,i,j) in the diagonal subband, the corresponding contrast sensitivity is the diagonal contrast sensitivity associated with the DWT scale with a DWT scale index of λ.

[0177] Notably, the CS-weighted source coefficients 422 and the CS-weighted restoration coefficients estimate the response of the HVS to the source frame 162(1) and the restored image, respectively, taking into account the PPD value 208. In some embodiments, the CS weighting engine 420 sets the source signal term 440 equal to a subset of the CS-weighted source coefficients 422 located in the central region of each subband with a margin factor (e.g., 0.1).

[0178] In some embodiments, the CS weighting engine 420 performs one or more contrast masking operations on the CS weighted restoration coefficients to mitigate any effect of the additive image on the visibility of the restored image to generate masked CS weighted restoration coefficients 424. The CS weighting engine 420 sets the restored signal term 450 equal to a subset of the masked CS weighted restoration coefficients 424 located in a central region of each subband having a boundary factor (e.g., 0.1).

[0179] As shown, in some embodiments, the PPD-aware DLM extractor 260(1) calculates the DLM value 490 according to the following equation:

[0180]

[0181] In equation (9), v, u, and n represent the source signal term 440, the restored signal term 450, and the noise term 480, respectively, and ∥∥3 represents a 3-norm operation. According to the summation sign in equation (9), the PPD-aware DLM extractor 260(1) independently sums the numerator and denominator over the vertical, horizontal, and diagonal subbands, and then sums over the four scales. As shown in the figure, the PPD-aware DLM extractor 260(1) outputs a DLM value 490 as the PPD_DLM value 168(1).

[0182] As will be appreciated by those skilled in the art, equation (9) is the ratio of the restored signal term 450 regularized by the noise term 480 to the source signal term 440 regularized by the noise term 480. Therefore, in order to be able to accurately detect the perceived visual information loss associated with the restored signal term 440 across different viewing parameters via equation (9), the PPD-aware DLM extractor 260(1) considers the NVD 124(3) when calculating the noise term 480. As shown, in some embodiments, the PPD-aware DLM extractor 260(1) calculates the noise factor 470 (denoted herein as k) based on the NVD 124(3) according to the following equation:

[0183] k = 0.000625 ( 3 / NVD ) 2 (12)

[0184] As shown, in the same or other embodiments, the PPD-aware DLM extractor 260(1) defines the noise term 480 as n=(kw λ h λ ) 1 / 3 In the noise term 480, w λ and h λ They represent the width and height of the DWT scale with the DWT scale index λ respectively, and k represents the noise coefficient 470.

[0185] Figure 5 is a flow chart of method steps for generating a trained perceptual quality model that takes into account NVD and display resolution according to various embodiments. Figure 1-4 Although the method steps are described with reference to a system, those skilled in the art will understand that any system configured to perform the method steps in any order falls within the scope of the embodiments.

[0186] As shown, method 500 begins at step 502, where training application 130 selects subjective quality score group 102(2). At step 504, training application 130 selects reconstructed video sequence(s), NVD, and DRH associated with the selected subjective quality score group.

[0187] At step 506 , for each selected reconstructed video sequence, the feature engine 140 calculates overall NVD-aware VIF feature values ​​at different spatial scales based on the reconstructed video sequence, the corresponding source video sequence, and the selected NVD.

[0188] At step 508, feature engine 140 calculates a PPD value based on the selected NVD and the selected DRH. At step 510, for each selected reconstructed video sequence, feature engine 14 calculates an overall PPD-aware DLM feature value based on the reconstructed video sequence, the corresponding source video sequence, the PPD value, and the selected NVD.

[0189] At step 512, the feature engine 140 calculates an overall temporal feature value for each selected reconstructed video sequence based on the corresponding source video sequence. At step 514, for each selected reconstructed video sequence, the feature engine 140 generates an overall feature vector specifying the associated overall feature value.

[0190] At step 516, the training application 130 determines whether the selected subjective quality score group is the last subjective quality score group. If at step 516, the training application 130 determines that the selected subjective quality score group is not the last subjective quality score group, the method 500 proceeds to step 518. At step 518, the training application 130 selects the next subjective quality score group, and the method 500 returns to step 504, where the training application 130 selects (one or more) reconstructed video sequences, NVDs, and DRHs associated with the selected subjective quality score group.

[0191] However, if at step 516, the training application 130 determines that the selected subjective quality score group is the final subjective quality score group, the method 500 proceeds directly to step 520. At step 520, the training engine 150 performs one or more supervised machine learning (ML) operations on the untrained ML model based on the subjective quality scores in the subjective quality core group(s) and the associated overall feature vectors to generate a trained perceptual quality model 158.

[0192] At step 522, the training engine 150 stores the trained perceptual quality model 158 in any amount of available memory and / or transmits it to any number of other software applications for future use. The method 500 then terminates.

[0193] Figure 6 is a flow chart of method steps for estimating the perceptual video quality of a reconstructed video using a trained perceptual quality model according to various embodiments. Figure 1-4 Although the method steps are described with reference to a system, those skilled in the art will understand that any system configured to perform the method steps in any order falls within the scope of the embodiments.

[0194] As shown, method 600 begins at step 602, where feature engine 140 calculates a plurality of VIF kernel sizes based on NVD. At step 604, for each reconstructed frame of the reconstructed video sequence, VIF extractor 250 calculates NVD-aware VIF feature values ​​based on the reconstructed frame, the corresponding source frame, and the VIF kernel size.

[0195] At step 606, the feature engine 140 calculates the PPD value based on NVD and DRH. At step 608, for each reconstructed frame, the PPD-aware DLM extractor 260 calculates different PPD-aware DLM feature values ​​based on the reconstructed frame, the corresponding source frame, the PPD value, the vertical / horizontal CSF, the diagonal CSF and the NVD.

[0196] At step 610, the TI extractor 270 computes TI features for the reconstructed frame based on the corresponding source frame. At step 612, for each reconstructed frame, the quality reasoning application 180 executes the trained perceptual quality model 158 on the associated NVD-perceptual VIF feature value, the associated PPD-perceptual DLM feature value, and the associated TI feature value to compute a perceptual quality score for the reconstructed frame.

[0197] At step 614, the quality reasoning application 180 calculates an overall perceptual quality score corresponding to the reconstructed video sequence, the NVD, and the DRH based on the perceptual quality scores. At step 616, the quality reasoning application 180 stores the overall perceptual quality score and / or one or more perceptual quality scores in any amount of available memory and / or sends them to any number of other software applications for future use. The method 600 then terminates.

[0198] In summary, the disclosed techniques can be used to accurately predict human perception of the quality of reconstructed video content on different combinations of DRH and NVD via a trained perceptual quality model. In some embodiments, a training application generates feature vectors corresponding to two sets of subjective quality scores associated with different combinations of DRH and NVD. Each subjective quality score is associated with one of a reconstructed video sequence, a source video sequence, and a combination of DRH and NVD.

[0199] The training application includes, but is not limited to, a feature engine, a feature pooling engine, and a training engine. For each subjective quality score, the feature engine generates a different feature vector for each reconstructed frame in the associated reconstructed video sequence based on the reconstructed frame, the corresponding source frame, the associated DRH, and the associated NVD. In some embodiments, each feature vector includes, but is not limited to, NVD_VIF1 value, NVD_VIF2 value, NVD_VIF3 value, NVD_VIF4 value, PPD_DLM value, and TI value.

[0200] For each reconstructed frame, the feature engine calculates four different VIF filter sizes based on the associated NVD. The feature engine then calculates NVD_VIF1, NVD_VIF2, NWD_VIF3, and NVD_VIF4 values ​​based on the VIF metrics based on the reconstructed frame, the corresponding source frame, and the different VIF filter sizes.

[0201] For each reconstructed frame, the feature engine calculates the PPD value based on the associated NVD and the associated DRH. The feature engine calculates the source DWT coefficients in the vertical, horizontal and diagonal subbands of the four DWT scales based on the source frame corresponding to the reconstructed frame. The feature engine calculates the restored DWT coefficients in the vertical, horizontal and diagonal subbands of the four DWT scales based on the reconstructed frame and the source frame. The restored DWT coefficients are associated with a restored image that includes the same detail loss as the reconstructed frame but does not include any additional impairments.

[0202] The feature engine weights the source DWT coefficients and restoration coefficients in the vertical and horizontal subbands of the four DWT scales based on the PPD value, vertical / horizontal CSF, and the four DWT scales. The feature engine weights the source DWT coefficients and restoration coefficients in the diagonal subbands of the four DWT scales based on the PPD value, diagonal CSF, and the four DWT scales. The feature extractor scales the noise term based on NVD. The feature engine then calculates the ratio of the weighted restoration coefficients regularized by the noise term to the weighted source coefficients normalized by the noise term to generate the PPD_DLM value.

[0203] In some embodiments, the feature engine sets each TI value equal to the average absolute pixel difference of the luminance component between adjacent pairs of corresponding reconstructed frames. For each reconstructed video sequence, the feature pooling engine calculates a single overall feature vector based on the feature vectors associated with the reconstructed frames in the reconstructed video sequence. In some embodiments, the training engine configures the SVR to train the SVM based on the overall feature vector and the corresponding subjective quality score group to generate a trained SVM, which is also referred to herein as a trained perceptual quality model.

[0204] In some embodiments, a quality reasoning application uses a trained perceptual quality model to calculate an overall perceptual quality score corresponding to a reconstructed video sequence, DRH, and NVD. The quality reasoning application includes, but is not limited to, a feature engine, a trained perceptual quality model, and a score pooling engine. The feature engine generates different feature vectors for each reconstructed frame of the reconstructed video sequence based on the reconstructed video sequence, the corresponding source video sequence, DRH, and NVD. For each reconstructed frame of the reconstructed video sequence, the quality reasoning application inputs the associated feature vector into the trained perceptual quality model. In response, the trained perceptual quality model outputs the perceptual quality score of the reconstructed frame. The score pooling engine calculates an overall perceptual quality score based on the perceptual quality score of the reconstructed frame. The overall perceptual quality score estimates the perceptual video quality of the reconstructed video sequence when viewed at a standardized viewing distance of NVD on a display having a digital resolution height equal to the DRH.

[0205] At least one technical advantage of the disclosed technology over the prior art is that, with the disclosed technology, a trained perceptual quality model can be used to predict more accurate perceptual visual quality scores for reconstructed video over a wider range of viewing parameters than can be achieved using prior art methods. Specifically, unlike features of conventional perceptual quality models, the disclosed technology incorporates features in a trained perceptual quality model that approximate a fundamental relationship between PPD values ​​and human perception of the quality of reconstructed video content. Thus, the trained perceptual quality model can be used to generate a different perceptual quality score for each different combination of display resolution and normalized viewing distance over a range of PPD values ​​typically associated with an actual viewing experience. Thus, implementing the trained perceptual quality model during the encoding process can improve the quality / bitrate tradeoff typically made when encoding video content, and thereby improve the overall streaming experience for the viewer. These technical advantages provide one or more technical advances over prior art methods.

[0206] 1. In some embodiments, a computer-implemented method for estimating perceptual video quality of a reconstructed video includes: calculating a first plurality of feature values ​​corresponding to a plurality of visual quality metrics based on a first reconstructed frame, a first source frame, a first display resolution, and a first normalized viewing distance; executing a trained perceptual quality model on the first plurality of feature values ​​to generate a first perceptual quality score, the first perceptual quality score indicating a first perceptual visual quality level of the first reconstructed frame; and performing one or more operations associated with an encoding process based on the first perceptual quality score.

[0207] 2. A computer-implemented method according to clause 1, wherein performing one or more operations associated with the encoding process includes: calculating an overall perceptual quality score of a first reconstructed video sequence including the first reconstructed frame based on the first perceptual quality score, wherein the overall perceptual quality score indicates an overall perceptual visual quality level of the first reconstructed video sequence; and calculating a trade-off between quality and bitrate based on the overall perceptual quality score when encoding a first source video sequence.

[0208] 3. A computer-implemented method according to clause 1 or 2, wherein performing one or more operations associated with the encoding process includes: evaluating at least one of an encoder, a decoder, or an adaptive bitrate streaming algorithm based on the first perceptual quality score.

[0209] 4. A computer-implemented method according to any of clauses 1-3, wherein the first perceived visual quality level approximates the quality of an actual viewing experience, in which a viewing device has the first display resolution and a viewer views the viewing device at the first standardized viewing distance.

[0210] 5. The computer-implemented method according to any one of clauses 1-4 further includes: calculating a second plurality of feature values ​​corresponding to the plurality of visual quality metrics based on the first reconstructed frame, the first source frame, the first display resolution, and a second standardized viewing distance; and executing the trained perceptual quality model on the second plurality of feature values ​​to generate a second perceptual quality score indicating a second perceptual visual quality level of the first reconstructed frame.

[0211] 6. The computer-implemented method according to any one of clauses 1-5 further includes: calculating a first unit angle pixel value based on the first display resolution and the first standardized viewing distance; and calculating an absolute perceptual quality score of a first reconstructed video sequence including the first reconstructed frame based on the first perceptual quality score and a first offset associated with the first unit angle pixel value, wherein the absolute perceptual quality score indicates an overall perceptual visual quality level of the first reconstructed video sequence.

[0212] 7. A computer-implemented method according to any of clauses 1-6, wherein the multiple visual quality metrics include at least one of a detail loss metric or a visual information fidelity index, the detail loss metric is modified to take into account different standardized viewing distances and different display resolutions, and the visual information fidelity index is modified to take into account different standardized viewing distances.

[0213] 8. A computer-implemented method according to any one of clauses 1-7, wherein calculating the first plurality of eigenvalues ​​includes calculating the first eigenvalue by the following steps: calculating a unit angle pixel value based on the first display resolution and the first standardized viewing distance; calculating a plurality of contrast sensitivities based on a plurality of scale indices, the unit angle pixel value and at least one contrast sensitivity function; and calculating the first eigenvalue based on the first reconstructed frame, the first source frame and the plurality of contrast sensitivities.

[0214] 9. A computer-implemented method according to any one of clauses 1-8, wherein calculating the first plurality of eigenvalues ​​includes calculating a first eigenvalue by the following steps: calculating a noise coefficient based on the first standardized viewing distance; and calculating the first eigenvalue based on the first reconstructed frame, the first source frame, the first display resolution, the first standardized viewing distance and the noise coefficient.

[0215] 10. A computer-implemented method according to any one of clauses 1-9, wherein calculating the first plurality of eigenvalues ​​includes calculating a first eigenvalue by the following steps: generating a Gaussian two-dimensional kernel based on the first standardized viewing distance; and calculating the first eigenvalue based on the Gaussian two-dimensional kernel, the first reconstructed frame and the first source frame.

[0216] 11. In some embodiments, one or more non-transitory computer-readable media include instructions that, when executed by one or more processors, cause the one or more processors to estimate the perceptual video quality of the reconstructed video by performing the following steps: calculating a first plurality of feature values ​​corresponding to a plurality of visual quality metrics based on a first reconstructed frame, a first source frame, a first display resolution, and a first standardized viewing distance; executing a trained perceptual quality model on the first plurality of feature values ​​to generate a first perceptual quality score, the first perceptual quality score indicating a first perceptual visual quality level of the first reconstructed frame; and performing one or more operations associated with an encoding process based on the first perceptual quality score.

[0217] 12. One or more non-transitory computer-readable media according to clause 11, wherein performing one or more operations associated with the encoding process includes: when encoding a first source video sequence including the first source frame, calculating a tradeoff between quality and bit rate based on the first perceptual quality score.

[0218] 13. One or more non-transitory computer-readable media according to clause 11 or 12, wherein performing one or more operations associated with the encoding process includes: calculating an overall perceptual quality score of a first reconstructed video sequence including the first reconstructed frame based on the first perceptual quality score, wherein the overall perceptual quality score indicates an overall perceptual visual quality level of the first reconstructed video sequence; and evaluating at least one of an encoder, a decoder, or an adaptive bitrate streaming algorithm based on the overall perceptual quality score.

[0219] 14. The one or more non-transitory computer-readable media of any of clauses 11-13, wherein the trained perceptual quality model comprises a trained support vector regression model, a trained artificial neural network, or a trained regression tree.

[0220] 15. One or more non-transitory computer-readable media according to any one of clauses 11-14, further comprising: calculating a second plurality of feature values ​​corresponding to the plurality of visual quality metrics based on a second reconstructed frame, a second source frame, a second display resolution, and a second standardized viewing distance; and executing the trained perceptual quality model on the second plurality of feature values ​​to generate a second perceptual quality score, the second perceptual quality score indicating a second perceptual visual quality level of the second reconstructed frame.

[0221] 16. One or more non-transitory computer-readable media according to any one of clauses 11-15, further comprising: calculating a first unit angle pixel value based on the first display resolution and the first standardized viewing distance; and calculating an absolute perceptual quality score of a first reconstructed video sequence including the first reconstructed frame based on the first perceptual quality score and a first offset associated with the first unit angle pixel value, wherein the absolute perceptual quality score indicates an overall perceptual visual quality level of the first reconstructed video sequence.

[0222] 17. One or more non-transitory computer-readable media according to any of clauses 11-16, wherein the multiple visual quality metrics include at least one of a detail loss metric or a visual information fidelity index, the detail loss metric is modified to account for different standardized viewing distances and different display resolutions, and the visual information fidelity index is modified to account for different standardized viewing distances.

[0223] 18. One or more non-transitory computer-readable media according to any one of clauses 11-17, wherein calculating the first plurality of eigenvalues ​​includes calculating the first eigenvalue by the following steps: calculating a unit angle pixel value based on the first display resolution and the first standardized viewing distance; calculating a plurality of contrast sensitivities based on a plurality of scale indices, the unit angle pixel value and at least one contrast sensitivity function; and calculating the first eigenvalue based on the first reconstructed frame, the first source frame and the plurality of contrast sensitivities.

[0224] 19. One or more non-transitory computer-readable media according to any one of clauses 11-18, wherein calculating the first plurality of eigenvalues ​​includes calculating a first eigenvalue by: generating a Gaussian two-dimensional kernel based on the first standardized viewing distance; and calculating the first eigenvalue based on the Gaussian two-dimensional kernel, the first reconstructed frame, and the first source frame.

[0225] 20. One or more non-transitory computer-readable media according to any of clauses 11-19, wherein the first reconstructed frame is in a high dynamic range format.

[0226] 21. In some embodiments, a system includes: one or more memories storing instructions; and one or more processors coupled to the one or more memories, wherein the one or more processors perform the following steps when executing the instructions: calculating a first plurality of feature values ​​corresponding to a plurality of visual quality metrics based on a first reconstructed frame, a first source frame, a first display resolution, and a first standardized viewing distance; executing a trained perceptual quality model on the first plurality of feature values ​​to generate a first perceptual quality score, the first perceptual quality score indicating a first perceptual visual quality level of the first reconstructed frame; and performing one or more operations associated with an encoding process based on the first perceptual quality score.

[0227] Any and all combinations of any claim elements recited in any claim and / or any elements described in the present application, in any manner, are within the scope and protection of the contemplated invention.

[0228] The description of the various embodiments is given for purposes of illustration, but is not intended to be exhaustive or limited to the embodiments disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.

[0229] Aspects of the embodiments herein may be embodied as systems, methods, or computer program products. Thus, aspects of the present disclosure may take the form of fully hardware embodiments, fully software embodiments (including firmware, resident software, microcode, etc.), or combined software and hardware embodiments, which may all be collectively referred to herein as "modules" or "systems." Additionally, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable media having computer-readable program code embodied thereon.

[0230] Any combination of one or more computer-readable media may be utilized. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of computer-readable storage media will include the following: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium may be any tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment.

[0231] Aspects of the present disclosure are described above with reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to the embodiments of the present disclosure. It will be understood that each square frame of the flowchart and / or block diagram and the combination of each square frame in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine. When the instruction is executed via the processor of a computer or other programmable data processing device, the function / action specified in one or more square frames of the flowchart and / or block diagram can be realized. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable gate array.

[0232] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a fragment or a part of a code, and the module, a fragment or a part of the code includes one or more executable instructions for implementing (one or more) specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box may also occur in an order different from that marked in the accompanying drawings. For example, depending on the functions involved, the two boxes shown in succession may actually be executed substantially simultaneously, or the boxes may sometimes be executed in reverse order. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs a specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0233] While the foregoing is directed to embodiments of the present disclosure, other and further embodiments of the disclosure may be devised without departing from the basic scope thereof, and the scope thereof is determined by the claims that follow.

Claims

1. A computer-implemented method for estimating perceptual video quality of a reconstructed video, the method comprising: calculating a first plurality of feature values ​​corresponding to a plurality of visual quality metrics based on the first reconstructed frame, the first source frame, the first display resolution, and the first normalized viewing distance; executing a trained perceptual quality model on the first plurality of feature values ​​to generate a first perceptual quality score indicating a first perceptual visual quality level of the first reconstructed frame; as well as One or more operations associated with an encoding process are performed based on the first perceptual quality score.

2. The computer-implemented method of claim 1 , wherein: Performing one or more operations associated with the encoding process includes: Calculating an overall perceptual quality score of a first reconstructed video sequence including the first reconstructed frame based on the first perceptual quality score, wherein the overall perceptual quality score indicates an overall perceptual visual quality level of the first reconstructed video sequence; and When encoding the first source video sequence, a trade-off between quality and bitrate is calculated based on the overall perceptual quality score.

3. The computer-implemented method of claim 1 , wherein: Performing one or more operations associated with the encoding process includes evaluating at least one of an encoder, a decoder, or an adaptive bitrate streaming algorithm based on the first perceptual quality score.

4. The computer-implemented method of claim 1 , wherein: The first perceived visual quality level approximates a quality of an actual viewing experience in which a viewing device has the first display resolution and a viewer views the viewing device at the first standardized viewing distance.

5. The computer-implemented method of claim 1 , further comprising: calculating a second plurality of feature values ​​corresponding to the plurality of visual quality metrics based on the first reconstructed frame, the first source frame, the first display resolution, and a second normalized viewing distance; as well as The trained perceptual quality model is performed on the second plurality of feature values ​​to generate a second perceptual quality score indicative of a second perceptual visual quality level of the first reconstructed frame.

6. The computer-implemented method of claim 1 , further comprising: Calculating a first unit angle pixel value based on the first display resolution and the first standardized viewing distance; as well as An absolute perceptual quality score of a first reconstructed video sequence including the first reconstructed frame is calculated based on the first perceptual quality score and a first offset associated with the first unit angle pixel value, wherein the absolute perceptual quality score indicates an overall perceptual visual quality level of the first reconstructed video sequence.

7. The computer-implemented method of claim 1 , wherein: The plurality of visual quality metrics include at least one of a detail loss metric modified to account for different normalized viewing distances and different display resolutions or a visual information fidelity index modified to account for different normalized viewing distances.

8. The computer-implemented method of claim 1 , wherein: Calculating the first plurality of eigenvalues ​​includes calculating a first eigenvalue by: Calculating a pixel value per unit angle based on the first display resolution and the first standardized viewing distance; calculating a plurality of contrast sensitivities based on a plurality of scaling indices, the unit angle pixel value, and at least one contrast sensitivity function; as well as The first eigenvalue is calculated based on the first reconstructed frame, the first source frame, and the plurality of contrast sensitivities.

9. The computer-implemented method of claim 1 , wherein: Calculating the first plurality of eigenvalues ​​includes calculating a first eigenvalue by: calculating a noise factor based on the first standardized viewing distance; as well as The first eigenvalue is calculated based on the first reconstructed frame, the first source frame, the first display resolution, the first normalized viewing distance, and the noise coefficient.

10. The computer-implemented method of claim 1, wherein: Calculating the first plurality of eigenvalues ​​includes calculating a first eigenvalue by: generating a Gaussian two-dimensional kernel based on the first normalized viewing distance; and The first eigenvalue is calculated based on the Gaussian two-dimensional kernel, the first reconstructed frame, and the first source frame.

11. One or more non-transitory computer-readable media comprising instructions that, when executed by one or more processors, cause the one or more processors to estimate a perceived video quality of a reconstructed video by performing the following steps: calculating a first plurality of feature values ​​corresponding to a plurality of visual quality metrics based on the first reconstructed frame, the first source frame, the first display resolution, and the first normalized viewing distance; executing a trained perceptual quality model on the first plurality of feature values ​​to generate a first perceptual quality score indicating a first perceptual visual quality level of the first reconstructed frame; as well as One or more operations associated with an encoding process are performed based on the first perceptual quality score.

12. The one or more non-transitory computer readable media of claim 11, wherein: Performing one or more operations associated with the encoding process includes calculating a tradeoff between quality and bitrate based on the first perceptual quality score when encoding a first source video sequence including the first source frame.

13. The one or more non-transitory computer readable media of claim 11, wherein: Performing one or more operations associated with the encoding process includes: Calculating an overall perceptual quality score of a first reconstructed video sequence including the first reconstructed frame based on the first perceptual quality score, wherein the overall perceptual quality score indicates an overall perceptual visual quality level of the first reconstructed video sequence; and At least one of an encoder, a decoder, or an adaptive bitrate streaming algorithm is evaluated based on the overall perceptual quality score.

14. The one or more non-transitory computer readable media of claim 11, wherein: The trained perceptual quality model includes a trained support vector regression model, a trained artificial neural network or a trained regression tree.

15. The one or more non-transitory computer-readable media of claim 11, further comprising: calculating a second plurality of feature values ​​corresponding to the plurality of visual quality metrics based on a second reconstructed frame, a second source frame, a second display resolution, and a second normalized viewing distance; as well as The trained perceptual quality model is performed on the second plurality of feature values ​​to generate a second perceptual quality score indicative of a second perceptual visual quality level of the second reconstructed frame.

16. The one or more non-transitory computer-readable media of claim 11, further comprising: Calculating a first unit angle pixel value based on the first display resolution and the first standardized viewing distance; as well as An absolute perceptual quality score of a first reconstructed video sequence including the first reconstructed frame is calculated based on the first perceptual quality score and a first offset associated with the first unit angle pixel value, wherein the absolute perceptual quality score indicates an overall perceptual visual quality level of the first reconstructed video sequence.

17. The one or more non-transitory computer readable media of claim 11, wherein: The plurality of visual quality metrics include at least one of a detail loss metric modified to account for different normalized viewing distances and different display resolutions or a visual information fidelity index modified to account for different normalized viewing distances.

18. The one or more non-transitory computer readable media of claim 11, wherein: Calculating the first plurality of eigenvalues ​​includes calculating a first eigenvalue by: Calculating a pixel value per unit angle based on the first display resolution and the first standardized viewing distance; calculating a plurality of contrast sensitivities based on a plurality of scaling indices, the unit angle pixel value, and at least one contrast sensitivity function; as well as The first eigenvalue is calculated based on the first reconstructed frame, the first source frame, and the plurality of contrast sensitivities.

19. The one or more non-transitory computer readable media of claim 11, wherein: Calculating the first plurality of eigenvalues ​​includes calculating a first eigenvalue by: generating a Gaussian two-dimensional kernel based on the first normalized viewing distance; and The first eigenvalue is calculated based on the Gaussian two-dimensional kernel, the first reconstructed frame, and the first source frame.

20. The one or more non-transitory computer readable media of claim 11, wherein: The first reconstructed frame is in a high dynamic range format.

21. A system comprising: One or more memories storing instructions; as well as One or more processors, coupled to the one or more memories, the one or more processors performing the following steps when executing the instructions: calculating a first plurality of feature values ​​corresponding to a plurality of visual quality metrics based on the first reconstructed frame, the first source frame, the first display resolution, and the first normalized viewing distance; executing a trained perceptual quality model on the first plurality of feature values ​​to generate a first perceptual quality score indicating a first perceptual visual quality level of the first reconstructed frame; as well as One or more operations associated with an encoding process are performed based on the first perceptual quality score.