Weighted PSNR quality metric for video coding data
By applying time and space high-pass filters to determine visual activity information in predetermined picture blocks of the video sequence, the video quality evaluation method is improved, the problem of inaccurate video quality evaluation in the prior art is solved, and higher correlation and lower algorithm complexity are achieved.
Patent Information
- Application Number
- CN202080073900.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-21
- Filing Date
- 2020-10-16
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-10-16
AI Technical Summary
Existing video quality evaluation methods, especially PSNR-based methods, cannot effectively reflect the subjective impression of video encoding quality, and the perceptual weighted PSNR (WPSNR) has poor correlation on video data.
By determining visual activity information in a predetermined picture block of a video sequence, using a temporal high-pass filter and a spatial high-pass filter, the video quality evaluation method is improved, specifically including receiving a predetermined picture block of a plurality of video frames, applying a high-pass filter to determine visual activity information, and adjusting the encoded quantization parameters based on this information.
The correlation with subjective average opinion score (MOS) data is achieved, which reduces the algorithm complexity of video quality evaluation and improves the evaluation accuracy of video encoding quality.
Smart Images

Figure CN114747214B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an apparatus and method for improved video quality assessment, and more particularly, to an apparatus and method for improved perceptual weighted PSNR (WPSNR; weighted peak signal-to-noise ratio) for video quality assessment. Background Art
[0002] It is well known that the objective PSNR metric correlates very poorly with the subjective impression of video coding quality. Therefore, several alternative metrics such as (MS-)SSIM and VMAF have been proposed.
[0003] In JVET-H0047 [6], a block-level perceptually weighted distortion metric was proposed as an extension of the PSNR metric, called WPSNR, which was improved in JVET-K0206 [7] and JVET-M0091 [8]. More recently, the WPSNR metric was found to correlate with subjective mean opinion score (MOS) data as well as (MS-)SSIM across at least several MOS-annotated static image databases [9], see Table 1. However, on video data, the correlation with MOS scores (e.g., the correlations provided in [4] or the results of the previous JVET call for proposals
[10] ) was found to be worse than that of (MS-)SSIM or VMAF, thus indicating the need for improvement. In the following, a summary of the block-level WPSNR metric is provided as well as a description of a low-complexity WPSNR extension for video coding that addresses the above shortcomings.
[0004]
[0005] Table 1 shows the average correlation between subjective MOS data and objective values on JPEG and JPEG 2000 compressed static images of the four databases. SROCC: Spearman rank order, PLCC: Pearson linear correlation [9].
[0006] Given the well-known inaccuracy of the Peak Signal-to-Noise Ratio (PSNR) in predicting the average subjective judgment of the visual coding quality for a given codec c and image or video stimulus s, several metrics that perform better have been developed over the past two decades. The most commonly used are the Structural Similarity Measure (SSIM) [1] and its multiscale extension MS-SSIM [2], and the more recent Video Multi-Method Assessment Fusion (VMAF), which combines a number of other metrics using machine learning [4]. VMAF methods have been found to be particularly useful for the assessment of video coding quality [4], but determining an objective VMAF score is algorithmically complex and requires two passes. More importantly, VMAF algorithms are not differentiable [5] and therefore cannot be used as a reference for perceptual bit allocation strategies during image or video encoding (such as PSNR-based or SSIM-based metrics).
[0007] As can be appreciated, it is desirable to provide improved video quality assessment concepts. Summary of the invention
[0008] It is an object of the present invention to provide improved concepts for improved video quality assessment.
[0009] A device for determining visual activity information for a predetermined picture block of a video sequence is provided, the video sequence comprising a plurality of video frames, the plurality of video frames comprising a current video frame and one or more temporally preceding video frames, wherein the one or more temporally preceding video frames temporally precede the current video frame. The device is configured to receive a predetermined picture block of each of the one or more temporally preceding video frames and a predetermined picture block of the current video frame. In addition, the device is configured to determine the visual activity information based on the predetermined picture block of the current video frame and based on the predetermined picture block of each of the one or more temporally preceding video frames and based on a temporal high-pass filter.
[0010] In addition, a device for determining visual activity information for a predetermined picture block of a video sequence is provided, the video sequence comprising a plurality of video frames, the plurality of video frames comprising a current video frame. The device is configured to receive a predetermined picture block of the current video frame. In addition, the device is configured to determine visual activity information according to the predetermined picture block of the current video frame and according to a spatial high-pass filter and / or a temporal high-pass filter. In addition, the device is configured to: downsample the predetermined picture block of the current video frame to obtain a downsampled picture block, and apply a spatial high-pass filter and / or a temporal high-pass filter to each of the plurality of picture samples of the downsampled picture block; or, the device is configured to: apply the spatial high-pass filter and / or the temporal high-pass filter only to the first group of the plurality of picture samples of the predetermined picture block but not to the second group of the plurality of picture samples of the predetermined picture block.
[0011] In addition, a device for changing a coding quantization parameter on a picture according to an embodiment is provided, the device comprising the device for determining visual activity information as described above. The device for changing a coding quantization parameter on a picture is configured to: determine a coding quantization parameter of a predetermined block according to the visual activity information.
[0012] Furthermore, an encoder for encoding a picture into a data stream is provided. The encoder comprises: an apparatus for varying a coding quantization parameter on a picture as described above, and an encoding stage configured to encode the picture into the data stream using the coding quantization parameter.
[0013] Furthermore, a decoder for decoding a picture from a data stream is provided. The decoder comprises: an apparatus for varying a coding quantization parameter on a picture as described above, and a decoding stage configured to decode the picture from the data stream using the coding quantization parameter. The decoding stage is configured to: decode a residual signal from the data stream, dequantize the residual signal using the coding quantization parameter, and use the residual signal and decode the picture from the data stream using predictive decoding.
[0014] In addition, a method for determining visual activity information for a predetermined picture block of a video sequence is provided, the video sequence comprising a plurality of video frames, the plurality of video frames comprising a current video frame and one or more temporally preceding video frames, wherein the one or more temporally preceding video frames temporally precede the current video frame. The method comprises:
[0015] - receiving a predetermined picture block of each of one or more temporally previous video frames and a predetermined picture block of a current video frame, and:
[0016] - determining visual activity information from a predetermined picture block of a current video frame and from a predetermined picture block of each of one or more temporally preceding video frames and from a temporal high pass filter.
[0017] In addition, a method for determining visual activity information for a predetermined picture block of a video sequence is provided, the video sequence comprising a plurality of video frames, the plurality of video frames comprising a current video frame. The method comprises:
[0018] - receiving a predetermined picture block of a current video frame, and:
[0019] - Determining visual activity information based on predetermined picture blocks of the current video frame and based on a spatial high pass filter and / or a temporal high pass filter.
[0020] The method includes: downsampling a predetermined picture block of a current video frame to obtain a downsampled picture block, and applying a spatial high-pass filter and / or a temporal high-pass filter to each of a plurality of picture samples of the downsampled picture block. Alternatively, the method includes: applying the spatial high-pass filter and / or the temporal high-pass filter to only a first group of a plurality of picture samples of the predetermined picture block but not to a second group of a plurality of picture samples of the predetermined picture block.
[0021] Furthermore, a method for varying an encoding quantization parameter across a picture according to an embodiment is provided. The method comprises the method for determining visual activity information as described above.
[0022] The method for varying a coding quantization parameter on a picture further includes determining a coding quantization parameter of a predetermined block according to visual activity information.
[0023] Furthermore, a coding method for encoding a picture into a data stream according to an embodiment is provided. The coding method comprises the method for varying a coding quantization parameter on a picture as described above.
[0024] The encoding method further includes: encoding the picture into a data stream using the encoding quantization parameter.
[0025] In addition, a decoding method for decoding a picture from a data stream according to an embodiment is provided. The decoding method comprises the method for varying the encoding quantization parameter on a picture as described above.
[0026] The decoding method further includes: decoding the picture from the data stream using the encoded quantization parameter.
[0027] Furthermore, a computer program is provided, comprising instructions which, when executed on a computer or a signal processor, cause the computer or the signal processor to perform one of the above methods.
[0028] Furthermore, there is provided a data stream having pictures encoded therein by an encoder as described above.
[0029] The examples show that, through a low-complexity extension of our previous work on perceptual weighted PSNR (WPSNR) proposed in JVET-H0047, JVET-K0206, JVET-M0091, a motion-aware WPSNR algorithm can be obtained that produces similar levels of correlation with the subjective mean opinion score as the above-mentioned state-of-the-art metrics, but with lower algorithmic complexity. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In the following, embodiments of the present invention will be described in more detail with reference to the accompanying drawings, in which:
[0031] Figure 1An apparatus for determining visual activity information according to an embodiment is shown.
[0032] Figure 2 The effect of different WPSNR averaging concepts over multiple frames is shown.
[0033] Figure 3 Sample-level high-pass filtering of s is shown without and with spatial downsampling of the input signal s during filtering.
[0034] Figure 4 A device for varying a coding quantization parameter QP across a picture is shown, which comprises a visual activity information determiner and a QP determiner.
[0035] Figure 5 A possible structure of the encoding stages is shown.
[0036] Figure 6 A possible decoder configured to decode a reconstructed version of a video and / or a picture from a data stream is shown. DETAILED DESCRIPTION
[0037] Figure 1 An apparatus for determining visual activity information for a predetermined picture block of a video sequence according to an embodiment is shown, wherein the video sequence includes a plurality of video frames, wherein the plurality of video frames include a current video frame and one or more temporally preceding video frames, wherein the one or more temporally preceding video frames temporally precede the current video frame.
[0038] The apparatus is configured to receive, for example, through the first module 110 , a predetermined picture block of each of one or more temporally preceding video frames and a predetermined picture block of a current video frame.
[0039] Furthermore, the apparatus is configured to determine visual activity information, for example by means of the second module 120 , from a predetermined picture block of the current video frame and from a predetermined picture block of each of one or more temporally preceding video frames and from a temporal high pass filter.
[0040] According to an embodiment, the temporal high pass filter may be, for example, a finite impulse response filter.
[0041] In an embodiment, the apparatus 100 may be configured to apply a temporal high pass filter by combining picture samples in a predetermined picture block of a current video frame with picture samples in a predetermined picture block of each of one or more temporally preceding video frames, for example.
[0042] According to an embodiment, each of the picture samples of the predetermined picture block of the current video frame and each of the picture samples of the predetermined picture block of each of the one or more temporally preceding video frames may be, for example, a luminance value. Alternatively, each of the picture samples of the predetermined picture block of the current video frame and each of the picture samples of the predetermined picture block of each of the one or more temporally preceding video frames may be, for example, a chrominance value. Alternatively, each of the picture samples of the predetermined picture block of the current video frame and each of the picture samples of the predetermined picture block of each of the one or more temporally preceding video frames may be, for example, a red value, a green value, or a blue value.
[0043] In an embodiment, the one or more temporally previous video frames are exactly one temporally previous video frame.The apparatus 100 may be configured to apply a temporal high pass filter by combining picture samples in a predetermined picture block of a current video frame with picture samples in a predetermined picture block of the exactly one temporally previous video frame.
[0044] According to an embodiment, the temporal high-pass filter may be defined, for example, according to the following formula:
[0045]
[0046] Where x is a first coordinate value of a sample position within a predetermined image block, where y is a second coordinate value of a sample position within a predetermined image block, where s i [x, y] indicates the picture sample at position (x, y) in a predetermined picture block of the current video frame, where s i-1 [x, y] indicates a picture sample at position (x, y) in a predetermined picture block of exactly one temporally previous video frame, where Indicates picture samples of a predetermined picture block generated by applying the temporal high pass filter.
[0047] In an embodiment, the one or more temporally preceding video frames are two or more temporally preceding video frames. The apparatus 100 may be configured, for example, to apply a temporal high pass filter by combining picture samples in a predetermined picture block of the current video frame with picture samples in a predetermined picture block of each of the two or more temporally preceding video frames.
[0048] According to an embodiment, the one or more temporally preceding video frames are exactly two temporally preceding video frames, a first temporally preceding video frame of the exactly two temporally preceding video frames is immediately before the current video frame in time, and wherein a second temporally preceding video frame of the exactly two temporally preceding video frames is immediately before the first temporally preceding video frame in time. The apparatus 100 may be configured, for example, to apply a temporal high pass filter by combining picture samples in a predetermined picture block of the current video frame with picture samples in a predetermined picture block of the first temporally preceding video frame and picture samples in a predetermined picture block of the second temporally preceding video frame.
[0049] In an embodiment, the temporal high-pass filter may be defined, for example, according to the following formula:
[0050]
[0051] Where x is a first coordinate value of a sample position within a predetermined image block, where y is a second coordinate value of a sample position within a predetermined image block, where s i [x, y] indicates the picture sample at position (x, y) in a predetermined picture block of the current video frame, where s i-1 [x, y] indicates a picture sample located at position (x, y) in a predetermined picture block of a first temporally previous video frame, where s i-2 [x,γ] indicates a picture sample located at position (x,y) in a predetermined picture block of a second temporally previous video frame, where Indicates picture samples of a predetermined picture block generated by applying the temporal high pass filter.
[0052] According to an embodiment, the device 100 can be configured, for example, to combine a spatially high-pass filtered version of a picture sample in a predetermined picture block of a current video frame with a temporally high-pass filtered picture sample, where the temporally high-pass filtered picture sample is generated by applying a temporal high-pass filter by combining the picture sample in the predetermined picture block of the current video frame with the picture samples in the predetermined picture block of each of one or more temporally previous video frames.
[0053] In an embodiment, the apparatus 100 may be configured to combine, for example, a spatial high-pass filtered version of a picture sample in a predetermined picture block of a current video frame with a temporal high-pass filtered picture sample according to the definition of the following formula:
[0054]
[0055] in, indicates a spatially high-pass filtered version of a picture sample at position (x, y) in a predetermined picture block of a current video frame, wherein γ indicates a constant, γ>0, and wherein, Indicates temporal high-pass filtered picture samples.
[0056] According to an embodiment, γ may be defined as γ=2, for example.
[0057] In an embodiment, in order to obtain a plurality of intermediate picture samples of a predetermined block, for each of the plurality of picture samples of the predetermined block, the apparatus 100 may be configured, for example, to determine an intermediate picture sample by combining a spatial high-pass filtered version of the picture sample in the predetermined picture block of the current video frame with a temporally high-pass filtered picture sample, the temporally high-pass filtered picture sample being generated by combining the picture sample of the predetermined picture block of the current video frame with the picture sample in each of the predetermined picture blocks of one or more temporally preceding video frames to apply a temporal high-pass filter. The apparatus 100 may be configured, for example, to determine the sum of the plurality of picture samples.
[0058] According to an embodiment, the apparatus 100 may be configured to determine the visual activity information according to the following formula:
[0059]
[0060] in, is an intermediate image sample located at (x, y) among multiple intermediate image samples, where B k Indicates a predetermined block of picture samples having N x N.
[0061] In an embodiment, the apparatus 100 may be configured to determine the visual activity information according to the following formula:
[0062]
[0063] in, indicates visual activity information, and wherein Indicates the minimum value greater than or equal to 0.
[0064] According to an embodiment, the apparatus 100 may be, for example, an apparatus for determining a visual quality value of a video sequence. The apparatus 100 may be, for example, configured to obtain a plurality of visual activity values by determining visual activity information for each of one or more picture blocks of a plurality of picture blocks of one or more video frames of a plurality of video frames of the video sequence. In addition, the apparatus 100 may be, for example, configured to determine a visual quality value based on the plurality of visual activity values.
[0065] In an embodiment, the apparatus 100 may be configured to obtain a plurality of visual activity values by determining visual activity information for each of a plurality of picture blocks of one or more video frames of a plurality of video frames of a video sequence, for example.
[0066] According to an embodiment, the apparatus 100 may be configured to obtain a plurality of visual activity values by determining visual activity information for each of a plurality of picture blocks of each video frame in a plurality of video frames of a video sequence, for example.
[0067] In an embodiment, the apparatus 100 may be configured to determine the visual quality value of the video sequence by determining a visual quality value for one or more video frames of a plurality of video frames of the video sequence, for example.
[0068] According to an embodiment, the apparatus 100 may be configured to define a visual quality value of a video frame in a plurality of video frames of a video sequence according to the following formula:
[0069]
[0070] Among them, WPSNR c,s Indicates a visual quality value of the video frame, where W is a width of a plurality of picture samples of the video frame, where H is a height of a plurality of picture samples of the video frame, where BD is the coded bit depth of each sample, and where s[x,y] is an original picture sample at (x,y), where s c [x, y] is a decoded picture sample at (x, y), which is generated by decoding the encoding of the original picture sample at (x, y), and where
[0071]
[0072] Among them, a k is the visual activity information of the image block, where a pic >0, and where 0<β<1.
[0073] In an embodiment, the apparatus 100 may be configured to determine the visual quality value of the video sequence by determining a visual quality value for each of a plurality of video frames of the video sequence, for example.
[0074] The apparatus 100 may be configured to determine the visual quality value of the video sequence according to the following formula, for example:
[0075]
[0076] Among them, WPSNR c Indicates the visual quality value of the video sequence,
[0077] Among them, s i indicating a video frame in a plurality of video frames of a video sequence,
[0078] in, Indicates the number of video frames in a video sequence represented by si indicating a visual quality value of the one video frame, and
[0079] Wherein, F indicates the number of multiple video frames of the video sequence.
[0080] According to an embodiment, For example, the WPSNR can be defined as above c,s .
[0081] In an embodiment, the apparatus 100 may be configured to determine the visual quality value of the video sequence by averaging frame-level weighted distortions of a plurality of video frames of the video sequence, for example.
[0082] According to an embodiment, the apparatus 100 may be configured to determine the visual quality value of the video sequence according to the following formula:
[0083]
[0084] Among them, WPSNR′ c indicates a visual quality value of a video sequence, wherein F indicates the number of a plurality of video frames of the video sequence, wherein W is a width of a plurality of picture samples of the video frame, wherein H is a height of a plurality of picture samples of the video frame, wherein BD is a coding bit depth of each sample, and wherein i is an index indicating a video frame among a plurality of video frames of the video sequence, wherein k is an index indicating a picture block among a plurality of picture blocks of a video frame of the plurality of video frames of the video sequence, wherein B k is the one picture block among the multiple picture blocks of one video frame in the multiple video frames of the video sequence, where s i [x,y] is the original image sample at (x,y), where s c,i [x, y] is the decoded picture sample at (x, y), which is generated by decoding the encoding of the original picture sample at (x, y), where
[0085]
[0086] Among them, a k is the picture block B k Visual activity information, where a pic >0, and where 0<β<1.
[0087] In an embodiment, the apparatus 100 may be configured to determine the visual quality value of the video sequence according to the following formula:
[0088]
[0089] Among them, WPSNR cindicates the visual quality value of the video sequence, where F indicates the number of multiple video frames of the video sequence, and WPSNR′ c As defined above, Defined as WPSNR above c,s , where 0<δ<1.
[0090] According to an embodiment, δ may be defined as δ=0.5, for example.
[0091] In an embodiment, the apparatus 100 may be configured to determine the visual quality value of the video sequence according to the following formula:
[0092]
[0093] in, indicates a visual quality value of a video sequence, wherein F indicates the number of a plurality of video frames of the video sequence, wherein W is a width of a plurality of picture samples of the video frame, wherein H is a height of a plurality of picture samples of the video frame, wherein BD is a coding bit depth of each sample, and wherein i is an index indicating a video frame among a plurality of video frames of the video sequence, wherein k is an index indicating a picture block among a plurality of picture blocks of a video frame of the plurality of video frames of the video sequence, wherein B k is the one picture block among the multiple picture blocks of one video frame in the multiple video frames of the video sequence, where s i [x,y] is the original image sample at (x,y), where s c,i [x,y] is the decoded picture sample at (x,y). The decoded picture sample is generated by decoding the encoding of the original picture sample at (x,y).
[0094]
[0095] Among them, a k is the picture block B k Visual activity information, where a pic >0, and where 0<β<1.
[0096] According to an embodiment, for example, it may be defined as β=0.5, and
[0097] or
[0098]
[0099] In an embodiment, the apparatus 100 may be configured to determine 120 the visual activity information based on a spatial high pass filter and / or a temporal high pass filter, for example.
[0100] According to an embodiment, the apparatus 100 may be configured, for example, to: downsample a predetermined picture block of a current video frame to obtain a downsampled picture block, and apply a spatial high-pass filter and / or a temporal high-pass filter to each of a plurality of picture samples in the downsampled picture block. Alternatively, the apparatus 100 may be configured, for example, to: apply a spatial high-pass filter and / or a temporal high-pass filter only to a first group of a plurality of picture samples of a predetermined picture block but not to a second group of a plurality of picture samples of a predetermined picture block.
[0101] Furthermore, an apparatus 100 for determining visual activity information for a predetermined picture block including a video sequence according to an embodiment is provided. The video sequence includes a plurality of video frames including a current video frame.
[0102] The apparatus is configured to receive 110 a predetermined picture block of a current video frame.
[0103] Furthermore, the apparatus 100 is configured to determine 120 visual activity information according to a predetermined picture block of the current video frame and according to a spatial high pass filter and / or a temporal high pass filter.
[0104] In addition, the apparatus 100 is configured to: downsample a predetermined picture block of the current video frame to obtain a downsampled picture block, and apply a spatial high-pass filter and / or a temporal high-pass filter to each of the multiple downsampled picture samples. Alternatively, the apparatus 100 is configured to: apply the spatial high-pass filter and / or the temporal high-pass filter only to a first group of multiple picture samples of the predetermined picture block but not to a second group of multiple picture samples of the predetermined picture block.
[0105] According to an embodiment, the device 100 can be configured, for example, to: apply a spatial high-pass filter and / or a temporal high-pass filter only to a first group of multiple picture samples of a predetermined picture block but not to a second group of multiple picture samples, the first group of multiple picture samples just including those picture samples of the multiple picture samples of the predetermined picture block that are located in rows with even row indices and in columns with even column indices, and the second group of multiple picture samples of the predetermined picture block just including those picture samples of the multiple picture samples of the predetermined picture block that are located in rows with odd row indices and / or in columns with odd column indices.
[0106] Alternatively, the device 100 can be configured, for example, to apply a spatial high-pass filter and / or a temporal high-pass filter only to a first group of multiple picture samples of a predetermined picture block but not to a second group of multiple picture samples, the first group of multiple picture samples just including those picture samples of the multiple picture samples of the predetermined picture block that are located in rows with odd row indices and in columns with odd column indices, and the second group of multiple picture samples of the predetermined picture block just including those picture samples of the multiple picture samples of the predetermined picture block that are located in rows with even row indices and / or in columns with even column indices.
[0107] Alternatively, the device 100 can be configured, for example, to apply a spatial high-pass filter and / or a temporal high-pass filter only to a first group of multiple picture samples of a predetermined picture block but not to a second group of multiple picture samples, the first group of multiple picture samples just including those picture samples of the multiple picture samples of the predetermined picture block that are located in rows with odd row indices and in columns with even column indices, and the second group of multiple picture samples of the predetermined picture block just including those picture samples of the multiple picture samples of the predetermined picture block that are located in rows with even row indices and / or in columns with odd column indices.
[0108] Alternatively, the device 100 can be configured, for example, to apply a spatial high-pass filter and / or a temporal high-pass filter only to a first group of multiple picture samples of a predetermined picture block but not to a second group of multiple picture samples, the first group of multiple picture samples just including those picture samples of the multiple picture samples of the predetermined picture block that are located in rows with even row indices and in columns with odd column indices, and the second group of multiple picture samples of the predetermined picture block just including those picture samples of the multiple picture samples of the predetermined picture block that are located in rows with odd row indices and / or in columns with even column indices.
[0109] In an embodiment, the spatial high-pass filter applied only to the first group of the plurality of picture samples may be defined, for example, according to the following formula:
[0110]
[0111] in, or
[0112]
[0113] Among them, s i [x,y] indicates the first set of picture samples.
[0114] According to an embodiment, the temporal high-pass filter can be based on Or according to is defined as, where x is a first coordinate value of a sample position within a predetermined image block, where y is a second coordinate value of a sample position within a predetermined image block, where Indicates the picture sample at position (x, y) in a predetermined picture block of the current video frame, where indicates a picture sample at position (x, y) in a predetermined picture block of a first temporally previous video frame, where s i-2 [x, y] indicates a picture sample located at position (x, y) in a predetermined picture block of a second temporally previous video frame, where Indicates picture samples of a predetermined picture block generated by applying the temporal high pass filter.
[0115] Before describing additional preferred embodiments, a review of the block-based WPNSR algorithm is provided.
[0116] Similar to PSNR, the WPSNR of a codec c and a video frame (or static image stimulus) s is c,s The value is given by:
[0117]
[0118] where W and H are the luma width and height of s, respectively, BD is the encoding bit depth per sample, and
[0119]
[0120] Denotes each block B of size N·N k The sensitivity weights are derived from the spatial activity of the block a k Exported from
[0121]
[0122] Select a pic So that on a larger image set w k ≈1. Note that for all k, if w k =1, then the PSNR is obtained. See [9],
[11] for details. For video, for frame-level WPSNR c,s The values are averaged to obtain the final output:
[0123]
[0124] Where F indicates the total number of frames in the video. Usually, the WPSNR of a high-quality video is c ≈40.
[0125] Hereinafter, an extension of WPSNR for moving pictures according to an embodiment is provided.
[0126] The spatially adaptive WPSNR algorithm introduced above can be used to introduce temporal adaptation into visual activity a k The calculation of s can be easily extended to motion picture signals i , where i represents the frame index in the video. i High-pass filtering, a k Determined as:
[0127]
[0128] where h s is to use convolution h s =s*H s The high-pass filtered signal obtained by convolution uses the spatial filter H s .
[0129] In an embodiment, for example, h after time high-pass filtering can be converted into t =s*H t Add to h s To incorporate time adaptation:
[0130]
[0131]
[0132] Formula (6) is visual activity information according to an embodiment. For example, it can be considered as temporal visual activity information.
[0133] In the embodiment, a k Replace with The above equations (1) to (4), especially equation (2), also apply to
[0134] In an embodiment, two temporal high pass filters are preferably used.
[0135] The first temporal high-pass filter is a first-order FIR (finite impulse response) filter for frame rates of 30 Hz or less (e.g., 24, 25, and 30 frames per second), and is given by:
[0136]
[0137] The second temporal high-pass filter is a second-order FIR filter for frame rates above 30 Hz (e.g., 48, 50, and 60 frames per second) and is given by:
[0138]
[0139] In other words, one or two previous frame inputs are used to determine each block B of each frame s over time. k A measure of temporal activity in a task.
[0140] The relative weight parameter γ is a constant that can be determined experimentally. For example, γ = 2. k Due to the introduction of |h t | and the resulting increase in sample variance, for example, w k Modified to:
[0141]
[0142] It is worth noting that the The temporal activity component in is a relatively crude (but very low complexity) approximation of the block-level motion estimation algorithm found in all modern video codecs. Naturally, more sophisticated (but computationally more complex) temporal activity measures can be designed that use the temporal filter h t Intra-block motion between frames i, i-1, and (if applicable) i-2 is taken into account before being applied to i
[12] ,
[13] . This extension is not used here due to its high algorithmic complexity.
[0143] In the following, changes to time-varying video quality according to embodiments are provided.
[0144] As already outlined, for video sequences, the conventional approach is to average the PSNR (or WPSNR) values of individual frames to obtain a single measure for the entire sequence. For compressed video material where the visual quality varies greatly over time, this form of averaging the output of frame-level metrics may not correlate well with the MOS values given by human observers (especially non-experts). Averaging the logarithmic (W)PSNR values seems to be particularly suboptimal on video content with high overall visual quality, where, however, some brief temporal segments exhibit low quality. Such scenarios are in fact not uncommon due to the introduction of rate-adaptive video streaming. It has been found in experiments that even if most frames of the compressed video appear to be of good quality to the eyes of non-expert viewers, in this case they assign relatively low scores during the video quality assessment task. Therefore, in this case, the logarithmic domain averaged WPSNR will usually overestimate the visual quality.
[0145] One solution to this problem is to c,s The frame-level weighted distortion determined during the calculation (i.e., the denominator in Eq. (1)) is not the WPSNR c,s The values themselves are averaged:
[0146]
[0147] Figure 2 The effect of different WPSNR averaging concepts on multiple frames (levels) is shown. Non-constant line: frame level WPSNR value, Constant line: traditional log-domain averaging: proposed linear-domain (distortion) averaging.
[0148] Specifically, Figure 2 The advantage of the above linear domain arithmetic averaging compared to the traditional log domain averaging (which is equivalent to the geometric mean over multiple frames in the linear domain) is shown. For sequences with relatively constant frame WPSNR, as shown on the left, the averaging methods produce very similar outputs. On videos with different frame quality, as shown on the right, the relatively low frame WPSNR (caused by relatively high frame distortion) dominates more in the linear domain averaging than in the log domain averaging. This results in slightly lower overall WPSNR values, and as expected, the overall WPSNR value is generally not overestimated.
[0149] The weighted average of the linear domain (arithmetic) and log domain (geometric) WPSNR averages can also be used to obtain the value between two output values (e.g., Figure 2 Specifically, the overall WPSNR average value can be given by the following formula:
[0150]
[0151] Where WPSNR′ c represents the linear domain average, and 0≤δ≤1 represents the linear-to-log weighting factor. This approach adds one more degree of freedom in the WPSNR calculation, which can be used to maximize WPSNR′ c The correlation between the ′ value and the experimental MOS results.
[0152] Another alternative is to export WPSNR′ c , using the “root mean square”
[14] distortion:
[0153]
[0154] The 20 at the beginning of the equation (instead of 10) "cancels" the square root of the power of 0.5. This form of calculating the average video WPSNR data produces a result that is between the log domain solution and the linear domain solution described above, and can be very close to WPSNR' c 'Results when preferably weight δ=0.5 or weight δ≈0.5.
[0155] Hereinafter, changes to ultra-high resolution video content according to embodiments are provided.
[0156] It has been observed that, especially for ultra-high-definition (UHD) video sequences with resolutions greater than 2048×1280 luminance samples, the original WPSNR methods in [6], [7], [8], [9], and
[11] still correlate poorly with subjective MOS data (e.g., on the JVET Call for Proposal dataset
[10] ). In this regard, WPSNR performs only slightly better than the traditional PSNR metric. One possible explanation is that UHD videos are often viewed on screen sizes similar to those of lower-resolution high-definition content with only 1920×1080 (HD) or 2048×1080 (2K) luminance samples. In short, samples of UHD videos are displayed as smaller than (upscaled) samples of HD or 2K videos, a fact that should be taken into account during the visual activity calculation in the WPSNR algorithm, as described above.
[0157] The solution to the above problem is to extend the spatial high-pass filter H s , which allows it to be extended to more adjacent samples across s[x,y]. Given this, in [7], [9],
[11] , for example:
[0158]
[0159] or a scaled version thereof (multiplied by 1 / 4 in [9]), one approach is to multiply H by a factor of 2 s Upsampling is performed, i.e., increasing its size from 3×3 to 6×6 or even 7×7. However, this will significantly increase the algorithmic complexity of the spatio-temporal visual activity computation. Therefore, an alternative solution was chosen, in which if the input image or video is larger than 2048×1280 luminance samples, then i–2 、s i–1 、s i Determine visual activity on a downsampled version of In other words, for s i For multiple samples of s i For each quadruple of samples, we can simply calculate A single value of , and optionally for video, one can calculate This approach has been applied to many quality metrics, most notably MS-SSIM [2]. However, it is worth noting that in the context of this study, by properly designing the high-pass filter, the downsampling operation and the high-pass operation can be unified into one process, thereby achieving minimal algorithmic complexity. For example, using the following filter:
[0160]
[0161] or
[0162]
[0163] or
[0164]
[0165] in represents downsampling, and
[0166]
[0167]
[0168] use The derived y needs to be determined only for even values of x and y (i.e., every fourth value of the input sample set s). (or for static images input is a k ) The specific benefits of the proposed downsampling high-pass operation are as follows Figure 3 Otherwise, as mentioned above, (or a k ) can remain the same (including the division by 4N 2 ).
[0169] It should be emphasized that only at the block level spatial-temporal visual activity (or a for a single static image k ,) is temporarily applied during the calculation of the WPSNR metric (i.e., )The sum of distortions evaluated is still determined at the input resolution without any downsampling, whether the input is UHD, HD or smaller.
[0170] Figure 3 Sample-level high-pass filtering of s (left) is shown without spatial downsampling of the input signal s (middle) and with spatial downsampling of the input signal s (right) during filtering. When downsampling is performed, the 4 inputs are mapped to one high-pass output.
[0171] In the following, further embodiments of determining a quantization parameter for video encoding are described.
[0172] Furthermore, a video encoder is provided for encoding a video sequence comprising a plurality of video frames according to a quantization parameter, wherein the quantization parameter is determined according to visual activity information. Furthermore, a corresponding decoder, a computer program and a data stream are provided.
[0173] A device for varying a coding quantization parameter on a picture according to an embodiment is provided, the device comprising the device 100 for determining visual activity information as described above.
[0174] The apparatus for changing the encoding quantization parameter on a picture is configured to determine the encoding quantization parameter of a predetermined block according to visual activity information.
[0175] In an embodiment, the means for varying the encoding quantization parameter may be configured to logarithmize the visual activity information when determining the encoding quantization parameter.
[0176] Furthermore, an encoder for encoding a picture into a data stream is provided. The encoder comprises: an apparatus for varying a coding quantization parameter on a picture as described above, and an encoding stage configured to encode the picture into a data stream using the coding quantization parameter.
[0177] In an embodiment, the encoder may, for example, be configured to encode the encoded quantization parameter into the data stream.
[0178] In an embodiment, the encoder may, for example, be configured to perform a two-dimensional median filter on the encoded quantization parameter.
[0179] In an embodiment, the encoding stage may be configured, for example, to use the picture and obtain a residual signal using predictive encoding, and encode the residual signal into a data stream using an encoded quantization parameter.
[0180] In an embodiment, the encoding stage may for example be configured to: encode a picture into a data stream using predictive coding to obtain a residual signal, quantize the residual signal using an encoded quantization parameter, and encode the quantized residual signal into the data stream.
[0181] In an embodiment, the encoding stage may be configured, for example, to adjust the Lagrangian rate-distortion parameter according to the encoding quantization parameter when encoding the picture into the data stream.
[0182] In an embodiment, the means for changing the encoding quantization parameter may be configured to: perform the change of the encoding quantization parameter based on an original version of the picture.
[0183] In an embodiment, the encoding stage may support, for example, one or more of the following:
[0184] Perform block-level switching between transform domain prediction residual coding and spatial domain prediction residual coding;
[0185] Block-level prediction residual coding with block sizes whose horizontal and vertical dimensions are multiples of 4;
[0186] Loop filter coefficients are determined and encoded into a data stream.
[0187] In an embodiment, the device for varying the coding quantization parameter may, for example, be configured to encode the coding quantization parameter into a data stream in a logarithmic domain, and the encoding engine may be configured to: when encoding a picture using the coding quantization parameter, apply the coding quantization parameter in the following manner: before quantization in a non-logarithmic domain, the coding quantization parameter is used as a divisor of the signal to be quantized.
[0188] Furthermore, a decoder for decoding a picture from a data stream is provided.
[0189] The decoder comprises: an apparatus for varying an encoding quantization parameter across a picture as described above, and a decoding stage configured to decode a picture from a data stream using the encoding quantization parameter.
[0190] The decoding stage is configured to decode a residual signal from the data stream, dequantize the residual signal using the encoded quantization parameter, and decode a picture from the data stream using the residual signal and using predictive decoding.
[0191] In an embodiment, the means for varying the encoding quantization parameter may be configured, for example, to perform the variation of the encoding quantization parameter based on a picture version reconstructed by a decoding stage from a data stream.
[0192] In an embodiment, the decoding stage may support, for example, one or more of the following:
[0193] Block between transform domain prediction residual coding and spatial domain prediction residual coding
[0194] Level switching;
[0195] Block-level prediction with block sizes that are multiples of 4 in both horizontal and vertical dimensions
[0196] Residual decoding;
[0197] Decode the loop filter coefficients from the data stream.
[0198] In an embodiment, the device for changing the coding quantization parameter can be configured, for example, to determine the coding quantization parameter based on the prediction deviation in the logarithmic domain, and the decoding engine is configured to: when decoding a picture using the coding quantization parameter, convert the coding quantization parameter from the logarithmic domain to the non-logarithmic domain by exponential operation, and apply the coding quantization parameter in the non-logarithmic domain as a factor for scaling the quantized signal sent by the data stream.
[0199] Furthermore, a data stream is provided, which has pictures encoded into the data stream by an encoder as described above.
[0200] Hereinafter, specific embodiments are described in more detail.
[0201] All contemporary perceptual image and video transform encoders apply a quantization parameter (QP) for rate control, where the quantization parameter (QP) is used as a divisor to normalize the transform coefficients before quantization in the encoder and to scale the quantized coefficient values for reconstruction in the decoder. In High Efficiency Video Coding (HEVC) as specified in [8], the QP value is encoded once per picture or once per N×N block, where N=8, 16, 32 or 64, with a step size of approximately 1 dB on a logarithmic scale.
[0202] Encoder: q = round(6log2(QP) + 4), decoder: QP' = 2 (q–4) / 6 , (18)
[0204] Where q is the encoded QP index and ' indicates reconstruction. Note that QP' is also used for encoder-side normalization to avoid any error propagation effects due to QP quantization. The present embodiment adjusts the QP locally for each 64×64 coding tree unit (CTU, i.e., N=64) in the case where images and videos have a resolution equal to or less than full high definition (FHD, 1920×1080 pixels), or locally for each 64×64 or 128×128 block in the case of a resolution greater than FHD (e.g., 3840×2160 pixels).
[0205] Now, the visual activity information determined above (e.g., determined according to equation (6)) is calculated over the entire picture (or, in the case of HEVC, on-chip). ) are averaged. For example, in an FHD picture, when N=64, for each B k 510 (per block) The values are averaged.
[0206] use
[0207] In HEVC, preferably, the constant c = 2 (19)
[0209] For the logarithmic transformation, which can be implemented efficiently using a table lookup (for a general algorithm, see e.g.
[16] ), the QP offset for each block k is –q <o b ≤51–q can finally be determined as:
[0210]
[0211] In HEVC, this CTU-level offset is added to the default slice-level QP index q, and the QP' of each CTU is obtained from (1).
[0212] Alternatively, assuming that the overall multiplier λ of a picture is associated with the overall QP of the picture, the QP allocation rule is obtained, for example, according to the following formula:
[0213]
[0214] where the half square brackets indicate rounding. At this point, note that the weighting factor w may be preferentially scaled in such a way that its average value over a picture or set of pictures or a video frame is close to 1, for example. k Then, the same relationship between the picture / set Lagrangian parameter λ and the picture / set QP as for the unweighted SSE distortion can be used.
[0215] Note that in order to slightly reduce the incremental QP auxiliary information rate, it is found that applying a two-dimensional median filter to q+o b The resulting matrix of the sum is advantageously sent to the decoder as part of the coded bitstream. In a preferred embodiment, a filter with a three-tap cross kernel is used, i.e., a high-pass filter similar to (1), which calculates the median of a value based on its immediate vertical and immediate horizontal neighbors. In addition, in each CTU, for example, the median of the values can be calculated based on q+o b Update rate distortion parameter λ b =λ k To maximize coding efficiency
[0216] Or in median filtering,
[0217] In
[15] , edge blocks are classified into separate classes and quantized using dedicated custom parameters to prevent a significant increase in quantization-induced ringing around straight lines or object boundaries. When the current embodiment is used in the context of HEVC, this effect is not observed even though no comparable classification is performed. This property is most likely due to the improved efficiency of HEVC in edge coding over the MPEG-2 standard used in
[15] . Most notably, HEVC supports smaller 4× 4 block, where optional Transform Skip to quantize directly in the spatial domain, and Shape Adaptive Offset (SAO) Post-filtering is performed to reduce sideband and ringing effects during decoding [8, 10].
[0218] Since the average of the images in (6) and It is combined that the average encoding bitrate does not increase significantly due to applying the QP adaptation proposal when measured on different sets of input material. In fact, for q=37 and similar nearby values, it is found that the average bitrate does not change at all when QP adaptation is employed. Therefore, this property can be regarded as a second advantage of the present embodiment, in addition to its low computational complexity.
[0219] It should be emphasized that the present embodiment can be easily extended to non-square coding blocks. It should be obvious to those skilled in the art that in (2-4), all occurrences of (here divided by) N can be replaced by (divided by) N1·N2. 2 To consider unequal horizontal and vertical block / CTU sizes, subscripts 1 and 2 denote the horizontal and vertical block sizes.
[0220] Having described a first embodiment of visual activity information for controlling the encoding quantization parameter of a block, reference is now made to Figure 4 Describe the corresponding embodiments, Figure 4 A device for varying or adjusting a coding quantization parameter on a picture and its possible application in an encoder for encoding a picture are shown, but the details presented above are generalized at this time and although Figure 4 Embodiments of may be implemented as a modification of the HEVC codec as was the case above, but this need not necessarily be the case as outlined in more detail below.
[0221] Figure 4 An apparatus 10 for varying a coding quantization parameter QP over a picture 12 is shown, comprising a visual activity information determiner (VAI determiner) 14 and a QP determiner 16. The visual activity information determiner determines visual activity information for a predetermined block of the picture 12. The visual activity information determiner 14 calculates the visual activity information using, for example, equation (6) Furthermore, as also described above, the visual activity information determiner 14 may first high pass filter the predetermined block and then determine the visual activity information 18. The visual activity information determiner 14 may alternatively use other equations instead of using equation (6) by changing some parameters used in equation (6).
[0222] The QP determiner 16 receives the visual activity information 18 and determines a quantization parameter QP according to the visual activity information 18. As described above, the QP determiner 16 may logarithmize the visual activity information received from the visual activity information determiner 14, for example as indicated in Equation 5, although any other conversion to the logarithmic domain may alternatively be used.
[0223] The QP determiner 16 may apply a logarithmization to the low pass filter domain visual activity information.For example, the determination by the QP determiner 16 may also involve rounding or quantization, ie, rounding of the visual activity information in the logarithmic domain.
[0224] The modes of operation of the visual activity information determiner 14 and the QP determiner 16 have been discussed above with respect to specific predetermined blocks of the picture 12. For example, Figure 4 Such a predetermined block is exemplarily indicated at 20a in FIG. In the manner just outlined, the determiners 14 and 16 act on each of the multiple blocks consisting of the picture 12, thereby implementing a QP change / adjustment on the picture 12, that is, adjusting the quantization parameter QP of the picture content so as to be suitable for, for example, the human visual system.
[0225] Due to this adjustment, the resulting quantization parameter can advantageously be used by the encoding stage 22 receiving the corresponding quantization parameter QP in order to encode the corresponding block of the picture 12 into the data stream 24. Figure 4 It is exemplarily shown how the apparatus 10 may be combined with an encoding stage 22 to form an encoder 26. The encoding stage 22 encodes the picture 12 into a data stream 24 and for this purpose uses a quantization parameter QP that is varied / adjusted by the apparatus 10 across the picture 12. That is, within each block constituting the picture 12, the encoding stage 22 uses a quantization parameter determined by the QP determiner 16.
[0226] For the sake of completeness, it should be noted that the quantization parameter used by the encoding stage 22 to encode the picture 12 may not be determined solely by the QP determiner 16. Some rate controls of the encoding stage 22 may cooperate, for example, to determine the QP q to determine the quantization parameter, and the contribution of the QP determiner 16 may ultimately enter the QP offset 0 b .like Figure 4 As shown, the encoding stage 22 may, for example, encode a quantization parameter into the data stream 24. As described above, the quantization parameter may be encoded into the data stream 24 for the corresponding block (e.g., block 20a) in the logarithmic domain. The encoding stage 22 may then apply the quantization parameter in the non-logarithmic domain, i.e., in order to normalize the signal to be encoded into the data stream 24 by using the quantization parameter in the non-logarithmic domain or the linear domain as a divisor to be applied to the corresponding signal. By this measure, the quantization noise generated by the quantization performed by the encoding stage 22 is controlled on the picture 12.
[0227] As discussed above, for example, for a picture 12 or slice thereof, the quantization parameter may be encoded into the data stream 24 as a difference from a globally determined larger base quantization parameter, i.e., as an offset of 0. b in the form of and the encoding may involve entropy coding and / or differential or predictive coding, merging or similar concepts.
[0228] Figure 5 A possible structure of the encoding stage 22 is shown. Specifically, Figure 4 Involved Figure 4 The encoder 26 is a video encoder, where the picture 12 is a picture in a video 28. Here, the encoding stage 22 uses hybrid video coding. Figure 5 The encoding stage 22 of the apparatus 10 comprises a subtractor 30 which subtracts the prediction signal 32 from the signal to be encoded (e.g., picture 12). In a cascade of optional transform stages 34, quantizers 36 and entropy encoders 38, they are connected to the output of the subtractor 30 in the order in which they are mentioned. The transform stage 34 is optional and may apply a transform such as a spectral decomposition transform to the residual signal output by the subtractor 30, and the quantizer 36 quantizes the residual signal in the transform domain or in the spatial domain based on a quantization parameter varied or adjusted by the apparatus 10. The residual signal thus quantized is entropy encoded into the data stream 24 by the entropy encoder 38. A cascade of inverse quantizers 42, followed by an optional inverter 44, inverts the transform and quantization of the modules 34 and 36 or performs the inverse operations of the transform and quantization of the modules 34 and 36 so as to reconstruct the residual signal output by the subtractor 30 except for the quantization error arising from the quantization of the quantizer 36. An adder 46 adds the reconstructed residual signal to the prediction signal 32 to obtain the reconstructed signal. A loop filter 48 may optionally be present in order to improve the quality of the fully reconstructed picture. The prediction stage 50 receives the reconstructed signal part, ie an already reconstructed part of the current picture and / or an already reconstructed previously encoded picture, and outputs the prediction signal 32.
[0229] therefore, Figure 5 It is clearly shown that the quantization parameter changed or adjusted by the device 10 can be used in the encoding stage 22 to quantize the prediction residual signal. The prediction stage 50 can support different prediction modes, such as intra-frame prediction mode and inter-frame prediction mode (e.g., motion compensation prediction mode), according to which the prediction block is spatially predicted from the already encoded part, and according to the inter-frame prediction mode, the prediction block is predicted based on the already encoded picture. It should be noted that, for example, the encoding stage 22 can support turning on / off the residual transformation performed by the transformation stage 34 and the corresponding inverse transformation performed by the inverse transformer 44 in units of residual blocks.
[0230] Furthermore, it should be noted that the block granularity mentioned may be different: the block for changing the prediction mode, the block for setting the prediction parameters for controlling the corresponding prediction mode and sending these prediction parameters in the data stream 24, the block for performing, for example, a separate spectral transform at the transform stage 34, and finally the blocks 20a and 20b for changing or adjusting the quantization parameters by the device 10 may be different from each other, or at least some of them may be different from each other. For example, and as illustrated in the above example with respect to HEVC, when the spectral transform may be, for example, a DCT, a DST, a KLT, a FFT or a Hadamard transform, the size of the blocks 20a and 20b for performing the quantization parameter change / adjustment by the device 10 may be more than four times larger than the smallest block size for which the transform stage 34 performs the transform in alignment. Alternatively, it may even be larger than eight times the smallest transform block size. As indicated above, the loop filter 48 may be an SAO filter
[17] . Alternatively, an ALF filter
[18] may be used. The filter coefficients of the loop filter may be encoded into the data stream 24.
[0231] Finally, as already indicated above, the QP as output of the device 10 may be encoded into the data stream in such a way that it has passed through some two-dimensional median filtering in order to reduce the necessary data rate.
[0232] Figure 6 A possible decoder 60 is shown, which is configured to decode a reconstructed version 62 of a video 28 and / or a picture 12 from a data stream 24. Internally, the decoder comprises an entropy decoder 64, at the input of which the data stream 24 enters, followed by the modules shown, and in a manner related to Figure 6 The manner shown is interconnected so that Figure 6 The same reference numerals have been used again in , but with a prime in order to indicate that they are present in the decoder 60 rather than in the encoder stage 22. That is, the reconstructed signal 62 is obtained at the output of the adder 46' or, alternatively, at the output of the loop filter 48'. In general, Figure 5 The encoding level 22 and Figure 6 The difference between the modules of the decoder 60 relies on the fact that the encoding stage 22 determines or sets the prediction parameters, the prediction mode, the switching between the spatial domain of the residual transform and the residual coding, etc. according to some optimization scheme using, for example, a Lagrangian cost function. Via the data stream 24, the quantizer 42' obtains a quantization parameter change / adjustment advantageously selected by the device 10. It uses the quantization parameter in the non-logarithmic domain as a factor in order to scale the quantized signal, i.e. the quantized residual signal obtained from the data stream 24 by the entropy decoder 64. The Lagrangian cost function just mentioned may involve a Lagrangian rate / distortion parameter, which is a factor applied to the coding rate, the corresponding product being added to the distortion to produce the Lagrangian cost function. This Lagrangian rate / distortion parameter may be adjusted by the encoding stage 22 according to the encoding quantization parameter.
[0233] It should be noted that above and below, the term "encoding" indicates source encoding of static or moving pictures. However, the present aspect of determining a visual encoding quality value according to the invention is equally applicable to other forms of encoding, most notably channel encoding that can result in visually similar forms of visible distortion (e.g., frame error concealment (FEC) artifacts caused by activation of an FEC algorithm in the event of network packet loss).
[0234] Although some aspects have been described in the context of an apparatus, it will be clear that these aspects also represent a description of a corresponding method, wherein a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step also represent a description of a feature of a corresponding block or item or a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware device (such as a microprocessor, a programmable computer, or an electronic circuit). In some embodiments, one or more of the most important method steps may be performed by such a device.
[0235] The data stream of the present invention may be stored on a digital storage medium, or may be transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium (eg, the Internet).
[0236] Depending on certain implementation requirements, embodiments of the present invention may be implemented in hardware or in software. Implementation may be performed using a digital storage medium (e.g., a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory) having electronically readable control signals stored thereon that cooperate (or are capable of cooperating) with a programmable computer system to perform the various methods. Thus, the digital storage medium may be computer readable.
[0237] Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
[0238] Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer.The program code may, for example, be stored on a machine readable carrier.
[0239] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
[0240] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0241] A further embodiment of the inventive method is therefore a data carrier (or a digital storage medium or a computer-readable medium) on which is recorded the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium is typically tangible and / or non-transitory.
[0242] Therefore, another embodiment of the inventive method is a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may, for example, be configured to be transmitted via a data communication connection (eg, via the Internet).
[0243] A further embodiment comprises a processing means, for example a computer or a programmable logic device, configured to or adapted to perform one of the methods described herein.
[0244] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0245] Another embodiment according to the invention comprises an apparatus or system configured to transmit a computer program to a receiver (e.g. electronically or optically), the computer program being used to perform one of the methods described herein. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, comprise a file server for transmitting the computer program to the receiver.
[0246] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array can collaborate with a microprocessor to perform one of the methods described herein. Typically, these methods are preferably performed by any hardware device.
[0247] The devices described herein may be implemented using hardware devices, or using computers, or using a combination of hardware devices and computers.
[0248] The apparatus described herein or any component of an apparatus described herein may be implemented at least partially in hardware and / or software.
[0249] The methods described herein may be performed using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0250] Any component of a method described herein or an apparatus described herein may be performed at least in part by hardware and / or by software.
[0251] The above embodiments are merely illustrative of the principles of the present invention. It should be understood that modifications and variations of the arrangements and details described herein will be apparent to other persons skilled in the art. Therefore, it is intended that the scope of the present invention be limited only by the scope of the appended patent claims and not by the specific details given by way of the description and explanation of the embodiments herein.
[0252] References
[0253] [1]Z.Wang,ACBovik,HRSheikh,andA.C.Bovik,"Image QualityAssessment:From Error Visibility to Structural Similarity,"IEEE Trans.ImageProcess.,vol.13,no.4,pp.600–612,Apr.2004.[2]Z.Wang,EPSimoncelli,andA.C.Bovik,"Multiscale Structural Similarity for Image Quality assessment,” in Proc.IEEE 37 th Asilomar Conf. on Signals, Systems, and Computers, Nov. 2003.
[0254] [3] Netflix, “VMAF–Video Multimethod Assessment Fusion,” online: https: / / github.com / Netflix / vmaf, https: / / medium.com / netflix-techblog / toward-a-practical-perceptual-vi deo-quality-metric-653f208b9652.
[0255] [4] P.Philippe, W.Hamidouche, J.Fournier, and JYAubié, "AHG4: Subjective comparison of VVC and HEVC," Joint Video Experts Team, doc.JVET-O0451, Gothenburg, July 2019.
[0256] [5]Z.Li,“VMAF:the Journey Continues,”in Proc.Mile High Videoworkshop,Denver,July 2019,online:http: / / mile-high.video / files / mhv2019 / pdf / day1 / 1_08_Li.pdf.
[0257] [6]S.Bosse,C.Helmrich,H.Schwarz,D.Marpe,T.Wiegand,“Perceptuallyoptimized QP adaptation and associated distortion measure,”doc.JVET-H0047,Macau,CN,Oct. / Dec.2017.
[0258] [7]C.Helmrich,H.Schwarz,D.Marpe,T.Wiegand,“AHG10:Improvedperceptually optimized QP adaptation and associated distortion measure,”doc.JVET-K0206,Ljubljana,July 2018.
[0259] [8]C.Helmrich,H.Schwarz,D.Marpe,T.Wiegand,“AHG10:Clean-up andfinalization of perceptually optimized QP adaptation method in VTM,”doc.JVET-M0091,Marrakech,Dec.2018.
[0260] [9]J.Erfurt,C.Helmrich,S.Bosse,H.Schwarz,D.Marpe,T.Wiegand,“A Studyof the Perceptually Weighted Peak Signal-to-Noise Ratio(WPSNR)for ImageCompression,”in Proc.IEEE Int.Conf.on Image Processing(ICIP),Taipei,CN,pp.2339–2343,Sep.2019.
[0261]
[10] V.Baroncini,“Results of Subjective Testing of Responses to theJoint CfP on Video Compression Technology with Capability beyond HEVC,”doc.JVET-J0080,San Diego,Apr.2018.
[0262]
[11] C.R.Helmrich,S.Bosse,M.Siekmann,H.Schwarz,D.Marpe,and T.Wiegand,“Perceptually Optimized Bit-Allocation and Associated Distortion Measure forBlock-Based Image or Video Coding,”in Proc.IEEE Data Compression Conf.(DCC),Snowbird,pp.172–181,Mar.2019.
[0263]
[12] M.Barkowsky,J.Bialkowski,B.Eskofier,R.Bitto,and A.Kaup,“TemporalTrajectory Aware Video Quality Measure,”IEEE J.Selected Topics in SignalProcessing,vol.3,no.2,pp.266–279,Apr.2009.
[0264]
[13] K.Seshadrinatan and A.C.Bovik,“Motion Tuned Spatio-TemporalQuality Assessment of Natural Videos,”IEEE Trans.Image Processing,vol.19,no.2,pp.335–350,Feb.2010.
[0265]
[14] D.McK.Kerslake,The Stress of Hot Environments,p.37,CambridgeUniversity Press,1972,online:https: / / books.google.de / books?id=FQo9AAAAIAAJ&pg=PA37&lpg=PA37&dq=%22square+mean+root%22&q=%22square%20mean%20root%22&f=false#v=snippet&q=%22square%20mean%20root%22&f=false.
[0266]
[15] W.Osberger,S.Hammond,and N.Bergmann,“An MPEG EncoderIncorporating Perceptually Based Quantisation,”in Proc.IEEE AnnualConf.Speech&Image Technologies for Comput.&Telecomm.,Brisbane,vol.2,pp.731–734,1997.
[0267]
[16] S.E.Anderson,“Bit Twiddling Hacks,”Stanford University,2005.http: / / graphics.stanford.edu / ~seander / bithacks.html
[0268]
[17] C.-M.Fu,E.Alshina,A.Alshin,Y.Huang,C.Chen,C.Tsai,C.Hsu,S.Lei,J.Park,and W.-J.Han,“Sample Adaptive Offset in the HEVC Standard,”IEEETrans.Circuits&Syst.for Video Technology,vol.22,no.12,pp.1755–1764,Dec.2012.
[0269]
[18] C.-Y.Tsai,C.-Y.Chen,T.Yamakage,I.S.Chong,Y.-W.Huang,C.-M.Fu,T.Itoh,T.Watanabe,T.Chujoh,M.Karczewicz,and S.-M.Lei,“Adaptive Loop Filteringfor Video Coding,”IEEE J.Selected Topics in Signal Process.,vol.7,no.6,pp.934–945,Dec.2013。
Claims
1. A device (100) for determining visual activity information for a predetermined picture block of a video sequence, the video sequence comprising a plurality of video frames, the plurality of video frames comprising a current video frame and one or more temporally preceding video frames, wherein: The one or more temporally preceding video frames temporally precede the current video frame, wherein the apparatus is configured to: receiving (110) a predetermined picture block of each of the one or more temporally preceding video frames and a predetermined picture block of the current video frame, determining (120) the visual activity information based on a predetermined picture block of the current video frame and based on a predetermined picture block of each of the one or more temporally preceding video frames and based on a temporal high pass filter, The apparatus (100) is configured to obtain a plurality of visual activity values by determining visual activity information for each of one or more picture blocks of a plurality of picture blocks of one or more video frames of the plurality of video frames of the video sequence, wherein the apparatus (100) is configured to determine a visual quality value for a video frame of one or more of the plurality of video frames of the video sequence based on the plurality of visual activity values, The apparatus (100) is configured to define a visual quality value of the video frame among the plurality of video frames of the video sequence according to the following formula: in indicating a visual quality value of the video frame, wherein W is a width of a plurality of picture samples of the video frame, wherein H is a height of the plurality of picture samples of the video frame, wherein BD is a coding bit depth of each sample, and wherein s indicates the video frame, k is an index indicating one of a plurality of picture blocks of the video frame, wherein is the original image sample at (x, y), where is a decoded picture sample at (x, y), the decoded picture sample being generated by decoding the encoding of the original picture sample at (x, y), and wherein represents the weighting factor, ,in is the visual activity information of the picture block, where It is obtained based on W, H and BD. , and among them ,in Indicates the picture block having N x N picture samples, where N indicates a positive integer value.
2. The device (100) according to claim 1, wherein: The apparatus (100) is further configured to determine a visual quality value for the video sequence by determining a visual quality value for a video frame of one or more of the plurality of video frames of the video sequence, The apparatus (100) is configured to determine a visual quality value of the video sequence by determining a visual quality value for each of the plurality of video frames of the video sequence, wherein the apparatus (100) is configured to determine the visual quality value of the video sequence according to the following formula: , in indicates a visual quality value of the video sequence, wherein indicates a video frame of the plurality of video frames of the video sequence, wherein Indicating that the plurality of video frames of the video sequence are A visual quality value of the one video frame indicated by F indicates the number of the multiple video frames of the video sequence, and i indicates a time index.
3. The device (100) according to claim 1, in, ,and in, or 。 4. A device (100) for determining visual activity information for a predetermined picture block of a video sequence, the video sequence comprising a plurality of video frames, the plurality of video frames comprising a current video frame and one or more temporally preceding video frames, wherein: The one or more temporally preceding video frames temporally precede the current video frame, wherein the apparatus is configured to: receiving (110) a predetermined picture block of each of the one or more temporally preceding video frames and a predetermined picture block of the current video frame, determining (120) the visual activity information based on a predetermined picture block of the current video frame and based on a predetermined picture block of each of the one or more temporally preceding video frames and based on a temporal high pass filter, The device (100) is a device for determining a visual quality value of the video sequence. The apparatus (100) is configured to obtain a plurality of visual activity values by determining visual activity information for each of one or more picture blocks of a plurality of picture blocks of one or more video frames of the plurality of video frames of the video sequence, wherein the apparatus (100) is configured to determine the visual quality value based on the plurality of visual activity values, The apparatus (100) is configured to determine the visual quality value of the video sequence according to the following formula: , in indicating a visual quality value of the video sequence, wherein F indicates the number of the plurality of video frames of the video sequence, wherein W is a width of a plurality of picture samples of the video frame, wherein H is a height of the plurality of picture samples of the video frame, wherein BD is a coding bit depth of each sample, and wherein i is an index indicating one of the plurality of video frames of the video sequence, wherein k is an index indicating one of the plurality of picture blocks of one of the plurality of video frames of the video sequence, wherein is the one picture block of the plurality of picture blocks of one video frame of the plurality of video frames of the video sequence, having N x N picture samples, where N indicates a positive integer value, wherein is the original image sample at (x, y), where is a decoded picture sample at (x, y), the decoded picture sample being generated by decoding the encoding of the original picture sample at (x, y), where represents the weighting factor, ,in is the picture block Visual activity information, where It is obtained based on W, H and BD. , and among them .
5. The device (100) according to claim 4, in, ,and in, or 。 6. A device (100) for determining visual activity information for a predetermined picture block of a video sequence, the video sequence comprising a plurality of video frames, the plurality of video frames comprising a current video frame and one or more temporally preceding video frames, wherein: The one or more temporally preceding video frames temporally precede the current video frame, wherein the apparatus is configured to: receiving (110) a predetermined picture block of each of the one or more temporally preceding video frames and a predetermined picture block of the current video frame, determining (120) the visual activity information based on a predetermined picture block of the current video frame and based on a predetermined picture block of each of the one or more temporally preceding video frames and based on a temporal high pass filter, The device (100) is a device for determining a visual quality value of the video sequence. The apparatus (100) is configured to obtain a plurality of visual activity values by determining visual activity information for each of one or more picture blocks of a plurality of picture blocks of one or more video frames of the plurality of video frames of the video sequence, wherein the apparatus (100) is configured to determine the visual quality value based on the plurality of visual activity values, The apparatus (100) is configured to determine the visual quality value of the video sequence according to the following formula: in indicating a visual quality value of the video sequence, wherein F indicates the number of the plurality of video frames of the video sequence, wherein W is a width of a plurality of picture samples of the video frame, wherein H is a height of the plurality of picture samples of the video frame, wherein BD is a coding bit depth of each sample, and wherein i is an index indicating one of the plurality of video frames of the video sequence, wherein k is an index indicating one of the plurality of picture blocks of one of the plurality of video frames of the video sequence, wherein is the one picture block of the plurality of picture blocks of one video frame of the plurality of video frames of the video sequence, having N x N picture samples, where N indicates a positive integer value, wherein is the original image sample at (x, y), where is a decoded picture sample at (x, y), the decoded picture sample being generated by decoding the encoding of the original picture sample at (x, y), where represents the weighting factor, ,in is the picture block Visual activity information, where It is obtained based on W, H and BD. , and among them .
7. The device (100) according to claim 6, in, ,and in, or 。 8. An apparatus (100) for determining visual activity information for a predetermined picture block of a video sequence, the video sequence comprising a plurality of video frames, the plurality of video frames comprising a current video frame and one or more temporally preceding video frames, wherein: The one or more temporally preceding video frames temporally precede the current video frame, wherein the apparatus is configured to: receiving (110) a predetermined picture block of each of the one or more temporally preceding video frames and a predetermined picture block of the current video frame, determining (120) the visual activity information based on a predetermined picture block of the current video frame and based on a predetermined picture block of each of the one or more temporally preceding video frames and based on a temporal high pass filter, The device (100) is a device for determining a visual quality value of the video sequence. The apparatus (100) is configured to obtain a plurality of visual activity values by determining visual activity information for each of one or more picture blocks of a plurality of picture blocks of one or more video frames of the plurality of video frames of the video sequence, wherein the apparatus (100) is configured to determine the visual quality value based on the plurality of visual activity values, The apparatus (100) is configured to determine the visual quality value of the video sequence according to the following formula: in indicates a visual quality value of the video sequence, wherein As defined in claim 4, for indicating a first visual quality value of the video sequence, wherein Indicates a video frame of a plurality of video frames of the video sequence, wherein As defined in claim 1, for indicating a plurality of video frames of the video sequence by The visual quality value of the one video frame indicated by indicates the number of the plurality of video frames of the video sequence, wherein i indicates a time index, wherein .
9. The device (100) according to claim 8, wherein .
10. A method for determining visual activity information for a predetermined picture block of a video sequence, the video sequence comprising a plurality of video frames, the plurality of video frames comprising a current video frame and one or more temporally preceding video frames, wherein: The one or more temporally preceding video frames temporally precede the current video frame, wherein the method comprises: receiving (110) a predetermined picture block of each of the one or more temporally preceding video frames and a predetermined picture block of the current video frame, determining (120) the visual activity information based on a predetermined picture block of the current video frame and based on a predetermined picture block of each of the one or more temporally preceding video frames and based on a temporal high pass filter, The method comprises: obtaining a plurality of visual activity values by determining visual activity information for each of one or more picture blocks of a plurality of picture blocks of one or more video frames of the plurality of video frames of the video sequence, wherein a visual quality value is determined for a video frame of one or more of the plurality of video frames of the video sequence according to the plurality of visual activity values, The method comprises: defining a visual quality value of the video frame among the plurality of video frames of the video sequence according to the following formula: in indicating a visual quality value of the video frame, wherein W is a width of a plurality of picture samples of the video frame, wherein H is a height of the plurality of picture samples of the video frame, wherein BD is a coding bit depth of each sample, and wherein s indicates the video frame, k is an index indicating one of a plurality of picture blocks of the video frame, wherein is the original image sample at (x, y), where is a decoded picture sample at (x, y), the decoded picture sample being generated by decoding the encoding of the original picture sample at (x, y), and wherein represents the weighting factor, ,in is the visual activity information of the picture block, where It is obtained based on W, H and BD. , and among them ,in Indicates the picture block having N x N picture samples, where N indicates a positive integer value.
11. The method according to claim 10, wherein: The method further comprises determining a visual quality value of the video sequence by determining a visual quality value for a video frame of one or more of the plurality of video frames of the video sequence, The method comprises: determining a visual quality value of the video sequence by determining a visual quality value for each of the plurality of video frames of the video sequence, wherein the method comprises: determining the visual quality value of the video sequence according to the following formula: , in indicates a visual quality value of the video sequence, wherein indicates a video frame of the plurality of video frames of the video sequence, wherein Indicating that the plurality of video frames of the video sequence are A visual quality value of the one video frame indicated by F indicates the number of the multiple video frames of the video sequence, and i indicates a time index.
12. A method for determining visual activity information for a predetermined picture block of a video sequence, the video sequence comprising a plurality of video frames, the plurality of video frames comprising a current video frame and one or more temporally preceding video frames, wherein: The one or more temporally preceding video frames temporally precede the current video frame, wherein the method comprises: receiving (110) a predetermined picture block of each of the one or more temporally preceding video frames and a predetermined picture block of the current video frame, determining (120) the visual activity information based on a predetermined picture block of the current video frame and based on a predetermined picture block of each of the one or more temporally preceding video frames and based on a temporal high pass filter, wherein the method is a method for determining a visual quality value of the video sequence, The method comprises: obtaining a plurality of visual activity values by determining visual activity information for each of one or more picture blocks of a plurality of picture blocks of one or more video frames of the plurality of video frames of the video sequence, wherein determining the visual quality value is performed based on the plurality of visual activity values, The method comprises: determining the visual quality value of the video sequence according to the following formula: , in indicating a visual quality value of the video sequence, wherein F indicates the number of the plurality of video frames of the video sequence, wherein W is a width of a plurality of picture samples of the video frame, wherein H is a height of the plurality of picture samples of the video frame, wherein BD is a coding bit depth of each sample, and wherein i is an index indicating one of the plurality of video frames of the video sequence, wherein k is an index indicating one of the plurality of picture blocks of one of the plurality of video frames of the video sequence, wherein is the one picture block of the plurality of picture blocks of one video frame of the plurality of video frames of the video sequence, having N x N picture samples, where N indicates a positive integer value, wherein is the original image sample at (x, y), where is a decoded picture sample at (x, y), the decoded picture sample being generated by decoding the encoding of the original picture sample at (x, y), where represents the weighting factor, ,in is the picture block Visual activity information, where It is obtained based on W, H and BD. , and among them .
13. A method for determining visual activity information for a predetermined picture block of a video sequence, the video sequence comprising a plurality of video frames, the plurality of video frames comprising a current video frame and one or more temporally preceding video frames, wherein: The one or more temporally preceding video frames temporally precede the current video frame, wherein the method comprises: receiving (110) a predetermined picture block of each of the one or more temporally preceding video frames and a predetermined picture block of the current video frame, determining (120) the visual activity information based on a predetermined picture block of the current video frame and based on a predetermined picture block of each of the one or more temporally preceding video frames and based on a temporal high pass filter, wherein the method is a method for determining a visual quality value of the video sequence, The method comprises: obtaining a plurality of visual activity values by determining visual activity information for each of one or more picture blocks of a plurality of picture blocks of one or more video frames of the plurality of video frames of the video sequence, wherein determining the visual quality value is performed based on the plurality of visual activity values, The method comprises: determining the visual quality value of the video sequence according to the following formula: in indicating a visual quality value of the video sequence, wherein F indicates the number of the plurality of video frames of the video sequence, wherein W is a width of a plurality of picture samples of the video frame, wherein H is a height of the plurality of picture samples of the video frame, wherein BD is a coding bit depth of each sample, and wherein i is an index indicating one of the plurality of video frames of the video sequence, wherein k is an index indicating one of the plurality of picture blocks of one of the plurality of video frames of the video sequence, wherein is the one picture block of the plurality of picture blocks of one video frame of the plurality of video frames of the video sequence, having N x N picture samples, where N indicates a positive integer value, wherein is the original image sample at (x, y), where is a decoded picture sample at (x, y), the decoded picture sample being generated by decoding the encoding of the original picture sample at (x, y), where represents the weighting factor, ,in is the picture block Visual activity information, where It is obtained based on W, H and BD. , and among them .
14. A method for determining visual activity information for a predetermined picture block of a video sequence, the video sequence comprising a plurality of video frames, the plurality of video frames comprising a current video frame and one or more temporally preceding video frames, wherein: The one or more temporally preceding video frames temporally precede the current video frame, wherein the method comprises: receiving (110) a predetermined picture block of each of the one or more temporally preceding video frames and a predetermined picture block of the current video frame, determining (120) the visual activity information based on a predetermined picture block of the current video frame and based on a predetermined picture block of each of the one or more temporally preceding video frames and based on a temporal high pass filter, wherein the method is a method for determining a visual quality value of the video sequence, The method comprises: obtaining a plurality of visual activity values by determining visual activity information for each of one or more picture blocks of a plurality of picture blocks of one or more video frames of the plurality of video frames of the video sequence, wherein determining the visual quality value is performed based on the plurality of visual activity values, The method comprises: determining the visual quality value of the video sequence according to the following formula: in indicates a visual quality value of the video sequence, wherein As defined in claim 12, for indicating a first visual quality value of the video sequence, wherein Indicates a video frame of a plurality of video frames of the video sequence, wherein As defined in claim 10, for indicating a plurality of video frames of the video sequence by The visual quality value of the one video frame indicated by indicates the number of the plurality of video frames of the video sequence, wherein i indicates a time index, wherein .
15. A non-transitory computer readable medium comprising computer readable instructions which, when executed on a computer or a signal processor, cause the computer or signal processor to perform the method according to any one of claims 10 to 14.
Citation Information
Patent Citations
Image anti-shake in digital cameras
CN101159813A