Method for Reference-Free Video Quality Prediction

By using neural networks and video processing blocks on the client side, high-level features are extracted and neural networks are trained, the universality problem of reference-free video quality prediction is solved, and adaptability and high-accuracy video quality prediction are achieved for different video compression standards and decoding structures.

CN115700750BActive Publication Date: 2025-06-03AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210849455.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-07-21
Filing Date
2022-07-19
Publication Date
2025-06-03
Estimated Expiration
2042-07-19

AI Technical Summary

Technical Problem

The prior art is difficult to implement reference-free video quality prediction on the client side, especially under different video compression standards and decoding structures, and lacks a general method for video quality prediction.

Method used

Using a neural network combined with video processing block method, the neural network is trained to predict video quality by receiving input bitstreams and extracting advanced features by training data.

Benefits of technology

It realizes video quality prediction without reference on the client side, can adapt to different video compression standards and decoding structures, and improves the accuracy and versatility of video quality prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115700750B_ABST
    Figure CN115700750B_ABST
Patent Text Reader

Abstract

This application is directed to a method for reference - free video quality prediction. A system for reference - free video quality prediction includes: a video processing block configured to receive an input bitstream and generate a first vector; and a neural network configured to provide a predicted quality vector after being trained using training data. The training data includes the first vector and a second vector, and the elements of the first vector include high - level features extracted according to high - level syntax processing of the input bitstream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This description generally relates to video processing, and more particularly, to methods for no-reference video quality prediction. Background Art

[0002] For remote monitoring of client-side video quality, no-reference video quality prediction has become increasingly important. Using no-reference video quality prediction, the video quality can be estimated without having to view the received video or require the original video content. By being able to automatically diagnose video quality problems reported by end-users, no-reference video quality prediction can help reduce customer support costs. A common practice is to perform video quality analysis on the decoded video sequence in the pixel domain. More accurate methods can use not only pixel-domain information but also bitstream characteristics measured at different decoding stages.

[0003] In the past few decades, several video compression standards have been developed, such as the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) Moving Picture Experts Group (MPEG) and the International Telecommunication Union (ITU-)T joint international standard MPEG-2 / H.262, Advanced Video Coding (AVC) / H.264, High Efficiency Video Coding (HEVC) / H.265, and Versatile Video Coding (VVC) / H.266, as well as the industry standards VP8, VP9, and Alliance for Open Media Video 1 (AV1). End-users can receive video content compressed in a variety of video formats. Although these standards offer different levels of compression efficiency and differ in details from each other, all of these standards use a common block-based hybrid coding structure. The common coding structure makes it possible to develop a general method for no-reference video quality prediction on the client side. For example, VVC (the latest video compression standard from MPEG / ITU-T) still employs a block-based hybrid coding structure. In VVC, a picture is partitioned into coding tree units (CTUs), which can be up to 128x128 pixels in size. The CTUs are further decompressed into coding units (CUs) of different sizes by using a so-called quadtree plus binary and ternary tree (QTBTT) recursive block partitioning structure. The CUs can have a four-way split by using quadtree partitioning, a two-way split by adapting horizontal or vertical binary tree partitioning, or a three-way split by using horizontal or vertical ternary tree partitioning. The CUs can be as large as the CTUs and as small as a 4x4 pixel block size. Summary of the Invention

[0004] In one aspect, the present application is directed to a system for reference - free video quality prediction, the system comprising: a video processing block that receives an input bitstream and generates a first vector; and a neural network configured to provide a predicted quality vector after being trained using training data, wherein: the training data includes the first vector and a second vector, and the elements of the first vector include high - level features extracted according to high - level syntax processing of the input bitstream.

[0005] In another aspect, the present application is directed to a method for reference - free video quality prediction, the method comprising: receiving a video data stream; generating a feature vector by decoding the video data stream and extracting features; and configuring a neural network to provide a predicted quality vector after being trained using training data, wherein: the training data includes the feature vector and a ground - truth video quality vector, and generating the feature vector includes high - level syntax processing of the video data stream to extract high - level feature elements.

[0006] In another aspect, the present application is directed to a method for training a neural network for reference - free video quality prediction, the method comprising: compressing an input data stream using an encoder - decoder chain to generate reconstructed data; calculating a first vector using the reconstructed data and the input data stream; decoding the output of the last encoder of the encoder - decoder chain using a decoder for high - level and block - level feature extraction to generate a feature vector; and training neural network parameters by processing a loss function. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The specific features of the technology are set forth in the appended claims. However, for explanatory purposes, several embodiments of the technology are set forth in the following figures.

[0008] Figure 1 is a high - level diagram illustrating an example of a neural - network - based reference - free video quality prediction system in accordance with various aspects of the present technology.

[0009] Figure 2 is a diagram illustrating an example of a versatile video codec (VVC) decoder in accordance with various aspects of the present technology.

[0010] Figure 3 is a diagram illustrating an example of a hierarchical decoding structure in accordance with various aspects of the present technology.

[0011] Figure 4 is a schematic diagram illustrating an example of a neural network for video quality prediction in accordance with various aspects of the present technology.

[0012] Figure 5 is a diagram illustrating an example of a process for training data generation and network training in accordance with various aspects of the present technology.

[0013] Figure 6 is a flowchart illustrating a method for no-reference video quality prediction in accordance with various aspects of the present technology.

[0014] Figure 7 is a block diagram illustrating an electronic system in which one or more aspects of the present technology may be implemented. Detailed Description

[0015] The detailed description set forth below is intended as a description of various configurations of the present technology and is not intended to represent the only configurations in which the present technology may be practiced. The accompanying drawings are incorporated into and constitute a part of the embodiments, which include specific details for providing a thorough understanding of the present technology. However, the present technology is not limited to the specific details set forth herein and may be practiced without these specific details. In some instances, structures and components are shown in block diagram form to avoid obscuring the concepts of the present technology.

[0016] The present technology is directed to methods and systems for no-reference video quality prediction. The disclosed technology implements no-reference video quality prediction by using a neural network that is trained to predict the root mean square error (RMSE) value between a reconstructed picture and an original picture after an in-loop filter, as described in more detail below. The RMSE value can be converted into a video quality score, such as a peak signal-to-noise ratio (PSNR) value.

[0017] Figure 1 is a high-level diagram illustrating an example of a neural-network-based no-reference video quality prediction system 100 in accordance with various aspects of the present technology. The neural-network-based no-reference video quality prediction system 100 (hereinafter referred to as system 100) includes a video processing block 110 and a neural network 120. The video processing block 110 is a decoding and feature extraction block, and some aspects of which (video decoding aspects) are discussed below with respect to Figure 2 The video processing block 110 provides a feature vector x(t) from an input bitstream corresponding to a picture. The elements of the feature vector x(t) are partitioned into two categories, namely, high-level features extracted from high-level syntax processing and block-level features obtained from block-level decoding processing.

[0018] Advanced features may include a transcoding indicator, codec type, picture decoding type, picture resolution, frame rate, bit depth, chroma format, compressed picture size, advanced quantization parameter (qp), average temporal distance, and temporal layer ID. The transcoding indicator determines whether to transcode the current picture. Transcoding means that the video can first be compressed and decompressed in one format (e.g., AVC / H.264) and then recompressed into the same or a different format (e.g., HEVC / H.265). This information is typically not available in the bitstream but can be communicated from the service area to the client via an external component. The codec type may include VVC / H.266, HEVC / H.265, AVC / H.264, VP8, VP9, AV1, etc. Each codec type may be assigned a codec ID. The picture decoding type may include I-, B-, and P-pictures, and each picture type may be assigned an ID. For example, the picture resolution may be 4K UHD, 1080p HD, 720p HD, etc. An ID may be assigned based on the luminance samples in the picture. Examples of frame rates may include 60, 50, 30, 20 frames per second. The frame rate is normalized to, for example, 120 frames per second. For example, the bit depth may be 8-bit or 10-bit and is normalized to 10-bit. For example, the chroma format may be 4:2:0, and each chroma format may be assigned an ID, e.g., 0 for the 4:2:0 chroma format. The compressed picture size is normalized by the luminance picture size to produce a bits per pixel (bbp) value. The advanced quantization parameter (qp) is the average qp of the picture obtained by parsing the quantization parameter in the slice header of the picture. The list0 average temporal distance represents the average temporal distance between the current picture and its forward (i.e., list0) reference picture, which is obtained by parsing the slice-level reference picture list (RPL) of the current picture. If the list0 reference picture does not exist, it is set to 0. The list1 average temporal distance represents the average temporal distance between the current picture and its backward (i.e., list1) reference picture, which is obtained by parsing the slice-level RPL of the current picture. If the list1 reference picture does not exist, it is set to 0. The temporal layer ID corresponds to the current picture. As discussed below, the temporal ID of the picture is assigned based on the hierarchical decoding structure.

[0019] The neural network 120 provides a predicted quality vector p(t), which is a neural network-based inference that enables prediction of the video quality of a picture. The predicted video quality can be measured using any suitable video metric, such as PSNR, Structural Similarity Index Measure (SSIM), Multi-Scale Structural Similarity Index Measure (MS-SSIM), Video Multimethod Assessment Fusion (VMAF), and Mean Opinion Score (MOS), depending on the video quality selected for neural network training. The predicted video quality of consecutive pictures can also be combined to produce a video quality prediction for a video segment.

[0020] Figure 2 FIG. is a diagram illustrating an example of a Versatile Video Coding (VVC) decoder 200 (an example of a video decoding block) in accordance with various aspects of the present technology. The VVC decoder 200 (hereinafter referred to as decoder 200) includes an advanced syntax processing 202 and a block-level processing 204, and the block-level processing 204 includes an entropy decoding engine 210, an inverse quantization block 220, an inverse transform block 230, an intra prediction mode reconstruction block 240, an intra prediction block 250, an in-loop filter block 260, an inter prediction block 270, and a motion data reconstruction block 280.

[0021] The advanced syntax processing block 202 includes suitable logic and buffer circuits to receive an input bitstream 202 and parse the advanced syntax elements to generate advanced features 203, which include transcoding indicators, codec types, picture coding types, picture resolutions, frame rates, bit depths, chroma formats, compressed picture sizes, advanced qps, average temporal distances, and temporal layer IDs, as discussed above with respect to Figure 1 The input bitstream 202 is composed of the output of the last encoder in the encode-decode chain (not shown for clarity). The advanced syntax elements may include Sequence Parameter Sets (SPSs), Picture Parameter Sets (PPSs), Video Parameter Sets (VPSs), Picture Headers (PHs), Slice Headers (SHs), Adaptive Parameter Sets (APSs), Supplemental Enhancement Information (SEI) messages, and so on. The decoded advanced information is then used to configure the decoder 200 to perform the block-level decoding process.

[0022] At the block level, the entropy decoding engine 210 decodes the incoming bitstream 202 and delivers decoded symbols that include quantized transform coefficients 212 and control information 214. The control information includes the intra prediction mode delta (relative to the most probable mode), the inter prediction mode, the motion vector difference (MVD, relative to the motion vector predictor), the merge index (merge_idx), the quantization scale, and the in-loop filter parameters 216. The intra prediction reconstruction block 240 reconstructs the intra prediction mode 242 of the coding unit (CU) by deriving the most probable mode (MPM) list and using the decoded intra prediction mode delta. The motion data reconstruction block 280 reconstructs the motion data 282 (e.g., motion vectors, reference indices (plural)) by deriving the advanced motion vector predictor (AMVP) list or the merge / skip list and using the MVD. The decoded motion data 282 of the current picture can serve as the temporal motion vector predictor (TMVP) 274 for decoding future pictures and is stored in the decoded picture buffer (DPB).

[0023] The quantized transform coefficients 212 are delivered to the inverse quantization block 220 and then to the inverse transform block 230 to reconstruct the residual block 232 of the CU. Based on the signaled intra or inter prediction mode, the decoder 200 can perform intra prediction or inter prediction (i.e., motion compensation) to generate the prediction block 282 of the CU. The residual block 232 is then added to the prediction block 282 to produce the reconstructed CU before the in-loop filter. The in-loop filter 260 performs in-loop filtering on the reconstructed block, such as deblocking filtering, sample adaptive offset (SAO) filtering, and adaptive loop filtering (ALF), to produce the reconstructed CU after the in-loop filter 262. The reconstructed picture 264 is stored in the DPB to serve as a reference picture for motion compensation of future pictures and is also sent to the display.

[0024] The block-based nature of video decoding processing enables it to extract features on the decoder side without incurring additional processing latency or increasing memory bandwidth consumption. The features extracted at the block level help improve the video quality prediction accuracy when compared to pixel-domain-only prediction methods.

[0025] Referring to the block-level processing 204, the block-level features can include the following: 1) the percentage of intra-coded blocks in the current picture delivered by the entropy decoding engine 210; 2) the percentage of inter-coded blocks in the current picture delivered by the entropy decoding engine 210; 3) the average block-level qp of the current picture delivered by the entropy decoding engine 210; 4) the maximum block-level qp of the current picture delivered by the entropy decoding engine 210; and 5) the minimum block-level qp of the current picture delivered by the entropy decoding engine 210. The block-level features can also include the standard deviation of the horizontal motion vectors of the current picture calculated in the motion data reconstruction block 280. For example, mvx0(i), i = 0, 1,..., mv cnt0-1 and mvx1(i), where i = 0, 1, ..., mv cnt1 -1 is set as the horizontal motion vectors of list0 and list1 for reconstructing the current picture, and mv cnt0 and mv cnt1 are respectively set as the numbers of block vectors of picture list0 and list1, and the vectors are normalized at the block level by using the temporal distance between the predicted reference blocks of the current prediction unit (PU). In this case, the standard deviation sd of the horizontal motion vectors of the current picture is calculated by the following equation max :

[0026]

[0027] Another feature that the block-level feature can include is the average motion vector magnitude of the current picture calculated in the motion data block 280. For example, (mvx0(i), mvy0(i)), where i = 0, 1, ..., mv cnt0 -1 and (mvx1(i), mvy1(i)), where i = 0, 1, ..., mv cnt1 -1 are set as the motion vectors of list0 and list1 for reconstructing the current picture, mv cnt0 and mv cnt1 are respectively set as the numbers of block vectors of picture list0 and list1, and the vectors are normalized at the block level by using the temporal distance between the predicted reference blocks of the current PU. In this case, the average motion vector magnitude avg is calculated by the following equation mv :

[0028]

[0029] The block-level feature can also include the average absolute magnitude of the low-frequency inverse quantization transform coefficients of the current picture calculated in the inverse quantization block 220. For example, if the transform unit (TU) size is W*H, then the coefficients are defined as low-frequency coefficients when their indices in the TU are less than W*H / 2 in the scanning order (i.e., the coefficient decoding order in the bitstream). The absolute magnitude is obtained by averaging the Y, U, and V components of the picture. Of course, the individual magnitudes can be calculated separately for the Y, U, and V components.

[0030] Another possible feature of the block-level feature is the average absolute magnitude of the high-frequency inverse quantized transform coefficients of the current picture calculated in the inverse quantization block 220. For example, if the TU size is W*H, then the coefficients are defined as high-frequency coefficients when their indices in the TU are greater than or equal to W*H / 2 in the scanning order (or the coefficient decoding order in the bitstream). The absolute magnitude is obtained by averaging the Y, U, and V components of the picture. Of course, the individual magnitudes can be calculated separately for the Y, U, and V components.

[0031] The block-level feature may further include the standard deviation of the prediction residuals of the current picture calculated separately by the inverse transform block 230 for the Y, U, and V components. Let resid(i, j), for i = 0, 1, ..., picHeight-1, j = 0, 1, ..., picWidth-1 be the prediction residual picture of the Y, U, or V component, and calculate the standard deviation sd of the prediction residuals of the component through the following equation resid :

[0032]

[0033] Another feature that the block-level feature may include is the root mean square error (RMSE) value between the reconstructed pictures before and after the in-loop filter calculated separately by the in-loop filter block 260 for the Y, U, and V components. For example, if the codec (e.g., MPEG-2) does not have an in-loop filter or the in-loop filter is turned off, then set the RMSE to 0 for the picture. Let dec(i, j) and rec(i, j), for i = 0, 1, ..., picHeight-1, j = 0, 1, ..., picWidth-1 be the reconstructed Y, U, or V component pictures before and after the in-loop filter respectively. Then, calculate the RMSE rmse of the component through the following equation

[0034]

[0035] The block-level feature may further include the standard deviation of the reconstructed picture after the in-loop filter calculated separately by the in-loop filter block for the Y, U, and V components. For example, let rec(i, j), for i = 0, 1, ..., picHeight-1, j = 0, 1, ..., picWidth-1 be the reconstructed Y, U, or V component picture after the in-loop filter. Then, calculate the standard deviation sd of the reconstructed component picture through the following equation rec :

[0036]

[0037] Another feature that may be included in the block-level feature is the edge sharpness of the reconstructed picture after the in-loop filter that can be calculated by the in-loop filter for the Y, U, and V components. For example, let rec(i, j), G x (i, j) and G y (i, j) for i = 0, 1, ..., picHeight-1, j = 0, 1, ..., picWidth-1

[0038] Are respectively set as the Y, U or V component pictures after the in-loop filter and their corresponding horizontal / vertical edge sharpness maps. Then, the edge sharpness edge of the reconstructed component pictures is calculated by the following equation sharpness :

[0039]

[0040] Where the edge sharpness map G can be calculated by the following equation (for example, using a Sobel filter) x (i, j) and G y (i, j) for i = 0, 1,..., picHeight - 1, j = 0, 1,..., picWidth - 1:

[0041]

[0042] It should be noted that in the above equations, the reconstructed picture samples used to calculate G x (i, j) and G y (i, j) can exceed the picture boundary, and the closest picture boundary samples can be filled in for the unavailable samples. Another solution is to simply avoid calculating G x (i, j) and G y (i, j) along the picture boundary and set it to 0, that is,

[0043]

[0044] Figure 3 Is a diagram illustrating an example of a hierarchical decoding structure 300 according to various aspects of the present technology. The vertical columns show the temporal ID (Tid) of the pictures, which is related to temporal scalable decoding and in some aspects is assigned based on Figure 3 The hierarchical decoding structure 300 shown in. Box 302 shows the original decoding order of the pictures, as received in the bitstream. The blocks (1, 2, 3... 16) shown in this diagram represent the pictures in the display order and the arrows indicate the prediction dependencies of the pictures. For example, arrows 0-8 and 16-8 indicate that picture 8 depends on pictures 0 and 16, and arrows 8-4 and 8-12 show the dependencies of pictures 4 and 12 on picture 8. The Tid values divide the pictures into several (e.g., 4) subsets. Pictures in the higher Tid value subsets are less efficient in decoding. For example, when a picture belongs to the least significant subset, a less powerful decoder can filter out pictures with Tid = 4 (pictures numbered 1, 3, 5, 7, 9, 11, 13, and 15).

[0045] Figure 4FIG. is a schematic diagram illustrating an example architecture of a neural network 400 for video quality prediction in accordance with various aspects of the present technology. The neural network 400 can be used for reference - free video quality prediction. The neural network 400 includes an input layer 410, a hidden layer 420, and an output layer 430. In some aspects, the input layer 410 includes a number of input nodes 412. For example, the hidden layer 420 consists of five fully - connected hidden layers 422 of 256, 128, 64, 32, and 16 neurons, respectively.

[0046] The input layer 410 takes as input the feature vectors extracted according to the decoding of the current picture. Since the quality metric used in this example is PSNR, the output layer produces the RMSE of the Y, U, and V components. In one or more aspects, the total number of network parameters is approximately 51,747. The activation function used is the rectified linear unit (ReLU). To convert the predicted RMSE into a PSNR value, the following equation can be used:

[0047]

[0048] Figure 5FIG. 500 illustrates a process for training data generation and network training in accordance with various aspects of the present technology. A neural network is represented by network parameters θ and an activation function g(). Training or test data vectors are associated with decoded pictures, which consist of a feature vector x(t) and a ground truth video quality vector q(t). Framework 500 is used to generate training data. Process 500 begins at process step 510, where the original sequence 502 is encoded and decoded using a selected compression standard (format), decoding structure (e.g., all intra, random access, and low latency configurations), bit rate, etc. Although typically the sequence is encoded and decoded once, in some use cases (e.g., transcoding and downscaling), the sequence may be encoded and decoded multiple times using cascades of encoding and decoding stages with different compression formats and bit rates. For example, the sequence may first be encoded and decoded using AVC / H.264 and then transcoded into the HEVC / H.265 format. In all cases involving transcoding and / or downscaling, at process step 520, a ground truth video quality vector q(t) of the decoded pictures between the original sequence 502 and the reconstructed sequence 514 is calculated using the reconstructed sequence 514. Any suitable quality metric (e.g., PSNR, SSIM, MS-SSIM, Video Multimethod Fusion (VMF), and Mean Opinion Score (MOS)) may be employed to represent the ground truth and predicted video quality vectors. Finally, the resulting bitstream 512 (i.e., the output of the last encoder in the encoding / decoding chain) is fed into a decoder for advanced and block-level feature extraction at process step 530 to form a feature vector x(t) of the sequence. Given a labeled training set {(x(0), q(0)), (x(1), q(1)),..., (x(T-1), q(T-1))}, the neural network parameters θ can be trained at process step 550 by processing (minimizing) a loss function J (plus some regularization term with respect to the parameter θ).

[0049]

[0050] The supervised training step includes calculating a predicted quality vector p(t) using the feature vector x(t) at inference step 558; calculating a prediction loss between the predicted quality vector p(t) and the ground truth quality vector q(t) at process step 552. At process step 554, backpropagation is used to calculate the partial derivatives (gradients) of each network layer. At process step 556, Stochastic Gradient Descent (SGD) is used to update the parameter θ and the updated parameter θ is fed into Figure 4 neural network 400. The above steps are repeated until the training criterion is met.

[0051] A feasibility study is performed on the neural network 400. A total of 444,960 training vectors and 49,440 test vectors are used in the study. Commercial AVC / H.264 and HEVC / H.265 encoders with four typical bitrate points and constant bitrate (CBR) control are used to generate the first vector set. The second vector set simulates a transcoding / downscaling environment where a test sequence is first compressed with an AVC / H.264 encoder and then recompressed with an HEVC / H.265 encoder (i.e., transcoding) and an AVC / H.264 encoder (i.e., downscaling) on the reconstructed sequence. As described above, here, the baseline true RMSE in the transcoding / downscaling case is calculated for the original sequence, rather than for the reconstructed sequence after the first-pass AVC / H.264 encoding.

[0052] After training for 2,000 maximum training epochs with mean absolute error as the loss function, the average PSNR (Y, U, V) prediction errors (in dB) and failure rates are (0.20, 0.16, 0.17) / 0.96% and (0.59, 0.41, 0.39) / 11.68% for the training set and the test set, respectively. It should be noted that here the prediction failure rate is the percentage of training / test vectors for which the average YUV PSNR prediction error (i.e., the average absolute PSNR difference between the prediction and the baseline true Y, U, V PSNR) is greater than 1 dB.

[0053] In some embodiments, instead of using the full-size input feature vector x(t), a feature subset can be used. For example, a less complex network (e.g., having a reduced count of hidden layers and / or neurons) can use an input feature vector that contains only high-level features for video quality prediction. High-level features can typically be extracted using firmware without the need to change the block-level decoder hardware / software. Decoders without block-level feature extraction can deploy non-complex or less complex neural networks for video quality prediction, while other decoders with full feature extraction capabilities can deploy more complex networks. The neural networks can have different network parameters and may or may not have the same network architecture. To share the same architecture with a more complex neural network, a less accurate network can still use the full-size input feature vector but set the block-level features to zero in the input vector. In one or more embodiments, the decoded pictures can be classified into different content categories (e.g., natural video, screen content, etc.) by analyzing the bitstream characteristics and / or the decoded pictures, or the classification information can be communicated by a server, and the network for video prediction can be switched at the picture level based on the content classification information. In some aspects, the classification information can be added as an additional feature to the input feature vector, thus avoiding the need to switch the network at the picture level.

[0054] In some embodiments, a user may be able to report a difference between a predicted video quality and an observed video quality. The deployed network may be refined to improve prediction accuracy by leveraging user feedback. To reduce the additional burden of updating the video quality prediction network, in some aspects, only a subset of network layers or parameters may be refined and updated.

[0055] Figure 6 FIG. 4 is a flow chart illustrating a method 600 for no-reference video quality prediction in accordance with various aspects of the present technology. Method 600 includes receiving a video data stream (610) and generating a feature vector by decoding the video data stream and extracting features (620). Method 600 further includes configuring a neural network to provide a predicted quality vector after being trained using training data (630). The training data includes the feature vector and a ground truth video quality vector, and generating the feature vector consists of high-level syntax processing of the video data stream to extract high-level feature elements and block-level processing to extract block-level feature elements.

[0056] Figure 7 FIG. 7 is a block diagram illustrating an electronic system in which one or more aspects of the present technology may be implemented. The electronic system 700 may be a communication device, such as, for example, a smart phone, a smart watch or a tablet computer, a desktop computer, a laptop computer, a wireless router, a wireless access point (AP), a server, or other electronic device. The electronic system 700 may include various types of computer-readable media, as well as interfaces for various other types of computer-readable media. The electronic system 700 includes a bus 708, one or more processors 712, a system memory 704 (and / or buffer), a read-only memory (ROM) 710, a permanent storage device 702, an input device interface 714, an output device interface 706, and one or more network interfaces 716, or subsets and variations thereof.

[0057] The bus 708 collectively represents a bus of all the systems, peripherals, and chip sets that communicatively connect the numerous internal devices of the electronic system 700. In one or more embodiments, the bus 708 communicatively connects the one or more processors 712 with the ROM 710, the system memory 704, and the permanent storage device 702. From these various memory units, the one or more processors 712 retrieve instructions to be executed and data to be processed in order to perform the processes of the present invention. In different embodiments, the one or more processors 712 may be a single processor or a multi-core processor.

[0058] The ROM 710 stores static data and instructions required by one or more processors 712 and other modules of the electronic system 700. On the other hand, the permanent storage device 702 can be a read and write memory device. The permanent storage device 702 can be a non-volatile memory unit that stores instructions and data even when the electronic system 700 is turned off. In one or more embodiments, a mass storage device (e.g., a magnetic disk or optical disk and its corresponding disk drive) can be used as the permanent storage device 702.

[0059] In one or more embodiments, a removable storage device (e.g., a flash drive and its corresponding disk drive) can be used as the permanent storage device 702. Like the permanent storage device 702, the system memory 704 can be a read and write memory device. However, unlike the permanent storage device 702, the system memory 704 can be a volatile read and write memory, e.g., random access memory. The system memory 704 can store any of the instructions and data that one or more processors 712 may need during runtime. In one or more embodiments, the processors of the present invention are stored in the system memory 704, the permanent storage device 702, and / or the ROM 710. From these various memory units, one or more processors 712 retrieve the instructions to be executed and the data to be processed in order to perform the processes of one or more embodiments.

[0060] The bus 708 is also connected to input and output device interfaces 714 and 706. The input device interface 714 enables a user to transfer information to the electronic system 700 and select commands for the electronic system 700. For example, input devices that can be used in conjunction with the input device interface 714 can include an alphanumeric keyboard and a pointing device (also referred to as a "cursor control device"). For example, the output device interface 706 can enable the display of images generated by the electronic system 700. For example, output devices that can be used in conjunction with the output device interface 706 can include a printer and a display device, such as a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, a flexible display, a flat panel display, a solid state display, a projector, or any other device for outputting information. One or more embodiments can include a device that serves as both an input and output device, such as a touch screen. In these embodiments, the feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and the input received from the user can be in any form, including sound, voice, or tactile input.

[0061] Finally, as Figure 7As shown, bus 708 also couples electronic system 700 to one or more networks and / or one or more network nodes via one or more network interfaces 716. In this way, electronic system 700 can be part of a computer network, such as, for example, a local area network (LAN), a wide area network (WAN), or an intranet, or a network of networks such as the Internet. Any or all components of electronic system 700 can be used in conjunction with the present invention, although the disclosed techniques can also be implemented using a distributed system, for example, a distributed processing and storage system.

[0062] Embodiments within the scope of the present invention can be implemented in part or in whole using a tangible computer-readable storage medium (or one or more types of multiple tangible computer-readable storage media) encoding one or more instructions. The tangible computer-readable storage medium can also be non-transitory in nature.

[0063] A computer-readable storage medium can be a storage medium readable, writable, or otherwise accessible by a general or special purpose computing device, including any processing electronic device and / or processing circuit capable of executing instructions. By way of example and without limitation, the computer-readable medium can include any volatile semiconductor memory, such as RAM, DRAM, SRAM, T-RAM, Z-RAM, or TTRAM. The computer-readable medium can also include any non-volatile semiconductor memory, such as, for example, ROM, PROM, EPROM, EEPROM, NVRAM, flash memory, nvSRAM, FeRAM, FeTRAM, MRAM, PRAM, CBRAM, SONOS, RRAM, NRAM, racetrack memory, FJG, or millipede memory.

[0064] In addition, the computer-readable storage medium can include any non-semiconductor memory, such as, for example, optical disk storage devices, disk storage devices, magnetic tape, other magnetic storage devices, or any other medium capable of storing one or more instructions, such as a distributed storage system. In one or more embodiments, the tangible computer-readable storage medium can be directly coupled to the computing device, while in other embodiments, the tangible computer-readable storage medium can be indirectly coupled to the computing device, for example, via one or more wired connections, one or more wireless connections, or any combination thereof.

[0065] Instructions may be directly executable or may be used to generate executable instructions. For example, instructions may be implemented as executable or non-executable machine code, or may be implemented in a high-level language as instructions that can be compiled to produce executable or non-executable machine code. Additionally, instructions may be implemented as data or may contain data. Computer-executable instructions may also be organized in any format, including routines, subroutines, programs, data structures, objects, modules, applications, applets, functions, and so on. As will be appreciated by those skilled in the art, details include but are not limited to number, structure, sequence, and the organization of instructions may vary significantly without changing the underlying logic, functionality, processing, and output.

[0066] Although the foregoing discussion has primarily referred to microprocessors or multi-core processors that execute software, one or more embodiments are executed by one or more integrated circuits, such as an ASIC or FPGA. In one or more embodiments, these integrated circuits execute instructions stored on the circuit itself.

[0067] Those skilled in the art will appreciate that the various illustrative blocks, modules, elements, components, memory systems, and algorithms described herein may be implemented as electronic hardware, computer software, or a combination of both. To illustrate this interchangeability of hardware and software, the various illustrative blocks, modules, elements, components, memory systems, and algorithms have generally been described above in terms of their functionality. Whether this functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application. The various components and blocks may be arranged differently (e.g., arranged in a different order or partitioned in a different way), all of which are within the scope of the present technology.

[0068] It should be understood that any particular order or hierarchy of blocks in the disclosed processes is illustrative of example methods. Based on design preferences, it should be understood that the particular order or hierarchy of blocks in the processes may be rearranged, or that not all of the illustrated blocks may be performed. Any one of the blocks may be performed simultaneously. In one or more embodiments, multitasking and parallel processing may be advantageous. Additionally, the separation of the various system components described in the embodiments above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged into multiple software products (e.g., cloud-based applications) or multiple devices in a distributed system.

[0069] As used in this specification and in any claims of this application, the terms "base station", "receiver", "computer", "server", "processor", and "memory" all refer to electronic or other technical devices. These terms do not include a human or a group of humans. For purposes of the specification, the term "display" or "displaying" means displaying on an electronic device.

[0070] As used herein, the phrase "at least one of" before a list of items (where any of the items in the list are separated by the term "and" or "or") modifies the list as a whole, rather than each member of the list (i.e., each item). The phrase "at least one of" does not necessarily require selection of at least one of each of the listed items, but rather the phrase allows the meaning of including at least one of any one of the items and / or at least one of any combination of the items and / or at least one of each of the items. By way of example, the phrases "at least one of A, B, and C" and "at least one of A, B, or C" each mean only A, only B, or only C; any combination of A, B, and C; and / or at least one of each of A, B, and C.

[0071] The predicate words "configured to", "operable to", and "programmed to" do not imply any particular tangible or intangible modification of an object, but are intended to be used interchangeably. In one or more embodiments, for example, a processor configured to monitor and control an operation or a component may also mean a processor programmed to monitor and control the operation or a processor operable to monitor and control the operation. Similarly, a processor configured to execute code may be regarded as a processor programmed to execute code or operable to execute code.

[0072] Phrases such as "aspect", "the aspect", "another aspect", "some aspects", "one or more aspects", "embodiment", "the embodiment", "another embodiment", "some embodiments", "one or more embodiments", "example", "the example", "another example", "some examples", "one or more examples", "configuration", "the configuration", "another configuration", "some configurations", "one or more configurations", "the present technology", "the present disclosure", "the present invention", and their various variations, etc., are for convenience purposes and do not imply that the disclosure related to such phrases is essential to the present technology or that such disclosure applies to all configurations of the present technology. The disclosure related to such phrases may apply to all configurations or one or more configurations. The disclosure related to such phrases provides one or more examples. The phrase "aspect" or "some aspects", for example, may mean one or more aspects and vice versa, and this similarly applies to other foreign language phrases.

[0073] The term "exemplary" is used herein to mean "serving as an example, instance, or illustration". Embodiments described herein as "exemplary" or "instances" are not necessarily to be construed as preferred or advantageous over other embodiments. Further, to the extent that the terms "comprising", "having", etc. are used in the detailed description or the claims, these terms are intended to be inclusive in a manner similar to the way the term "including" is interpreted when "including" is used as a transitional word in a claim.

[0074] All structural and functional equivalents of the elements of the various aspects described throughout this disclosure that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and intended to be covered by the claims. In addition, nothing disclosed herein is dedicated to the public, whether or not the disclosure is explicitly recited in the claims. No claim element is to be construed under the provisions of 35 U.S.C. § 112(f), unless the element is expressly recited using the phrase "means for" or, in the case of a claim element of a memory system, the element is recited using the phrase "step for".

[0075] The foregoing description is provided to enable any person skilled in the art to practice the various aspects described herein. Those skilled in the art will readily recognize various modifications to these aspects, and the general principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein, but are to be accorded the full scope consistent with the language of the claims, where the reference to an element in the singular is not intended to mean "one and only one" (unless specifically so stated) but rather "one or more". The term "some", unless specifically stated otherwise, means one or more. Masculine pronouns (e.g., "his") include feminine and neuter genders (e.g., "her" and "its"), and vice versa. Headings and subheadings, if any, are used for convenience only and do not limit the invention.

Claims

1. A system, which comprises: a video decoder configured to receive and decode an input bitstream to reconstruct a picture from the input bitstream and generate a first vector, the first vector including features extracted by the video decoder from the input bitstream and features determined by the video decoder during reconstruction of the picture; and a neural network configured to generate a second vector based on the first vector, the second vector including one or more metrics representing the predicted quality of the picture reconstructed by the video decoder, wherein one or more parameters of the neural network are trained by: calculating a first training vector based on a decoded test bitstream, using the video decoder to generate a second training vector, the second training vector including features extracted from the test bitstream and determined by the video decoder, using the neural network to generate a third training vector based on the second training vector, calculating a prediction loss between the first training vector and the third training vector, and updating the one or more parameters based on the prediction loss.

2. The system according to claim 1, wherein the video decoder is configured to extract the features from the input bitstream by parsing syntax elements in the input bitstream.

3. The system according to claim 2, wherein the syntax elements include at least one of a sequence parameter set, a picture parameter set, a video parameter set, a picture header, a slice header, an adaptive parameter set, or a supplementary enhancement information message.

4. The system according to claim 2, wherein the features extracted from the input bitstream include at least one of the following: a transcoding indicator, a codec type, a picture coding type, a picture resolution, a frame rate, a bit depth, a chroma format, a compressed picture size, a high-level quantization parameter, an average temporal distance, or a temporal layer identifier.

5. The system according to claim 1, wherein the features determined by the video decoder during reconstruction of the picture include at least one of the following: the percentage of intra-coded blocks in the picture, the percentage of inter-coded blocks in the picture, the average block-level quantization parameter of the picture, the maximum block-level quantization parameter of the picture, the minimum block-level quantization parameter of the picture, the standard deviation of the horizontal motion vectors of the picture, the average motion vector magnitude of the picture, the average absolute magnitude of the low-frequency inverse quantization transform coefficients of the picture, the average absolute magnitude of the high-frequency inverse quantization transform coefficients of the picture, the standard deviation of the prediction residuals of the picture, the root mean square error value between the reconstructed pictures before and after in-loop filtering, the standard deviation of the reconstructed picture after in-loop filtering, or the edge sharpness of the reconstructed picture after in-loop filtering.

6. The system according to claim 1, wherein the one or more metrics representing the predicted quality of the reconstructed picture include at least one of the following: peak signal-to-noise ratio, structural similarity index measure, multi-scale structural similarity index measure, video multi-method assessment fusion, or mean opinion score.

7. The system according to claim 6, wherein an output layer of the neural network is configured to produce a root mean square error value.

8. The system according to claim 7, wherein the root mean square error value is converted into a peak signal-to-noise ratio value.

9. A method, which comprises: receiving and decoding, by a video decoder, an input bitstream to reconstruct a picture from the input bitstream; generating, by the video decoder, a first vector, the first vector comprising features extracted by the video decoder from the input bitstream and determined by the video decoder during reconstruction of the picture; using a neural network to generate a second vector, the second vector comprising one or more metrics representing a predicted quality of the reconstructed picture, wherein the second vector is based on the first vector, wherein one or more parameters of the neural network are trained by: calculating a first training vector based on a decoded test bitstream, using the video decoder to generate a second training vector, the second training vector comprising features extracted from the test bitstream and determined by the video decoder, using the neural network to generate a third training vector based on the second training vector, calculating a prediction loss between the first training vector and the third training vector, and updating the one or more parameters based on the prediction loss.

10. The method according to claim 9, wherein the features are extracted from the input bitstream by parsing syntax elements in the input bitstream.

11. The method according to claim 10, wherein the syntax elements comprise at least one of a sequence parameter set, a picture parameter set, a video parameter set, a picture header, a slice header, an adaptive parameter set, or a supplementary enhancement information message.

12. The method according to claim 10, wherein the features extracted from the input bitstream comprise at least one of the following: a transcoding indicator, a codec type, a picture coding type, a picture resolution, a frame rate, a bit depth, a chroma format, a compressed picture size, a high-level quantization parameter, an average temporal distance, or a temporal layer identifier.

13. The method according to claim 9, wherein the features determined by the video decoder during reconstruction of the picture comprise at least one of the following: a percentage of intra-coded blocks in the picture, a percentage of inter-coded blocks in the picture, an average block-level quantization parameter of the picture, a maximum block-level quantization parameter of the picture, a minimum block-level quantization parameter of the picture, a standard deviation of horizontal motion vectors of the picture, an average motion vector magnitude of the picture, an average absolute magnitude of low-frequency inverse quantization transform coefficients of the picture, an average absolute magnitude of high-frequency inverse quantization transform coefficients of the picture, a standard deviation of prediction residuals of the picture, a root mean square error value between the reconstructed pictures before and after an in-loop filter, a standard deviation of the reconstructed picture after the in-loop filter, or an edge sharpness of the reconstructed picture after the in-loop filter.

14. The method according to claim 9, wherein the one or more metrics representing the predicted quality of the reconstructed picture comprise at least one of the following: peak signal-to-noise ratio, structural similarity index measure, multi-scale structural similarity index measure, video multi-method assessment fusion, or mean opinion score.

15. The method according to claim 14, wherein an output layer of the neural network is configured to generate a root mean square error value.

16. The method according to claim 15, wherein the root mean square error value is converted into a peak signal-to-noise ratio value.

Citation Information

Patent Citations

  • No-Reference Video / Image Quality Measurement with Compressed Domain Features

    US20130293725A1

  • Fast encoding loss metric

    US20180084280A1