Systems and methods for signaling and deriving quantization parameters for frame level interpolation prediction mode

By conditionally signaling the quantization parameters in video decoding, the problem of low notification efficiency of quantization parameters in the prior art is solved according to the activation state of the frame-level interpolation mode, and higher accuracy and bandwidth efficiency are achieved.

CN120283403APending Publication Date: 2025-07-08TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480005149.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-28
Filing Date
2024-04-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the intra-interpolation prediction mode of existing video decoding technology, the notification method of quantization parameters is inefficient, resulting in high bandwidth consumption and insufficient accuracy.

Method used

The method of conditionally signaling the quantization parameters is adopted, and the quantization parameters are derived or parsed through frame interpolation according to the activation state of the frame-level interpolation mode to reduce unnecessary bitstream notifications.

Benefits of technology

It improves the accuracy and bandwidth efficiency of quantization parameters, reduces the amount of notifications in the bitstream, and improves the efficiency and quality of video decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120283403A_ABST
    Figure CN120283403A_ABST
Patent Text Reader

Abstract

An example method of video coding includes receiving a video bitstream including a plurality of encoded pictures. The method further includes deriving a reconstructed picture of an encoded picture of the plurality of encoded pictures using frame interpolation. The method further includes determining whether one or more quantization parameters for encoding the picture are signaled in the video bitstream based on the signaled indicator in the video bitstream. When the signaled indicator indicates that the one or more quantization parameters are signaled in the video bitstream, the one or more quantization parameters are parsed from the video bitstream. When the signaled indicator indicates that the one or more quantization parameters are not signaled in the video bitstream, the one or more quantization parameters are derived at the decoder.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the priority of U.S. Provisional Patent Application No. 63 / 596,907, entitled "Signaling and Derivation of Quantization Parameters for Frame - Level Interpolation Prediction Mode", filed on November 7, 2023, and this application is a continuation of and claims the priority of U.S. Patent Application No. 18 / 620,927, entitled "Systems and Methods for Signaling and Derivation of Quantization Parameters for a Frame - Level Interpolation Prediction Mode", filed on March 28, 2024. Technical field

[0003] The disclosed embodiments generally relate to video coding, including but not limited to systems and methods for implementing frame - level interpolation prediction mode and for signaling and / or deriving quantization parameters for such a mode. Background art

[0004] Digital video is supported by various electronic devices such as digital televisions, laptop or desktop computers, tablet computers, digital imaging devices, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. Electronic devices send and receive or otherwise convey digital video data across communication networks and / or store digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video coding can be used to compress video data according to one or more video coding standards before transmitting or storing the video data. Video coding can be performed by hardware and / or software on an electronic / client device or a server providing cloud services.

[0005] Video coding typically utilizes prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.) that exploit the redundancy inherent in video data. Video coding aims to compress video data into a form that uses a lower bit rate while avoiding or minimizing the degradation of video quality. A variety of video codec standards have been developed. For example, High-Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H (Moving Picture Experts Group - HEVC, MPEG-H) project. The ITU-T (International Telecommunication Union - Telecommunication Standardization Sector, ITU-T) and ISO / IEC (International Organization for Standardization / International Electrotechnical Commission, ISO / IEC) released the HEVC / H.265 standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Versatile Video Coding (VVC / H.266) is a video compression standard intended as a successor to HEVC. The ITU-T and ISO / IEC released the VVC / H.266 standard in 2020 (version 1) and 2022 (version 2). AOMedia Video 1 (Alliance for Open Media Video 1, AV1) is an open video coding format designed as an alternative to HEVC. On January 8, 2019, the verified version 1.0.0 of the specification with errata 1 was released. Summary of the Invention

[0006] The present disclosure describes a set of methods for video (image) compression, and more particularly, relates to in-loop filtering. Some embodiments include conditionally signaling and / or deriving quantization parameters (e.g., based on whether an intra-frame interpolation mode is active). Conditionally signaling quantization parameters can improve accuracy compared to systems that always (e.g., at the decoder) derive quantization parameters, and can improve bandwidth efficiency (reduce the number of bits signaled) compared to systems that always (e.g., in the video bitstream) signal quantization parameters.

[0007] According to some embodiments, a method for video decoding includes: (i) receiving a video bitstream including a plurality of coded pictures; (ii) using intra interpolation to obtain reconstructed pictures of the coded pictures among the plurality of coded pictures; (iii) determining whether one or more quantization parameters for a coded picture are signaled in the video bitstream based on an indicator signaled in the video bitstream; (iv) when the signaled indicator indicates that one or more quantization parameters (e.g., having a first value) are signaled in the video bitstream, parsing the one or more quantization parameters from the video bitstream; and (v) when the signaled indicator indicates that one or more quantization parameters are not signaled in the video bitstream (e.g., having a second value), obtaining the one or more quantization parameters.

[0008] According to some embodiments, a method for video decoding includes: (i) receiving a video bitstream including a plurality of coded pictures; (ii) using intra interpolation to obtain reconstructed pictures of the coded pictures among the plurality of coded pictures; (iii) determining whether an incremental value of a quantization parameter for a coded picture is signaled in the video bitstream based on an indicator signaled in the video bitstream, where the incremental value represents a difference between a value of a reference quantization parameter and a value of the quantization parameter; (iv) when the signaled indicator has a first value, parsing the incremental value of the quantization parameter from the video bitstream; and (v) when the signaled indicator has a second value, obtaining one or more quantization parameters without parsing the incremental value.

[0009] According to some embodiments, a method for video encoding includes: (i) receiving video data including a plurality of pictures; (ii) encoding a first picture among the plurality of pictures according to a first type of intra interpolation; (iii) determining whether one or more quantization parameters for the first picture to be encoded are to be signaled in the video bitstream; (iv) transmitting the encoded first picture via the video bitstream; and (v) signaling a first indicator via the video bitstream, where the first indicator is used to indicate whether one or more quantization parameters for the first picture to be encoded are signaled in the video bitstream.

[0010] According to some embodiments, a method for processing visual media data includes: (i) obtaining a source video sequence; and (ii) performing a conversion between the source video sequence and a bitstream of visual media data, where the bitstream includes: (a) a plurality of coded pictures, including a first coded picture encoded according to a first type of intra interpolation; and (b) a first indicator, which is used to indicate whether one or more quantization parameters for the first coded picture are signaled in the video bitstream.

[0011] According to some embodiments, a computing system is provided, such as a streaming system, a server system, a personal computer system, or other electronic devices. The computing system includes control circuitry and a memory storing one or more sets of instructions. The one or more sets of instructions include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and a decoder component (e.g., a transcoder). According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more sets of instructions executable by the computing system. The one or more sets of instructions include instructions for performing any of the methods described herein.

[0012] Accordingly, devices and systems having methods for encoding and decoding video are disclosed. Such methods, devices, and systems may supplement or replace conventional methods, devices, and systems for encoding / decoding video. The features and advantages described in the specification are not necessarily all inclusive, and in particular, given the figures, specification, and claims provided in this disclosure, some additional features and advantages will be apparent to those of ordinary skill in the art. Further, it should be noted that the language used in the specification has been principally selected for readability and instructional purposes and not necessarily for the purpose of delineating or limiting the subject matter described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] For a more detailed understanding of the present disclosure, reference may be made to the features of various embodiments, some of which are illustrated in the accompanying figures. However, the figures only illustrate relevant features of the present disclosure and are therefore not necessarily considered limiting, as the specification may allow other effective features that will be understood by those skilled in the art upon reading the present disclosure.

[0014] Figure 1 is a block diagram showing an example communication system according to some embodiments.

[0015] Figure 2A is a block diagram showing example elements of an encoder component according to some embodiments.

[0016] Figure 2B is a block diagram showing example elements of a decoder component according to some embodiments.

[0017] Figure 3 is a block diagram showing an example server system according to some embodiments.

[0018] Figures 4A to 4C shows an example prediction block, residual block, and reconstruction block according to some embodiments.

[0019] Figure 5A Shows an example Temporal Interpolated Prediction (TIP) mode according to some embodiments.

[0020] Figure 5B Shows an example in-loop filter stage according to some embodiments.

[0021] Figure 6A Shows an example video decoding process according to some embodiments.

[0022] Figure 6B Shows an example video encoding process according to some embodiments.

[0023] By convention, the various features shown in the drawings are not necessarily drawn to scale, and throughout the specification and drawings, like reference numerals may be used to represent like features. Detailed Description

[0024] The present disclosure describes video / image compression techniques that include signaling and deriving quantization parameters for prediction modes based on frame interpolation. The present disclosure includes a description of conditionally signaling quantization parameters and / or derivation methods based on whether a frame-level intra-frame interpolation mode is enabled for an encoded picture. For example, when the TIP frame-level mode (e.g., the tip_frame_as_output mode) is selected, the luminance and chrominance quantization indices of the AC coefficients can be derived from a reference frame rather than signaled from the encoder side to the decoder side. As an example, the quantization index can be derived as an average of the quantization indices from the reference frame.

[0025] Example systems and devices

[0026] Figure 1 Is a block diagram showing a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m) communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is, for example, a streaming system for use with video-enabled applications such as video conferencing applications, digital TV applications, and media storage and / or distribution applications.

[0027] The source device 102 includes a video source 104 (e.g., a camera device component or a media storage device) and an encoder component 106. In some embodiments, the video source 104 is a digital camera device (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates one or more encoded video bitstreams based on the video stream. The video stream from the video source 104 may be of high data volume compared to the encoded video bitstreams generated by the encoder component 106. Since the encoded video bitstreams 108 are of lower data volume (less data) compared to the video stream from the video source, the encoded video bitstreams 108 require less bandwidth to transmit and less storage space to store. In some embodiments, the source device 102 does not include the encoder component 106 (e.g., configured to transmit uncompressed video data to the network 110).

[0028] One or more networks 110 represent any number of networks for transmitting information between the source device 102, the server system 112, and / or the electronic devices 120, including, for example, wired (wired) and / or wireless communication networks. One or more networks 110 may exchange data in circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0029] One or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is a streaming server or includes a streaming server (e.g., configured to store and / or distribute video content such as the encoded video stream from the source device 102). The server system 112 includes a decoder component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the decoder component 114 includes an encoder component and / or a decoder component. In various embodiments, the decoder component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, the decoder component 114 is configured to decode the encoded video bitstreams 108 and re-encode the video data using different coding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings based on the encoded video bitstreams 108. In some embodiments, the server system 112 serves as a Media-Aware Network Element (MANE). For example, the server system 112 may be configured to trim the encoded video bitstreams 108 to customize potentially different bitstreams for one or more of the electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.

[0030] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to generate an outgoing video stream that can be presented on a display or other type of rendering device. In some embodiments, one or more of the electronic devices 120 do not include a display component (e.g., are communicatively coupled to an external display device and / or include a media storage device). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.

[0031] The source device and / or the plurality of electronic devices 120 are sometimes referred to as "terminal devices" or "user devices". In some embodiments, one or more of the electronic devices 120 and / or the source device 102 are examples of a server system, a personal computer, a portable device (e.g., a smart phone, a tablet computer, or a laptop computer), a wearable device, a video conferencing device, and / or other types of electronic devices.

[0032] In an example operation of the communication system 100, the source device 102 transmits the encoded video bitstream 108 to the server system 112. For example, the source device 102 may encode a picture stream captured by the source device. The server system 112 receives the encoded video bitstream 108 and may use the decoder component 114 to decode and / or encode the encoded video bitstream 108. For example, the server system 112 may apply an encoding that is more optimized for network transmission and / or storage to the video data. The server system 112 may transmit the encoded video data 116 (e.g., one or more decoded video bitstreams) to one or more of the electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display the video pictures.

[0033] Figure 2AFIG. 0 is a block diagram showing example elements of an encoder component 106 according to some embodiments. The encoder component 106 receives video data (e.g., a source video sequence) from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives the video sequence from a remote video source (e.g., a video source that is a component of a device different from the encoder component 106). The video source 104 may provide the source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8 bits, 10 bits, or 12 bits), any color space (e.g., BT.601 Y CrCb or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, the video source 104 is a storage device that stores previously captured / prepared video. In some embodiments, the video source 104 is a camera device that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. The pictures themselves may be organized as a spatial pixel array, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. A person of ordinary skill in the art can easily understand the relationship between pixels and samples.

[0034] The encoder component 106 is configured to decode and / or compress the pictures of the source video sequence into a decoded video sequence 216 in real time or under other time constraints required by the application. In some embodiments, the encoder component 106 is configured to perform a conversion between the source video sequence and a visual media data bitstream (e.g., a video bitstream). Enforcing an appropriate decoding speed is a function of the controller 204. In some embodiments, the controller 204 controls the other functional units as described below and is functionally coupled to the other functional units. The parameters set by the controller 204 may include rate control related parameters (e.g., picture skip, quantizer, and / or the λ value of rate distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. A person of ordinary skill in the art can easily identify other functions of the controller 204, as such functions may belong to the encoder component 106 optimized for a specific system design.

[0035] In some embodiments, the encoder component 106 is configured to operate in a decoding loop. In a simplified example, the decoding loop includes a source decoder 202 (e.g., responsible for creating symbols such as a symbol stream based on an input picture and reference pictures to be decoded) and a (local) decoder 210. The decoder 210 reconstructs the symbols in a manner similar to a (remote) decoder to create sample data (in the case where the compression between the symbols and the decoded video bitstream is lossless). The reconstructed sample stream (sample data) is input to the reference picture memory 208. Since the decoding of the symbol stream produces bit-exact results regardless of the decoder location (local or remote), the content in the reference picture memory 208 is also bit-exact between the local encoder and the remote encoder. In this way, the reference picture samples interpreted by the prediction part of the encoder will be the same as the sample values that the decoder will interpret when using prediction during decoding.

[0036] The operation of the decoder 210 can be the same as that of a remote decoder such as the decoder component 122 described in detail below in connection with Figure 2B However, briefly referring to Figure 2B , since the symbols are available and the encoding of the symbols into the decoded video sequence by the entropy decoder 214 and the decoding of the symbols by the parser 254 can be lossless, the entropy decoding part including the buffer memory 252 and the parser 254 of the decoder component 122 may not be fully implemented in the local decoder 210.

[0037] Except for parsing / entropy decoding, the decoder techniques described herein may exist in a corresponding encoder in a form with substantially the same functionality. For this reason, the disclosed subject matter focuses on decoder operations. Additionally, the description of encoder techniques can be simplified because encoder techniques can be inverse to decoder techniques.

[0038] As part of the operation of the source decoder 202, the source decoder 202 may perform motion compensated predictive decoding that predictively decodes an input frame by referring to one or more previously decoded frames designated as reference frames from a video sequence. In this way, the decoding engine 212 decodes the difference between a pixel block of the input frame and a pixel block of the reference frame, and the reference frame can be selected as a prediction reference for the input frame. The controller 204 can manage the decoding operations of the source decoder 202, including, for example, setting parameters and subgroup parameters for encoding video data.

[0039] The decoder 210 decodes the decoded video data of a frame that can be designated as a reference frame based on the symbols created by the source decoder 202. The operation of the decoding engine 212 can advantageously be a lossy process. When the decoded video data is in a video decoder ( Figure 2AWhen decoded at the (not shown), the reconstructed video sequence can be a copy of the source video sequence with some errors. The decoder 210 repeats the decoding process that can be performed by a remote video decoder on the reference frames, and can cause the reconstructed reference frames to be stored in the reference picture memory 208. In this way, the encoder component 106 locally stores a copy of the reconstructed reference frames, which has the same content (without transmission errors) as the reconstructed reference frames that will be obtained by the remote video decoder.

[0040] The predictor 206 can perform a prediction search for the decoding engine 212. That is, for a new frame to be decoded, the predictor 206 can search in the reference picture memory 208 for sample data (as a candidate reference pixel block) or some metadata such as reference picture motion vectors, block shapes, etc. that can be used as an appropriate prediction reference for the new picture. The predictor 206 can operate on a per-pixel block basis of the sample blocks to find an appropriate prediction reference. As determined by the search results obtained by the predictor 206, the input picture can have prediction references taken from multiple reference pictures stored in the reference picture memory 208.

[0041] The outputs of all the foregoing functional units can be subjected to entropy coding in the entropy coder 214. The entropy coder 214 converts the symbols into a coded video sequence by losslessly compressing the symbols generated by the various functional units according to techniques known to those of ordinary skill in the art (e.g., Huffman coding, variable length coding, and / or arithmetic coding).

[0042] In some embodiments, the output of the entropy coder 214 is coupled to a transmitter. The transmitter can be configured to buffer the coded video sequence created by the entropy coder 214 in preparation for transmission via a communication channel 218, which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter can be configured to merge the decoded video data from the source decoder 202 with other data to be transmitted, such as decoded audio data and / or an auxiliary data stream (source not shown). In some embodiments, the transmitter can transmit additional data along with the encoded video. The source decoder 202 can include such data as part of the decoded video sequence. The additional data can include temporal / spatial / SNR (Signal-to-Noise Ratio, SNR) enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, etc.

[0043] The controller 204 can manage the operation of the encoder component 106. During decoding, the controller 204 can assign a specific decoded picture type to each decoded picture, which may affect the decoding technique applied to the corresponding picture. For example, a picture can be assigned as an intra picture (I picture), a predictive picture (P picture), or a bi-predictive picture (B picture). An intra picture can be decoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of those variations of I pictures and their corresponding applications and characteristics, and thus will not be repeated here. A predictive picture can be decoded and decoded using inter prediction or intra prediction that uses at most one motion vector and a reference index to predict the sample values of each block. A bi-predictive picture can be decoded and decoded using inter prediction or intra prediction that uses at most two motion vectors and a reference index to predict the sample values of each block. Similarly, a multi-predictive picture can use more than two reference pictures and associated metadata for reconstructing a single block.

[0044] Source pictures can generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples respectively), and decoded on a block-by-block basis. These blocks can be predictively decoded with reference to other (already decoded) blocks, which are determined by the decoding assignment applied to the corresponding picture of the block. For example, blocks of an I picture can be non-predictively decoded, or predictively decoded (spatial prediction or intra prediction) with reference to already decoded blocks of the same picture. Pixel blocks of a P picture can be non-predictively decoded with reference to one previously decoded reference picture via temporal prediction or via spatial prediction. Blocks of a B picture can be non-predictively decoded with reference to one or two previously decoded reference pictures via temporal prediction or via spatial prediction.

[0045] Video can be captured as a plurality of source pictures (video pictures) in a time series. Intra picture prediction (commonly abbreviated as intra prediction) exploits the spatial correlation within a given picture, while inter picture prediction exploits the (temporal or other) correlation between pictures. In an example, a specific picture in encoding / decoding, which is referred to as the current picture, is segmented into blocks. In the case where a block in the current picture is similar to a reference block in a previously decoded and still buffered reference picture in the video, the block in the current picture can be decoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.

[0046] The encoder component 106 may perform decoding operations according to any of the predetermined video decoding techniques or standards such as those described herein. In the operation of the encoder component 106, the encoder component 106 may perform various compression operations, including predictive decoding operations that utilize temporal redundancy and spatial redundancy in the input video sequence. Thus, the decoded video data may conform to the syntax specified by the video decoding technique or standard being used.

[0047] Figure 2B is a block diagram showing example elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 in is coupled to the channel 218 and the display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to the loop filter 256 and configured to transmit data (e.g., via a wired connection or a wireless connection) to the display 124.

[0048] In some embodiments, the decoder component 122 includes a receiver coupled to the channel 218 and configured to receive data (e.g., via a wired connection or a wireless connection) from the channel 218. The receiver may be configured to receive one or more decoded video sequences to be decoded by the decoder component 122. In some embodiments, the decoding of each decoded video sequence is independent of other decoded video sequences. Each decoded video sequence may be received from the channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive the encoded video data as well as other data, such as decoded audio data and / or auxiliary data streams, which may be forwarded to their respective consuming entities (not depicted). The receiver may separate the decoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data along with the encoded video. The additional data may be included as part of the (one or more) decoded video sequences. The additional data may be used by the decoder component 122 to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0049] According to some embodiments, the decoder component 122 includes a buffer memory 252, a parser 254 (sometimes also referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra picture prediction unit 262, a motion compensation prediction unit 260, an aggregator 268, a loop filter unit 256, a reference picture memory 266, and a current picture memory 264. In some embodiments, the decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuitry systems. The decoder component 122 may be implemented at least partially in software.

[0050] The buffer memory 252 is coupled between the channel 218 and the parser 254 (e.g., to counter network jitter). In some embodiments, the buffer memory 252 is separate from the decoder component 122. In some embodiments, a separate buffer memory is provided between the output of the channel 218 and the decoder component 122. In some embodiments, in addition to the buffer memory 252 inside the decoder component 122 (e.g., which is configured to handle playback timing), a separate buffer memory is provided outside the decoder component 122 (e.g., to counter network jitter). When receiving data from a store-and-forward device with sufficient bandwidth and controllability or from an isochronous synchronization network, the buffer memory 252 may not be needed, or the buffer memory 252 may be smaller. In order to use a packet network such as the Internet as much as possible, the buffer memory 252 may be needed. The buffer memory 252 may be relatively large and / or have an adaptive size, and may be implemented at least partially in the operating system or a similar element outside the decoder component 122.

[0051] The parser 254 is configured to reconstruct symbols 270 according to the decoded video sequence. The symbols may include, for example, information for managing the operation of the decoder component 122, and / or information for controlling a rendering device such as the display 124. The control information for the rendering device may be in the form of, for example, a supplementary enhancement information (SEI) message or a video usability information (VUI) parameter set fragment (not depicted). The parser 254 parses (entropy decodes) the decoded video sequence. The decoding of the decoded video sequence may be performed according to video decoding techniques or standards, and may follow principles well known to those skilled in the art, including: variable length decoding, Huffman decoding, arithmetic decoding with or without context sensitivity, etc. The parser 254 may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the decoded video sequence based on at least one parameter corresponding to the subgroup. The subgroup may include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser 254 may also extract information from the decoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0052] Depending on the type of the decoded video picture or a part thereof (e.g., an inter-picture and an intra-picture, an inter-block and an intra-block) and other factors, the reconstruction of the symbols 270 may involve multiple different units. Which units are involved and the way they are involved may be controlled by the parser 254 through subgroup control information parsed from the decoded video sequence. For clarity, this subgroup control information flow between the parser 254 and the multiple units below is not depicted.

[0053] The decoder component 122 can be conceptually subdivided into multiple functional units, and in some implementations, these units interact closely with each other and can be at least partially integrated with each other. However, for clarity, the conceptual subdivision of the functional units is maintained herein.

[0054] The scaler / inverse transform unit 258 receives, from the parser 254, the quantized transform coefficients as symbols 270 and control information (such as which transform to use, block size, quantization factor, and / or quantization scaling matrix). The scaler / inverse transform unit 258 can output a block including sample values, which can be input into the aggregator 268. In some cases, the output samples of the scaler / inverse transform unit 258 belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information can be provided by the intra-picture prediction unit 262. The intra-picture prediction unit 262 can generate a block of the same size and shape as the block being reconstructed using the surrounding reconstructed information obtained from the current (partially reconstructed) picture from the current picture memory 264. The aggregator 268 can add the predictive information that the intra-picture prediction unit 262 has generated to the output sample information provided by the scaler / inverse transform unit 258 on a per-sample basis.

[0055] In other cases, the output samples of the scaler / inverse transform unit 258 belong to an inter-coded and potentially motion-compensated block. In such a case, the motion compensation prediction unit 260 can access the reference picture memory 266 to obtain samples for prediction. After motion-compensating the obtained samples according to the symbols 270 belonging to the block, these samples can be added by the aggregator 268 to the output of the scaler / inverse transform unit 258 (referred to as residual samples or residual signals in this case) to generate output sample information. The addresses within the reference picture memory 266 can be controlled by motion vectors, and the motion compensation prediction unit 260 obtains prediction samples from the reference picture memory 266. The motion vectors can be available to the motion compensation prediction unit 260 in the form of symbols 270, which can have, for example, an X component, a Y component, and a reference picture component. Motion compensation can also include, for example, interpolation of sample values obtained from the reference picture memory 266 when using sub-sampled accurate motion vectors, a motion vector prediction mechanism.

[0056] The output samples of aggregator 268 can undergo various loop filtering techniques in loop filter unit 256. Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the decoded video bitstream and available to loop filter unit 256 as symbols 270 from parser 254, but video compression techniques can also respond to meta-information obtained during decoding of previous (in decoding order) portions of the decoded picture or decoded video sequence, and to previously reconstructed and loop-filtered sample values. The output of loop filter unit 256 can be a sample stream that can be output to a rendering device such as display 124, and stored in reference picture memory 266 for use in future inter-picture prediction.

[0057] Once reconstructed, some decoded pictures can be used as reference pictures for future prediction. Once a decoded picture is reconstructed and the decoded picture has been identified (e.g., by parser 254) as a reference picture, the current reference picture can become part of reference picture memory 266, and a new current picture memory can be reallocated before starting to reconstruct subsequent decoded pictures.

[0058] Decoder component 122 can perform decoding operations according to predetermined video compression techniques that can be recorded in any standard such as those described herein. In the sense that the decoded video sequence follows the syntax of the video compression technique or standard, the decoded video sequence can conform to the syntax specified by the video compression technique or standard used, as specified in the video compression technique document or standard and particularly in the profile thereof. Additionally, to conform to some video compression techniques or standards, the complexity of the decoded video sequence can be within the range defined by the levels of the video compression technique or standard. In some cases, the levels limit the maximum picture size, maximum frame rate, maximum reconstruction sample rate (measured, for example, in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the levels can be further restricted by the Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the decoded video sequence.

[0059] Figure 3is a block diagram showing a server system 112 according to some embodiments. The server system 112 includes control circuitry 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuitry 302 includes one or more processors (e.g., a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and / or a DPU (Data Processing Unit)). In some embodiments, the control circuitry includes a field programmable gate array, a hardware accelerator, and / or an integrated circuit (e.g., an application specific integrated circuit).

[0060] The network interface 304 may be configured to interface with one or more communication networks (e.g., a wireless network, a wired network, and / or an optical network). The communication network may be local, wide area, metropolitan area, vehicular and industrial, real-time, delay tolerant, etc. Examples of communication networks include: local area networks such as Ethernet, wireless LAN (Local Area Network); cellular networks including GSM (Global System for Mobile Communications), 3G (the Third Generation), 4G (the Fourth Generation), 5G (the Fifth Generation), LTE (Long Term Evolution), etc.; TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; vehicular and industrial networks including CANBus (Controller Area Network - BUS), etc. Such communication may be only unidirectional reception (e.g., broadcast TV), only unidirectional transmission (e.g., CAN bus to certain CAN bus devices), or bidirectional (e.g., to other computer systems using local digital networks or wide area digital networks). Such communication may include communication to one or more cloud computing networks.

[0061] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input device 310 may include one or more of the following: a keyboard, a mouse, a touchpad, a touch screen, a data glove, a joystick, a microphone, a scanner, a camera device, etc. The output device 308 may include one or more of the following: an audio output device (e.g., a speaker), a visual output device (e.g., a display or a monitor), etc.

[0062] The memory 314 may include high-speed random access memory (such as DRAM (Dynamic Random Access Memory), SRAM (Static Random Access Memory), DDR RAM (Double Data Rate Random Access Memory), and / or other random access solid-state memory devices) and / or non-volatile memory (such as one or more disk storage devices, optical disc storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). The memory 314 optionally includes one or more storage devices remote from the control circuitry 302. The memory 314, or alternatively, the non-volatile solid-state memory device within the memory 314, includes a non-transitory computer-readable storage medium. In some embodiments, the memory 314 or the non-transitory computer-readable storage medium of the memory 314 stores the following programs, modules, instructions, and data structures, or subsets or supersets thereof:

[0063] ● An operating system 316 that includes procedures for handling various basic system services and for performing hardware-related tasks;

[0064] ● A network communication module 318 for connecting the server system 112 to other computing devices via one or more network interfaces 304 (e.g., via

[0065] wired connections and / or wireless connections); ● A decoding module 320 for performing various functions related to encoding and / or decoding data such as video data. In some embodiments, the decoding module 320 is an instance of the decoder component 114. The decoding module 320 includes, but is not limited to, one or more of the following:

[0066] o A decoding module 322 for performing various functions related to decoding encoded data, such as those previously described with respect to the decoder component 122; and

[0067] o An encoding module 340 for performing various functions related to encoding data, such as those previously described with respect to the encoder component 106; and ● A picture memory 352 for storing pictures and picture data, e.g., for use with the decoding module 320. In some embodiments, the picture memory 352 includes one or more of the following: a reference picture memory 208, a buffer memory 252, a current picture memory 264, and a reference picture memory 266.

[0068] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform the various functions previously described with respect to the parser 254), a transform module 326 (e.g., configured to perform the various functions previously described with respect to the scaling / inverse transform unit 258), a prediction module 328 (e.g., configured to perform the various functions previously described with respect to the motion compensation prediction unit 260 and / or the intra picture prediction unit 262), and a filter module 330 (e.g., configured to perform the various functions previously described with respect to the loop filter 256).

[0069] In some embodiments, the encoding module 340 includes a code module 342 (e.g., configured to perform the various functions previously described with respect to the source decoder 202 and / or the decoding engine 212) and a prediction module 344 (e.g., configured to perform the various functions previously described with respect to the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 includes Figure 3 a subset of the modules shown. For example, a shared prediction module is used by both the decoding module 322 and the encoding module 340.

[0070] Each of the modules identified above stored in the memory 314 corresponds to an instruction set for performing the functions described herein. The modules identified above (e.g., the instruction sets) need not be implemented as separate software programs, procedures, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, the decoding module 320 optionally does not include separate decoding and encoding modules, but instead uses the same set of modules to perform two sets of functions. In some embodiments, the memory 314 stores a subset of the modules and data structures identified above. In some embodiments, the memory 314 stores additional modules and data structures not described above.

[0071] Although Figure 3 a server system 112 is shown in accordance with some embodiments, Figure 3 it is intended more as a functional description of the various features that may exist in one or more server systems rather than a structural schematic of the embodiments described herein. In practice, the items shown separately may be combined, and some items may be separated. For example, Figure 3 some of the items shown separately in may be implemented on a single server, and a single item may be implemented by one or more servers. The actual number of servers used to implement the server system 112 and how the features are distributed among the servers will vary depending on the implementation, and optionally, it depends in part on the amount of data traffic processed by the server system during peak usage periods and during average usage periods.

[0072] Example decoding techniques

[0073] As discussed above, some codecs (e.g., AV1) operate on pixel blocks. Each pixel block can be processed in a predictive transform coding scheme where intra-frame reference pixels, inter-frame motion compensation, or some combination of both are used to obtain a prediction. The residual from the prediction can undergo a transform (e.g., a 2-D unitary transform) to further remove spatial correlation, and the transform coefficients are quantized. Then, arithmetic coding can be used to entropy code both the predicted syntax elements and the quantized transform coefficient indices.

[0074] Figure 4A The calculation of a predicted block according to some embodiments is shown. In Figure 4A the example, intra-frame prediction is performed on the current block 402 to generate a predicted block 404. The current block 402 includes a set of samples (e.g., a pixel block), and the predicted block 404 includes a set of predictions corresponding to the set of samples. Figure 4B The calculation of a residual block according to some embodiments is shown. As Figure 4B shown, the predicted block 404 is subtracted from the current block 402 to generate a residual block 406 that includes a set of residuals. For example, the corresponding difference between each sample and the corresponding prediction is calculated. Figure 4C The calculation of a reconstructed block according to some embodiments is shown. As Figure 4C shown, the residual block 406 undergoes one or more transforms and quantization to generate a set of residual coefficients. The set of residual coefficients can be transmitted from an encoder component to a decoder component. The set of residual coefficients undergoes inverse quantization and inverse transform to generate a reconstructed residual block 408. The reconstructed residual block 408 is combined with the predicted block 404 (e.g., the reconstructed residuals of the reconstructed residual block 408 are added to the predictions of the predicted block 404) to generate a reconstructed block 410 corresponding to the current block 402.

[0075] To reduce redundancy in the residual signal, various residual prediction techniques have been developed. These techniques predict the residual signal and encode the refined residual. Residual Difference Pulse Code Modulation (RDPCM) requires sample-based differential pulse code modulation along the horizontal or vertical axis. By doing so, each residual row in the horizontal mode (or column in the case of a vertical orientation) can be reconstructed at the decoder by summing the scaled differential pulse code modulation residual levels along the corresponding row (or column). RDPCM can be of an explicit type or an implicit type. The explicit type requires supplementary signaling of the direction and its application is limited to inter-prediction blocks. On the other hand, the implicit type does not require signaling of the direction and can only be applied to intra-prediction blocks where the prediction direction is associated with the intra-prediction mode. Block-based Differential Pulse Code Modulation (BDPCM) performs sample-based differential pulse code modulation on the reconstructed samples rather than the residual samples. The use of the second mode indication occurs during the prediction mode reconstruction process. This signaling involves two syntax elements, each for both luminance and chrominance. For example, the initial syntax element flag indicates its utilization, while the second syntax element flag specifies the horizontal or vertical direction. Thus, the decoder can receive video data from the video bitstream, which includes a plurality of blocks and a plurality of residual coefficients, and the plurality of blocks includes a first block.

[0076] Figure 5A FIG. shows a TIP mode according to some embodiments. In Figure 5A the example, interpolation processing is used to combine the information in reference frames 504-1 and 504-2 and project it to the same temporal instance as the current frame 502. In some embodiments, multiple TIP modes are supported. In a first example TIP mode, the interpolated frame 506 can be used as an additional reference frame. The decoded blocks of the current frame 502 can directly reference the interpolated frame 506 and thus utilize information from two different references, with only the overhead cost of a single inter-prediction mode. In another example TIP mode, the interpolated frame 506 is directly assigned as the decoded frame 508, as the output of the decoding process of the current frame 502 (e.g., skipping other conventional decoding steps such as generating residual blocks). This mode can provide significant decoding and simplification benefits, especially for low bitrate applications. Other techniques can be used to interpolate frames between two reference frames, such as Frame Rate Up Conversion (FRUC).

[0077] The exemplary TIP mode includes generating an interpolated frame 506 corresponding to the current frame 502. The interpolated frame 506 can then be used as an additional reference frame for the current frame 502 or be directly specified as the reconstructed output of the decoder for the current frame 502. On the decoder side, the blocks decoded in the TIP mode can be generated on the fly, such that it is not necessary to create the entire interpolated frame 506 at the decoder, thus saving decoding time and processing. Syntax elements can be used to indicate the frame-level TIP mode. Table 1 below shows examples of the modes indicated by the values of the tip_frame_mode parameter.

[0078] tip_frame_mode Meaning 0 Disable the TIP mode in this frame 1 Use the TIP frame as an additional reference frame 2 Directly output the TIP frame without decoding the current frame

[0079] Table 1 - Exemplary TIP Mode

[0080] Exemplary interpolation methods for interpolating an intermediate frame between two frames can reuse motion vectors from available references. The same motion vectors can also be used, with slight modification, for temporal motion vector predictor (TMVP) processing. For example, a rough motion vector field can be created for the TIP frame by projection of a modified TMVP field. In this example, the rough motion vector field is refined by filling holes and using a smoothing operation. In this example, the refined motion vector field is used to generate the TIP frame. On the decoder side, the blocks decoded using the TIP mode can be generated on the fly without creating the entire TIP frame. However, other suitable interpolation methods can be substituted in combination with other features discussed in this disclosure.

[0081] Figure 5B shows an exemplary in-loop filtering stage according to some embodiments. In Figure 5B the example, the in-loop filtering stage applied to the decoded frame 508 includes a deblocking filter 512, a constrained direction enhancement filter (CDEF) 514, and a loop restoration filter 516. In some embodiments, the filtered output frame is used as a reference frame for subsequent frames (e.g., stored in the reference frame buffer 520). In some embodiments, a normative film grain synthesis stage is also applied to generate the corresponding display picture 518. Different from the in-loop filter stage, the result of the film grain synthesis stage (e.g., out-of-loop filter) does not affect the prediction of subsequent frames. The loop filtering method can include any filtering process applied to the reconstructed samples (e.g., after adding the residual to the prediction), and the filtering process includes Wiener loop filtering, cross-component filtering, and a constrained direction enhancement filter (CDEF).

[0082] The cross-component filtering method may use co-located reconstructed samples and neighboring reconstructed samples from a first color component as inputs to perform filtering on a current reconstructed sample of a second color component. The cross-component offset filtering method may use co-located reconstructed samples and their neighboring reconstructed samples from a first color component as inputs to derive an offset value that is added to the current sample of the second color component to adjust its reconstructed value. The first color component may refer to a luminance color component, and the second color component may refer to a chrominance color component. The first color component and the second color component may be the same color component (e.g., the luminance component).

[0083] The deblocking filter 512 may be applied across transform block boundaries to remove block artifacts caused by quantization errors. In some embodiments, the filter length is determined based on the minimum transform block size on both sides. In some embodiments, the deblocking filter 512 uses a finite impulse response (FIR) filter (e.g., a low-pass filter). Edge detection may be used to disable the deblocking filter at transitions containing high variance signals (e.g., to avoid blurring actual edges in the original image). In this way, the deblocking filtering method may be applied to reconstructed samples located near block boundaries. Block boundaries may include transform block boundaries, motion compensation block boundaries, decoding block boundaries, and / or fixed block size boundaries.

[0084] The CDEF 514 applies a non-linear deringing filter along a specific (e.g., tilted) direction. The CDEF 514 may operate on the output of the deblocking filter 512. The CDEF 514 may operate in 8×8 units. In some embodiments, eight preset directions are defined by rotating and reflecting a template in a preset direction. The decoder may use the reconstructed pixels to select a general direction index. A primary filter may be applied along the selected direction, and a secondary filter may be applied along an offset direction (e.g., oriented 45° away from the primary direction). In some embodiments, up to eight sets of filter parameters are signaled (e.g., in a frame header). The filter parameter sets may include primary filter strength indices and secondary filter strength indices for luminance and chrominance components. The CDEF may apply filtering to the reconstructed samples by identifying the direction of each block and then performing adaptive filtering with a high degree of control over the filter strength along and across that direction.

[0085] In some embodiments, the loop restoration filter 516 is applied to the reconstructed pixels after any previous in-loop filtering stage (e.g., the deblocking filter 512 and / or the CDEF 514). The loop restoration filter 516 can be applied to a loop restoration unit (LRU), such as a 64×64, 128×128, and / or 256×256 pixel block. Bypass filtering, a Wiener filter (e.g., the Wiener loop filtering method), and / or a self-guided filter can be independently applied to each LRU. The Wiener loop filtering method can use a linear weighted sum of the current reconstructed sample and multiple spatially adjacent reconstructed samples as an input to obtain a modified value of the current reconstructed sample as an output.

[0086] As described above, the intra-frame interpolation method can obtain prediction samples by interpolating the current picture using one or more reference pictures (e.g., 1, 2, or 3 reference pictures) and directly obtaining the prediction samples from the interpolated picture. The intra-frame interpolation method can include a frame-level mode (e.g., tip_frame_mode = 2 in Table 1), which directly uses the interpolated picture as the reconstructed picture without sending any residuals. The intra-frame interpolation method can include a block-level mode that uses the interpolated picture as an additional reference frame and can further signal the motion vector and the residual. When applying the frame-level mode of the intra-frame interpolation method, a deblocking filtering process can be applied to the reconstructed picture, e.g., to reduce block artifacts caused by the block-based intra-frame interpolation process. To perform deblocking, some parameters related to the quantization process need to be provided to control the intensity of deblocking.

[0087] Sub-block based inter-prediction techniques such as TIP and optical flow motion vector refinement (OPFL) may introduce block artifacts during the prediction process. These artifacts may be difficult to remove by the deblocking filter. In some embodiments, a prediction enhancement filter (PEF) is employed during the prediction stage (e.g., to improve the visual quality with a minor impact on the encoding and decoding implementation).

[0088] In sub - block - based inter - frame prediction, such as for TIP and OPFL modes, a prediction unit (PU) can be split into smaller motion compensation units (MCUs). Each MCU can have its own motion vector pointing to a reference frame. When the motion information of an MCU is different from its neighbors, block artifacts may appear along the MCU boundaries (e.g., when there is no residual due to a low bit - rate budget). The de - block filter 512 discussed above can only process PU and TU boundaries. Therefore, when the MCU boundary does not align with the PU or TU boundary, the de - block filter may not process the MCU boundary.

[0089] In the TIP mode, a TIP reference frame can be generated in 8×8 MCU units. Then, the current block refers to the TIP frame via a motion vector. Since the motion vector can have arbitrary values, the block artifacts caused by using the TIP mode in the final prediction and / or reconstruction may not align the 8×8 grid of the PU with the 8×8 grid of the reconstructed frame. Additionally, the TIP reference may already contain block artifacts along the boundaries of each MCU.

[0090] In some embodiments, PEF is applied during the prediction stage to reduce the block artifacts caused by TIP and / or OPFL prediction processing and thus improve visual quality. For example, since the position of the block artifacts may not align with the 8×8 grid, two parameters can be derived based on the value of the motion vector to identify the position of the block artifacts. Then, PEF can be applied to the predicted samples at the internal MCU boundaries to reduce the block artifacts. When the TIP reference frame is used as the direct output, a filter can be applied to the 8×8 grid of the TIP frame. In the OPFL mode, the size of the MCU can be equal to 8×8 or 4×4, and the position of the block artifacts can align with the 8×8 or 4×4 grid. Filtering can be applied to the predicted samples located on the internal MCU boundaries to reduce the block artifacts.

[0091] PEF can include multiple steps. First, it is determined whether the MCU - level filter is on or off. For example, the motion vector difference on both sides of the MCU boundary is checked. When the motion vector difference is less than a threshold, filtering of the boundary can be skipped. For filtering on TIP, the TMVP motion vector can be used to check the motion vector difference; for filtering on OPFL, the OPFL - refined motion vector can be used alternatively. Next, a filter on / off decision can be made at the sample level. For example, a mask can be derived based on the samples near the boundary (e.g., the logic can be a simplified version of the logic used in the de - block filter 512). Next, an increment value is derived (similar to the de - block filter logic), and an offset is derived and applied to each of the samples to be filtered.

[0092] In some embodiments, when using sub-block motion compensation, a prediction filtering method (e.g., PEF) applies filtering to a prediction block. In some embodiments, the sub-block motion method performs motion compensation based on sub-blocks, e.g., when there are multiple sub-blocks within a decoded block. In an example, optical flow-based prediction can be used to refine the motion compensation for each sub-block within a given decoded block using an optical flow function, and the optical flow-based prediction is an example of the above sub-block motion method. Additionally, the prediction filtering method can also be applied to block units that perform intra interpolation in the frame-level mode of the intra interpolation method.

[0093] In some embodiments, the TIP frame-level mode is modified by using an implicit quantization index. When the interpolated frame obtained through the TIP mode is directly assigned as the output for the decoding process of the current frame (e.g., TIP frame-level mode), the luminance and chrominance quantization indices of the current frame are not signaled but are implicitly derived from the quantization indices of the reference frames. In this way, the decoding bits consumed for signaling the quantization index are saved. In some embodiments, a sequence-level flag is used to switch between implicit and explicit frame-level signaling of the luminance and chrominance quantization parameters (QP) for the TIP frame-level mode.

[0094] As discussed above, in the TIP mode, an intermediate frame is generated by interpolating using the motion vector fields of the forward and backward reference frames. In the first TIP mode (e.g., TIP_FRAME_AS_REF, sometimes also referred to as the TIP block-level mode), the interpolated frame is used as an additional reference frame for the current frame. In the second TIP mode (e.g., TIP_FRAME_AS_OUTPUT, sometimes also referred to as the TIP frame-level mode), the interpolated frame is directly output as the reconstruction of the current frame. As described above, when using the TIP mode, PEF can be applied to the interpolated frame to remove block artifacts. PEF requires the luminance and chrominance quantization indices of the AC coefficients for deblocking, and the quantization indices may need to be signaled from the encoder side to the decoder side. However, in some existing codec standards, only the quantization index associated with the luminance AC coefficients is signaled, while the quantization index associated with the chrominance AC coefficients is not signaled. When a bitstream is generated by an encoder using a non-CTC configuration, this may lead to an encoder-decoder mismatch because the knowledge of the quantization index associated with the chrominance AC coefficients is missing for the decoder but is used by the encoder to perform PEF filtering.

[0095] The quantization index for the TIP frame-level mode can be used to perform PEF filtering and is not used for coefficient decoding because there are no signaled residuals in the TIP frame-level mode. Therefore, when the residuals are decoded, the cost of signaling the quantization index is relatively higher compared to other frames. The method described below solves this problem.

[0096] When the TIP frame level mode is selected, the luminance and chrominance quantization indices of the AC coefficients can be derived from the reference frames instead of being signaled from the encoder side to the decoder side. For example, the quantization index can be derived as the average of the quantization indices from the reference frames. The terms base_q_idx, DeltaQUAc, and DeltaQVAc can represent the quantization index of the luminance AC coefficients and the incremental quantization indices of the Cb and Cr AC coefficients relative to base_q_idx, respectively. In this way, the quantization indices of the current interpolated frame can be represented as shown in Equation Set 1:

[0097]

[0098] DeltaQUAc cur =(DeltaQUAc ref1 +DeltaQUAc ref2 +1)>>1

[0099] DeltaQVAc cur =(DeltaQVAc ref1 +DeltaQVAc ref2 +1)>>1

[0100] Equation Set 1 - Quantization Indices

[0101] In Equation Set 1, the subscript "cur" corresponds to the current interpolated frame, and the subscripts "ref1" and "ref2" correspond to the two reference frames. Additionally, a sequence-level flag can be used to switch between the above-described implicit QP derivation scheme for the luminance and chrominance QPs for the TIP frame level mode and explicit frame-level signaling.

[0102] Figure 6A is a flowchart showing a method 600 for decoding video according to some embodiments. Method 600 can be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuitry and a memory storing instructions for execution by the control circuitry. In some embodiments, method 600 is performed by executing instructions stored in the memory (e.g., memory 314) of the computing system.

[0103] The system receives (602) a video bitstream including a plurality of encoded pictures. The system uses in - frame interpolation to derive (604) a reconstructed picture of an encoded picture among the plurality of encoded pictures. The system determines (606) whether one or more quantization parameters for the encoded picture are signaled in the video bitstream based on a signaled indicator in the video bitstream. When the signaled indicator indicates that one or more quantization parameters are signaled in the video bitstream, the system parses (608) out one or more quantization parameters from the video bitstream. When the signaled indicator indicates that one or more quantization parameters are not signaled in the video bitstream, the system derives (610) one or more quantization parameters.

[0104] In some embodiments, a first flag (e.g., a first indicator) is signaled to indicate whether to apply the frame - level mode of the in - frame interpolation method, and a second flag (e.g., a second indicator) is signaled to indicate whether to apply the prediction filtering method. In some embodiments, a third flag (e.g., a third indicator) is conditionally signaled to indicate whether to apply an explicit or implicit quantization parameter derivation method (e.g., whether one or more quantization parameters are signaled) on top of the frame - level mode of the in - frame interpolation method. In some embodiments, the condition depends on the signaling of the values of the first and second flags. As used herein, a "flag" may refer to a syntax having a binary value or a syntax having more than two value options.

[0105] In some embodiments, the first flag and / or the second flag are signaled in an advanced syntax. An example implementation is shown below in Example Syntax 1.

[0106]

[0107] Example Syntax 1 - General Sequence Header OBU Syntax

[0108] In Example Syntax 1, enable_tip is the first flag that signals whether to apply the in - frame interpolation method, and enable_tip equal to 1 indicates that the frame - level mode of the in - frame interpolation method is being applied. In Example Syntax 1, enable_pef is the second flag that signals whether to apply the prediction filtering method, and enable_tip_explicit_qp is the third flag that signals whether to apply an explicit or implicit quantization parameter derivation method on top of the frame - level mode of the in - frame interpolation method, where the third flag (enable_tip_explicit_qp) is signaled based on the condition that enable_tip is equal to 1 and enable_pef is non - zero.

[0109] In some embodiments, the third flag is signaled after the first and second flags. For example, the third flag is signaled immediately after the second flag. In another example, the third flag is signaled immediately after the first flag.

[0110] In some embodiments, when the third flag is signaled using a value derived from applying an explicit quantization parameter on top of a frame-level mode of an intra interpolation method, the following flags may be signaled partially or together. The fourth flag is signaled to specify a luminance-related quantization parameter. If the current frame has more than one color component, the fifth flag is signaled to specify whether the other two chrominance color components share the same delta quantization parameter relative to the luminance quantization parameter. If the fifth syntax is signaled using a value indicating that the two chrominance color components share the same delta quantization parameter, a single delta quantization parameter value is signaled. Otherwise, two delta quantization parameter values for the Cb and Cr color components are signaled separately. If the current frame has only a luminance component (e.g., is monochrome), the delta quantization parameter values for the Cb and Cr color components are not signaled and are set to a default value, e.g., 0. In some embodiments, the above flags / syntax are signaled using an advanced syntax.

[0111] An example frame header is shown below in Example Syntax 2.

[0112]

[0113]

[0114] Example Syntax 2 - Uncompressed Header Syntax

[0115] In Example Syntax 2, base_q_idx is the fourth flag, diff_uv_delta is the fifth flag, and DeltaQUAc and DeltaQVAc are the delta quantization parameters for Cb and Cr relative to the luminance quantization parameter.

[0116] In some embodiments, when the third flag is signaled using a value derived from applying an explicit quantization parameter on top of a frame-level mode of an intra interpolation method, the following flags may be signaled partially or together. The fourth flag is signaled to specify a luminance-related quantization parameter. If the current frame has more than one color component, two delta quantization parameter values for the Cb and Cr color components are signaled separately. If the current frame has only a luminance component (monochrome), the delta quantization parameter values for the Cb and Cr color components are set to a default value, e.g., 0.

[0117] In some embodiments, a first syntax (e.g., an indicator signaled in 606) is signaled to indicate whether the quantization parameter used in the frame-level mode of the frame interpolation method is implicitly derived or explicitly signaled. In some embodiments, the first syntax is a high-level syntax. In some embodiments, when the first syntax is signaled using a first value, at least one (or all) of the syntax related to the quantization parameter is explicitly signaled (e.g., at the frame level). In some embodiments, when the first syntax is signaled using a second value, at least one (or all) of the syntax related to the quantization parameter is not signaled but implicitly derived.

[0118] In some embodiments, the first syntax is signaled using more than N candidate values, and each candidate value specifies a predefined selection of whether to explicitly signal or implicitly derive the relevant quantization parameter. For example, the first syntax may include a plurality of flags, where each flag indicates whether to explicitly signal or implicitly derive the quantization parameter for each color (e.g., Y, Cb, and Cr) or color group (e.g., Cb and Cr together).

[0119] In some embodiments, the incremental value between the reference quantization parameter and the actual quantization parameter used for the current frame is signaled, rather than directly signaling the quantization parameter for the frame-level mode of the frame interpolation method. In some embodiments, the reference quantization parameter is derived using the quantization parameter of one or more reference pictures in the reference pictures of the current frame decoded in the frame-level mode of the frame interpolation method. In some embodiments, the incremental value may be signaled separately or jointly for different color components.

[0120] In some embodiments, when the quantization parameter used in the frame-level mode of the frame interpolation method is implicitly derived (e.g., loop filter processing), a weighted average of the quantization parameters associated with the reference pictures of the current frame decoded in the frame-level mode of the frame interpolation method is used. For example, for a selected reference frame, the weight is non-zero, while for other reference frames, the weight is zero. As another example, the implicitly derived quantization parameter is the minimum (or maximum) of the quantization parameters associated with all reference frames.

[0121] In some embodiments, when implicitly deriving quantization parameters used in the frame-level mode of an intra prediction method, a weighted average of quantization parameters associated with a particular block in a reference picture of the current frame decoded in the frame-level mode of the intra prediction method is used. For example, quantization parameters associated with a block located at predefined coordinates in a reference picture of the current frame decoded in the frame-level mode of the intra prediction method are used. Examples of the predefined coordinates include the four corner positions and the middle position of the reference picture. As an example, a weighted sum of quantization parameters associated with one or more blocks located at predefined coordinates in the reference picture of the current frame is used to derive the quantization parameters for the current frame. In some embodiments, the maximum (or minimum, or median, or average, or average of the maximum and minimum) of the quantization parameters associated with selected (or all) blocks in one or more reference frames is used to derive the quantization parameters for the current frame.

[0122] In some embodiments, when implicitly deriving quantization parameters used in the frame-level mode of an intra prediction method, the derivation is performed on a block-by-block basis. The block basis may include the largest decoded block, decoded block, transform block, and / or a block having a predefined block size. In some embodiments, for each current block, quantization parameters associated with one or more reference blocks in a reference frame are used to derive the quantization parameters for applying the frame-level mode of the intra prediction method to the current block. In this way, the quantization parameters used in the current frame may vary for different blocks.

[0123] In some embodiments, a signal is used to indicate a flag (e.g., an index) to indicate whether one or more of the loop filter methods are applied on the reconstructed samples obtained through the frame-level mode of the intra prediction method. For example, a signal is used to indicate a flag to indicate whether one or more of the following are applied: deblocking filter processing, cross-component sample offset (CCSO) loop filter method, Wiener loop filter processing, and CDEF processing. As an example, the flag is signaled to indicate whether multiple loop filter methods are applied simultaneously or disabled simultaneously. Example loop filter methods include deblocking filter processing, CCSO loop filter processing, Wiener loop filter processing, and CDEF filter processing. In some embodiments, a combined flag (e.g., an index) is signaled to indicate whether the loop filter methods are all disabled or all enabled. For example, a combined flag is signaled to indicate the application order of the loop filter methods.

[0124] In some embodiments, depending on the enabling of one or more loop filter methods, the reconstructed frame obtained through the frame-level mode of the intra prediction method is conditionally allowed to be used as a reference frame for another picture.

[0125] In some embodiments, parameters used in loop filtering processing performed on reconstructed samples obtained by a frame-level mode of an intra prediction method are different from parameters used in loop filtering processing on reconstructed samples that are not obtained by a frame-level mode. In some embodiments, at least some of the parameters used in loop filtering processing on reconstructed samples obtained by a frame-level mode are not signaled, but are derived using relevant loop filtering parameters used in reference pictures used to obtain interpolated frames.

[0126] In some embodiments, the maximum number of bands used in a cross-component filtering method for an intra prediction method is different from the maximum number of bands used in a cross-component filtering method for other methods. For example, the maximum number of bands used in a cross-component filtering method for an intra prediction method is less than the maximum number of bands used for other methods.

[0127] In some embodiments, a flag indicating whether one or more loop filtering methods are applied to reconstructed samples obtained by a frame-level mode of an intra prediction method is not signaled, but is implicitly derived using decoding information known to both an encoder and a decoder. For example, one or more loop filtering methods are always applied to reconstructed samples obtained by a frame-level mode without signaling a flag or index indicating an on / off selection. For example, the decoding information may include a quantization parameter, a temporal distance between a current picture and a reference picture used to perform intra prediction, a motion vector used to perform intra prediction, and / or whether a hole filling process is applied to the reconstructed samples to be loop filtered.

[0128] In some embodiments, a flag or index is signaled at a high level (e.g., frame level) to indicate whether one or more simplified versions of an existing loop filter are applied to a prediction mode of an intra prediction mode. Simplification methods may be predefined. For example, a simplified filter may be a simplified cross-component loop filter that is restricted to be applied to only one component and / or has fewer offset categories.

[0129] Figure 6B FIG. 650 is a flow chart illustrating a method 650 of encoding video according to some embodiments. Method 650 may be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuitry and a memory storing instructions for execution by the control circuitry. In some embodiments, method 650 is performed by executing instructions stored in a memory (e.g., memory 314) of the computing system.

[0130] The system receives (652) video data including a plurality of pictures. The system encodes (654) a first picture among the plurality of pictures according to a first type of intra prediction. The system determines (656) whether to signal in the video bitstream one or more quantization parameters for the first picture encoded. The system transmits (658) the encoded first picture via the video bitstream. The system signals (660) via the video bitstream a first indicator for indicating whether to signal in the video bitstream one or more quantization parameters for the first picture encoded. As described above, the encoding process may mirror the decoding process described herein. For the sake of brevity, these details are not repeated here.

[0131] Although Figure 6A and Figure 6B a number of logical stages are shown in a particular order, stages that are not order dependent can be reordered and other stages can be combined or split. Some reorderings or other groupings not specifically mentioned will be apparent to those of ordinary skill in the art, and so the orderings and groupings presented herein are not exhaustive. Additionally, it should be recognized that the stages may be implemented in hardware, firmware, software, or any combination thereof.

[0132] Turning now to some example embodiments.

[0133] (A1)In one aspect, some embodiments include a method of video decoding (e.g., method 600). In some embodiments, the method is performed at a computing device having a memory and one or more processors. The method includes: (i) receiving a video bitstream including a plurality of coded pictures; (ii) using intra interpolation to obtain reconstructed pictures of the coded pictures in the plurality of coded pictures; (iii) determining whether one or more quantization parameters for a coded picture are signaled in the video bitstream based on a signaled indicator in the video bitstream; (iv) when the signaled indicator has a first value (e.g., indicating that one or more quantization parameters are signaled in the video bitstream), parsing one or more quantization parameters from the video bitstream; and (v) when the signaled indicator has a second value (e.g., indicating that one or more quantization parameters are not signaled in the video bitstream), obtaining one or more quantization parameters. For example, a signaled indicator (e.g., a first syntax) is signaled to indicate whether the quantization parameters used in the frame-level mode of the intra interpolation mode are implicitly obtained or explicitly signaled. In some embodiments, one or more quantization parameters are parsed from the video bitstream according to determining that the signaled indicator indicates that one or more quantization parameters are signaled in the video bitstream. In some embodiments, one or more quantization parameters are obtained based on decoding information according to determining that the signaled indicator indicates that one or more quantization parameters are not signaled in the video bitstream.

[0134] (A2)In some embodiments according to A1, the signaled indicator is a high-level syntax element. For example, the signaled indicator (e.g., a first syntax) is a high-level syntax, which may be signaled at the sequence level, picture level, sub-picture level, slice level, tile level, or maximum decoding block level.

[0135] (A3)In some embodiments according to A1 or A2, the first value for the signaled indicator indicates that at least one of one or more quantization parameters is explicitly signaled in the video bitstream. For example, when the first syntax is signaled using the first value, at least one (or all) of the quantization parameter-related syntax is explicitly signaled (e.g., at the frame level).

[0136] (A4)In some embodiments according to any one of A1 to A3, the second value for the signaled indicator indicates that at least one of one or more quantization parameters will be obtained by the computing system. For example, when the first syntax is signaled using the second value, at least one (or all) of the quantization parameter-related syntax is implicitly obtained without being signaled.

[0137] (A5) In some embodiments according to any one of A1 to A4: (i) one or more quantization parameters include a plurality of quantization parameters, and (ii) the method further includes: when the signaled indicator has a third value, parsing a first subset of the plurality of quantization parameters from the video bitstream and deriving a second subset of the plurality of quantization parameters. For example, more than N candidate values may be used to signal the first syntax, and each candidate value specifies a predefined choice of whether the associated quantization parameter is signaled explicitly or derived implicitly.

[0138] (A6) In some embodiments according to any one of A1 to A5, the signaled indicator includes a set of flags, and each flag in the set of flags indicates whether one or more corresponding quantization parameters for a corresponding color component are to be parsed from the video bitstream or derived by the computing system. For example, the first syntax may include a plurality of flags, each flag indicating for each color component or group of color components (e.g., Cb and Cr together) whether the quantization parameter is signaled explicitly or derived implicitly.

[0139] (A7) In some embodiments according to any one of A1 to A6, deriving one or more quantization parameters includes using a weighted average of the corresponding quantization parameters from a plurality of reference pictures for an encoded picture. For example, when implicitly deriving the quantization parameter used in the intra prediction method's frame-level mode, a weighted average of the quantization parameters associated with the reference pictures of the current frame decoded in the frame-level mode is used.

[0140] (A8) In some embodiments according to A7, non-zero weights are applied to the selected reference pictures among the plurality of reference pictures, and zero weights are applied to the unselected reference pictures among the plurality of reference pictures. For example, for one selected reference frame, the weighting is non-zero, while for other reference frames, the weighting is zero.

[0141] (A9) In some embodiments according to any one of A1 to A8, deriving one or more quantization parameters includes selecting a minimum value or a maximum value from the corresponding quantization parameters of a plurality of reference pictures for an encoded picture. For example, the implicitly derived quantization parameter is the minimum value (or maximum value) of the quantization parameters associated with all reference frames.

[0142] (A10)In some embodiments according to any one of A1 to A9, obtaining one or more quantization parameters includes using a weighted average of corresponding quantization parameters of a plurality of blocks from one or more reference pictures for an encoded picture. For example, when implicitly obtaining the quantization parameter used in the frame-level mode of an intra prediction method, a weighted average of the quantization parameters associated with specific blocks in the reference picture of the current frame decoded in the frame-level mode is used.

[0143] (A11)In some embodiments according to A10, the plurality of blocks include blocks at predefined coordinates in one or more reference pictures. For example, the quantization parameters associated with the blocks located at the predefined coordinates in the reference picture of the current frame decoded in the frame-level mode of an intra prediction method are used. Examples of the predefined coordinates include, but are not limited to, the four corner positions or the middle position of the reference picture.

[0144] (A12)In some embodiments according to any one of A1 to A11, obtaining one or more quantization parameters includes using a weighted sum of corresponding quantization parameters of a plurality of blocks from one or more reference pictures for an encoded picture. For example, the quantization parameter for the current frame is obtained using a weighted sum of the quantization parameters associated with one or more blocks located at predefined coordinates in the reference picture of the current frame decoded in the frame-level mode of an intra prediction method.

[0145] (A13)In some embodiments according to any one of A1 to A12, obtaining one or more quantization parameters includes selecting a minimum value or a maximum value from the corresponding quantization parameters of a plurality of blocks from one or more reference pictures for an encoded picture. For example, the maximum value (or minimum value, or median value, or average value, or average of the maximum and minimum values) of the quantization parameters associated with the selected (or all) blocks in one or more reference frames is used to obtain the quantization parameter for the current frame decoded using the frame-level mode.

[0146] (A14)In some embodiments according to any one of A1 to A13, obtaining one or more quantization parameters includes obtaining one or more block-level quantization parameters. For example, when implicitly obtaining the quantization parameter used in the frame-level mode of an intra prediction method, the obtaining is performed on a block-by-block basis. The block basis includes, but is not limited to, the largest decoded block, decoded block, transform block, block with a predefined block size.

[0147] (A15)In some embodiments according to A14, deriving one or more block-level quantization parameters includes using corresponding quantization parameters from multiple reference blocks for an encoded picture. For example, for each current block, quantization parameters associated with one or more reference blocks in a reference frame are used to derive quantization parameters for a frame-level mode for applying an intra prediction method to the current block. In this way, the quantization parameters used in the current frame can vary for different blocks.

[0148] (B1)On the other hand, some embodiments include a method of video encoding (e.g., method 650). In some embodiments, the method is performed at a computing device having a memory and one or more processors. The method includes: (i) receiving video data including multiple pictures; (ii) encoding a first picture of the multiple pictures according to a first type of intra prediction; (iii) determining whether to signal in a video bitstream one or more quantization parameters for the first picture being encoded; (iv) transmitting the encoded first picture via the video bitstream; and (v) signaling via the video bitstream a first indicator for indicating whether one or more quantization parameters for the first picture being encoded are signaled in the video bitstream.

[0149] (B2)In some embodiments according to B1, the one or more quantization parameters include multiple quantization parameters, and the first indicator indicates that only a subset of the multiple quantization parameters is signaled.

[0150] (B3)In some embodiments according to B1 or B2, the first indicator includes a set of flags, and each flag in the set of flags indicates whether one or more corresponding quantization parameters for a corresponding color component are to be parsed from the video bitstream or derived by the computing system.

[0151] (B4)In some embodiments according to any one of B1 to B3, the method further includes: deriving at least one quantization parameter using a weighted average or weighted sum of corresponding quantization parameters from multiple reference pictures for an encoded picture.

[0152] (B5)In some embodiments according to B4, non-zero weights are assigned to selected reference pictures among the multiple reference pictures, and zero weights are assigned to unselected reference pictures among the multiple reference pictures.

[0153] (B6)In some embodiments according to any one of B1 to B5, the method further includes: deriving at least one quantization parameter by selecting a minimum value or a maximum value from corresponding quantization parameters of multiple reference pictures for an encoded picture.

[0154] (C1)In another aspect, some embodiments include a method of video decoding. In some embodiments, the method is performed at a computing device having a memory and one or more processors. The method includes: (i) receiving a video bitstream including a plurality of coded pictures; (ii) using intra interpolation to obtain a reconstructed picture of a coded picture among the plurality of coded pictures; (iii) determining whether a delta value of a quantization parameter for the coded picture is signaled in the video bitstream based on a signaled indicator in the video bitstream; (iv) when the signaled indicator has a first value, parsing the delta value of the quantization parameter from the video bitstream, wherein the delta value represents a difference between a value of a reference quantization parameter and a value of the quantization parameter; and (v) when the signaled indicator has a second value, obtaining one or more quantization parameters without parsing the delta value. For example, instead of directly signaling the quantization parameter for the frame level mode of Method A, the delta value between the reference quantization parameter and the actual quantization parameter for the current frame is signaled. In some embodiments, the delta value of the quantization parameter is parsed from the video bitstream according to determining that the signaled indicator has the first value. In some embodiments, according to determining that the signaled indicator has the second value, the decoder uses decoding information to obtain the delta value of the quantization parameter.

[0155] (C2)In some embodiments according to C1, the reference quantization parameter is obtained using a corresponding quantization parameter of one or more reference pictures for the coded picture. For example, the reference quantization parameter is obtained using the quantization parameter of one or more reference pictures in the reference pictures of the current frame decoded in the frame level mode of the intra interpolation method.

[0156] (C3)In some embodiments according to C1 or C2, the corresponding delta values for different color components of the coded picture are jointly signaled in the video bitstream. For example, the delta values can be signaled individually or jointly for different color components.

[0157] (D1)In another aspect, some embodiments include a method of visual media data processing. In some embodiments, the method is performed at a computing device having a memory and one or more processors. The method includes: (i) obtaining a source video sequence; and (ii) performing a conversion between the source video sequence and a bitstream of visual media data, wherein the bitstream includes: (a) a plurality of coded pictures, including a first coded picture encoded according to a first type of intra interpolation; and (b) a first indicator for indicating whether one or more quantization parameters for the first picture to be encoded are signaled in the video bitstream.

[0158] In another aspect, some embodiments include a computing system (e.g., server system 112) that includes control circuitry (e.g., control circuitry 302) and a memory (e.g., memory 314) coupled to the control circuitry, the memory storing one or more sets of instructions configured to be executed by the control circuitry, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 to A15, B1 to B6, C1 to C3, and D1 above). In yet another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more sets of instructions for execution by control circuitry of a computing system, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 to A15, B1 to B6, C1 to C3, and D1 above).

[0159] Unless otherwise specified, any of the syntax elements described herein can be High-Level Syntax (HLS). As used herein, HLS is signaled at a level higher than the block level. For example, HLS can correspond to the sequence level, the frame level, the slice level, or the tile level. As another example, HLS elements can be signaled in a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), a Picture Parameter Set (PPS), an Adaptation Parameter Set (APS), a slice header, a picture header, a tile header, and / or a CTU header.

[0160] It will be understood that although the terms “first,” “second,” etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” are also intended to include the plural forms. It will also be understood that the term “and / or” as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will also be understood that when used in this specification, the terms “comprises” and / or “comprising” specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0161] As used herein, the term "if" can be interpreted, depending on the context, to mean "when the precondition is true" or "after the precondition is true" or "in response to determining that the precondition is true" or "in accordance with determining that the precondition is true" or "in response to detecting that the precondition is true". Similarly, the phrases "if it is determined that [the precondition is true]" or "if [the precondition is true]" or "when [the precondition is true]" can be interpreted, depending on the context, to mean "after determining that the precondition is true" or "in response to determining that the precondition is true" or "in accordance with determining that the precondition is true" or "after detecting that the precondition is true" or "in response to detecting that the precondition is true".

[0162] For purposes of illustration, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best illustrate the principles of operation and practical application, thereby enabling others skilled in the art to implement them.

Claims

1. A method for video decoding performed at a computing system, the computing system having a memory and one or more processors, the method comprising: Receiving a video bitstream including a plurality of coded pictures; Deriving reconstructed pictures of the coded pictures in the plurality of coded pictures using intra interpolation; Determining whether one or more quantization parameters for the coded pictures are signaled in the video bitstream based on a signaled indicator in the video bitstream; When the signaled indicator indicates that the one or more quantization parameters are signaled in the video bitstream, parsing the one or more quantization parameters from the video bitstream; And When the signaled indicator indicates that the one or more quantization parameters are not signaled in the video bitstream, deriving the one or more quantization parameters.

2. The method according to claim 1, wherein, The one or more quantization parameters include a plurality of quantization parameters, and the method further comprises: when the signaled indicator indicates that only a subset of the plurality of quantization parameters is signaled in the video bitstream, parsing a first subset of the plurality of quantization parameters from the video bitstream and deriving a second subset of the plurality of quantization parameters.

3. The method according to claim 1, wherein, The signaled indicator includes a set of flags, and each flag in the set of flags indicates whether one or more corresponding quantization parameters for a corresponding color component are to be parsed or derived by the computing system from the video bitstream.

4. The method according to claim 1, wherein Deriving the one or more quantization parameters includes using a weighted average of corresponding quantization parameters from a plurality of reference pictures for the coded picture.

5. The method according to claim 4, wherein, Non-zero weights are applied to selected reference pictures among the plurality of reference pictures, and zero weights are applied to unselected reference pictures among the plurality of reference pictures.

6. The method according to claim 1, wherein Deriving the one or more quantization parameters includes selecting a minimum value or a maximum value from corresponding quantization parameters of a plurality of reference pictures for the coded picture.

7. The method according to claim 1, wherein, Deriving the one or more quantization parameters includes using a weighted average of corresponding quantization parameters from a plurality of blocks in one or more reference pictures for the coded picture.

8. The method according to claim 7, wherein The plurality of blocks includes blocks at predefined coordinates in the one or more reference pictures.

9. The method according to claim 1, wherein Deriving the one or more quantization parameters includes using a weighted sum of corresponding quantization parameters from a plurality of blocks in one or more reference pictures for the coded picture.

10. The method according to claim 1, wherein Deriving the one or more quantization parameters includes selecting a minimum value or a maximum value from corresponding quantization parameters of a plurality of blocks in one or more reference pictures for the coded picture.

11. The method according to claim 1, wherein, Deriving the one or more quantization parameters includes deriving one or more block-level quantization parameters.

12. The method according to claim 11, wherein, Deriving the one or more block-level quantization parameters includes using corresponding quantization parameters from a plurality of reference blocks for the coded picture.

13. The method according to claim 1, wherein, The signaled indicator includes high-level syntax elements.

14. A computing system, comprising: Control circuitry; Memory; And One or more sets of instructions stored in the memory and configured to be executed by the control circuitry, the one or more sets of instructions including instructions for the following operations: Receiving video data including a plurality of pictures; Encoding a first picture of the plurality of pictures according to a first type of intra prediction; Determining whether to signal in a video bitstream one or more quantization parameters for the encoded first picture; Transmitting the encoded first picture via the video bitstream; And Signaling via the video bitstream a first indicator for indicating whether one or more quantization parameters for the encoded first picture are signaled in the video bitstream.

15. The system according to claim 14, wherein The one or more quantization parameters include a plurality of quantization parameters, and wherein the first indicator indicates that only a subset of the plurality of quantization parameters is signaled.

16. The system according to claim 14, wherein, The first indicator includes a set of flags, and wherein each flag in the set of flags indicates whether one or more corresponding quantization parameters for a corresponding color component are to be parsed or derived by the computing system from the video bitstream.

17. The system according to claim 14, further comprising: Deriving the one or more quantization parameters using a weighted average or weighted sum of corresponding quantization parameters from a plurality of reference pictures for the encoded picture.

18. The system according to claim 17, wherein, Assigning non-zero weights to selected reference pictures of the plurality of reference pictures, and wherein zero weights are assigned to unselected reference pictures of the plurality of reference pictures.

19. The system according to claim 14, further comprising: Deriving the one or more quantization parameters by selecting a minimum value or a maximum value from corresponding quantization parameters of a plurality of reference pictures for the encoded picture.

20. A non-transitory computer-readable storage medium storing one or more sets of instructions configured to be executed by a computing device having control circuitry and a memory, the one or more sets of instructions including instructions for the following operations: Obtain the source video sequence; And Performing a conversion between the source video sequence and a bitstream of visual media data, wherein The bitstream includes: a plurality of encoded pictures including a first encoded picture encoded according to a first type of intra prediction; And a first indicator for indicating whether one or more quantization parameters for the encoded first picture are signaled in the bitstream.