Systems and methods for derivation of quantization parameters for frame interpolation
By deriving the quantization parameter set from the reference frame and performing loop filtering after reconstructing the image, the problems of insufficient accuracy and bandwidth efficiency in video decoding during frame interpolation are solved, achieving a more efficient video decoding effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT AMERICA LLC
- Filing Date
- 2024-04-08
- Publication Date
- 2026-05-29
AI Technical Summary
Existing video decoding technologies struggle to effectively utilize the quantization parameters of reference frames during frame interpolation, resulting in insufficient video decoding accuracy and bandwidth efficiency.
By deriving a set of quantization parameters from a reference frame and performing loop filtering after reconstructing the image, the accuracy and bandwidth efficiency of video decoding are improved.
It improves the accuracy and bandwidth efficiency of video decoding, and enhances video quality and transmission efficiency.
Smart Images

Figure CN122122895A_ABST
Abstract
Description
Related applications
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 546,908, filed November 1, 2023, entitled “Derivation of Quantization Parameters for Frame Interpolation,” and is a continuation of U.S. Patent Application No. 18 / 620,933, filed March 28, 2024, entitled “System and Method for Derivation of Quantization Parameters for Frame Interpolation,” and claims priority to that U.S. Patent Application. Technical Field
[0002] The disclosed implementations generally relate to video decoding, including but not limited to systems and methods for deriving quantization parameters for frame interpolation. Background Technology
[0003] Digital video is supported by various electronic devices such as digital televisions, laptop or desktop computers, tablet computers, digital camera devices, digital recording devices, digital media players, video game consoles, smartphones, video conferencing equipment, video streaming devices, etc. Electronic devices send and receive or otherwise transmit digital video data across communication networks, and / or store digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video decoding can be used to compress video data according to one or more video decoding standards before transmission or storage. Video decoding can be performed by hardware and / or software on electronic / client devices or servers providing cloud services.
[0004] Video decoding typically uses prediction methods that leverage the inherent redundancy in video data (e.g., inter-frame prediction, intra-frame prediction, etc.). Video decoding aims to compress video data into a form using a lower bitrate while avoiding or minimizing video quality degradation. Several video codec standards have been developed. For example, High-Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. The ITU-T and ISO / IEC released the HEVC / H.265 standard in 2013 (Revision 1), 2014 (Revision 2), 2015 (Revision 3), and 2016 (Revision 4). Versatile Video Coding (VVC / H.266) is a video compression standard designed to succeed HEVC. The ITU-T and ISO / IEC released the VVC / H.266 standard in 2020 (Revision 1) and 2022 (Revision 2). AOMedia Video 1 (AV1) is an open video decoding format designed as an alternative to HEVC. On January 8, 2019, a validated version 1.0.0 of the specification with errata table 1 was released. Summary of the Invention
[0005] In addition, this disclosure describes deriving a set of quantization parameters for reconstructing the image, which is derived from a reference set of quantization parameters used to encode the image. Loop filtering is then performed on the reconstructed image using this set of quantization parameters. Deriving quantization parameters for each color component can improve video decoding accuracy and also improve bandwidth efficiency (e.g., compared to systems that represent quantization parameters as signals).
[0006] According to some implementations, a video decoding method includes: (i) receiving a video bitstream comprising a plurality of coded images; (ii) obtaining a reconstructed image corresponding to one of the coded images; (iii) deriving a set of quantization parameters for the reconstructed image, the set of quantization parameters being derived from a reference set of quantization parameters for the coded image; and (iv) performing loop filtering on the reconstructed image using the set of quantization parameters.
[0007] According to some implementations, a video encoding method includes: (i) receiving video data comprising a plurality of images; (ii) encoding a first image among the plurality of images according to a frame-level mode; (iii) determining, based on the frame-level mode, one or more quantization parameters to be derived for the encoded first image; (iv) transmitting the encoded first image via a video bitstream; and (v) discarding the one or more quantization parameters for the encoded first image signaled via the video bitstream.
[0008] According to some embodiments, a method for processing visual media data includes: (i) obtaining a source video sequence; and (ii) performing a conversion between the source video sequence and a bitstream of visual media data, wherein the bitstream includes: (a) a plurality of coded images, the plurality of coded images including a first coded image encoded according to a frame-level mode; and (b) a first indicator for indicating whether one or more quantization parameters for the first coded image are signaled in the video bitstream.
[0009] According to some embodiments, a computing system, such as a streaming system, server system, personal computer system, or other electronic device, is provided. The computing system includes a control circuitry system and a memory storing one or more sets of instructions. The one or more sets of instructions include instructions for performing any of the methods described herein. In some embodiments, the computing system includes encoder components and decoder components (e.g., a transcoder). According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more sets of instructions executable by the computing system. The one or more sets of instructions include instructions for performing any of the methods described herein.
[0010] Therefore, apparatus and systems utilizing methods for encoding and decoding video are disclosed. Such methods, apparatus, and systems can supplement or replace conventional methods, apparatus, and systems for encoding / decoding video. The features and advantages described in the specification are not necessarily exhaustive, and in particular, some additional features and advantages will be apparent to those skilled in the art in light of the accompanying drawings, specification, and claims provided in this disclosure. Furthermore, it should be noted that the language used in the specification has been chosen primarily for readability and instructional purposes and is not necessarily intended to depict or limit the subject matter described herein. Attached Figure Description
[0011] To provide a more detailed understanding of this disclosure, a more specific description can be made by referring to the features of various embodiments, some of which are illustrated in the accompanying drawings. However, the drawings only illustrate relevant features of this disclosure and are therefore not intended to be limiting, as those skilled in the art will understand upon reading this disclosure that other valid features may be permitted in this description.
[0012] Figure 1 This is a block diagram illustrating an example communication system according to some implementations.
[0013] Figure 2A This is a block diagram illustrating example elements of an encoder component according to some embodiments.
[0014] Figure 2B This is a block diagram illustrating example elements of a decoder component according to some embodiments.
[0015] Figure 3 This is a block diagram illustrating an example server system according to some implementation methods.
[0016] Figures 4A to 4C Example prediction blocks, residual blocks, and reconstruction blocks are shown according to some implementation methods.
[0017] Figure 5A An example of a Temporal Interpolation Prediction (TIP) pattern according to some implementations is shown.
[0018] Figure 5B An example in-loop filtering stage according to some implementations is shown.
[0019] Figure 6A An example video decoding process according to some implementation methods is shown.
[0020] Figure 6B An example video encoding process according to some implementation methods is shown.
[0021] By convention, the various features shown in the accompanying drawings are not necessarily drawn to scale, and similar reference numerals may be used to indicate similar features throughout the specification and the drawings. Detailed Implementation
[0022] This disclosure describes a video / image compression technique that conditionally derives quantization parameters using reference quantization parameters from a reference frame derived from the current frame. The derived quantization parameters can then be used when performing subsequent loop filtering operations on the reconstructed image. In some implementations, the quantization parameters are derived from the current frame using frame-level interpolation (e.g., via TIP mode). For example, the quantization index can be derived from a reference frame as the average of the quantization indices. Example systems and devices
[0023] Figure 1 This is a block diagram illustrating a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to 120-m) that are communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is, for example, a streaming system for use with video-enabled applications such as video conferencing applications, digital TV applications, and media storage and / or distribution applications.
[0024] Source device 102 includes a video source 104 (e.g., a camera device component or media storage device) and an encoder component 106. In some embodiments, the video source 104 is a digital camera device (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates one or more encoded video bitstreams based on the video stream. The video stream from video source 104 can have a high data volume compared to the encoded video bitstream 108 generated by encoder component 106. Because the encoded video bitstream 108 has a lower data volume (less data) compared to the video stream from video source 104, it requires less bandwidth to transmit and less storage space to store compared to the video stream from video source 104. In some embodiments, source device 102 does not include encoder component 106 (e.g., configured to transmit uncompressed video to network 110).
[0025] One or more networks 110 represent any number of networks that transmit information between source device 102, server system 112, and / or electronic device 120, including, for example, wired (connected) and / or wireless communication networks. One or more networks 110 may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks (LANs), wide area networks (WANs), and / or the Internet.
[0026] One or more networks 110 include server systems 112 (e.g., distributed / cloud computing systems). In some embodiments, server system 112 is a streaming server (e.g., configured to store and / or distribute video content, such as encoded video streams from source device 102) or includes such a streaming server. Server system 112 includes decoder components 114 (e.g., configured to encode and / or decode video data). In some embodiments, decoder components 114 include encoder components and / or decoder components. In various embodiments, decoder components 114 are instantiated as hardware, software, or a combination thereof. In some embodiments, decoder components 114 are configured to decode encoded video bitstream 108 and re-encode video data using different encoding standards and / or methods to generate encoded video data 116. In some embodiments, server system 112 is configured to generate multiple video formats and / or encodings based on encoded video bitstream 108. In some embodiments, server system 112 serves as a Media-Aware Network Element (MANE). For example, server system 112 can be configured to trim encoded video bitstream 108 to tailor potentially different bitstreams for one or more of the electronic devices 120. In some implementations, MANE is provided separately from server system 112.
[0027] Electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode encoded video data 116 to generate an outgoing video stream that can be displayed on a display or other type of presentation device. In some embodiments, one or more of the electronic devices 120 do not include a display component (e.g., communicatively coupled to an external display device and / or includes a media storage device). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access a server system 112 to obtain encoded video data 116.
[0028] The source device and / or multiple electronic devices 120 are sometimes referred to as “terminal devices” or “user devices”. In some implementations, one or more of the electronic devices 120 and / or the source device 102 are instances of server systems, personal computers, portable devices (e.g., smartphones, tablets, or laptops), wearable devices, video conferencing equipment, and / or other types of electronic devices.
[0029] In an example operation of communication system 100, source device 102 transmits encoded video bitstream 108 to server system 112. For example, source device 102 may decode a stream of images captured by the source device. Server system 112 receives encoded video bitstream 108 and may decode and / or encode encoded video bitstream 108 using decoder component 114. For example, server system 112 may apply more optimized decoding to video data for network transmission and / or storage. Server system 112 may transmit encoded video data 116 (e.g., one or more encoded video bitstreams) to one or more electronic devices 120. Each electronic device 120 may decode encoded video data 116 and optionally display video images.
[0030] Figure 2AThis is a block diagram illustrating example elements of an encoder component 106 according to some embodiments. The encoder component 106 receives video data (e.g., a source video sequence) from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives the video sequence from a remote video source (e.g., a video source that is a component of a device different from the encoder component 106). The video source 104 can provide the source video sequence in the form of a digital video sample stream, which can have any suitable bit depth (e.g., 8-bit, 10-bit, or 12-bit), any color space (e.g., BT.601YCrCb or RGB), and any suitable sampling structure (e.g., YCrCb 4:2:0 or YCrCb 4:4:4). In some embodiments, the video source 104 is a storage device storing previously captured / prepared video. In some embodiments, the video source 104 is a camera device that captures local image information as a video sequence. Video data can be provided as multiple individual images that are given motion when viewed sequentially. Each image can be organized as a spatial array of pixels, where, depending on the sampling structure, color space, etc., each pixel may include one or more samples. The relationship between pixels and samples will be readily understood by those skilled in the art.
[0031] Encoder component 106 is configured to decode and / or compress images of a source video sequence into a decoded video sequence 216 in real time or under other time constraints required by the application. In some embodiments, encoder component 106 is configured to perform a conversion between the source video sequence and a bitstream of visual media data (e.g., a video bitstream). Implementing an appropriate decoding speed is a function of controller 204. In some embodiments, controller 204 controls and is functionally coupled to other functional units as described below. Parameters set by controller 204 may include rate control-related parameters (e.g., image skipping, quantizer and / or rate-distortion optimization techniques with λ values), image size, group of pictures (GOP) layout, maximum motion vector search range, etc. Other functions of controller 204 can be readily identified by those skilled in the art, as these functions may belong to encoder component 106 optimized for a particular system design.
[0032] In some implementations, encoder component 106 is configured to operate within a decoding loop. In a simplified example, the decoding loop includes a source decoder 202 (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and a reference image) and a (local) decoder 210. Decoder 210 reconstructs the symbols in a manner similar to that of the (remote) decoder to create sample data (assuming lossless compression between the symbols and the decoded video bitstream). The reconstructed sample stream (sample data) is input to reference image memory 208. Since decoding of the symbol stream produces bit-accurate results regardless of the decoder's location (local or remote), the contents of reference image memory 208 are also bit-accurate between the local and remote encoders. In this way, the encoder's prediction portion interprets the same sample values as the sample values that the decoder interprets during prediction as reference image samples.
[0033] The operation of decoder 210 can be combined with a remote decoder, for example, as shown below. Figure 2B The operation of decoder component 122 is the same as described in the detailed description. However, a brief reference is provided. Figure 2B Since the symbols are available and the encoding / decoding of the symbols into a decoded video sequence by the entropy decoder 214 and the parser 254 can be lossless, the entropy decoding part of the decoder component 122, including the buffer memory 252 and the parser 254, does not need to be fully implemented in the local decoder 210.
[0034] Aside from parsing / entropy decoding, the decoder techniques described in this paper can exist in their corresponding encoders with essentially the same functionalities. For this reason, the subject matter focuses on decoder operations. Furthermore, the descriptions of encoder techniques can be simplified, as they can be the opposite of decoder techniques.
[0035] As part of its operation, source decoder 202 can perform motion-compensated predictive decoding, referencing one or more previously decoded frames in the video sequence designated as reference frames, to predictively decode the input frame. In this way, decoding engine 212 decodes the differences between pixel blocks of the input frame and pixel blocks of the reference frame, which can be selected as the prediction reference for the input frame. Controller 204 can manage the decoding operations of source decoder 202, including, for example, setting parameters and subgroup parameters for encoding video data.
[0036] Decoder 210 decodes decoded video data based on symbols created by source decoder 202, which can be designated as reference frames. The operation of decoding engine 212 can advantageously handle lossy processing. When decoded video data is in the video decoder (… Figure 2AWhen decoded at a location (not shown), the reconstructed video sequence can be a copy of the source video sequence with some errors. Decoder 210 can replicate the decoding process performed on the reference frame by a remote video decoder, and the reconstructed reference frame can be stored in reference image memory 208. In this way, encoder component 106 locally stores a copy of the reconstructed reference frame that shares common content (no transmission errors) with the reconstructed reference frame to be obtained by the remote video decoder.
[0037] Predictor 206 can perform a prediction search against decoding engine 212. That is, for a new frame to be decoded, predictor 206 can search the reference image memory 208 for sample data (as candidate reference pixel blocks) or specific metadata such as reference image motion vectors, block shapes, etc., that can be used as appropriate prediction references for the new image. Predictor 206 can operate pixel-by-pixel based on the sample blocks to find appropriate prediction references. As determined by the search results obtained by predictor 206, the input image can have prediction references obtained from multiple reference images stored in reference image memory 208.
[0038] The outputs of all the aforementioned functional units can undergo entropy decoding in entropy decoder 214. Entropy decoder 214 converts the symbols, such as those generated by the various functional units, into a decoded video sequence by lossless compression according to techniques known to those skilled in the art (e.g., Huffman decoding, variable-length decoding, and / or arithmetic decoding).
[0039] In some implementations, the output of entropy decoder 214 is coupled to a transmitter. The transmitter can be configured to buffer, for example, the decoded video sequence created by entropy decoder 214, in preparation for transmission via communication channel 218, which can be a hardware / software link to a storage device storing the encoded video data. The transmitter can be configured to combine decoded video data from source decoder 202 with other data to be transmitted, such as decoded audio data and / or auxiliary data streams (source not shown). In some implementations, the transmitter can transmit additional data along with the encoded video. Source decoder 202 can include such data as part of the decoded video sequence. Additional data can include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant images and slices, Supplementary Enhancement Information (SEI) messages, fragments of Visual Usability Information (VUI) parameter sets, etc.
[0040] Controller 204 can manage the operation of encoder component 106. During decoding, controller 204 can assign a specific decoding image type to each decoding image, which may affect the decoding technique applied to the corresponding image. For example, images can be assigned as intra-frame images (I-images), predictive images (P-images), or bidirectional predictive images (B-images). Intra-frame images can be decoded and decoded without using any other frames in the sequence as prediction sources. Some video codecs allow different types of intra-frame images, including, for example, Independent Decoder Refresh (IDR) images. Those skilled in the art are familiar with those variations of I-images and their corresponding applications and characteristics, and therefore will not be repeated here. Predictive images can be decoded and decoded using inter-frame prediction or intra-frame prediction that uses at most one motion vector and reference index to predict sample values for each block. Bidirectional predictive images can be decoded and decoded using inter-frame prediction or intra-frame prediction that uses at most two motion vectors and reference indexes to predict sample values for each block. Similarly, multiple predictive images can use more than two reference images and associated metadata to reconstruct a single block.
[0041] The source image can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and decoded on a block-by-block basis. Predictive decoding can be performed on these blocks with reference to other (already decoded) blocks, determined by the decoding assignment of the corresponding images applied to the blocks. For example, blocks of image I can be unpredictably decoded, or blocks of image I can be predictively decoded (spatial prediction or intra-frame prediction) with reference to already decoded blocks of the same image. Pixel blocks of image P can be unpredictably decoded via spatial prediction or via temporal prediction with reference to a previously decoded reference image. Blocks of image B can be unpredictably decoded via spatial prediction or via temporal prediction with reference to one or two previously decoded reference images.
[0042] Video can be captured as multiple source images (video pictures) in a time series. Intra-frame picture prediction (often simplified to intra-frame prediction) utilizes spatial correlations within a given image, while inter-frame picture prediction utilizes (temporal or other) correlations between images. In the example, a specific image in the encoding / decoding process—referred to as the current image—is segmented into blocks. Where a block in the current image is similar to a reference block in a previously decoded and still buffered reference image in the video, the block in the current image can be decoded using a vector called a motion vector. The motion vector points to the reference block in the reference image, and when using multiple reference images, the motion vector can have a third dimension that identifies the reference images.
[0043] Encoder component 106 can perform decoding operations according to any predetermined video decoding technology or standard described herein. In operation, encoder component 106 can perform various compression operations, including predictive decoding operations utilizing temporal and spatial redundancy in the input video sequence. Therefore, the decoded video data can conform to the syntax specified by the video decoding technology or standard used.
[0044] Figure 2B This is a block diagram illustrating example elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 is coupled to channel 218 and display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to loop filter 256 and configured to transmit data to display 124 (e.g., via a wired or wireless connection).
[0045] In some embodiments, decoder component 122 includes a receiver coupled to channel 218 and configured to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more decoded video sequences to be decoded by decoder component 122. In some embodiments, the decoding of each decoded video sequence is independent of other decoded video sequences. Each decoded video sequence may be received from channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive encoded video data as well as other data, such as decoded audio data and / or auxiliary data streams, which may be forwarded to their respective user entities (not depicted). The receiver may separate the decoded video sequences from other data. In some embodiments, the receiver receives additional (redundant) data along with the decoded video. The additional data may be included as part of the encoded video sequence. The decoder component 122 may use the additional data to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0046] According to some embodiments, the decoder component 122 includes a buffer memory 252, a parser 254 (sometimes also referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra-frame image prediction unit 262, a motion compensation prediction unit 260, an aggregator 268, a loop filter unit 256, a reference image memory 266, and a current image memory 264. In some embodiments, the decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuit systems. The decoder component 122 can be implemented at least partially in software.
[0047] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to combat network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, a separate buffer memory is provided outside decoder component 122 (e.g., to combat network jitter) in addition to buffer memory 252 inside decoder component 122 (e.g., configured to handle broadcast timing). Buffer memory 252 may not be necessary or may be small when receiving data from a store / forward device with sufficient bandwidth and controllability or from an isochronous network. Buffer memory 252 may be required to make the best use of packet networks such as the Internet; buffer memory 252 may be relatively large and / or have an adaptive size and may be implemented at least partially in an operating system or similar component outside decoder component 122.
[0048] Parser 254 is configured to reconstruct symbols 270 from the decoded video sequence. Symbols may include, for example, information for managing the operation of decoder component 122, and / or information for controlling presentation devices such as display 124. Control information for the presentation device may be in the form of, for example, supplementary enhancement information (SEI) messages or fragments of video availability information (VUI) parameter sets (not depicted). Parser 254 parses (entropy decodes) the decoded video sequence. Decoding of the decoded video sequence may be performed according to video decoding techniques or standards and may follow principles known to those skilled in the art, including variable-length decoding, Huffman decoding, arithmetic decoding with or without context sensitivity, etc. Parser 254 may extract a subgroup parameter set from the decoded video sequence for at least one subgroup of pixels in the subgroups used in the video decoder, based on at least one parameter corresponding to a group. Subgroups may include picture groups (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser 254 can also extract information such as transform coefficients, quantizer parameter values, and motion vectors from the decoded video sequence.
[0049] Depending on the type of the decoded video picture or a portion thereof (e.g., inter-frame and intra-frame pictures, inter-frame and intra-frame blocks) and other factors, the reconstruction of symbol 270 may involve multiple different units. Which units are involved and how they are involved can be controlled by subgroup control information parsed by parser 254 from the decoded video sequence. For simplicity, such subgroup control information flow between parser 254 and the following multiple units is not depicted.
[0050] The decoder component 122 can be conceptually subdivided into multiple functional units, and in some implementations, these units interact closely with each other and can be at least partially integrated with each other. However, for the sake of brevity, this paper retains the conceptual subdivision of functional units.
[0051] The scaler / inverse transform unit 258 receives quantized transform coefficients as symbols 270 and control information (such as which transform to use, block size, quantization factor, and / or quantization scaling matrix) from the parser 254. The scaler / inverse transform unit 258 can output blocks comprising sample values, which can be input to the aggregator 268. In some cases, the output samples of the scaler / inverse transform unit 258 belong to intra-coded blocks; that is, blocks that do not use predictive information from previously reconstructed images, but can use predictive information from previously reconstructed portions of the current image. Such predictive information can be provided by the intra-picture prediction unit 262. The intra-picture prediction unit 262 can use surrounding reconstructed information obtained from the current (partially reconstructed) image in the current image memory 264 to generate blocks of the same size and shape as the blocks in the reconstruction. The aggregator 268 can add the prediction information already generated by the intra-picture prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258 based on each sample.
[0052] In other cases, the output samples of the scaler / inverse transform unit 258 belong to an inter-frame decoding block and potentially to a motion-compensated block. In such cases, the motion-compensated prediction unit 260 can access the reference image memory 266 to obtain samples for prediction. After motion compensation of the obtained samples according to the symbols 270 belonging to the block, these samples can be added by the aggregator 268 to the output of the scaler / inverse transform unit 258 (referred to in this case as residual samples or residual signals) to generate output sample information. The address in the reference image memory 266 from which the motion-compensated prediction unit 260 obtains the predicted samples can be controlled by motion vectors. Motion vectors can be used by the motion-compensated prediction unit 260 in the form of symbols 270, which can have, for example, X components, Y components, and reference image components. Motion compensation can also include, for example, interpolation of sample values obtained from the reference image memory 266 when using subsampled precise motion vectors, and motion vector prediction mechanisms.
[0053] The output samples of aggregator 268 can undergo various loop filtering techniques in loop filter unit 256. Video compression techniques may include in-loop filtering techniques controlled by parameters included in the decoded video bitstream and available to loop filter unit 256 as symbols 270 from parser 254. However, video compression techniques may also respond to metadata obtained during decoding of previous (in decoding order) portions of the decoded picture or decoded video sequence, and to sample values obtained from previous reconstruction and loop filtering. The output of loop filter unit 256 may be a sample stream that can be output to a presentation device such as display 124 and stored in reference picture memory 266 for future inter-frame picture prediction.
[0054] Once reconstructed, certain decoded images can be used as reference images for future predictions. Once a decoded image has been reconstructed and has been identified as a reference image (e.g., by parser 254), the current reference image can become part of the reference image memory 266, and a new current image memory can be reallocated before reconstructing subsequent decoded images begins.
[0055] Decoder component 122 can perform decoding operations according to a predetermined video compression technique that may be recorded in any standard such as those described herein. As specified in a video compression technique document or standard, and particularly in a configuration file therein, the decoded video sequence may conform to the syntax specified by the video compression technique or standard used, in the sense that the decoded video sequence follows the syntax of the video compression technique or standard. Furthermore, to conform to some video compression techniques or standards, the complexity of the decoded video sequence may be within a range defined by the hierarchy of the video compression technique or standard. In some cases, the hierarchy limits the maximum image size, maximum frame rate, maximum reconstruction sample rate (measured, for example, megasamples per second), maximum reference image size, etc. In some cases, the limitations set by the hierarchy can be further limited by the Hypothetical Reference Decoder (HRD) specification and metadata managed by the HRD buffers used to represent signals in the decoded video sequence.
[0056] Figure 3This is a block diagram illustrating a server system 112 according to some embodiments. The server system 112 includes a control circuitry system 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuitry system 302 includes one or more processors (e.g., CPU, GPU, and / or DPU). In some embodiments, the control circuitry system includes a field-programmable gate array, a hardware accelerator, and / or an integrated circuit (e.g., an application-specific integrated circuit).
[0057] Network interface 304 can be configured to interface with one or more communication networks (e.g., wireless networks, wired networks, and / or optical networks). Communication networks can be local, wide area, metropolitan area, vehicle and industrial, real-time, latency-tolerant, etc. Examples of communication networks include: local area networks such as Ethernet; wireless LANs (LANs); cellular networks including GSM (Global System for Mobile Communication), 3G (the Third Generation), 4G (the Fourth Generation), 5G (the Fifth Generation), LTE (Long Term Evolution), etc.; wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; vehicle and industrial networks including CANbus, etc. Such communication can be one-way receiving (e.g., broadcast TV), one-way transmitting (e.g., to a CAN bus of some CAN bus device), or bidirectional (e.g., to other computer systems using local or wide area digital networks). Such communication can include communication to one or more cloud computing networks.
[0058] User interface 306 includes one or more output devices 308 and / or one or more input devices 310. Input devices 310 may include one or more of the following: keyboard, mouse, touchpad, touch screen, data glove, joystick, microphone, scanner, camera device, etc. Output devices 308 may include one or more of the following: audio output devices (e.g., speakers), visual output devices (e.g., monitors or displays), etc.
[0059] Memory 314 may include high-speed random access memory (e.g., DRAM (Dynamic Random Access Memory), SRAM (Static Random Access Memory), DDR RAM (Double Data Rate Random Access Memory), and / or other random access solid-state memory devices) and / or non-volatile memory (e.g., one or more disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state memory devices). Memory 314 may optionally include one or more storage devices remote from the control circuitry system 302. Memory 314, or alternatively, the non-volatile solid-state memory devices within memory 314 include non-transitory computer-readable storage media. In some embodiments, memory 314 or the non-transitory computer-readable storage media of memory 314 stores programs, modules, instructions, and data structures, or subsets or supersets thereof: ● Operating system 316, which includes procedures for handling various basic system services and for performing hardware-related tasks; ● Network communication module 318, which is used via one or more network interfaces 304 (For example, via wired and / or wireless connection) connect server system 112 to other computing devices; ● Decoding module 320 performs various functions related to encoding and / or decoding data, such as video data. In some embodiments, decoding module 320 is an example of decoder component 114. Decoding module 320 includes, but is not limited to, one or more of the following: Decoding module 322, which performs various functions related to decoding encoded data, such as those previously described with respect to decoder component 122; and The encoding module 340 performs various functions related to encoding data, such as those previously described with respect to encoder component 106; and ● Image memory 352 is used to store images and image data, for example, for use with decoding module 320. In some embodiments, image memory 352 includes one or more of the following: reference image memory 208, buffer memory 252, current image memory 264, and reference image memory 266.
[0060] In some implementations, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions previously described with respect to the parser 254), a transformation module 326 (e.g., configured to perform various functions previously described with respect to the scaler / inverse transformation unit 258), a prediction module 328 (e.g., configured to perform various functions previously described with respect to the motion compensation prediction unit 260 and / or the intra-frame image prediction unit 262), and a filter module 330 (e.g., configured to perform various functions previously described with respect to the loop filter 256).
[0061] In some embodiments, the encoding module 340 includes a code module 342 (e.g., configured to perform the various functions previously described with respect to source decoder 202 and / or decoding engine 212) and a prediction module 344 (e.g., configured to perform the various functions previously described with respect to predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 includes... Figure 3 A subset of the modules shown. For example, a shared prediction module is used by both the decoding module 322 and the encoding module 340.
[0062] Each of the modules identified above and stored in memory 314 corresponds to a set of instructions for performing the functions described herein. The modules identified above (e.g., instruction sets) do not need to be implemented as separate software programs, processes, or modules, and therefore various subsets of these modules can be combined or otherwise rearranged in various embodiments. For example, decoding module 320 may optionally not include separate decoding and encoding modules, but instead use the same set of modules to perform both sets of functions. In some embodiments, memory 314 stores a subset of the modules and data structures identified above. In some embodiments, memory 314 stores additional modules and data structures not described above.
[0063] Although Figure 3 A server system 112 according to some embodiments is shown, but Figure 3 This is intended more as a functional description of various features that can exist in one or more server systems than as a structural diagram of the implementation described herein. In practice, items shown individually may be combined, and some items may be separate. For example, Figure 3 Some items shown individually can be implemented on a single server, and a single item can be implemented by one or more servers. The actual number of servers used to implement server system 112 and how features are distributed among said servers will vary depending on the implementation method, and optionally, it depends in part on the amount of data traffic processed by the server system during peak usage periods and during average usage periods. Example Decoding Techniques
[0064] As described above, some codecs (e.g., AV1) operate on pixel blocks. Each pixel block can be processed in a predictive transform decoding scheme, where prediction is obtained using intra-frame reference pixels, inter-frame motion compensation, or some combination of both. The residual from the prediction can undergo a transform (e.g., a 2-D unit transform) to further remove spatial correlations, and the transform coefficients are quantized. Then, arithmetic decoding can be used to entropy decode the predictive syntax elements and the quantized transform coefficient indices.
[0065] Figure 4A The calculation of a prediction block according to some implementations is shown. Figure 4A In the example, intra-frame prediction is performed on the current block 402 to generate prediction block 404. The current block 402 includes a set of samples (e.g., pixel blocks), and prediction block 404 includes a prediction set corresponding to that set of samples. Figure 4B The calculation of residual blocks according to some implementation methods is shown. For example... Figure 4B As shown, prediction block 404 is subtracted from the current block 402 to generate residual block 406, which includes the residual set. For example, the corresponding difference between each sample and the corresponding prediction is calculated. Figure 4C The calculation of the reconstructed block according to some implementation methods is shown. For example... Figure 4C As shown, residual block 406 undergoes one or more transforms and quantizations to generate a set of residual coefficients. This set of residual coefficients can be transferred from the encoder component to the decoder component. The set of residual coefficients undergoes inverse quantization and inverse transform to generate a reconstructed residual block 408. The reconstructed residual block 408 is combined with prediction block 404 (e.g., the reconstructed residuals of the reconstructed residual block 408 are added to the predictions of prediction block 404) to generate a reconstructed block 410 corresponding to the current block 402.
[0066] To reduce redundancy in residual signals, various residual prediction techniques have been developed. These techniques predict the residual signal and encode the refined residual. Residual Difference Pulse Code Modulation (RDPCM) requires sample-based differential pulse code modulation along either the horizontal or vertical axis. By doing so, each residual row (or column in the case of vertical orientation) in the horizontal mode can be reconstructed at the decoder by summing the scaled differential pulse code modulation residual levels along the corresponding row (or column). RDPCM can be explicit or implicit. Explicit types require directional supplementary signaling and their application is limited to inter-frame prediction blocks. Implicit types, on the other hand, do not require directional signaling and can be applied only to intra-frame prediction blocks, where the prediction direction is associated with the intra-frame prediction mode. Block-based Differential Pulse Code Modulation (BDPCM) performs sample-based differential pulse code modulation on the reconstructed samples rather than the residual samples. Indication of the use of the second mode occurs during the prediction mode reconstruction process. This signaling involves two syntax elements, each for both luma and chroma. For example, the initial syntax element flag indicates its utilization, while the second syntax element flag specifies the horizontal or vertical direction. Therefore, the decoder can receive video data from the video bitstream comprising multiple blocks and multiple residual coefficients, said multiple blocks including the first block.
[0067] Figure 5A TIP patterns according to some implementation methods are shown. Figure 5A In the example, interpolation is used to combine information from reference frames 504-1 and 504-2 and project that information onto the same time instance as the current frame 502. In some implementations, multiple TIP modes are supported. In a first example TIP mode, the interpolated frame 506 is used as an additional reference frame. The decoded block of the current frame 502 can directly reference the interpolated frame 506, thus utilizing information from two different references at the overhead of only a single inter-frame prediction mode. In another example TIP mode, the interpolated frame 506 is directly designated as the decoded frame 508, i.e., the output of the decoding process of the current frame 502 (e.g., skipping other conventional decoding steps, such as generating residual blocks). This mode can provide considerable decoding and simplification advantages, especially for low bit-rate applications. Other techniques for interpolating frames between two reference frames can be used, such as Frame Rate Up Conversion (FRUC).
[0068] Example TIP modes include generating an interpolated frame 506 corresponding to the current frame 502. The interpolated frame 506 can then be used as an additional reference frame for the current frame 502, or directly designated as the reconstructed output of the decoder for the current frame 502. On the decoder side, blocks decoded in TIP mode can be dynamically generated, eliminating the need to create the entire interpolated frame 506 at the decoder, saving decoding time and processing. Frame-level TIP modes can be indicated using syntax elements. Examples of modes indicated by the value of the `tip_frame_mode` parameter are shown in Table 1 below. tip_frame_mode meaning 0 Disable TIP mode in this frame 1 Use TIP frames as additional reference frames. 2 Output the TIP frame directly without decoding the current frame. Table 1—Example TIP Patterns
[0069] Example interpolation methods for interpolating intermediate frames between two frames can reuse motion vectors from available references. The same motion vectors, with slight modifications, can also be used for temporal motion vector predictor (TMVP) processing. For example, a coarse motion vector field can be created for a TIP frame by projecting a modified TMVP field. In this example, the coarse motion vector field is refined by filling holes and using a smoothing operation. In this example, the TIP frame is generated using the refined motion vector field. On the decoder side, blocks decoded using the TIP mode can be dynamically generated without creating the entire TIP frame. However, other suitable interpolation methods can be used instead, in conjunction with other features discussed in this disclosure.
[0070] Figure 5B An example in-loop filtering stage according to some implementations is shown. Figure 5B In the example, the in-loop filtering stage applied to decoded frame 508 includes a deblocking filter 512, a Constrained Direction Enhancement Filter (CDEF) 514, and a loop recovery filter 516. In some embodiments, the filtered output frame is used as a reference frame for subsequent frames (e.g., stored in a reference frame buffer 520). In some embodiments, a canonical film grain synthesis stage is also applied to generate the corresponding display image 518. Unlike the in-loop filtering stages, the results of the film grain synthesis stage (e.g., an out-of-loop filter) do not affect the prediction of subsequent frames. Loop filtering methods can include any filtering applied to the reconstructed samples (e.g., after adding residuals to the prediction), including Wiener loop filtering, cross-component filtering, and the Constrained Direction Enhancement Filter (CDEF).
[0071] Cross-component filtering methods can use co-located reconstructed samples and neighboring reconstructed samples from a first color component as input to perform filtering on the current reconstructed sample of the second color component. Cross-component offset filtering methods can use co-located reconstructed samples from a first color component and their neighboring reconstructed samples as input to derive an offset value, which is added to the current sample of the second color component to adjust its reconstructed value. The first color component can refer to the luminance color component, and the second color component can refer to the chrominance color component. The first and second color components can be the same color component (e.g., the luminance component).
[0072] The deblocking filter 512 can be applied across transform block boundaries to remove block artifacts caused by quantization errors. In some implementations, the filter length is determined based on the minimum transform block size on either side. In some implementations, the deblocking filter 512 uses a finite impulse response (FIR) filter (e.g., a low-pass filter). Edge detection can be used to disable the deblocking filter at transitions containing high-variance signals (e.g., to avoid blurring actual edges in the original image). In this way, the deblocking filtering method can be applied to reconstructed samples located near block boundaries. Block boundaries can include transform block boundaries, motion-compensated block boundaries, decoding block boundaries, and / or fixed block size boundaries.
[0073] CDEF 514 applies a nonlinear deringing filter along a specific (e.g., tilted) direction. CDEF 514 can operate on the output of deblocking filter 512. CDEF 514 can operate in 8×8 units. In some implementations, eight preset directions are defined by rotating and reflecting templates in preset directions. The decoder can use reconstructed pixels to select a popular direction index. A primary filter can be applied along the selected direction, and a secondary filter can be applied along an offset direction (e.g., 45° away from the primary direction). In some implementations, up to eight sets of filter parameters are represented as signals (e.g., in the frame header). The filter parameter sets can include primary and secondary filter strength indices for the luma and chroma components. CDEF can apply filtering to reconstructed samples by identifying the direction of each block and then adaptively filtering along and across the direction with highly controlled filter strength.
[0074] In some implementations, a loop restoration filter 516 is applied to the reconstructed pixels after any prior in-loop filtering stage (e.g., deblocking filter 512 and / or CDEF 514). The loop restoration filter 516 can be applied to loop restoration units (LRUs), such as 64×64 pixel blocks, 128×128 pixel blocks, and / or 256×256 pixel blocks. Bypass filtering, Wiener filtering (e.g., the Wiener loop filtering method), and / or self-guided filters can be applied independently to each LRU. The Wiener loop filtering method can use the current reconstructed sample and a linear weighted sum of multiple spatially adjacent reconstructed samples as input to derive a modified value for the current reconstructed sample as output.
[0075] As described above, frame interpolation methods derive prediction samples by interpolating the current image using one or more reference images (e.g., 1, 2, or 3 reference images) and directly obtaining prediction samples from the interpolated images. Frame interpolation methods can include frame-level modes (e.g., tip_frame_mode = 2 in Table 1), which directly use the interpolated image as the reconstructed image without sending any residuals. Frame interpolation methods can also include block-level modes, which use the interpolated image as additional reference frames and can further represent motion vectors and residuals as signals. When applying the frame-level mode of a frame interpolation method, deblocking filtering can be applied to the reconstructed image, for example, to mitigate block artifacts caused by block-based frame interpolation. To perform deblocking, some parameters related to quantization processing are required to control the intensity of deblocking.
[0076] Sub-block-based inter-frame prediction techniques such as TIP and Optical Flow Motion Vector Refinement (OPFL) can introduce block artifacts during prediction processing. These artifacts can be difficult to remove using deblocking filters. In some implementations, a Prediction Enhancement Filter (PEF) is employed during the prediction phase (e.g., to improve visual quality with minimal encoding and decoding impact).
[0077] In sub-block-based inter-frame prediction, such as for TIP and OPFL modes, the prediction unit (PU) can be broken down into smaller motion compensation units (MCUs). Each MCU can have its own motion vector pointing to the reference frame. Block artifacts may occur along the MCU boundaries when the MCU's motion information differs from its neighbors (e.g., when there is no residual due to a low bit rate budget). The deblocking filter 512 discussed above can address only the PU and TU boundaries. Therefore, when the MCU boundary is not aligned with the PU or TU boundary, the MCU boundary may not be processed by the deblocking filter.
[0078] In TIP mode, TIP reference frames can be generated in 8×8 MCU units. The TIP frames are then referenced by the current block via motion vectors. Since motion vectors can have arbitrary values, block artifacts resulting from the use of TIP mode in the final prediction and / or reconstruction may prevent the 8×8 grid of the PU from aligning with the 8×8 grid of the reconstructed frame. Furthermore, the TIP references may already contain block artifacts along the boundaries of each MCU.
[0079] In some implementations, a PEF is applied during the prediction phase to reduce block artifacts caused by TIP prediction processing and / or OPFL prediction processing, thereby improving visual quality. For example, since the location of block artifacts may not align with the 8×8 grid, two parameters can be derived based on the values of the motion vectors to identify the location of the block artifacts. The PEF can then be applied to the prediction samples at the internal MCU boundaries to reduce block artifacts. When the TIP reference frame is used as the direct output, a filter can be applied to the 8×8 grid of the TIP frame. In OPFL mode, the MCU size can be equal to 8×8 or 4×4, and the location of the block artifacts can be aligned with the 8×8 or 4×4 grid. Filtering can be applied to the prediction samples located on the internal MCU boundaries to reduce block artifacts.
[0080] PEF can include several steps. First, it is determined whether the MCU-level filter is on or off. For example, the motion vector difference between the two sides of the MCU boundary is examined. When the motion vector difference is less than a threshold, the boundary filtering can be skipped. For filtering on TIP, the TMVP motion vector can be used to examine the motion vector difference; for filtering on OPFL, the OPFL-refined motion vector can be used instead. Next, the decision to turn the filter on / off can be made at the sample level. For example, a mask can be derived based on samples near the boundary (e.g., this logic could be a simplified version of the logic used in the deblocking filter 512). Next, the increment value is derived (similar to the deblocking filter logic), and the offset is derived and applied to each sample in the samples to be filtered.
[0081] In some implementations, when sub-block motion compensation is used, a prediction filtering method (e.g., PEF) applies filtering to the prediction block. In some implementations, such as when multiple sub-blocks exist within a decoding block, the sub-block motion method performs motion compensation on a sub-block basis. In the example, optical flow-based prediction can be used to refine the motion compensation for each sub-block within a given decoding block using an optical flow function, and optical flow-based prediction is an example of the sub-block motion method described above. Furthermore, the prediction filtering method can also be applied to block units where frame interpolation is performed in frame-level mode of the frame interpolation method.
[0082] In some implementations, the TIP frame-level mode is modified using implicit quantization indexes. When an interpolated frame derived from the TIP mode is directly designated as the output of the decoding process for the current frame (e.g., in TIP frame-level mode), the luma quantization index and chroma quantization index of the current frame are not represented by signals, but are implicitly derived from the quantization index of the reference frame. In this way, the decoding bits consumed by representing the quantization index by signals are saved. In some implementations, sequence-level flags are used to switch between implicit and explicit frame-level signaling for the luma quantization parameter (QP) and chroma quantization parameter (QP) in the TIP frame-level mode.
[0083] As discussed above, in TIP mode, intermediate frames are generated by interpolation using the motion vector fields of forward and backward reference frames. In the first TIP mode (e.g., TIP_FRAME_AS_REF, sometimes also called TIP block-level mode), the interpolated frame is used as an additional reference frame for the current frame. In the second TIP mode (e.g., TIP_FRAME_AS_OUTPUT, sometimes also called TIP frame-level mode), the interpolated frame is directly output as a reconstruction of the current frame. As mentioned above, when using TIP mode, PEF can be applied to the interpolated frame to remove block artifacts. PEF requires luma quantization indices and chroma quantization indices for deblocking AC coefficients, which may need to be represented by signals from the encoder side to the decoder side. However, in some existing codec standards, only the quantization index associated with the luma AC coefficients is represented by signals, and the quantization index associated with the chroma AC coefficients is not represented by signals. When the bitstream is generated by an encoder using a non-CTC configuration, this can lead to an encoder-decoder mismatch because knowledge about the quantization index associated with the chroma AC coefficients is missing for the decoder, but is used by the encoder to perform PEF filtering.
[0084] The quantization index used in TIP frame-level mode can be used to perform PEF filtering, but not for coefficient decoding, because there are no residuals represented by signals in TIP frame-level mode. Therefore, the signaling of the quantization index is relatively more expensive than in other frames when decoding the residuals. The method described below addresses this problem.
[0085] When the TIP frame-level mode is selected, the luminance and chrominance quantization indices of the AC coefficients can be derived from the reference frame, rather than being represented as signals from the encoder side to the decoder side. For example, the quantization index can be derived from the reference frame as the average of the quantization indices. The terms base_q_idx, DeltaQUAc, and DeltaQVAc can represent the quantization index for the luminance AC coefficients, and the incremental quantization index relative to base_q_idx for the Cb and Cr AC coefficients, respectively. In this way, the quantization index of the current interpolation frame can be represented as shown in Equation 1: DeltaQUAc cur =(DeltaQUAc ref1 +DeltaQUAc ref2 +1)>>1 DeltaQVAc cur =(DeltaQVAc ref1 +DeltaQVAc ref2 +1)>>1 Equation System 1—Quantization Index
[0086] In Equation 1, the subscript "cur" corresponds to the current interpolation frame, and the subscripts "ref1" and "ref2" correspond to the two reference frames. Furthermore, sequence-level flags can be used to switch between the implicit QP derivation scheme and explicit frame-level signaling for the Luminance QP and Chroma QP in TIP frame-level mode.
[0087] Figure 6A This is a flowchart illustrating a method 600 for decoding video according to some embodiments. Method 600 can be executed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuitry and memory storing instructions for execution by the control circuitry. In some embodiments, method 600 is executed by executing instructions stored in the computing system's memory (e.g., memory 314).
[0088] The system receives (602) a video bitstream comprising multiple coded images. The system obtains (604) a reconstructed image corresponding to one of the coded images. For example, the system derives the reconstructed image using a frame interpolation method. The system derives (606) a set of quantization parameters for the reconstructed image, which is derived from a reference quantization parameter set used for the coded image. The system performs (608) loop filtering on the reconstructed image using this set of quantization parameters. For example, the system performs one or more loop filtering techniques on the reconstructed image.
[0089] In some implementations, a first flag (e.g., a first indicator) is signaled to indicate whether a frame-level mode of the frame interpolation method has been applied, and a second flag (e.g., a second indicator) is signaled to indicate whether a predictive filtering method has been applied. In some implementations, a third flag (e.g., a third indicator) is conditionally signaled to indicate whether an explicit or implicit quantization parameter derivation method (e.g., whether one or more quantization parameters are signaled) has been applied above the frame-level mode of the frame interpolation method. In some implementations, the signaling depends on the values of the first and second flags. As used herein, a "flag" can refer to a syntax with binary values or a syntax with more than two value options.
[0090] In some implementations, the first flag and / or the second flag are represented by signals using high-level syntax. An example implementation is shown below in Example Syntax 1. Example Syntax 1—General Sequence Header OBU Syntax
[0091] In Example Syntax 1, `enable_tip` is a first flag indicating whether the frame interpolation method is applied, and `enable_tip` equal to 1 indicates the frame-level mode in which the frame interpolation method is applied. In Example Syntax 1, `enable_pef` is a second flag indicating whether the predictive filtering method is applied, and `enable_tip_explicit_qp` is a third flag indicating whether an explicit or implicit quantization parameter derivation method is applied above the frame-level mode of the frame interpolation method. The third flag (`enable_tip_explicit_qp`) is represented by a signal based on the conditions that `enable_tip` equals 1 and `enable_pef` is non-zero.
[0092] In some implementations, the third flag is indicated by a signal following the first and second flags. For example, the third flag is indicated by a signal immediately after the second flag. In another example, the third flag is indicated by a signal immediately after the first flag.
[0093] In some implementations, when the third flag is represented by a signal using the value derived from the explicit quantization parameter method applied over the frame-level mode of the frame interpolation method, the following flags may be represented partially or together by signals. A fourth flag is represented by a signal to specify a luminance-related quantization parameter. If the current frame has more than one color component, a fifth flag is represented by a signal to specify whether the other two chroma color components share the same incremental quantization parameter relative to the luminance quantization parameter. If the fifth syntax is represented by a value indicating that two chroma color components share the same incremental quantization parameter, a single incremental quantization parameter value is represented by a signal. Otherwise, two incremental quantization parameter values are represented by signals for the Cb and Cr color components, respectively. If the current frame has only a luminance component (e.g., is monochromatic), the incremental quantization parameter values for the Cb and Cr color components are not signaled and are set to a default value, such as 0. In some implementations, the above flags / syntax are represented by signals using a high-level syntax.
[0094] The example frame header is shown below in Example Syntax 2. Example Syntax 2—Uncompressed Header Syntax
[0095] In example syntax 2, base_q_idx is the fourth flag, diff_uv_delta is the fifth flag, and DeltaQUAc and DeltaQVAc are incremental quantization parameters for Cb and Cr relative to the luminance quantization parameters.
[0096] In some implementations, quantization-related parameters are derived and used in the frame-level mode of the frame interpolation method. In some implementations, the quantization-related parameters for each color component (including luma and chroma) are represented as signals, and these quantization-related parameters are used to perform loop filtering on the reconstructed image in the frame-level mode of the frame interpolation method. In some implementations, quantization parameters related to the chroma color component are represented as signals on top of the frame-level mode of the frame interpolation method. For example, quantization parameters related to the quantization of chroma AC coefficients (e.g., u_ac_delta_q and v_ac_delta_q) can be represented as signals in the video bitstream. As another example, quantization parameters related to the quantization of chroma DC coefficients can be represented as signals. In some implementations, the frame-level mode is such as a full-frame skip mode, direct output prediction of global motion, and / or temporal interpolation prediction.
[0097] In some implementations, only a selected subgroup of quantization parameters for the image is represented by a signal, and a portion of the quantization parameters is implicitly derived but not represented by a signal. In some implementations, for the image, only the quantization parameters related to the quantization applied to the luminance color component are represented by a signal, and the quantization parameters related to the chrominance color component are implicitly derived. As an example, the chrominance quantization parameter is restricted to be the same as the luminance parameter, such that the chrominance quantization-related parameters are not represented by a signal but are implicitly derived. As another example, the chrominance increment quantization parameter (used to indicate the difference between chrominance and luminance or between the chrominance AC coefficient and the chrominance DC coefficient) is restricted to zero, such that no signaling for the chrominance increment quantization parameter is required.
[0098] In some implementations, quantization parameters associated with one or more reference images (used to derive the interpolated frame of the current frame using a frame interpolation method) are used to derive quantization parameters for performing loop filtering on the interpolated frame of the current frame using a frame interpolation method.
[0099] In some implementations, a weighted sum of quantization parameters associated with a reference image (used to derive an interpolated frame for the current frame using a frame interpolation method) is calculated to derive quantization parameters for performing loop filtering on the interpolated frame for the current frame using the frame interpolation method.
[0100] In some implementations, the difference between the quantization parameters for the chroma AC coefficients and other quantization parameters represented by signals (e.g., u_ac_delta_q and v_ac_delta_q) is derived using the following set of equations 1 through the corresponding syntax used in the reference frames (e.g., u_ac_delta_qref1 and u_ac_delta_qref2, v_ac_delta_qref1 and v_ac_delta_qref2, where ref1 and ref2 refer to the first and second reference images). u_ac_delta_q=(w0*u_ac_delta_q ref1 +w1*u_ac_delta_q ref2 +r0) / (w0+w1) v_ac_delta_q=(v0*v_ac_delta_q ref1 +v1*v_ac_delta_q ref2 +r1) / (v0+v1) Equation set 2—Derivation of chromaticity AC increment
[0101] Example values for weighting factors w0, w1, v0, and v1 include, but are not limited to: w0 = w1 = v0 = v1 = 1, where r0 = (w0 + w1) / 2 and r1 = (v0 + v1) / 2.
[0102] In some implementations, the quantization parameters for the chroma color components (e.g., base_qindex+u_ac_delta_q and base_qindex+v_ac_delta_q) are derived using the following set of equations 2 from the quantization parameters used in the reference frame (base_qindexref1+u_ac_delta_qref1, base_qindexref2+u_ac_delta_qref2, base_qindexref1+v_ac_delta_qref1 and base_qindexref2+v_ac_delta_qref2). base_qindex+u_ac_delta_q =(w0*(base_qindex) ref1 +u_ac_delta_q ref1 )+w1*(base_qindex ref2 +u_ac_delta_q ref2 )+r0) / (w0+w1) base_qindex+v_ac_delta_q =(v0*(base_qindex) ref1 +v_ac_delta_q ref1 )+v1*(base_qindex ref2 +v_ac_delta_q ref2 )+r1) / (v0+v1) Equation set 3—Derivation of chromaticity AC coefficients
[0103] In Equation 2, ref1 and ref2 refer to the first reference image and the second reference image, respectively. In some implementations, the weighting factor is implicitly derived using decoding information, such as the time distance from the reference frame to the current frame, the time layer associated with the reference frame, the decoding mode used in the reference frame, whether the reference frame is an intra-frame-only frame, and / or whether global motion is used in the reference frame.
[0104] In some implementations, all quantization parameters for performing loop filtering on frame reconstruction in a frame-level mode using the frame interpolation method are implicitly derived. In some implementations, quantization parameters associated with a reference image (used to derive the interpolated frame of the current frame using the frame interpolation method) are used to derive quantization parameters for performing loop filtering on the interpolated frame of the current frame using the frame interpolation method.
[0105] In some implementations, a weighted sum of quantization parameters associated with a reference image (used to derive an interpolated frame for the current frame using a frame interpolation method) is calculated to derive quantization parameters for performing loop filtering on the interpolated frame for the current frame using the frame interpolation method.
[0106] In some implementations, quantization parameters for luminance and / or chrominance associated with a reference frame are used to calculate quantization parameters for luminance and / or chrominance. For example, the luminance AC coefficient QP (quantization parameter) is represented as the base QP, and the difference between luminance ACQP and chrominance AC QP is represented as u_ac_delta_q and v_ac_delta_q for the Cb color component and Cr color component, respectively. base_qindex = (k0 * base_qindex) ref1 +k1*base_qindex ref2 +r0) / (k0+k1) u_ac_delta_q=(w0*u_ac_delta_q ref1 +w1*u_ac_delta_q ref2 +r1) / (w0+w1) v_ac_delta_q=(v0*v_ac_delta_q ref1 +v1*v_ac_delta_q ref2 +r2) / (v0+v1) Equation set 4—Luminance basis AC derivation and chromaticity increment AC derivation
[0107] In equation set 3, example values of weighting factors w0, w1, v0 and v1 include, but are not limited to: w0 = w1 = v0 = v1 = 1, where r0 = (w0 + w1) / 2 and r1 = (v0 + v1) / 2.
[0108] In some implementations, the quantization parameters for the chroma color components (e.g., base_qindex+u_ac_delta_q and base_qindex+v_ac_derta_q) are derived using the following set of equations 4 from the quantization parameters used in the reference frame (base_qindexref1+u_ac_delta_qref1, base_qindexref2+u_ac_delta_qref2, base_qindexref1+v_ac_delta_qref1 and base_qindexref2+v_ac_delta_qref2). base_qindex = (k0 * base_qindex) ref1+k1*base_qindex ref2 +r0) / (k0+k1) base_qindex+u_ac_delta_q =(w0*(base_qindex) ref1 +u_ac_delta_q ref1 )+w1*(base_qindex ref2 +u_ac_delta_q ref2 )+r0) / (w0+w1) base_qindex+v_ac_delta_q =(v0*(base_qindex) ref1 +v_ac_delta_q ref1 )+v1*(base_qindex ref2 +v_ac_delta_q ref2 )+r1) / (v0+v1) Equation set 5—Derivation of luminance and chromaticity AC coefficients
[0109] In equation system 4, ref1 and ref2 refer to the first reference image and the second reference image.
[0110] In some implementations, the weighting factor is implicitly derived from decoding information, such as the time distance from the reference frame to the current frame, the time layer associated with the reference frame, the decoding mode used in the reference frame, whether the reference frame is an intra-frame-only frame, and / or whether global motion is used in the reference frame.
[0111] Figure 6B This is a flowchart illustrating a method 650 for encoding video according to some embodiments. Method 650 can be executed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuitry and memory storing instructions for execution by the control circuitry. In some embodiments, method 650 is executed by executing instructions stored in the computing system's memory (e.g., memory 314).
[0112] The system receives (652) video data comprising multiple images. The system encodes (654) the first image among the multiple images. The system determines (656) whether to signal one or more quantization parameters for the first image based on whether frame-level interpolation mode is used to encode the first image and whether a prediction enhancement filter (PEF) is enabled for the first image. When frame-level interpolation mode is used to encode the first image and PEF is enabled, the system signals (658) the base and incremental quantization parameter sets for the first encoded image. When frame-level interpolation mode is used to encode the first image and PEF is disabled, the system abandons (660) signaling the base and incremental quantization parameter sets for the first encoded image. The system transmits (662) the encoded first image via a video bitstream.
[0113] As mentioned earlier, the encoding process can reflect the decoding process described in this article. For the sake of brevity, these details will not be repeated here.
[0114] although Figure 6A and Figure 6B Some logical levels are shown in a specific order, but levels that are not dependent on the order can be reordered, and other levels can be combined or decomposed. Some reorderings or other groupings not specifically mentioned will be obvious to those skilled in the art, and therefore the orderings and groupings presented herein are not exhaustive. Furthermore, it should be recognized that the levels can be implemented in hardware, firmware, software, or any combination thereof.
[0115] Now let's turn to some example implementations.
[0116] (A1) In one aspect, some implementations include a method for video decoding (e.g., method 600). In some implementations, the method is performed at a computing system having memory and one or more processors. The method includes: (i) receiving a video bitstream comprising a plurality of coded images; (ii) obtaining a reconstructed image corresponding to one of the coded images among the plurality of coded images; (iii) deriving a set of quantization parameters for the reconstructed image, the set of quantization parameters being derived from a reference set of quantization parameters used for encoding the image; and (iv) performing loop filtering on the reconstructed image using the set of quantization parameters. For example, all quantization parameters used to perform loop filtering on frame reconstruction in a frame-level mode using a frame interpolation mode are implicitly derived. In some implementations, the set of quantization parameters is used to encode each color component (e.g., luma and chroma components) in the color components of the image. As another example, selected quantization parameters are represented only by signals, and a portion of the quantization parameters are implicitly derived but not represented by signals. In some implementations, the reconstructed image corresponds to a frame-level mode applied to the coded image.
[0117] (A2) In some implementations of A1, the set of quantization parameters used to reconstruct the image is derived when a frame-level interpolation mode is enabled for the coded image. In some implementations, the frame-level mode is a direct output prediction and / or temporal interpolation prediction, such as a full-frame skip mode, a global motion mode, or a global motion mode. In some implementations, the set of quantization parameters is derived based on the frame-level mode being applied.
[0118] (A3) In some embodiments of A2, the method further includes: (i) determining whether to derive or parse a set of quantization parameters from the video bitstream based on a first indicator represented by a signal in the video bitstream, wherein the set of quantization parameters for reconstructing the image is derived according to a first value of the first indicator represented by a signal; and (ii) parsening the set of quantization parameters from the video bitstream when the first indicator represented by a signal has a second value. In some embodiments, the set of quantization parameters is derived according to the first indicator represented by a signal having a first value, and the set of quantization parameters is parsed from the video bitstream according to the first indicator represented by a signal having a second value. For example, quantization-related parameters for each color component (including luminance and chrominance) are represented by signals, and these quantization-related parameters are used to perform loop filtering on the reconstructed image in the frame-level mode of the frame interpolation method. In some embodiments, quantization parameters related to the chrominance color component are represented by signals on the frame-level mode of the frame interpolation method. For example, signals can be used to represent quantization parameters related to the quantization of chromaticity AC coefficients (e.g., u_ac_delta_q and v_ac_delta_q as defined in AVM). As another example, signals can be used to represent quantization parameters related to the quantization of chromaticity DC coefficients.
[0119] (A4) In some implementations of A2 or A3, the frame-level frame interpolation mode is a temporal interpolation prediction (TIP) mode. In some implementations, the set of quantization parameters used to reconstruct the image is derived based on a determination that enables the TIP mode for the encoded image.
[0120] (A5) In some embodiments of any of A2 to A4, the method further includes: determining whether a frame-level interpolation mode is enabled for the encoded picture based on a second indicator represented by a signal in the video bitstream. For example, a flag (e.g., indicating enable_tip) equal to a first value (e.g., 1) indicates that the TIP mode is enabled for the encoded picture.
[0121] (A6) In some implementations of any of A1 to A5, the reference quantization parameter set corresponds to one or more reference images of the encoded image. For example, quantization parameters associated with the reference images (e.g., for deriving the interpolated frame of the current frame using a frame interpolation method) are used to derive quantization parameters for performing loop filtering on the interpolated frame of the current frame using the frame interpolation method.
[0122] (A7) In some implementations of any of A1 to A6, the quantization parameter set includes AC chromaticity coefficients. For example, the quantization parameter set includes a basic quantization index, incremental quantization values for the U channel, and incremental quantization values for the V channel.
[0123] (A8) In some embodiments of any of A1 to A7, the method further includes: deriving a second set of quantization parameters for reconstructing the image based on the decoded information. For example, only the quantization parameters related to the quantization applied to the luminance color component are represented by signals, while these parameters are implicitly derived for the quantization parameters related to the chrominance color component.
[0124] (A9) In some implementations of A8, the second set of quantization parameters is the same as the set of quantization parameters. For example, the chromaticity quantization parameters are restricted to be the same as the luminance quantization parameters, such that parameters related to chromaticity quantization are not represented by signals but are implicitly derived. As another example, the chromaticity increment quantization parameters (e.g., the difference between quantization parameters indicating chromaticity and luminance or between chromaticity AC coefficients and chromaticity DC coefficients) are restricted to zero, such that the chromaticity increment quantization parameters do not need to be represented by signals.
[0125] (A10) In some embodiments of A8 or A9, the quantization parameter set corresponds to a first color component, and the second quantization parameter set corresponds to a second color component. For example, the quantization parameter set corresponds to the luminance component, and the second quantization parameter set corresponds to the chrominance component.
[0126] (A11) In some implementations of any of A1 to A10, the quantization parameter set is derived from a weighted sum of a reference quantization parameter set. For example, a weighted sum of quantization parameters associated with a reference image is calculated to derive quantization parameters for performing loop filtering on an interpolated frame of the current frame using method A. As another example, quantization parameters for luminance and / or chrominance associated with a reference frame are used to calculate quantization parameters for luminance and / or chrominance. For example, the luminance AC coefficient QP (quantization parameter) is represented as base QP, and the difference between luminance AC QP and chrominance AC QP is represented as u_ac_delta_q and v_ac_delta_q for the Cb color component and Cr color component, respectively, as shown in Equation 4. As another example, quantization parameters for the chrominance color component (e.g., base_qindex+u_ac_delta_q and base_qindex+v_ac_delta_q) are obtained by using the quantization parameter (base_qindex) used in the reference frame. ref1 +u_ac_delta_q ref1 base_qindex ref2 +u_ac_delta_q ref2 base_qindex ref1+v_ac_delta_q ref1 and base_qindex ref2 +v_ac_delta_q ref2 The result is shown in equation system 5.
[0127] (A12) In some implementations of A11, the set of weighting factors used for the weighted sum is derived from decoding information. For example, the weighting factors are implicitly derived from decoding information such as the time distance from the reference frame to the current frame, the time layer associated with the reference frame, the decoding mode used in the reference frame, whether the reference frame is an intra-frame-only frame, and / or whether global motion is used in the reference frame.
[0128] (A13) In some implementations of any of A1 to A12, the quantization parameter set is derived from a reference quantization parameter set using one or more AC increment parameters. For example, the difference between the quantization parameters for the chroma AC coefficients and other quantization parameters represented by signals (e.g., u_ac_delta_q and v_ac_delta_q) is derived using the corresponding syntax used in the reference frame (e.g., u_ac_delta_qref1 and u_ac_delta_qref2, v_ac_delta_qref1 and v_ac_delta_qref2, where ref1 and ref2 refer to the first and second reference images), as shown in Equation 2.
[0129] (A14) In some implementations of any of A1 to A13, the quantization parameter set is derived from a reference quantization parameter set using one or more basic quantization parameters. For example, quantization parameters for the chroma color components (e.g., base_qindex+u_ac_delta_q and base_qindex+v_ac_delta_q) are derived from quantization parameters used in the reference frame (base_qindexref1+u_ac_delta_qref1, base_qindexref2+u_ac_delta_qref2, base_qindexref1+v_ac_delta_qref1, and base_qindexref2+v_ac_delta_qref2), as shown in Equation 3.
[0130] (B1) In another aspect, some implementations include a method for video encoding (e.g., method 650). In some implementations, the method is performed on a computing system having memory and one or more processors. The method includes: (i) receiving video data comprising a plurality of images; (ii) encoding a first image among the plurality of images; (iii) determining whether to signal one or more quantization parameters for the first image based on whether a frame-level interpolation mode is used to encode the first image and whether a prediction enhancement filter (PEF) is enabled for the first image; (iv) when the first image is encoded using a frame-level interpolation mode and PEF is enabled, signaling a set of base and incremental quantization parameters for the first encoded image; (v) when the first image is encoded using a frame-level interpolation mode and PEF is disabled, abandoning the signaling of the set of base and incremental quantization parameters for the first encoded image; and (vi) transmitting the encoded first image via a video bitstream.
[0131] (B2) In some implementations of B1, the frame-level frame interpolation mode is the temporal interpolation prediction (TIP) mode.
[0132] (B3) In some implementations of B1 or B2, the basic and incremental quantization parameter sets include AC coefficients.
[0133] (B4) In some implementations of any of B1 to B3, the base and incremental quantization parameter sets correspond to two or more color components.
[0134] (C1) In another aspect, some embodiments include a method for processing visual media data. In some embodiments, the method is performed on a computing system having memory and one or more processors. The method includes: (i) obtaining a source video sequence; and (ii) performing a conversion between the source video sequence and a bitstream of visual media data, wherein the bitstream includes: (a) a plurality of coded pictures, the plurality of coded pictures including a first coded picture encoded according to a first type of frame interpolation; (b) a first indicator indicating whether a temporal interpolation mode (TIP) is enabled for the first coded picture; (c) a second indicator indicating whether a prediction enhancement filter (PEF) is enabled for the first coded picture; and (d) a set of base and incremental quantization parameters for the first coded picture when the first indicator indicates that the TIP mode is enabled and the second indicator indicates that the PEF is enabled; and wherein, when the first indicator indicates that the TIP mode is enabled and the second indicator indicates that the PEF is disabled, the bitstream does not include the set of base and incremental quantization parameters for the first coded picture.
[0135] (C2) In some implementations of C1, the bitstream also includes a reference base and a set of incremental quantization parameters for one or more reference images for the first coded image.
[0136] In another aspect, some embodiments include a computing system (e.g., server system 112) comprising a control circuitry system (e.g., control circuitry system 302) and a memory (e.g., memory 314) coupled to the control circuitry system. The memory stores one or more sets of instructions configured to be executed by the control circuitry system, including instructions for performing any of the methods described herein (e.g., A1 to A14, B1 to B4, and C1 to C2 above). In yet another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more sets of instructions for execution by the control circuitry system of the computing system, including instructions for performing any of the methods described herein (e.g., A1 to A14, B1 to B4, and C1 to C2 above).
[0137] As used in this article, N refers to a variable number. Unless explicitly stated otherwise, different instances of N may refer to the same number (e.g., the same integer value, such as the number 2) or different numbers.
[0138] Unless otherwise stated, any syntax element described herein can be a High-Level Syntax (HLS). As used herein, HLS is represented by signals at a level higher than the block level. For example, HLS can correspond to the sequence level, frame level, slice level, or tile level. As another example, HLS elements can be represented by signals in a Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Adaptation Parameter Set (APS), slice header, picture header, tile header, and / or CTU header.
[0139] It should be understood that while the terms “first,” “second,” etc., may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the claims. As used in the description of embodiments and appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or,” as used herein, refers to and covers any and all possible combinations of one or more of the associated listed items. It will also be understood that, when used in this specification, the terms “comprising” and / or “including” specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0140] As used herein, the term "if" may be interpreted, depending on the context, as meaning "when the prerequisite is true," "after the prerequisite is true," "in response to determining that the prerequisite is true," "based on determining that the prerequisite is true," or "in response to detecting that the prerequisite is true." Similarly, the phrases "if it is determined [that the prerequisite is true]," "if [that the prerequisite is true]," or "when [that the prerequisite is true]" may be interpreted, depending on the context, as meaning "after determining that the prerequisite is true," "in response to determining that the prerequisite is true," "based on determining that the prerequisite is true," "after detecting that the prerequisite is true," or "in response to detecting that the prerequisite is true."
[0141] For illustrative purposes, the foregoing description has been described with reference to specific embodiments. However, the illustrative discussion above is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described to best illustrate the operating principles and practical applications, thereby enabling others skilled in the art to implement them.
Claims
1. A method for video decoding, the method being performed at a computing system having a memory and one or more processors, the method comprising: Receives a video bitstream that includes multiple encoded images; Obtain the reconstructed image corresponding to the coded image among the plurality of coded images; Export a set of quantization parameters for the reconstructed image, the set of quantization parameters being derived from a reference set of quantization parameters for the encoded image; as well as Loop filtering is performed on the reconstructed image using the quantization parameter set.
2. The method according to claim 1, wherein, The quantization parameter set for the reconstructed image is exported when frame-level frame interpolation mode is enabled for the encoded image.
3. The method according to claim 2, further comprising: Whether to derive or parse the quantization parameter set from the video bitstream is determined based on a first indicator represented by a signal in the video bitstream; Wherein, the quantization parameter set for the reconstructed image is derived with a first value based on the first indicator represented by the signal; and When the first indicator represented by a signal has a second value, the quantization parameter set is parsed from the video bitstream.
4. The method according to claim 2, wherein, The frame-level frame interpolation mode is the temporal interpolation prediction (TIP) mode.
5. The method according to claim 2, further comprising: The frame-level frame interpolation mode is enabled for the encoded image based on a second indicator represented by a signal in the video bitstream.
6. The method according to claim 1, wherein, The reference quantization parameter set corresponds to one or more reference images for the encoded image.
7. The method according to claim 1, wherein, The quantization parameter set includes AC chromaticity coefficients.
8. The method according to claim 1, further comprising: A second set of quantization parameters for the reconstructed image is derived based on the decoded information.
9. The method according to claim 8, wherein, The second set of quantization parameters is the same as the set of quantization parameters.
10. The method according to claim 8, wherein, The quantization parameter set corresponds to the first color component, and the second quantization parameter set corresponds to the second color component.
11. The method according to claim 1, wherein, The quantization parameter set is derived from a weighted sum of the reference quantization parameter set.
12. The method according to claim 11, wherein, The set of weighting factors for the weighted sum is derived from the decoding information.
13. The method according to claim 1, wherein, The quantization parameter set is derived from the reference quantization parameter set using one or more AC increment parameters.
14. The method according to claim 1, wherein, The quantization parameter set is derived from the reference quantization parameter set using one or more basic quantization parameters.
15. A computing system, comprising: Control circuit system; Memory; as well as A set of one or more sets of instructions, stored in the memory and configured to be executed by the control circuitry system, the set of one or more sets of instructions including instructions for: Receive video data including multiple images; Encode the first image among the plurality of images; Whether to use a signal to represent one or more quantization parameters for the first image is determined based on whether frame-level interpolation mode is used to encode the first image and whether a prediction enhancement filter (PEF) is enabled for the first image. When the first image is encoded using the frame-level interpolation mode and the PEF is enabled, the base and incremental quantization parameter sets for the first encoded image are represented by signals. When the first image is encoded using the frame-level interpolation mode and the PEF is disabled, the signal representation of the base and incremental quantization parameter sets for the first encoded image is abandoned; and The first image, encoded, is transmitted via video bitstream.
16. The system according to claim 15, wherein, The frame-level interpolation mode is the temporal interpolation prediction (TIP) mode.
17. The system according to claim 15, wherein, The set of basic and incremental quantization parameters includes AC coefficients.
18. The system according to claim 15, wherein, The basic and incremental quantization parameter sets correspond to two or more color components.
19. A non-transitory computer-readable storage medium storing one or more sets of instructions configured for execution by a computing device having a control circuitry system and a memory, the one or more sets of instructions including instructions for: Obtain the source video sequence; as well as Perform the conversion between the source video sequence and the bitstream of visual media data, wherein, The bit stream includes: Multiple encoded images, the multiple encoded images including a first encoded image encoded according to a first type of frame interpolation; A first indicator indicates whether time interpolation mode (TIP) is enabled for the first coded image; A second indicator indicates whether a prediction enhancement filter (PEF) is enabled for the first coded image; and When the first indicator indicates that the TIP mode is enabled and the second indicator indicates that the PEF is enabled, the set of base and incremental quantization parameters for the first encoded image; and Wherein, when the first indicator indicates that the TIP mode is enabled and the second indicator indicates that the PEF is disabled, the bitstream does not include the base and incremental quantization parameter set for the first encoded image.
20. The non-transitory computer-readable storage medium according to claim 19, wherein, The bitstream also includes a reference base and an incremental quantization parameter set for one or more reference images of the first coded image.