System and method for extended multi-residual block coding

By adopting the multi-residual block encoding mode in video encoding and decoding technology, and using independent parameter sets to generate and combine multiple residual blocks, the problems of poor encoding efficiency and video quality in the prior art are solved, and more efficient video compression and better reconstruction quality are achieved.

CN119948872APending Publication Date: 2025-05-06TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480004178.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-29
Filing Date
2024-05-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies are difficult to effectively utilize multiple residual block encoding modes when compressing video data, resulting in poor encoding efficiency and video quality.

Method used

Using a multi-residual block encoding mode, multiple residual blocks are obtained by using independent parameter sets and combined to generate the final residual block for reconstructing the video block. This method involves using different transform sizes, transform types and encoding modes to improve encoding efficiency.

Benefits of technology

Improve the quality and coding efficiency of video reconstruction, reduce encoding overhead, and achieve better compression ratio and more accurate video reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948872A_ABST
    Figure CN119948872A_ABST
Patent Text Reader

Abstract

Various implementations described herein include methods and systems for encoding and decoding video. In one aspect, a method includes receiving a video bitstream including a plurality of encoded blocks, the plurality of encoded blocks including a current block. The method includes obtaining a first residual block of the current block using a first parameter set. The method includes obtaining a second residual block of the current block using a second set of parameters, where the second set of parameters is independent of the first set of parameters. The method includes obtaining a combined residual block using the first residual block and the second residual block. The method further includes reconstructing the current block using the combined residual block.
Need to check novelty before this filing date? Find Prior Art

Description

Related Applications

[0001] This application claims priority to (i) U.S. Provisional Patent Application No. 63 / 524,549, entitled “Multiple Residual Block Coding Mode,” filed on June 30, 2023, and (ii) U.S. Provisional Patent Application No. 63 / 611,089, entitled “Extended Multi-Residue Block Coding,” filed on December 15, 2023, and is a continuation of and claims priority to U.S. Patent Application No. 18 / 677,729, entitled “Systems and Methods for Extended Multi-Residue Block Coding,” filed on May 29, 2024, each of which is incorporated herein by reference in its entirety. Technical Field

[0002] The disclosed embodiments relate generally to video coding and decoding, including but not limited to systems and methods for obtaining a combined residual block for encoding and decoding a coded block. Background Art

[0003] Digital video is supported by a variety of electronic devices, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. Electronic devices send and receive or otherwise communicate digital video data over a communication network, and / or store digital video data on a storage device. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video encoding may be used to compress video data according to one or more video encoding standards before transmitting or storing the video data. Video encoding may be performed by hardware and / or software on an electronic / client device or a server providing a cloud service.

[0004] Video coding typically utilizes prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.) that exploit the redundancy inherent in video data. Video coding aims to compress video data into a form that uses a lower bit rate while avoiding or minimizing the degradation of video quality. A variety of video codec standards have been developed. For example, High Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. ITU-T and ISO / IEC released the HEVC / H.265 standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Versatile Video Coding (VVC / H.266) is a video compression standard intended to be a successor to HEVC. ITU-T and ISO / IEC released the VVC / H.266 standard in 2020 (version 1) and 2022 (version 2). Alliance for Open Media (AOMedia) Video 1 (AV1) is an open video coding format designed as an alternative to HEVC. On January 8, 2019, the validated version 1.0.0 of the specification (with Errata 1) was released. Summary of the invention

[0005] The present disclosure describes a set of methods for video (image) compression, including residual coding techniques for encoding video blocks. The residual coding techniques include multi-residual block coding modes. In some embodiments, in order to reconstruct the current block, multiple residual blocks (e.g., N partial residual blocks) are encoded and used to obtain a combined residual block. For example, the values ​​of the residual blocks can be added to generate a combined residual block, and then the combined residual block is added to the prediction block to obtain a reconstructed block. In some embodiments, the residual block has different parameters, such as transform size, transform type and / or coding mode (e.g., lossless coding mode or lossy coding mode). The advantage of using multiple residual coding techniques and combined residual blocks is that the quality of the corresponding reconstructed (decoded) video can be improved. In addition, the use of multiple residual coding techniques and combined residual blocks can reduce coding overhead (e.g., more efficiently encoding and / or decoding the residual of the video).

[0006] According to some embodiments, a video decoding method includes: (i) receiving a video code stream including multiple coding blocks, the multiple coding blocks including a current block; (ii) using a first parameter set to obtain a first residual block (sometimes also referred to as a residual block) of the current block; (iii) using a second parameter set to obtain a second residual block of the current block, wherein the second parameter set is independent of the first parameter set; (iv) using the first residual block and the second residual block to obtain a combined residual block; and (v) reconstructing the current block using the combined residual block.

[0007] According to some embodiments, a video encoding method includes: (i) receiving video data including multiple blocks, the multiple blocks including a current block; (ii) obtaining a first residual block of the current block using a first parameter set; (iii) obtaining a second residual block of the current block using a second parameter set, wherein the second parameter set is independent of the first parameter set; and (iv) representing the first residual block and the second residual block with a signal in a video bitstream.

[0008] According to some embodiments, a bitstream conversion method includes: (i) obtaining a source video sequence including multiple frames; and (ii) performing conversion between the source video sequence and a video bitstream of visual media data, wherein the video bitstream includes: (a) multiple encoded blocks corresponding to the multiple frames, the multiple encoded blocks including a current block; (b) a first residual block for the current block, wherein the first residual block is generated using a first parameter set; and (c) a second residual block for the current block, wherein the second residual block is generated using a second parameter set.

[0009] According to some embodiments, a computing system is provided, such as a streaming system, a server system, a personal computer system, or other electronic device. The computing system includes a control circuit and a memory storing one or more instruction sets. The one or more instruction sets include instructions for performing any method described herein. In some embodiments, the computing system includes an encoder component and a decoder component (e.g., a transcoder).

[0010] According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more instruction sets executed by a computing system. The one or more instruction sets include instructions for performing any of the methods described herein.

[0011] Thus, devices and systems using methods for encoding and decoding video are disclosed. Such methods, devices and systems may supplement or replace conventional methods, devices and systems for video encoding / decoding. The features and advantages described in the specification are not necessarily all-inclusive, and in particular, some additional features and advantages will be apparent to one of ordinary skill in the art in view of the drawings, description and claims provided by the present disclosure. Furthermore, it should be noted that the language used in the specification is primarily selected for readability and instructional purposes, and is not necessarily selected to describe or limit the subject matter described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to be able to understand the present disclosure in more detail, a more specific description may be made with reference to the features of various embodiments, some of which are illustrated in the accompanying drawings. However, the accompanying drawings only illustrate the relevant features of the present disclosure and are therefore not necessarily to be considered limiting, as those skilled in the art will understand upon reading the present disclosure that the description may contain other effective features.

[0013] Figure 1 is a block diagram illustrating an exemplary communication system in accordance with some embodiments.

[0014] Figure 2A is a block diagram illustrating exemplary elements of an encoder assembly according to some embodiments.

[0015] Figure 2B is a block diagram illustrating exemplary elements of a decoder component according to some embodiments.

[0016] Figure 3 is a block diagram illustrating an exemplary server system in accordance with some embodiments.

[0017] 4A to 4D An exemplary coding tree structure according to some embodiments is shown.

[0018] FIG. 5A to FIG. 5E Exemplary prediction blocks, residual blocks, and reconstructed blocks according to some embodiments are shown.

[0019] Fig. 6A An exemplary video decoding process according to some embodiments is shown.

[0020] Figure 6B An exemplary video encoding process according to some embodiments is shown.

[0021] According to common practice, the various features shown in the drawings are not necessarily drawn to scale, and like reference numerals may be used to denote like features throughout the specification and drawings. DETAILED DESCRIPTION

[0022] The present disclosure describes a video / image compression technique including multi-residual block encoding. The disclosed technique includes: obtaining a first residual block of a current block using a first parameter set; obtaining a second residual block of the current block using an independent second parameter set; and obtaining a combined residual block using the first residual block and the second residual block. The combined residual block can then be used to reconstruct the corresponding coded block. Using independent parameter sets for the first residual block and the second residual block can reduce coding overhead (e.g., by limiting the options for the second residual block). For example, the parameter set may include a corresponding transform core index and a transform index. In this example, the second parameter set may indicate a transform core index for the following transform set, which has fewer options and therefore requires fewer bits for the corresponding transform index. Similarly, limiting the options for the second residual block by independent encoding can be applied to secondary transforms and / or quantization indices. In this way, coding overhead can be reduced for the second residual block, while the multi-residual block technique maintains or improves the quality (e.g., accuracy and precision) of the reconstructed video. As described below, some multiple residual block encoding techniques use a lossy transform technique for a first residual block and a lossless transform technique for a second residual block, which allows for better compression ratios than lossless-only methods and more accurate / precise reconstruction than lossless-only methods. Exemplary Systems and Apparatus

[0023] Figure 1 1 is a block diagram illustrating a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic device 120-1 to electronic device 120-m), which are communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, for example, which is used with video-enabled applications, such as video conferencing applications, digital television applications, media storage and / or distribution applications.

[0024] Source device 102 includes a video source 104 (e.g., a camera component or a media storage) and an encoder component 106. In some embodiments, video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). Encoder component 106 generates one or more encoded video streams from the video stream. The video stream from video source 104 can be high data volume compared to the encoded video stream 108 generated by encoder component 106. Because the encoded video stream 108 has a lower data volume (less data) than the video stream from the video source, the encoded video stream 108 requires less bandwidth for transmission and less storage space for storage than the video stream from video source 104. In some embodiments, source device 102 does not include encoder component 106 (e.g., configured to transmit uncompressed video to one or more networks 110).

[0025] One or more networks 110 represent any number of networks that transmit information between source device 102, server system 112, and / or electronic device 120, including, for example, wired and / or wireless communication networks. One or more networks 110 can exchange data in circuit switching channels and / or packet switching channels. Representative networks include telecommunication networks, local area networks, wide area networks, and / or the Internet.

[0026] One or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is a streaming server or includes a streaming server (e.g., configured to store and / or distribute video content, such as an encoded video stream from the source device 102). The server system 112 includes a codec component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the codec component 114 includes an encoder component and / or a decoder component. In various embodiments, the codec component 114 is instantiated as hardware, software, or a combination of hardware and software. In some embodiments, the codec component 114 is configured to decode the encoded video stream 108 and re-encode the video data using different encoding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings from the encoded video stream 108. In some embodiments, the server system 112 acts as a media-aware network element (MANE). For example, the server system 112 may be configured to prune the encoded video stream 108 to tailor a potentially different stream for one or more electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.

[0027] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to generate an output video stream that can be presented on a display or other type of presentation device. In some embodiments, one or more electronic devices 120 do not include a display component (e.g., communicatively coupled to an external display device and / or include a media storage device). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.

[0028] The source device and / or multiple electronic devices 120 are sometimes referred to as “end devices” or “user devices.” In some embodiments, the source device 102 and / or one or more electronic devices 120 are examples of server systems, personal computers, portable devices (e.g., smartphones, tablets, or laptops), wearable devices, video conferencing devices, and / or other types of electronic devices.

[0029] In an exemplary operation of the communication system 100, the source device 102 transmits an encoded video stream 108 to a server system 112. For example, the source device 102 may encode a picture stream captured by the source device. The server system 112 receives the encoded video stream 108 and may decode and / or encode the encoded video stream 108 using a codec component 114. For example, the server system 112 may apply encoding to the video data, which is better for network transmission and / or storage. The server system 112 may transmit the encoded video data 116 (e.g., one or more encoded video streams) to one or more electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display the video picture.

[0030] Figure 2A 1 is a block diagram illustrating exemplary elements of an encoder component 106 according to some embodiments. The encoder component 106 receives video data (e.g., a source video sequence) from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., a transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives the video sequence from a remote video source (e.g., a video source that is a component of a device different from the encoder component 106). The video source 104 may provide a source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8 bits, 10 bits, or 12 bits), any color space (e.g., BT.601Y CrCB, or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0, or Y CrCb 4:4:4). In some embodiments, the video source 104 is a storage device that stores previously captured / prepared videos. In some embodiments, the video source 104 is a camera that captures local image information as a video sequence. The video data may be provided as a plurality of separate pictures that are given motion when viewed sequentially. The picture itself may be organized into a spatial array of pixels, where each pixel may include one or more samples depending on the sampling structure used, the color space, etc. The relationship between pixels and samples may be readily understood by one of ordinary skill in the art.

[0031] The encoder component 106 is configured to encode and / or compress the pictures of the source video sequence into the encoded video sequence 216 in real time or under other time constraints required by the application. In some embodiments, the encoder component 106 is configured to perform conversion between the source video sequence and a code stream of visual media data (e.g., a video code stream). Implementing the appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls other functional units as described below and is functionally coupled to the other functional units. The parameters set by the controller 204 may include rate control related parameters (e.g., lambda values ​​for picture skipping, quantizer, and / or rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Other functions of the controller 204 can be easily identified by those of ordinary skill in the art, as these functions may involve the encoder component 106 optimized for a certain system design.

[0032] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. In a simplified example, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and a reference picture) and a (local) decoder 210. The decoder 210 reconstructs the symbols to create sample data in a manner similar to that of the (remote) decoder (when the compression between the symbols and the encoded video code stream is lossless). The reconstructed sample stream (sample data) is input to the reference picture memory 208. Since the decoding of the symbol stream produces bit-accurate results that are independent of the decoder location (local or remote), the contents of the reference picture memory 208 are also bit-accurately corresponding between the local encoder and the remote encoder. In this way, the prediction part of the encoder interprets the reference picture samples as the same sample values ​​as the decoder will interpret when using prediction during decoding.

[0033] The operation of decoder 210 may be combined with, for example, Figure 2B The remote decoder of the decoder assembly 122 described in detail is the same. However, briefly refer to Figure 2B , when symbols are available and the entropy encoder 214 and the parser 254 are capable of losslessly encoding / decoding the symbols into an encoded video sequence, the entropy decoding portion of the decoder component 122 including the buffer memory 252 and the parser 254 may not be fully implemented in the local decoder 210.

[0034] The decoder techniques described herein (except for parsing / entropy decoding) can be present in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter focuses on the decoder operation. In addition, the description of the encoder techniques can be simplified because the encoder techniques can be mutually inverse with the decoder techniques.

[0035] As part of its operation, source encoder 202 may perform motion compensated predictive coding, which predictively encodes an input frame by referencing one or more previously encoded frames from a video sequence designated as reference frames. In this manner, encoding engine 212 encodes the differences between pixel blocks of an input frame and pixel blocks of a reference frame that may be selected as a prediction reference for the input frame. Controller 204 may manage the encoding operations of source encoder 202, including, for example, setting parameters and subgroup parameters for encoding video data.

[0036] The decoder 210 decodes the encoded video data of the frame that can be designated as the reference frame based on the symbols created by the source encoder 202. The operation of the encoding engine 212 can advantageously be a lossy process. When the encoded video data is in the video decoder ( Figure 2A When the video sequence is decoded at a remote video decoder (not shown), the reconstructed video sequence may be a replica of the source video sequence with some errors. The decoder 210 replicates the decoding process that may be performed by the remote video decoder on the reference frame and may cause the reconstructed reference frame to be stored in the reference picture memory 208. In this way, the encoder component 106 locally stores a copy of the reconstructed reference frame that has common content (absent transmission errors) with the reconstructed reference frame that will be obtained by the remote video decoder.

[0037] The predictor 206 may perform a prediction search for the encoding engine 212. That is, for a new frame to be encoded, the predictor 206 may search the reference picture memory 208 for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may be used as appropriate prediction references for the new picture. The predictor 206 may operate on a pixel block by pixel block basis to find a suitable prediction reference. As determined by the search results obtained by the predictor 206, the input picture may have prediction references taken from a plurality of reference pictures stored in the reference picture memory 208.

[0038] The outputs of all the above functional units may be entropy encoded in the entropy encoder 214. The entropy encoder 214 performs lossless compression on the symbols generated by the various functional units according to techniques known to those skilled in the art (e.g., Huffman coding, variable length coding, and / or arithmetic coding) to convert the symbols into a coded video sequence.

[0039] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter may be configured to buffer the encoded video sequence created by the entropy encoder 214, thereby preparing for transmission through a communication channel 218, which may be a hardware / software link to a storage device that can store the encoded video data. The transmitter may be configured to merge the encoded video data from the source encoder 202 with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown). In some embodiments, the transmitter may transmit additional data when transmitting the encoded video. The source encoder 202 may include such data as part of the encoded video sequence. The additional data may include time / space / signal noise ratio (SNR) enhancement layers, other forms of redundant data such as redundant pictures and slices, supplementary enhancement information (SEI) messages, video usability information (VUI) parameter set fragments, etc.

[0040] The controller 204 may manage the operation of the encoder component 106. During encoding, the controller 204 may assign a certain encoded picture type to each encoded picture, but this may affect the encoding technique applied to the corresponding picture. For example, a picture may be assigned as an intra picture (I picture), a predictive picture (P picture), or a bidirectional predictive picture (B picture). An intra picture may be encoded and decoded without using any other frame in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh (IDR) pictures. Those of ordinary skill in the art are aware of these variants of I pictures and their corresponding applications and features, so they are not discussed here. Predictive pictures may be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and reference index to predict the sample values ​​of each block. Bidirectional predictive pictures may be encoded and decoded using intra prediction or inter prediction, which uses at most two motion vectors and reference indexes to predict the sample values ​​of each block. Similarly, multiple predictive pictures may use more than two reference pictures and associated metadata for reconstructing a single block.

[0041] The source picture may typically be spatially subdivided into blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples) and coded block by block. These blocks may be predictively coded with reference to other (already coded) blocks, which are determined by the coding allocation applied to the corresponding picture of the block. For example, blocks of an I picture may be non-predictively coded, or blocks of an I picture may be predictively coded (spatial prediction or intra-frame prediction) with reference to already coded blocks of the same picture. Blocks of pixels of a P picture may be non-predictively coded by spatial prediction with reference to one previously coded reference picture or by temporal prediction. Blocks of a B picture may be non-predictively coded by spatial prediction with reference to one or two previously coded reference pictures or by temporal prediction.

[0042] The captured video may be taken as a plurality of source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often simplified to intra-prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In one example, a particular picture being encoded / decoded is divided into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that was previously encoded in the video and is still buffered, the block in the current picture may be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case where multiple reference pictures are used, the motion vector may have a third dimension that identifies the reference picture.

[0043] The encoder component 106 may perform encoding operations according to any predetermined video encoding technique or standard such as described herein. In operation, the encoder component 106 may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.

[0044] Figure 2B is a block diagram illustrating exemplary elements of decoder component 122 according to some embodiments. Figure 2B The decoder component 122 in is coupled to the channel 218 and the display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to the loop filter 256 and configured to transmit data to the display 124 (eg, via a wired or wireless connection).

[0045] In some embodiments, the decoder component 122 includes a receiver coupled to the channel 218 and configured to receive data from the channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more encoded video sequences to be decoded by the decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. Each encoded video sequence may be received from the channel 218, which may be a hardware / software link leading to a storage device storing the encoded video data. The receiver may receive encoded video data and other data that may be forwarded to their respective use entities (not depicted), such as encoded audio data and / or auxiliary data streams. The receiver may separate the encoded video sequence from other data. In some embodiments, the receiver receives additional (redundant) data when receiving the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the decoder component 122 to decode the data and / or more accurately reconstruct the original video data. The additional data may take the form of, for example, time, space or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0046] According to some embodiments, the decoder component 122 includes a buffer memory 252, a parser 254 (sometimes also referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra picture prediction unit 262, a motion compensated prediction unit 260, an aggregator 268, a loop filter unit 256, a reference picture memory 266, and a current picture memory 264. In some embodiments, the decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. The decoder component 122 is at least partially implemented in software.

[0047] The buffer memory 252 is coupled between the channel 218 and the parser 254 (e.g., to prevent network jitter). In some embodiments, the buffer memory 252 is separate from the decoder component 122. In some embodiments, a separate buffer memory is provided between the output of the channel 218 and the decoder component 122. In some embodiments, in addition to the buffer memory 252 located inside the decoder component 122 (e.g., configured to handle playout timing), a separate buffer memory is provided outside the decoder component 122 (e.g., to prevent network jitter). When receiving data from a store / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory 252 may not be required or may be made smaller. For use on a best-effort network such as the Internet, the buffer memory 252 may be required, the buffer memory 252 may be relatively large, and / or may have an adaptive size, and may be implemented at least in part in an operating system or similar element external to the decoder component 122.

[0048] The parser 254 is configured to reconstruct symbols 270 from the encoded video sequence. These symbols may include, for example, information for managing the operation of the decoder component 122, and / or information for controlling a rendering device such as the display 124. The control information for the rendering device may be in the form of, for example, a supplementary enhancement information (SEI) message or a video usability information (VUI) parameter set fragment (not depicted). The parser 254 parses (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence may be performed according to a video coding technique or standard, and may follow various principles known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 254 may extract a subgroup parameter set for at least one of the subgroups of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the group. The subgroup may include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser 254 may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0049] Depending on the type of the coded video picture or a portion of the coded video picture (e.g., inter-frame and intra-frame pictures, inter-frame blocks and intra-frame blocks) and other factors, the reconstruction of the symbol 270 may involve multiple different units. Which units are involved and how they are involved can be controlled by the parser 254 through subgroup control information parsed from the coded video sequence. For clarity, such subgroup control information flow between the parser 254 and the multiple units below is not depicted.

[0050] The decoder component 122 may be conceptually subdivided into multiple functional units, which in some implementations closely interact with each other and may be at least partially integrated with each other. However, for the sake of clarity, the conceptual subdivision of multiple functional units is maintained herein.

[0051] The sealer / inverse transform unit 258 receives the quantized transform coefficients as symbols 270 from the parser 254, as well as control information (e.g., which transform to use, block size, quantization factor, and / or quantization scaling matrix). The sealer / inverse transform unit 258 may output a block including sample values, which may be input into the aggregator 268. In some cases, the output samples of the sealer / inverse transform unit 258 belong to intra-coded blocks, i.e., blocks that do not use predictive information from previously reconstructed pictures, but may use predictive information from previously reconstructed portions of the current picture. Such predictive information may be provided by the intra-picture prediction unit 262. The intra-picture prediction unit 262 may use surrounding reconstructed information extracted from the current (partially reconstructed) picture from the current picture memory 264 to generate a block of the same size and shape as the block being reconstructed. The aggregator 268 may add the predictive information generated by the intra-picture prediction unit 262 to the output sample information provided by the sealer / inverse transform unit 258 on a per-sample basis.

[0052] In other cases, the output samples of the scaler / inverse transform unit 258 belong to a block that is inter-coded and potentially motion compensated. In this case, the motion compensated prediction unit 260 may access the reference picture memory 266 to extract samples for prediction. After the extracted samples are motion compensated according to the symbols 270 belonging to the block, these samples may be added to the output of the scaler / inverse transform unit 258 (in this case, referred to as residual samples or residual signals) by the aggregator 268, thereby generating output sample information. The address at which the motion compensated prediction unit 260 extracts the prediction samples from the reference picture memory 266 may be controlled by a motion vector. The motion vector may be provided to the motion compensated prediction unit 260 for use in the form of a symbol 270, which may have, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include, for example, interpolation of sample values ​​extracted from the reference picture memory 266 when using sub-sample accurate motion vectors, motion vector prediction mechanisms, and the like.

[0053] The output samples of the aggregator 268 may be employed by various loop filtering techniques in the loop filter unit 256. The video compression techniques may include in-loop filter techniques controlled by parameters included in the encoded video bitstream and available to the loop filter unit 256 as symbols 270 from the parser 254, and may also be responsive to meta-information obtained during decoding of a previous (in decoding order) portion of an encoded picture or encoded video sequence, and to previously reconstructed and loop filtered sample values. The output of the loop filter unit 256 may be a sample stream that may be output to a presentation device (e.g., the display 124) and stored in the reference picture memory 266 for future inter-picture prediction.

[0054] Once reconstructed, certain coded pictures may be used as reference pictures for future prediction. Once a coded picture is reconstructed, and the coded picture is identified as a reference picture (e.g., by parser 254), the current reference picture may become part of reference picture memory 266, and a new current picture memory may be reallocated before starting reconstruction of a subsequent coded picture.

[0055] The decoder component 122 may perform decoding operations according to a predetermined video compression technique that may be recorded in a standard (e.g., any standard described herein). In the sense that the encoded video sequence follows the syntax of the video compression technique or standard, the encoded video sequence may conform to the syntax specified by the video compression technique or standard used (as specified in the video compression technology document or standard, particularly in the profile of the video compression technology or standard). In addition, in order to comply with some video compression techniques or standards, the complexity of the encoded video sequence may be within the range defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, mega samples per second), maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further defined by the hypothetical reference decoder (HRD) specification and metadata of the HRD buffer management signaled in the encoded video sequence.

[0056] Figure 3 is a block diagram illustrating a server system 112 according to some embodiments. The server system 112 includes a control circuit 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuit 302 includes one or more processors (e.g., a CPU (central processing unit), a GPU (graphics processing unit), and / or a DPU (data processing units)). In some embodiments, the control circuit includes one or more field-programmable gate arrays (FPGAs), hardware accelerators, and / or one or more integrated circuits (e.g., application-specific integrated circuits).

[0057] The network interface 304 may be configured to connect to one or more communication networks (e.g., wireless networks, wired networks, and / or optical networks). The communication network may be a local network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of communication networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Such communications may be one-way receiving only (e.g., broadcast television), one-way sending only (e.g., CANBus connected to certain CANBus devices), or bidirectional (e.g., using a local area network or wide area network digital network to connect to other computer systems). Such communications may include communications to one or more cloud computing networks.

[0058] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input devices 310 may include one or more of a keyboard, a mouse, a trackpad, a touch screen, a data glove, a joystick, a microphone, a scanner, a camera, etc. The output devices 308 may include one or more of an audio output device (e.g., a speaker), a visual output device (e.g., a display or monitor), etc.

[0059] The memory 314 may include a high-speed random access memory (e.g., DRAM (dynamic random access memory), SRAM (static random access memory), DDR RAM (double data rate random access memory), and / or other solid-state random access memory devices) and / or a non-volatile memory (e.g., one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). Optionally, the memory 314 includes one or more storage devices arranged away from the control circuit 302. The memory 314, or alternatively, the non-volatile solid-state memory device within the memory 314, includes a non-transitory computer-readable storage medium. In some embodiments, the memory 314 or the non-transitory computer-readable storage medium of the memory 314 stores the following programs, modules, instructions, and data structures, or a subset or superset thereof: An operating system 316, which includes programs for handling various basic system services and for performing hardware-related tasks; A network communications module 318 for connecting the server system 112 to other computing devices via one or more network interfaces 304 (eg, via wired and / or wireless connections); ● Codec module 320, which is used to perform and encode data (eg, video data) and / or In some embodiments, the codec module 320 is an instance of the codec component 114. The codec module 320 includes, but is not limited to, one or more of the following: ○ A decoding module 322 for performing various functions associated with decoding encoded data, such as those previously described with respect to the decoder component 122; and ○ an encoding module 340 for performing various functions associated with encoding data, such as those previously described with respect to the encoder component 106; and Picture memory 352 for storing pictures and picture data, e.g., for use with encoding module 320. In some embodiments, picture memory 352 includes one or more of reference picture memory 208, buffer memory 252, current picture memory 264, and reference picture memory 266.

[0060] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform the various functions previously described for the parser 254), a transform module 326 (e.g., configured to perform the various functions previously described for the scaler / inverse transform unit 258), a prediction module 328 (e.g., configured to perform the various functions previously described for the motion compensated prediction unit 260 and / or the intra-picture prediction unit 262), and a filter module 330 (e.g., configured to perform the various functions previously described for the loop filter 256).

[0061] In some embodiments, the encoding module 340 includes an encoding module 342 (e.g., configured to perform the various functions previously described for the source encoder 202 and / or the encoding engine 212) and a prediction module 344 (e.g., configured to perform the various functions previously described for the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 includes Figure 3 For example, the decoding module 322 and the encoding module 340 both use a shared prediction module.

[0062] Each module in the modules identified above stored in the memory 314 corresponds to an instruction set for performing the functions described herein. The modules identified above (e.g., instruction sets) do not need to be implemented as separate software programs, processes, or modules, so each subset of these modules can be combined or otherwise rearranged in various embodiments. For example, optionally, the encoding module 320 does not include a separate decoding module and an encoding module, but uses the same set of modules to perform these two groups of functions. In some embodiments, the memory 314 stores a subset of the modules and data structures identified above. In some embodiments, the memory 314 stores additional modules and data structures not described above.

[0063] Although Figure 3 A server system 112 is shown according to some embodiments, but Figure 3 The present invention is intended more as a functional description of various features that may exist in one or more server systems rather than as a schematic diagram of the structure of the embodiments described herein. In practice, items shown separately may be combined, and some items may be separated. For example, Figure 3 Some of the items shown individually in the figure can be implemented on a single server, and a single item can be implemented by one or more servers. The actual number of servers used to implement the server system 112, and how the features are distributed among the servers, will vary from implementation to implementation, and optionally, will depend in part on the amount of data traffic that the server system handles during peak usage periods and during average usage periods. Exemplary Encoding Techniques

[0064] The encoding processes and techniques described below may be performed in the devices and systems described above (e.g., source device 102, server system 112, and / or electronic device 120). In the following, a block may refer to a coding tree block, a maximum coding block, a predetermined fixed block size, a coding block, a prediction block, a residual block, or a transform block. In the following, a partial residual block is a residual block of a coding block that is not directly combined with a corresponding prediction block (e.g., further processing is required, such as combining multiple partial residual blocks together).

[0065] According to some embodiments, to reconstruct a block, a plurality of residual blocks (e.g., N residual blocks, where N is greater than 1) are encoded and used to obtain a final (e.g., combined) residual block. For example, the values ​​of the residual blocks may be added to generate a final residual block, which is then added to the predicted block to obtain a reconstructed block.

[0066] In some instances, the quantization step size may be different for the partial residual block. For example, for the second partial residual block, when the quantization step size is a predetermined default value (e.g., 1), the second partial residual block is encoded in a lossless mode. Otherwise, for the second partial residual block, when the quantization step size is not a predetermined default value (e.g., greater than 1), the second partial residual block is encoded in a lossy mode. In some embodiments, after the partial residual block with a lossless transform and a predetermined quantization step size, there are no further partial residual blocks (e.g., because the lossless transform does not have any distortion).

[0067] In some instances, in existing lossless coding schemes, lossy transforms cannot be applied, although they provide significant compression ratios. Some embodiments disclosed herein propose methods of using lossy transforms to achieve better compression ratios and lossless reconstruction at the same time. In the following, identity transforms refer to transform processes whose input and output are the same or scaled versions of each other.

[0068] Now, onto block partitioning. 4A to 4D An exemplary coding tree structure according to some embodiments is shown. Figure 4A As shown in the first coding tree structure (400) in , some coding methods use a 4-way partition tree starting from a 64×64 level and reducing to a 4×4 level, for example, which has some additional restrictions for 8×8 blocks. Figure 4A In , partitioning specified as "R" is recursive partitioning because the same partition tree is repeated in smaller sizes until the minimum level is reached. Figure 4B As shown in the exemplary coding tree structure (402) in FIG. 1 , some coding methods expand the partition tree to a 10-way structure and increase the maximum size (e.g., sometimes called a superblock) to start at 128×128. The second coding tree structure includes 4:1 / 1:4 rectangular partitions that are not in the first coding tree structure. Figure 4B The partition type with 3 sub-partitions in the second row of is called T-type partition. In addition to the coding block size, a coding tree depth may also be defined to indicate the partition depth from the root node.

[0069] As an example, a coding tree unit (CTU) can be divided into coding units (CUs) using a quadtree structure represented as a coding tree to accommodate various local characteristics. In some embodiments, the decision on whether to encode a picture area using inter-picture (temporal) prediction or intra-picture (spatial) prediction is made at the CU level. Depending on the prediction unit (PU) partition type, each CU can be further divided into one, two, or four PUs. Within a PU, the same prediction process is applied, and relevant information is transmitted to the decoder based on the PU. After obtaining the residual block by applying a prediction process based on the PU partition type, the CU can be divided into transform units (TUs) according to another quadtree structure (similar to the coding tree for the CU).

[0070] A quadtree with nested multi-type trees using binary and ternary partitioning segment structures can be used to replace the concept of multiple partition unit types. In the coding tree structure, a CU can have a square or rectangular shape. First, the CTU is partitioned by a four-level tree structure. The leaf nodes of the four-level tree can be further partitioned by a multi-type tree structure. Figure 4C As shown in the third coding tree structure (404) in FIG. 4 , the multi-type tree structure includes four types of partitioning. The leaf nodes of the multi-type tree are called CUs, unless the CU is too large for the maximum transform length. This means that CUs, PUs, and TUs have the same block size in a quadtree with a nested multi-type tree coding block structure. An example of block partitioning of a CTU (406) is shown in FIG. Figure 4D As shown, Figure 4D An exemplary quadtree is shown.

[0071] The coding tree scheme supports that luminance and chrominance can have separate block tree structures (for example, in VTM7). In some cases, for P slices and B slices, the luminance CTB and chrominance CTB in one CTU share the same coding tree structure. However, for I slices, luminance and chrominance may have separate block tree structures. When a separate block tree mode is applied, the luminance CTB is divided into CUs by one coding tree structure, and the chrominance CTB is divided into chrominance CUs by another coding tree structure. This means that the CU in the I slice may include a coding block of the luminance component or a coding block of two chrominance components, or that the CU in the I slice may consist of a coding block of the luminance component or a coding block of two chrominance components, and the CU in the P slice or B slice may always include coding blocks of all three color components or consist of coding blocks of all three color components, unless the video is monochrome.

[0072] Now, turning to transform blocks, to support extended coding block partitioning, multiple transform sizes (e.g., ranging from 4 to 64 points for each dimension) and transform shapes (e.g., square or rectangular with width / height ratios of 2:1 / 1:2 and 4:1 / 1:4) can be used.

[0073] The two-dimensional transform process may involve the use of a hybrid transform kernel (e.g., composed of a different one-dimensional transform for each dimension of the coded residual block). The main one-dimensional transform may include at least one of the following: a) 4-point, 8-point, 16-point, 32-point, 64-point discrete cosine transform DCT-2; b) 4-point, 8-point, 16-point asymmetric discrete sine transform (DST-4, DST-7) and its flipped version; or c) 4-point, 8-point, 16-point, 32-point identity transform. For example, the basis functions of DCT-2 and asymmetric DST used in AV1 are listed in Table 1. Table 1 - Exemplary primary transform basis functions

[0074] The availability of hybrid transform kernels may be based on transform block size and prediction mode. Exemplary dependencies are listed in Table 2 below, where "→" and "↓" represent horizontal and vertical dimensions, and "√" and "×" represent the availability of kernels for that block size and prediction mode. IDTX (or IDT) stands for Identity Transform. Table 2 - Availability of hybrid transform kernels

[0075] For chroma components, transform type selection may be performed in an implicit manner. For intra prediction residuals, the transform type may be selected based on the intra prediction mode, e.g., as specified in Table 3. For inter prediction residuals, the transform type may be selected based on the transform type selection of the co-located luma block. Therefore, for chroma components, it may not be necessary to signal the transform type in the bitstream. Intra prediction Vertical Transformation Horizontal Transform DC_PRED DCT DCT V_PRED ADST DCT H_PRED DCT ADST D45_PRED DCT DCT D135_PRED ADST ADST D113_PRED ADST DCT D157_PRED DCT ADST D203_PRED DCT ADST D67_PRED ADST DCT SMOOTH_PRED ADST ADST SMOOTH_V_PRED ADST DCT SMOOTH_H_PRED DCT ADST PAETH_PRED ADST ADST Table 3 - Transform type selection for chroma intra prediction residual

[0076] Now, turning to exemplary encoding and decoding using prediction blocks and residual blocks, Figure 5A Calculation of a prediction block according to some embodiments is shown. Figure 5A In the example of , intra prediction is performed on the current block 502 to generate a prediction block 504. In some embodiments, inter prediction is performed to generate a prediction block. The current block 502 includes a set of samples (eg, pixel blocks), and the prediction block 504 includes a prediction set corresponding to the set of samples. Figure 5B Calculation of the residual block according to some embodiments is shown. Figure 5B As shown, the prediction block 504 is subtracted from the current block 502 to generate a residual block 506 including a residual set. For example, a corresponding difference is calculated between each sample and the corresponding prediction. Figure 5C Calculation of a reconstructed block according to some embodiments is shown. Figure 5CAs shown, the residual block 506 undergoes one or more transformations and quantizations to generate a residual coefficient set. The residual coefficient set can be transmitted from the encoder component to the decoder component. The residual coefficient set undergoes inverse quantization and inverse transformation to generate a reconstructed residual block 508. The reconstructed residual block 508 is combined with the prediction block 504 (e.g., the reconstructed residual of the reconstructed residual block 508 is added to the prediction of the prediction block 504) to generate a reconstructed block 510 corresponding to the current block 502.

[0077] Figure 5D An example of generating multiple residual blocks for a current block is shown. Figure 5D , a first residual block 514 (e.g., a first partial residual block) of the current block 512 is obtained using a first parameter set, a second residual block 516 (e.g., a second partial residual block) of the current block 512 is obtained using a second parameter set, and a third residual block 518 (e.g., a third partial residual block) of the current block 512 is obtained using a third parameter set. In some embodiments, the first residual block 514, the second residual block 516, and the third residual block 520 are combined (e.g., by weighted sum) to form a combined residual block 520. Although Figure 5D An example with three partial residual blocks is shown, but in some embodiments, more or fewer partial residual blocks are used (eg, 2, 4, or 5 partial residual blocks).

[0078] Figure 5E Calculation of a reconstructed block according to some embodiments is shown. Figure 5E As shown, the combined residual block 520 undergoes one or more transformations and quantizations to generate a set of residual coefficients. The set of residual coefficients can be transmitted from the encoder component to the decoder component. The set of residual coefficients undergoes inverse quantization and inverse transformation to generate a reconstructed residual block 524. The reconstructed residual block 524 is combined with the prediction block 526 to generate a reconstructed block 528 corresponding to the current block 512. In some embodiments, the residual coefficients of the combined residual block 520 are not signaled, but the residual coefficients of the residual blocks 514, 516 and 518 are signaled. For example, the reconstructed residual block 524 is generated by reconstructing the residual blocks 514, 516 and 518 (at the decoder) from the coefficients represented by the signal, and then combining the reconstructed blocks. In some embodiments, at least one subset of the residual coefficients is not quantized (for example, the corresponding residual block transformed by the lossless method is not quantized).

[0079] Fig. 6A1 is a flow chart illustrating a method 600 of decoding a video according to some embodiments. The method 600 may be performed in a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuit and a memory storing instructions executed by the control circuit. In some embodiments, the method 1000 is performed by executing instructions stored in a memory (e.g., memory 314) of the computing system.

[0080] The system receives (602) a video bitstream including a plurality of coded blocks, the plurality of coded blocks including a current block. The system obtains (604) a first residual block of the current block using a first parameter set. The system obtains (606) a second residual block of the current block using a second parameter set, wherein the second parameter set is independent of the first parameter set. The system obtains (608) a combined residual block using the first residual block and the second residual block. The system reconstructs (610) the current block using the combined residual block. In some embodiments, for a block (e.g., a transform block), a plurality of residual blocks are encoded and used to obtain a combined residual block (e.g., a final residual block) for reconstruction of the block. Parameters / syntax (e.g., transform size, transform kernel type, and quantization index) used to obtain each partial residual block can be independently obtained or signaled, and related syntax elements can be parsed or inferred at the encoder side and the decoder side.

[0081] In some embodiments, among all partial residual blocks used to obtain a final (e.g., combined) residual block to reconstruct the current block, a first transform block size is used for a first partial residual block and a second transform block size is used for a second partial residual block. The first partial residual block and the second partial residual block may be different residual blocks. The first transform block size and the second transform block size may be different. For example, the transform block size of the first partial residual block is one of 4×4, 8×8, 16×16, and 32×32. The transform block size of the second partial residual block is also one of 4×4, 8×8, 16×16, and 32×32.

[0082] In some embodiments, for the lossless coding mode, the transform size of the first partial residual block can be fixed (e.g., 4×4). The transform type of the second partial residual block is also fixed (e.g., Hadamard transform). In some embodiments, for the lossless coding mode, if the transform type is an identity transform, the transform size for the last partial residual block can be 4×4, 8×8, 16×16, or 32×32. For example, for the first partial residual block, the transform block size of the first partial residual block (e.g., 4×4) can be fixed, and the transform block size of the second partial residual block (e.g., 4×4, 8×8, 16×16, or 32×32) can be selected from a predetermined size set. The selection can be signaled or obtained implicitly.

[0083] In some embodiments, among the partial residual blocks used to obtain a final (e.g., combined) residual block to reconstruct a current block, a first transform kernel type is used for a first partial residual block and a second transform kernel type is used for a second partial residual block. The first partial residual block and the second partial residual block may be different residual blocks. The first transform kernel type and the second transform kernel type may be different. For example, the transform kernel type of one partial residual block may be DCT-DCT, while the transform block type of another partial residual block may be DST-DST. Using different transform kernel types for different partial residual blocks can improve coding efficiency.

[0084] In some embodiments, different transform kernel sets may be used for multiple partial residual blocks. For example, the transform kernel set may include multiple transform kernels. In one example, for a first partial residual block, the transform kernel set is DCT-DCT, DST-DST, and identity transform. For another partial residual block, the set may be DCT-DCT and DST-DST.

[0085] In some embodiments, the primary transform kernel may be signaled for all partial residual blocks. In some embodiments, the primary transform kernel is predefined for at least one partial residual block.

[0086] In some embodiments, for the lossless coding mode, the transform kernel type of at least one partial residual block is signaled as a Hadamard transform or an identity transform to ensure that the lossless requirement is met.

[0087] In some embodiments, among the partial residual blocks encoded using a quantization step size greater than a predetermined default value (e.g., 1), a first secondary transform type is used for the first partial residual block and a second secondary transform type is used for the second partial residual block. The first partial residual block and the second partial residual block may be different residual blocks. The first secondary transform type and the second secondary transform type may be different. In some embodiments, a secondary transform is applied after the main transform for each lossy partial residual block. In some embodiments, a secondary transform is applied to a subset of the partial residual blocks. In some embodiments, a secondary transform is not applied to a partial residual block with a predetermined quantization step size.

[0088] In some embodiments, among the partial residual blocks used to obtain a final (e.g., combined) residual block to reconstruct a current block, a first quantization index is used for a first partial residual block and a second quantization is used for a second partial residual block. The first partial residual block and the second partial residual block may be different residual blocks. The first quantization index and the second quantization index may be different.

[0089] In some embodiments, the quantization index is received at the decoder side for use in the dequantization process. In some embodiments, the basic quantization index may be received at the block, frame, sequence, or GOP (Group of Pictures) level, while the quantization index increment is received at the block level. The quantization index is recovered by adding the basic quantization index and the quantization index increment. The recovered quantization index is used for the dequantization process at the decoder side.

[0090] In some embodiments, a base quantization index is signaled at the coded block level and a quantization index increment is signaled at the transform block level. In some embodiments, a base quantization index is signaled at the frame level and a quantization index increment is signaled at the coded block level. In some embodiments, for lossless coding, at least one partial residual block does not require a quantization index.

[0091] In some embodiments, the quantization index is not signaled, but is implicitly obtained based on any coded information known to both the encoder and the decoder, such coded information including but not limited to the coding block size, the temporal layer, the picture resolution size, the coding mode. In some embodiments, the context for signaling and parsing syntax elements related to the multiple residual block mode may depend on the coded information of the left block, the coded information of the upper block, or the coded information of the left block and the upper block combined to encode (including but not limited to) the multiple residual block mode, the transform size of the partial residual block, the transform type of the partial residual block, the quantization index of the partial residual block, and the mode in the secondary transform mode. In some embodiments, the signaling of the multiple residual block mode can be bypass coded, which may not use any context.

[0092] In some embodiments, the various parameters (e.g., indexes) described above are signaled using a high-level syntax, including but not limited to signaling at a sequence level, a picture level, a sub-picture level, a slice level, a tile level, a tile group level, and a maximum coding block level.

[0093] In some embodiments, to reconstruct a block, multiple residual blocks (each residual block is referred to as a partial residual block) are encoded and used to obtain a final (e.g., combined) residual block, which is added to the prediction block to obtain the reconstructed block. In some embodiments, the final (e.g., combined) residual block is obtained as the sum or weighted sum of multiple partial residual blocks in the frequency domain or the spatial domain.

[0094] In some embodiments, different quantization methods are applied to multiple residual blocks. Quantization method refers to any parameter or operation involved in quantization. The quantization method can be applied in the quantization process at the encoder, or in the dequantization process at the encoder and / or decoder. In some embodiments, different quantization step sizes are applied to partial residual blocks. In some embodiments, quantization is not applied to at least one of the partial residual blocks (or the quantization step size is 1).

[0095] In some embodiments, different transform coding methods are applied to multiple residual blocks. Transform coding methods refer to any parameters or operations involved in the transform process. The transform method can be applied in the forward transform process at the encoder, or in the inverse transform process at the encoder and / or decoder. In some embodiments, different transform kernels are applied to partial residual blocks. In some embodiments, no transform is applied to at least one of the partial residual blocks (or the transform kernel is an identity transform).

[0096] In some embodiments, different entropy coding methods are applied to the multiple residual blocks. Entropy coding methods refer to any parameters or operations involved in entropy coding. Entropy coding methods can be applied in an encoder, where the encoder converts residual samples (or transform coefficients) into entropy-coded binary bits (bins), or can also be applied during the parsing process of a decoder, where the decoder converts entropy-coded bins back to residual samples (or transform coefficients). In some embodiments, different binarizations (or from syntax values ​​to symbol values) are applied to some residual blocks during entropy coding.

[0097] In some embodiments, different coefficient (or residual) scanning orders are applied to these partial residual blocks in entropy coding. In some embodiments, different contexts are applied to the coefficients (or residuals) of these partial residual blocks in entropy coding.

[0098] In some embodiments, multiple (more than one) residual blocks may be encoded only for lossless coding mode. For lossy coding mode, only a single residual block may be encoded. In some embodiments, N (where N>1) residual blocks may be encoded only for lossless coding mode. For lossy coding mode, less than N and at least one residual block may be encoded.

[0099] In some embodiments, the block sizes of the plurality of partial residual blocks are different. In some embodiments, the block size of at least one partial residual block is the same as the block size of the currently encoded block. In some embodiments, the block size of at least one partial residual block is smaller than the block size of the currently encoded block. In one example, the block width and / or height of a partial residual block is half the block width and / or height of the currently encoded block.

[0100] In some embodiments, a flag is signaled to indicate whether to encode a single or multiple residual blocks. In some embodiments, the flag is signaled at the sequence level, frame level, slice level, super block, coding tree unit, or coded block level. In some embodiments, for lossless mode, the flag is inferred to encode all residual blocks.

[0101] In some embodiments, multiple reconstructed blocks may be generated for the current block using different partial residual blocks. In some embodiments, the reconstructed samples used as reference samples for another block and the reconstructed samples used for display (or further loop filtering) may be different. In some embodiments, a reconstructed block is obtained using partial residual blocks associated with a lossy quantization and / or dequantization process, and another reconstructed block is obtained using all partial residual blocks that provide lossless reconstruction of the current block. In some embodiments, for lossless reconstruction of the current block, the partial residual blocks are added to the reconstruction. In some embodiments, the reconstructed samples used as reference samples for another block in the same picture may be different from the reconstructed samples used as reference samples for another block in a different picture.

[0102] Figure 6B 6 is a flow chart illustrating a video encoding method 650 according to some embodiments. The method 650 may be performed in a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuit and a memory storing instructions executed by the control circuit. In some embodiments, the method 650 is performed by executing instructions stored in a memory (e.g., memory 314) of the computing system.

[0103] The system receives (652) video data including a plurality of blocks, the plurality of blocks including a current block. The system obtains (654) a first residual block for the current block using a first parameter set. The system obtains (656) a second residual block for the current block using a second parameter set, wherein the second parameter set is independent of the first parameter set. The system signals (658) the first residual block and the second residual block in a video bitstream. As previously described, the encoding process may mirror the decoding process described herein (e.g., obtaining and / or using a plurality of residual blocks). For the sake of brevity, these details are not repeated here.

[0104] Although Fig. 6A and Figure 6B The various logical stages are shown in a particular order, but stages that are not order-dependent may be reordered, and other stages may be combined or decomposed. Some reordering or other groupings not specifically mentioned will be apparent to one of ordinary skill in the art, and thus the ordering and groupings presented herein are not exhaustive. Furthermore, it should be appreciated that the various stages may be implemented in hardware, firmware, software, or any combination thereof.

[0105] Now, turning to some exemplary embodiments:

[0106] (A1) In one aspect, some embodiments include a video decoding method (e.g., method 600). In some embodiments, the method is performed in a computing system (e.g., server system 112) having a memory and a control circuit. In some embodiments, the method is performed in an encoding module (e.g., encoding module 320). In some embodiments, the method is performed in a source encoding component (e.g., source encoder 202), an encoding engine (e.g., encoding engine 212), and / or an entropy encoder (e.g., entropy encoder 214). The method includes: (i) receiving a video code stream including a plurality of encoding blocks, the plurality of encoding blocks including a current block; (ii) obtaining a first residual block of the current block using a first parameter set; (iii) obtaining a second residual block of the current block using a second parameter set, wherein the second parameter set is independent of the first parameter set; (iv) obtaining a combined residual block using the first residual block and the second residual block; and (v) reconstructing the current block using the combined residual block. For example, for a block (e.g., a transform block), a plurality of residual blocks are encoded, and the plurality of residual blocks are used together to obtain a residual block for reconstruction of the block. Parameters / syntax (e.g., transform size, transform kernel type, and quantization index) for obtaining each partial residual block are independently obtained / signaled, and relevant syntax elements are parsed or inferred at the encoder side and the decoder side. In some embodiments, a first parameter set is obtained and a second parameter set is signaled in the video bitstream. In some embodiments, the first parameter set and / or the second parameter set include one or more of the transform size, transform kernel type, and quantization index. In some embodiments, at least a subset of the first parameter set and / or the second parameter set is signaled in a high-level syntax.

[0107] (A2) In some embodiments of A1, the first residual block corresponds to a first transform block size, and the second residual block corresponds to a second transform block size, and the second transform block size is different from the first transform block size. For example, among all partial residual blocks used together to obtain a final residual block to reconstruct a current block, for the first partial residual block, the first transform block size is used, and for the second partial residual block, the second transform block size is used. The first partial residual block and the second partial residual block may be different residual blocks, and the first transform block size and the second transform block size may be different.

[0108] (A3) In some embodiments of A2, the first transform block size and the second transform block size are selected from the group consisting of 4×4, 8×8, 16×16 and 32×32. For example, the transform block size of the first partial residual block is one of 4×4, 8×8, 16×16 and 32×32, and the transform block size of the second partial residual block is also one of 4×4, 8×8, 16×16 and 32×32.

[0109] (A4) In some embodiments of A2 or A3, the first transform block size is a fixed size, wherein the transform type of the first residual block is fixed. For example, for a lossless coding mode, the transform size of the first partial residual block may be fixed (e.g., 4×4), and the transform type of the first partial residual block is also fixed (e.g., Hadamard transform).

[0110] (A5) In some embodiments of any one of A2 to A4, the first transform block size is a fixed size, and the second transform block size is selected from a group of predetermined transform sizes. For example, for a first partial residual block, the transform block size of the first partial residual block is fixed, such as 4×4, and the transform block size of the second partial residual block can be selected from a plurality of sizes among 4×4, 8×8, 16×16 or 32×32, and the selection can be signaled or obtained implicitly. As an example, the selection can be signaled or obtained implicitly based on the encoded block size, for example, if the block size is 16×16, the transform size can be set to be the same as the block size. For another partial residual block, the transform size can always be, for example, 4×4, regardless of the encoded block size.

[0111] (A6) In some embodiments of any one of A1 to A5, the second residual block corresponds to a lossless coding mode. The method includes: according to the second residual block having an identity transform type, the second residual block corresponds to a first transform block size selected from the group consisting of 4×4, 8×8, 16×16 and 32×32. For example, for the lossless coding mode, if the transform type is an identity transform, then for the last partial residual block, the transform size can be 4×4, 8×8, 16×16 or 32×32.

[0112] (A7) In some embodiments of any one of A1 to A6, the first residual block corresponds to a first transform kernel type, and the second residual block corresponds to a second transform kernel type, and the second transform kernel type is different from the first transform kernel type. For example, among all partial residual blocks used together to obtain a final residual block to reconstruct the current block, the first transform kernel type is used for the first partial residual block, and the second transform kernel type is used for the second partial residual block. The first partial residual block and the second partial residual block may be different residual blocks, and the first transform kernel type and the second transform kernel type may be different. In some embodiments, the first transform kernel type is DCT-DCT, and the second transform kernel type is DST-DST. For example, the transform kernel type of one partial residual block is DCT-DCT, and the transform block type of another partial residual block is DST-DST.

[0113] (A8) In some embodiments of A7, the first transform kernel type is selected from a first transform kernel set, and the second transform kernel type is selected from a second transform kernel set, and the second transform kernel set is different from the first transform kernel set. For example, different transform kernel sets may be used for multiple partial residual blocks.

[0114] (A9) In some embodiments of A7 or A8, for each of the first residual block and the second residual block, a corresponding main transform kernel is predefined or signaled. For example, the main transform kernel may be signaled for all partial residual blocks. As another example, the main transform kernel may be predefined for at least one partial residual block.

[0115] (A10) In some embodiments of any one of A7 to A9, the second residual block corresponds to a lossless coding mode, and the second transform kernel type is limited to a lossless transform set. For example, for the lossless coding mode, the transform kernel type of at least one partial residual block is signaled as a Hadamard transform or an identity transform to ensure lossless requirements.

[0116] (A11) In some embodiments of any one of A1 to A10, the first residual block corresponds to a first quantization step size, and the second residual block corresponds to a second quantization step size, and the second quantization step size is different from the first quantization step size. For example, among all partial residual blocks encoded using a quantization step size greater than a predetermined default value (e.g., 1), for the first partial residual block, a first secondary transform type is used, and for the second partial residual block, a second secondary transform type is used. The first partial residual block and the second partial residual block may be different residual blocks, and the first secondary transform type and the second secondary transform type may be different.

[0117] (A12) In some embodiments of any one of A1 to A11, the first residual block corresponds to a first secondary transform type, and the second residual block corresponds to a second secondary transform type, the second secondary transform type being different from the first secondary transform type. For example, a secondary transform may be applied after a primary transform for each lossy portion residual block. In some embodiments, when the residual block has a quantization step size greater than a predetermined threshold (e.g., 1), the residual block is generated according to the secondary transform.

[0118] (A13) In some embodiments of any one of A1 to A12, the first residual block corresponds to application of a secondary transform and the second residual block corresponds to not applying a secondary transform. For example, the secondary transform is applied to a subset of the partial residual blocks. In some embodiments, the secondary transform is not applied to the partial residual blocks with a predetermined quantization step size.

[0119] (A14) In some embodiments of any one of A1 to A13, the first residual block corresponds to a first quantization index, and the second residual block corresponds to a second quantization index, and the second quantization index is different from the first quantization index. For example, among all partial residual blocks used to obtain a final residual block to reconstruct a current block, a first quantization index is used for the first partial residual block, and a second quantization is used for the second partial residual block. The first partial residual block and the second partial residual block may be different residual blocks, and the first quantization index and the second quantization index may be different. In some embodiments, the first quantization index and / or the second quantization index are represented by a signal in the video bitstream. For example, the quantization index can be received at the decoder side for use in the dequantization process.

[0120] (A15) In some embodiments of A14, the first quantization index is determined based on a basic quantization value signaled at a first coding level and an index increment signaled at a second coding level, the second coding level being smaller than the first coding level. For example, a basic quantization index may be received at a block, frame, sequence, or GOP (group of pictures) level, and a quantization index increment may be received at a block level. The quantization index is recovered by adding the basic quantization index and the quantization index increment. The recovered quantization index is used for a dequantization process at the decoder side. As another example, the basic quantization index is signaled at the coded block level, and the quantization index increment is signaled at the transform block level. As another example, the basic quantization index is signaled at the frame level, and the quantization index increment is signaled at the coded block level.

[0121] (A16) In some embodiments of A14 or A15, at least one of the first quantization index and the second quantization index is not represented by a signal in the video bitstream. For example, for lossless coding, at least one partial residual block may not require a quantization index. In some embodiments, at least one of the first quantization index and the second quantization index is obtained by a computing system based on encoded information. For example, the quantization index is not represented by a signal, but is implicitly obtained based on any encoded information known to both the encoder and the decoder, such encoded information including but not limited to the coding block size, time layer, picture resolution size, and coding mode.

[0122] (A17) In some embodiments of any one of A1 to A16, the method includes: (i) using encoded information to determine a context for entropy encoding a first parameter set, the encoded information including one or more of the following: a residual mode, a transform size, a transform type, a quantization index, and a secondary transform mode; and (ii) entropy encoding the first parameter set using the determined context. For example, the context for signaling and parsing syntax elements related to a multiple residual block mode may depend on encoded information of a left block, encoded information of an upper block, or encoded information of a combination of a left block and an upper block to encode (including but not limited to) a multiple residual block mode, a transform size of a partial residual block, a transform type of a partial residual block, a quantization index of a partial residual block, and a mode in a secondary transform mode. In some embodiments, a bypass mode is used to signal at least one of the first parameter set and the second parameter set. For example, the signaling of the multiple residual block mode may be bypass encoded, which may not use any context.

[0123] (B1) On the other hand, some embodiments include a video encoding method (e.g., method 650). In some embodiments, the method is performed in a computing system having a memory and one or more processors. The method includes: (i) receiving video data including a plurality of blocks, the plurality of blocks including a current block; (ii) obtaining a first residual block of the current block using a first parameter set; (iii) obtaining a second residual block of the current block using a second parameter set, wherein the second parameter set is independent of the first parameter set; and (iv) signaling the first residual block and the second residual block in a video bitstream.

[0124] (B2) In some embodiments of B1, the method includes: using a signal to represent at least one of the first parameter set and the second parameter set through a video bit stream.

[0125] (C1) On the other hand, some embodiments include a method for processing visual media data. In some embodiments, the method is performed in a computing system having a memory and one or more processors. The method includes: (i) obtaining a source video sequence including a plurality of frames; and (ii) performing conversion between the source video sequence and a video stream of visual media data, wherein the video stream includes: (a) a plurality of encoded blocks corresponding to the plurality of frames, the plurality of encoded blocks including a current block; (b) a first residual block for the current block, wherein the first residual block is generated using a first parameter set; and (c) a second residual block for the current block, wherein the second residual block is generated using a second parameter set.

[0126] On the other hand, some embodiments include a computing system (e.g., server system 112) that includes a control circuit (e.g., control circuit 302) and a memory (e.g., memory 314) coupled to the control circuit, the memory storing one or more instruction sets configured to be executed by the control circuit, the one or more instruction sets including instructions for performing any of the methods described herein (e.g., A1 to A17, B1 and B2, and C1 above).

[0127] On the other hand, some embodiments include a non-transitory computer-readable storage medium that stores one or more instruction sets executed by control circuitry of a computing system, the one or more instruction sets including instructions for performing any of the methods described herein (e.g., A1 to A17, B1 and B2, and C1 above).

[0128] Unless otherwise specified, any of the syntax elements described herein may be high-level syntax (HLS). As used herein, HLS is signaled at a level higher than a block level. For example, HLS may correspond to a sequence level, a frame level, a slice level, or a tile level. As another example, an HLS element may be signaled in a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, a picture header, a tile header, and / or a CTU header.

[0129] It should be understood that, although the terms "first", "second" etc. can be used in this article to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another element. The terms used herein are only for the purpose of describing a specific embodiment and are not intended to limit the claims. As used in the description of the embodiment and the attached claims, the singular forms "one", "one" and "the" are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "and / or" used herein refer to any and all possible combinations of one or more listed related items, and include any and all possible combinations of one or more listed related items. It should also be further understood that when used in this specification, the terms "include" and / or "comprising" specify the existence of stated features, integers, steps, operations, elements and / or components, but do not exclude the existence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups.

[0130] As used herein, the term "when" may be interpreted to mean "if" or "when" or "in response to determining" or "according to determining" or "in response to detecting" the stated precondition is true, depending on the context. Similarly, the phrase "if it is determined that [the stated precondition is true]" or "if [the stated precondition is true]" or "when [the stated precondition is true]" may be interpreted to mean "upon determining" or "in response to determining" or "according to determining" or "upon detecting" or "in response to detecting" the stated precondition is true, depending on the context. As used herein, N refers to a variable number. Unless explicitly stated, different instances of N may refer to the same number (e.g., the same integer value (e.g., the number 2)) or different numbers.

[0131] For purposes of explanation, the foregoing description has been described with reference to specific embodiments. However, the illustrative discussion above is not intended to be exhaustive or to limit the claims to the precise forms disclosed. In view of the above teachings, many modifications and variations are possible. The embodiments are chosen and described in order to best explain the principles of operation and practical application, thereby enabling others skilled in the art to implement.

Claims

1. A video decoding method, the method being executed in a computing system having a memory and one or more processors, the method comprising: Receiving a video stream including a plurality of coding blocks, wherein the plurality of coding blocks include a current block; Obtain a first residual block of the current block using a first parameter set; Obtaining a second residual block of the current block using a second parameter set, wherein the second parameter set is independent of the first parameter set; obtaining a combined residual block using the first residual block and the second residual block; and The current block is reconstructed using the combined residual block.

2. The method according to claim 1, wherein: The first residual block corresponds to a first transform block size, and wherein the second residual block corresponds to a second transform block size, the second transform block size being different from the first transform block size.

3. The method according to claim 2, wherein: The first transform block size and the second transform block size are selected from the group consisting of 4×4, 8×8, 16×16 and 32×32.

4. The method according to claim 2, wherein: The first transform block size is a fixed size, and wherein the transform type of the first residual block is fixed.

5. The method according to claim 2, wherein: The first transform block size is a fixed size, and wherein the second transform block size is selected from a predetermined set of transform sizes.

6. The method according to claim 1, wherein: The second residual block corresponds to a lossless coding method; and According to the second residual block having an identity transform type, the second residual block corresponds to a transform block size selected from the group consisting of 4×4, 8×8, 16×16 and 32×32.

7. The method according to claim 1, wherein: The first residual block corresponds to a first transform kernel type, and wherein the second residual block corresponds to a second transform kernel type, the second transform kernel type being different from the first transform kernel type.

8. The method according to claim 7, wherein: The first transform kernel type is selected from a first transform kernel set, and wherein the second transform kernel type is selected from a second transform kernel set, the second transform kernel set being different from the first transform kernel set.

9. The method according to claim 7, wherein: For each of the first residual block and the second residual block, a corresponding primary transform kernel is predefined or signaled.

10. The method according to claim 7, wherein: The second residual block corresponds to a lossless coding mode, and wherein the second transform kernel type is restricted to a lossless transform set.

11. The method according to claim 1, wherein: The first residual block corresponds to a first quantization step size, and wherein the second residual block corresponds to a second quantization step size, the second quantization step size being different from the first quantization step size.

12. The method according to claim 1, wherein: The first residual block corresponds to a first secondary transform type, and wherein the second residual block corresponds to a second secondary transform type, the second secondary transform type being different from the first secondary transform type.

13. The method according to claim 1, wherein: The first residual block corresponds to applying a secondary transform, and wherein the second residual block corresponds to not applying a secondary transform.

14. The method according to claim 1, wherein: The first residual block corresponds to a first quantization index, and wherein the second residual block corresponds to a second quantization index, the second quantization index being different from the first quantization index.

15. The method according to claim 14, wherein: The first quantization index is determined from a base quantization value signaled at a first coding level and an index increment signaled at a second coding level, the second coding level being smaller than the first coding level.

16. The method according to claim 14, wherein: At least one of the first quantization index and the second quantization index is not represented by a signal in the video bitstream.

17. The method according to claim 1, further comprising: Determining a context for entropy encoding the first parameter set using encoded information, the encoded information comprising one or more of: a residual mode, a transform size, a transform type, a quantization index, and a secondary transform mode; and The first parameter set is entropy encoded using the determined context.

18. A computing system comprising: Control circuit; Memory; as well as one or more instruction sets stored in the memory and configured to be executed by the control circuitry, the one or more instruction sets comprising instructions for: receiving video data including a plurality of blocks, the plurality of blocks including a current block; Obtain a first residual block of the current block using a first parameter set; Obtaining a second residual block of the current block using a second parameter set, wherein the second parameter set is independent of the first parameter set; as well as The first residual block and the second residual block are represented by a signal in a video bitstream.

19. The computing system of claim 18, further comprising signaling, via the video stream, at least one of the first parameter set and the second parameter set.

20. A non-transitory computer-readable storage medium storing one or more instruction sets configured to be executed by a computing device having control circuitry and memory, the one or more instruction sets comprising instructions for: Obtaining a source video sequence including a plurality of frames; as well as A conversion is performed between the source video sequence and a video code stream of visual media data, wherein: The video code stream includes: a plurality of coded blocks corresponding to the plurality of frames, the plurality of coded blocks including the current block; a first residual block for the current block, wherein the first residual block is generated using a first parameter set; and A second residual block for the current block, wherein the second residual block is generated using a second parameter set.