Systems and methods for applying inseparable transforms to inter prediction residuals

By transmitting flags in the video code stream to control the application of inseparable transform cores, the problems of encoding quality and efficiency in inter-frame prediction residual processing are solved, and higher quality and more efficient video encoding and decoding are achieved.

CN120391052APending Publication Date: 2025-07-29TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480005040.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-29
Filing Date
2024-05-31
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing video encoding technology is difficult to effectively utilize inseparable transformation in inter-frame prediction residual processing to improve encoding quality and encoding and decoding efficiency.

Method used

By transmitting the flags in the video code stream, it indicates whether the inseparable transform core is applied to the set of residual blocks of inter-frame mode, and decide whether the first inseparable transform core is applied based on the value of the flags, the quality of the reconstruction video is improved and the calculation cost is reduced.

Benefits of technology

Improve the accuracy and accuracy of reconstructed videos, while reducing calculation costs and improving encoding and decoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120391052A_ABST
    Figure CN120391052A_ABST
Patent Text Reader

Abstract

Various implementations described herein include methods and systems for video coding and decoding. In one aspect, a method includes receiving a video bitstream, the video bitstream including a set of inter-mode coding blocks and a corresponding set of transform coefficients. The method includes deriving a set of inter-mode residual blocks from a set of transform coefficients. The method includes determining whether to apply one or more inseparable transform kernels to an inter-mode residual block set according to a value of a first indicator in a video bitstream. The method includes applying a first inseparable transform kernel when the indicator has a first value, and abandoning application of the first inseparable transform kernel when the indicator has a second value. The method further includes reconstructing the set of video blocks using the set of inter-mode residual blocks and the corresponding set of prediction blocks.
Need to check novelty before this filing date? Find Prior Art

Description

Cross - Reference to Related Applications

[0001] This application claims the benefit of priority of U.S. Provisional Application No. 63 / 603,544, filed on November 28, 2023, entitled "Methods and Apparatus for Applying Non - Separable Transforms on Inter Prediction Residuals", and this application is a continuation - in - part of, and claims the priority of, U.S. Patent Application No. 18 / 677,733, filed on May 29, 2024, entitled "Methods and Apparatus for Applying Non - Separable Transforms on Inter Prediction Residuals". The entire contents of the above - mentioned provisional application and patent application are incorporated herein by reference. Technical Field

[0002] The disclosed embodiments generally relate to video coding and decoding, including but not limited to transform coding applied to prediction residuals. Background Art

[0003] Digital video is supported by various electronic devices, such as digital televisions, laptop or desktop computers, tablets, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video teleconferencing devices, video streaming devices, etc. Electronic devices send and receive or otherwise transfer digital video data through communication networks and / or store digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video coding can be used to compress video data according to one or more video coding standards before transferring or storing the video data. Video coding and decoding can be performed by hardware and / or software on an electronic / client device or a server providing cloud services.

[0004] Video coding typically utilizes prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.), and such prediction methods utilize the redundancy inherent in video data. Video coding aims to compress video data into a form using a lower bitrate while avoiding or minimizing the degradation of video quality. Multiple video codec standards have been developed. For example, High-Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. The HEVC / H.265 standard was released by ITU-T and ISO / IEC in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Versatile Video Coding (VVC / H.266) is a video compression standard intended as a successor to HEVC. The VVC / H.266 standard was released by ITU-T and ISO / IEC in 2020 (version 1) and 2022 (version 2). AV1 (AOMedia Video 1) is an open video coding format designed as an alternative to HEVC. On January 8, 2019, the verification version 1.0.0 (with errata 1) of this specification was released. Summary of the Invention

[0005] The present disclosure describes a set of methods for video (image) compression, including a method of applying an inseparable transform to an inter-frame mode block residual. According to some embodiments, a video bitstream includes a set of inter-frame mode coded blocks and a corresponding set of transform coefficients. A set of inter-frame mode residual blocks can be derived from the set of transform coefficients. The value of a flag transmitted in the video bitstream can indicate whether one or more inseparable transform kernels are applied to the set of inter-frame mode residual blocks. Depending on the value of the flag, a first inseparable transform kernel may or may not be applied to the set of inter-frame mode residual blocks. A set of video blocks can be reconstructed from the set of inter-frame mode residual blocks and a corresponding set of prediction blocks. The advantage of using inseparable transform coding and decoding techniques in this way is that the quality (e.g., accuracy and / or precision) of the corresponding reconstructed (decoded) video can be improved. In addition, signaling whether to use an inseparable transform kernel can reduce the computational cost (e.g., more efficiently encode and / or decode the residuals of the video).

[0006] According to some embodiments, a video decoding method includes (i) receiving a video bitstream that includes a set of inter-mode coded blocks and a corresponding set of transform coefficients; (ii) deriving a set of (one or more) inter-mode residual blocks from the set of transform coefficients; (iii) determining whether to apply one or more non-separable transform kernels to the set of inter-mode residual blocks according to the value of a first indicator in the video bitstream; (iv) when the indicator has a first value, applying a first non-separable transform kernel to the set of inter-mode residual blocks; (v) when the indicator has a second value, forgoing applying the first non-separable transform kernel to the set of inter-mode residual blocks; and (vi) using the set of inter-mode residual blocks and a corresponding set of prediction blocks to reconstruct a set of video blocks.

[0007] According to some embodiments, a video encoding method includes (i) receiving video data that includes a set of video blocks; (ii) deriving a set of inter-mode residual blocks from the set of video blocks; (iii) determining whether to apply one or more non-separable transform kernels to the set of inter-mode residual blocks; (iv) generating a set of transform coefficients from the set of inter-mode residual blocks according to whether to apply one or more non-separable transform kernels to the set of inter-mode residual blocks; (v) determining the value of a first indicator according to whether to apply one or more non-separable transform kernels to the set of inter-mode residual blocks; (vi) transmitting the first indicator in the video bitstream; and (vii) transmitting the set of transform coefficients in the video bitstream.

[0008] According to some embodiments, a method of bitstream conversion includes (i) obtaining a source video sequence that includes a plurality of pictures; and (ii) performing a conversion between the source video sequence and a bitstream of visual media data, where the bitstream includes: (a) a plurality of coded blocks corresponding to the plurality of pictures; (b) a set of transform coefficients corresponding to the plurality of coded blocks; and (c) a first indicator that indicates whether to apply one or more non-separable transform kernels to a set of inter-mode residual blocks corresponding to the set of transform coefficients.

[0009] According to some embodiments, there is provided a computing system, such as a streaming system, a server system, a personal computer system, or other electronic device. The computing system includes a control circuit and a memory that stores one or more instruction sets. The one or more instruction sets include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and a decoder component (e.g., a transcoder).

[0010] According to some embodiments, there is provided a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores one or more instruction sets executable by a computing system. The one or more instruction sets include instructions for performing any of the methods described herein.

[0011] Accordingly, apparatuses and systems using methods for encoding and decoding video are disclosed. Such methods, apparatuses, and systems may supplement or replace conventional methods, apparatuses, and systems for video encoding / decoding. The features and advantages described in the specification do not necessarily include all of them. In particular, considering the accompanying drawings, specification, and claims provided in this disclosure, some additional features and advantages will be apparent to those of ordinary skill in the art. In addition, it should be noted that the language used in the specification is mainly selected for readability and guiding purposes, and not necessarily for depicting or limiting the subject matter described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] To be able to understand the present disclosure in more detail, a more specific description may be made with reference to the features of various embodiments, some of which are illustrated in the accompanying drawings. However, the accompanying drawings only show the relevant features of the present disclosure and are not necessarily considered restrictive, because those skilled in the art will understand that the specification may include other valid features when reading this disclosure.

[0013] Figure 1 is a block diagram showing an exemplary communication system according to some embodiments.

[0014] Figure 2A is a block diagram showing exemplary elements of an encoder assembly according to some embodiments.

[0015] Figure 2B is a block diagram showing exemplary elements of a decoder assembly according to some embodiments.

[0016] Figure 3 is a block diagram showing an exemplary server system according to some embodiments.

[0017] Figures 4A - 4D shows an example coding tree structure according to some embodiments.

[0018] Figures 5A - 5C shows example prediction blocks, residual blocks, and reconstruction blocks according to some embodiments.

[0019] Figure 6 shows a prediction block with corner samples according to some embodiments.

[0020] Figure 7A shows an example video decoding process according to some embodiments.

[0021] Figure 7B shows an example video encoding process according to some embodiments.

[0022] According to convention, the various features shown in the accompanying drawings are not necessarily drawn to scale, and the same reference numerals may be used throughout the specification and the drawings to denote the same features. Detailed Implementation Manner

[0023] The present disclosure describes a set of methods for video (image) compression, including methods for applying a transform to a residual block. For example, a method for applying a non-separable transform to an inter-prediction mode block residual. One or more flags may be received in an encoded bitstream, the flags being associated with applying a non-separable transform kernel to an inter-prediction residual block. Applying a non-separable transform to an inter-prediction mode block residual can improve the quality of video encoding / decoding (e.g., improve accuracy and / or precision). Additionally, signaling the application of a non-separable transform on different blocks can improve the encoding / decoding efficiency (e.g., the decoder does not need to derive which transform to use and can apply the most suitable transform for a given block). Example Systems and Devices

[0024] Figure 1 is a block diagram showing a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m), and the source device 102 and the plurality of electronic devices 120 are communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, for example, which is used with video-supporting applications, such as video conferencing applications, digital television applications, media storage, and / or distribution applications.

[0025] The source device 102 includes a video source 104 (e.g., a camera assembly or a media memory) and an encoder component 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates one or more encoded video bitstreams based on the video stream. Compared with the video stream from the video source 104, the video stream from the video source 104 may have a high data volume. Since the data volume of the encoded video bitstream 108 is lower (less data) compared with the video stream from the video source, the encoded video bitstream 108 requires less bandwidth for transmission and less storage space for storage compared with the video stream from the video source 104. In some embodiments, the source device 102 does not include the encoder component 106 (e.g., configured to transmit uncompressed video to the network 110).

[0026] One or more networks 110 represent any number of networks for transmitting information between the source device 102, the server system 112, and / or the electronic devices 120, including, for example, wired (wired) and / or wireless communication networks. One or more networks 110 may exchange data in circuit-switched channels and / or packet-switched channels. Representative networks include telecommunication networks, local area networks, wide area networks, and / or the Internet.

[0027] One or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is a streaming server or includes a streaming server (e.g., configured to store and / or distribute video content, such as an encoded video stream from the source device 102). The server system 112 includes an encoder component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the encoder component 114 includes an encoder component and / or a decoder component. In various embodiments, the encoder component 114 is instantiated as hardware, software, or a combination of hardware and software. In some embodiments, the encoder component 114 is configured to decode the encoded video bitstream 108 and re-encode the video data using different encoding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings based on the encoded video bitstream 108. In some embodiments, the server system 112 serves as a media-aware network element (MANE). For example, the server system 112 may be configured to trim the encoded video bitstream 108 to customize potentially different bitstreams for one or more electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.

[0028] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to generate an output video stream that can be presented on a display or other type of presentation device. In some embodiments, one or more of the electronic devices 120 do not include a display component (e.g., are communicatively coupled to an external display device and / or include a media memory). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.

[0029] The source device and / or the multiple electronic devices 120 are sometimes referred to as "terminal devices" or "user devices". In some embodiments, the source device 102 and / or one or more of the electronic devices 120 are examples of server systems, personal computers, portable devices (e.g., smartphones, tablets, or laptops), wearable devices, video conferencing devices, and / or other types of electronic devices.

[0030] In an exemplary operation of communication system 100, source device 102 transmits encoded video bitstream 108 to server system 112. For example, source device 102 may encode a picture stream captured by the source device. Server system 112 receives encoded video bitstream 108 and may use encoder component 114 to decode and / or encode encoded video bitstream 108. For example, server system 112 may apply encoding to video data, which is more optimal for network transmission and / or storage. Server system 112 may transmit encoded video data 116 (e.g., one or more encoded video bitstreams) to one or more electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display video pictures.

[0031] Figure 2A is a block diagram showing exemplary elements of encoder component 106 according to some embodiments. Encoder component 106 receives video data (e.g., a source video sequence) from video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, encoder component 106 receives a video sequence from a remote video source (e.g., a video source that is a component of a device different from encoder component 106). Video source 104 may provide a source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8 bits, 10 bits, or 12 bits), any color space (e.g., BT.601 Y CrCb, or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0, or Y CrCb 4:4:4). In some embodiments, video source 104 is a storage device that stores previously captured / prepared video. In some embodiments, video source 104 is a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that, when viewed in sequence, are given motion. The pictures themselves may be organized into a spatial pixel array, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those of ordinary skill in the art can readily understand the relationship between pixels and samples.

[0032] The encoder component 106 is configured to encode and / or compress pictures of a source video sequence into an encoded video sequence 216 in real time or under other time constraints required by an application. In some embodiments, the encoder component 106 is configured to perform a conversion between a source video sequence and a bitstream of visual media data (e.g., a video bitstream). Enforcing an appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls and is functionally coupled to other functional units as described below. Parameters set by the controller 204 may include rate control related parameters (e.g., picture skip, quantizer, and / or lambda value of rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those of ordinary skill in the art can readily identify other functions of the controller 204 as these functions may relate to the encoder component 106 optimized for a particular system design.

[0033] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. In a simplified example, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols, such as a symbol stream, based on input pictures to be encoded and reference pictures) and a (local) decoder 210. The decoder 210 reconstructs the symbols to create sample data in a manner similar to how a (remote) decoder creates sample data (when the compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input to the reference picture memory 208. Since the decoding of the symbol stream produces bit-exact results independent of the decoder location (local or remote), the content in the reference picture memory 208 is also bit-exact corresponding between the local encoder and the remote encoder. Thus, the prediction part of the encoder interprets the reference picture samples as the same sample values as the decoder will interpret during decoding when using prediction.

[0034] The operation of the decoder 210 may be the same as that of the remote decoder of the decoder component 122 described in detail below. However, briefly referring to Figure 2B Since the symbols are available and the entropy encoder 214 and the parser 254 can encode / decode the symbols into the encoded video sequence losslessly, the entropy decoding part of the decoder component 122 including the buffer memory 252 and the parser 254 may not be fully implemented in the local decoder 210. Figure 2B

[0035] Except for parsing / entropy decoding, the decoder techniques described herein may exist in corresponding encoders in substantially the same functional form. For this reason, the disclosed subject matter focuses on decoder operations. Additionally, the description of encoder techniques may be simplified as encoder techniques may be inverse to decoder techniques.

[0036] As part of its operation, the source encoder 202 may perform motion compensated predictive coding. Referencing one or more previously encoded frames designated as reference frames in the video sequence, this motion compensated predictive coding predictively encodes the input frame. In this way, the encoding engine 212 encodes the difference between a pixel block of the input frame and a pixel block of the reference frame, which reference picture may be selected as the prediction reference for the input frame. The controller 204 may manage the encoding operations of the source encoder 202, including, for example, setting parameters and subgroup parameters for encoding the video data.

[0037] The decoder 210 decodes the encoded video data of the frame that may be designated as a reference frame, based on the symbols created by the source encoder 202. The operation of the encoding engine 212 may advantageously be a lossy process. When the encoded video data is decoded at the video decoder ( Figure 2A not shown), the reconstructed video sequence may be a replica of the source video sequence with some errors. The decoder 210 replicates the decoding process that may be performed by the remote video decoder on the reference frame and may cause the reconstructed reference frame to be stored in the reference picture memory 208. In this way, the encoder assembly 106 stores a copy of the reconstructed reference frame locally, which copy has common content (absent transmission errors) with the reconstructed reference frame that will be obtained by the remote video decoder.

[0038] The predictor 206 may perform a prediction search for the encoding engine 212. That is, for a new frame to be encoded, the predictor 206 may search the reference picture memory 208 for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may be used as an appropriate prediction reference for the new picture. The predictor 206 may operate on a per pixel block basis of sample blocks to find a suitable prediction reference. As determined by the search results obtained by the predictor 206, the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory 208.

[0039] The outputs of all the above functional units may be entropy encoded in the entropy encoder 214. The entropy encoder 214 performs lossless compression on the symbols generated by the various functional units according to techniques known to those of ordinary skill in the art (e.g., Huffman coding, variable length coding, and / or arithmetic coding), thereby converting the symbols into an encoded video sequence.

[0040] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter may be configured to buffer the encoded video sequence created by the entropy encoder 214 to prepare for transmission over a communication channel 218, which may be a hardware / software link to a storage device that can store the encoded video data. The transmitter may be configured to combine the encoded video data from the source encoder 202 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown). In some embodiments, the transmitter may transmit additional data when transmitting the encoded video. The source encoder 202 may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / signal-to-noise ratio (SNR) enhancement layers, other forms of redundant data such as redundant pictures and slices, supplementary enhancement information (SEI) messages, visual usability information (VUI) parameter set segments, and the like.

[0041] The controller 204 may manage the operation of the encoder components 106. During encoding, the controller 204 may assign a certain encoded picture type to each encoded picture, but this may affect the encoding technique applied to the corresponding picture. For example, a picture may be assigned as an intra picture (I picture), a predictive picture (P picture), or a bi-predictive picture (B picture). An intra picture can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including for example independent decoder refresh (IDR) pictures. Those of ordinary skill in the art are familiar with these variations of I pictures and their corresponding applications and characteristics, and thus will not be described in detail herein. Predictive pictures can be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict the sample values of each block. Bi-predictive pictures can be encoded and decoded using intra prediction or inter prediction, which uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures may use more than two reference pictures and associated metadata for reconstructing a single block.

[0042] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, which are determined by the encoding assignment of the corresponding picture applied to the block. For example, blocks of an I picture can be non-predictively encoded, or blocks of an I picture can be predictively encoded with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be non-predictively encoded with reference to a previously encoded reference picture through spatial prediction or through temporal prediction. Blocks of a B picture can be non-predictively encoded with reference to one or two previously encoded reference pictures through spatial prediction or through temporal prediction.

[0043] The video captured can be a plurality of source pictures (video pictures) in a time series. Intra picture prediction (commonly abbreviated to intra prediction) exploits the spatial correlation within a given picture, while inter picture prediction exploits the (temporal or other) correlation between pictures. In one example, a particular picture being encoded / decoded is divided into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.

[0044] The encoder component 106 can perform encoding operations according to a predetermined video coding technique or standard (such as any of the standards described herein). In operation, the encoder component 106 can perform various compression operations, including predictive coding operations that exploit the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard being used.

[0045] Figure 2B is a block diagram showing exemplary elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 in is coupled to the channel 218 and the display 124. In some embodiments, the decoder component 122 includes a transmitter that is coupled to the loop filter 256 and is configured to transmit data to the display 124 (e.g., via a wired or wireless connection).

[0046] In some embodiments, decoder component 122 includes a receiver that is coupled to channel 218 and configured to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more encoded video sequences to be decoded by decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. Each encoded video sequence may be received from channel 218, which may be a hardware / software link leading to a storage device storing the encoded video data. The receiver may receive encoded video data and other data, such as encoded audio data and / or auxiliary data streams, that may be forwarded to their respective consuming entities (not depicted). The receiver may separate the encoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data when receiving the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by decoder component 122 to decode the data and / or more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0047] According to some embodiments, decoder component 122 includes buffer memory 252, parser 254 (sometimes also referred to as an entropy decoder), scaler / inverse transform unit 258, intra picture prediction unit 262, motion compensation prediction unit 260, aggregator 268, loop filter unit 256, reference picture memory 266, and current picture memory 264. Decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. In some embodiments, decoder component 122 is at least partially implemented in software.

[0048] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to prevent network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 located inside decoder component 122 (e.g., configured to handle playback timing), a separate buffer memory is provided outside decoder component 122 (e.g., to prevent network jitter). When receiving data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous synchronous network, buffer memory 252 may not be needed, or buffer memory 252 may be smaller. For use on a traffic packet network such as the Internet, buffer memory 252 may be needed, buffer memory 252 may be relatively large and / or have an adaptive size, and may be at least partially implemented in an operating system or a similar element external to decoder component 122.

[0049] The parser 254 is configured to reconstruct symbols 270 from the encoded video sequence. These symbols may include, for example, information for managing the operation of the decoder component 122, and / or information for controlling a rendering device such as the display 124. The control information for the rendering device may be in the form of, for example, Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not depicted). The parser 254 parses (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence may be performed according to a video coding technology or standard, and may follow various principles known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 254 may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to a group. The subgroups may include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser 254 may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0050] Depending on the type of the encoded video picture or a portion of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols 270 may involve multiple different units. Which units are involved and the manner of involvement may be controlled by the parser 254 through subgroup control information parsed from the encoded video sequence. For clarity, such subgroup control information flows between the parser 254 and the multiple units below are not depicted.

[0051] The decoder component 122 may be conceptually divided into multiple functional units, and in some embodiments, these units interact closely with each other and may be at least partially integrated with each other. However, for clarity, the conceptually divided functional units are retained here.

[0052] The Scaler / Inverse Transform Unit 258 receives, from the Parser 254, the quantized transform coefficients as symbols 270, as well as control information (e.g., which transform to use, block size, quantization factor, and / or quantization scaling matrix). The Scaler / Inverse Transform Unit 258 may output a block including sample values, which may be input into the Aggregator 268. In some cases, the output samples of the Scaler / Inverse Transform Unit 258 belong to intra-coded blocks, i.e., blocks that do not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the Intra Picture Prediction Unit 262. The Intra Picture Prediction Unit 262 may generate a block having the same size and shape as the block being reconstructed, using the surrounding reconstructed information extracted from the current (partially reconstructed) picture from the Current Picture Memory 264. The Aggregator 268 may add, on a per-sample basis, the predictive information generated by the Intra Picture Prediction Unit 262 to the output sample information provided by the Scaler / Inverse Transform Unit 258.

[0053] In other cases, the output samples of the Scaler / Inverse Transform Unit 258 belong to inter-coded and potentially motion-compensated blocks. In such a case, the Motion Compensation Prediction Unit 260 may access the Reference Picture Memory 266 to extract samples for prediction. After motion compensation of the extracted samples according to the symbols 270 belonging to the block, these samples may be added by the Aggregator 268 to the output of the Scaler / Inverse Transform Unit 258 (in this case, referred to as residual samples or residual signal), thereby generating output sample information. The extraction of the prediction samples by the Motion Compensation Prediction Unit 260 from addresses within the Reference Picture Memory 266 may be controlled by a motion vector. The motion vector may be provided in the form of symbols 270 for use by the Motion Compensation Prediction Unit 260, and the symbols 270 may have, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include interpolation of sample values extracted from the Reference Picture Memory 266, a motion vector prediction mechanism, etc., when using sub-sample accurate motion vectors.

[0054] The output samples of the Aggregator 268 may be subject to various loop filtering techniques in the Loop Filter Unit 256. The video compression technique may include in-loop filter techniques, which are controlled by parameters included in the encoded video bitstream and available for use by the Loop Filter Unit 256 as symbols 270 from the Parser 254, and the video compression technique may also respond to meta-information obtained during the decoding of a previous (in decoding order) portion of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values. The output of the Loop Filter Unit 256 may be a sample stream, which may be output to a rendering device (e.g., the display 124) and stored in the Reference Picture Memory 266 for future inter-picture prediction.

[0055] Once reconstructed, some of the encoded pictures can be used as reference pictures for future prediction. Once an encoded picture is fully reconstructed and the encoded picture (by, for example, the parser 254) is identified as a reference picture, the current reference picture can become part of the reference picture memory 266, and a new current picture memory can be reallocated before starting to reconstruct subsequent encoded pictures.

[0056] The decoder component 122 can perform decoding operations according to a predetermined video compression technique that can be recorded in a standard (e.g., any standard described herein). In the sense that the encoded video sequence follows the syntax of the video compression technique or standard, the encoded video sequence can conform to the syntax specified by the video compression technique or standard used (as specified in the video compression technique document or standard, particularly in the profile of the video compression technique or standard). Additionally, to conform to some video compression techniques or standards, the complexity of the encoded video sequence can be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in mega samples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.

[0057] Figure 3 is a block diagram showing a server system 112 according to some embodiments. The server system 112 includes a control circuit 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuit 302 includes one or more processors (e.g., a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and / or a DPU (Data Processing Units)). In some embodiments, the control circuit includes one or more Field-Programmable Gate Arrays (FPGAs), hardware accelerators, and / or one or more integrated circuits (e.g., application-specific integrated circuits).

[0058] The network interface 304 can be configured to connect to one or more communication networks (e.g., wireless network, wired network, and / or optical network). The communication network can be a local area network, wide area network, metropolitan area network, vehicle and industrial network, real-time network, delay-tolerant network, etc. Examples of communication networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television cable or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Such communication can be only one-way reception (e.g., broadcast television), only one-way transmission (e.g., CANBus connected to certain CANBus devices), or two-way (e.g., connecting to other computer systems using a local area network or wide area digital network). Such communication can include communication to one or more cloud computing networks.

[0059] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input device 310 can include one or more of a keyboard, mouse, touchpad, touch screen, data glove, joystick, microphone, scanner, camera, etc. The output device 308 can include one or more of an audio output device (e.g., speaker), visual output device (e.g., display or monitor), etc.

[0060] The memory 314 can include high-speed random access memory (e.g., DRAM (Dynamic Random Access Memory), SRAM (Static Random Access Memory), DDR RAM (Double Data Rate Random Access Memory), and / or other solid-state random access memory devices) and / or non-volatile memory (e.g., one or more disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). Optionally, the memory 314 includes one or more storage devices arranged remotely from the control circuit 302. The memory 314, or alternatively, the non-volatile solid-state memory device within the memory 314, includes a non-transitory computer-readable storage medium. In some embodiments, the memory 314 or the non-transitory computer-readable storage medium of the memory 314 stores the following programs, modules, instructions, and data structures, or subsets or supersets thereof: · An operating system 316, which includes programs for handling various basic system services and for performing hardware-related tasks; · A network communication module 318 for connecting the server system 112 to other computing devices via one or more network interfaces 304 (e.g., via wired and / or wireless connections); · An encoding / decoding module 320 for performing various functions related to encoding and / or decoding data (e.g., video data). In some embodiments, the encoding / decoding module 320 is an instance of the encoder component 114. The encoding / decoding module 320 includes, but is not limited to, one or more of the following: ○ A decoding module 322 for performing various functions related to decoding encoded data, such as the functions previously described for the decoder component 122; and ○ An encoding module 340 for performing various functions related to encoding data, such as the functions previously described for the encoder component 106; and · A picture memory 352 for storing pictures and picture data, e.g., for use with the encoding / decoding module 320. In some embodiments, the picture memory 352 includes one or more of the reference picture memory 208, the buffer memory 252, the current picture memory 264, and the reference picture memory 266.

[0061] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform the various functions previously described for the parser 254), a transformation module 326 (e.g., configured to perform the various functions previously described for the scaler / inverse transform unit 258), a prediction module 328 (e.g., configured to perform the various functions previously described for the motion compensation prediction unit 260 and / or the intra-picture prediction unit 262), and a filter module 330 (e.g., configured to perform the various functions previously described for the loop filter 256).

[0062] In some embodiments, the encoding module 340 includes a code module 342 (e.g., configured to perform the various functions previously described for the source encoder 202 and / or the encoding engine 212) and a prediction module 344 (e.g., configured to perform the various functions previously described for the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 include Figure 3 A subset of the modules shown. For example, both the decoding module 322 and the encoding module 340 use a shared prediction module.

[0063] Each of the above-identified modules stored in the memory 314 corresponds to an instruction set for performing the functions described herein. The above-identified modules (e.g., instruction sets) need not be implemented as separate software programs, procedures, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, optionally, the codec module 320 does not include separate decoding and encoding modules, but uses the same set of modules to perform both sets of functions. In some embodiments, the memory 314 stores subsets of the above-identified modules and data structures. In some embodiments, the memory 314 stores additional modules and data structures not described above.

[0064] Although Figure 3 FIG. 112 shows a server system 112 according to some embodiments, Figure 3 it is more intended as a functional description of the various features that may exist in one or more server systems than as a structural schematic of the embodiments described herein. In practice, items shown separately may be combined, and some items may be separated. For example, Figure 3 some of the items shown separately in FIG. 112 may be implemented on a single server, and a single item may be implemented by one or more servers. The actual number of servers used to implement the server system 112, and how the features are distributed among the servers, will vary depending on the implementation, and optionally, will depend in part on the amount of data traffic processed by the server system during peak usage periods and during average usage periods. Example Coding and Decoding Techniques

[0065] The codec processes and techniques described below may be performed on the devices and systems described above (e.g., source device 102, server system 112, and / or electronic device 120). According to some embodiments, a method of applying a non-separable transform kernel to an inter-frame mode residual block is described. Hereinafter, a block may refer to a coding tree block, a maximum coding block, a predefined fixed block size, a coding block, a prediction block, a residual block, or a transform block. An inter-frame mode coding block refers to a block using an inter-frame prediction mode. Additionally, a non-separable transform may refer to a primary transform directly applied to a residual, or a secondary transform applied to a transform coefficient block generated by a primary transform. The non-separable transform kernels may be grouped into sets represented by a set index and a kernel index within the set.

[0066] First, block partitioning is introduced. Figures 4A - 4D FIG. 172 shows an example coding tree structure according to some embodiments. As Figure 4A shown in the first coding tree structure (400) in FIG. 172, some coding methods use a four-way partitioning tree starting from the 64×64 level down to the 4×4 level, e.g., with some additional restrictions for blocks of 8×8. In Figure 4AIn it, the partition designated as "R" is recursive because the same partitioning tree repeats at a lower ratio until the lowest level is reached. As Figure 4B shown in the example coding tree structure (402) in Figure 4B , some coding methods extend the partitioning tree to a 10-way structure and increase the maximum size (e.g., sometimes referred to as a superblock) starting from 128×128. The second coding tree structure includes 4:1 / 1:4 rectangular partitions not in the first coding tree structure. Figure 4B The partitioning type with 3 sub-partitions in the second row of Figure 4B is called a T-type partition. In addition to the coding block size, the coding tree depth can also be defined to indicate the splitting depth from the root node.

[0067] As an example, a coding tree unit (CTU) can be split into coding units (CUs) by using a quadtree structure represented as a coding tree to adapt to various local characteristics. In some embodiments, a decision is made at the CU level on whether to use inter-picture (temporal) or intra-picture (spatial) prediction to code a picture region. Depending on the PU partitioning type, each CU can be further split into one, two, or four prediction units (PUs). Inside the PU, the same prediction process can be applied, and relevant information can be sent to the decoder based on the PU. After obtaining the residual block by applying the prediction process based on the PU partitioning type, the CU can be partitioned into transform units (TUs) according to another quadtree structure similar to the coding tree of the CU.

[0068] A quadtree partitioning structure with a nested multi-type tree can be used to replace the concept of multiple partition unit types, and this nested multi-type tree uses binary and ternary splits. In the coding tree structure, a CU can have a square or rectangular shape. The CTU is first partitioned by the quadtree structure. The quadtree leaf nodes can be further partitioned by the multi-type tree structure. As Figure 4C shown in the third coding tree structure (404) in Figure 4C , the multi-type tree structure includes four splitting types. The multi-type tree leaf nodes are called CUs unless the CU is too large for the maximum transform length. This means that CUs, PUs, and TUs can have the same block size in the quadtree coding block structure with a nested multi-type tree. In Figure 4D an example of the block partitioning of a CTU (406) is shown, which shows an example quadtree.

[0069] For example, in VTM7, the coding tree scheme supports the ability for luminance and chrominance to have separate block tree structures. In some cases, for P and B slices, the luminance and chrominance CTBs in a CTU have the same coding tree structure. However, for I slices, luminance and chrominance can have separate block tree structures. When the separate block tree mode is applied, the luminance CTB is divided into CUs by one coding tree structure, and the chrominance CTB is divided into chrominance CUs by another coding tree structure. This means that the CUs in an I slice can include or consist of coded blocks of the luminance component or both chrominance components, and the CUs in a P or B slice can always include or consist of coded blocks of all three color components, unless the video is monochrome.

[0070] Now, the transform and transform block are introduced. Multiple transform sizes (e.g., ranging from 4 points to 64 points for each dimension) and transform shapes (e.g., squares or rectangles with width / height ratios of 2:1 / 1:2 and 4:1 / 1:4) can be utilized. It should be noted that when the encoder component applies a transform, the decoder component performs the inverse transform. Therefore, in the following description, the transform described in the context of the decoder component can be the inverse process of the transform applied on the encoder side.

[0071] The two-dimensional transform process can involve using a hybrid transform kernel (e.g., composed of different one-dimensional transforms for each dimension of the coded residual block). The primary one-dimensional transforms can include at least one of the following: a) 4-point, 8-point, 16-point, 32-point, 64-point discrete cosine transform DCT-2; b) 4-point, 8-point, 16-point asymmetric discrete sine transforms (DST-4, DST-7) and their flipped versions; or c) 4-point, 8-point, 16-point, 32-point identity transform. The basis functions of DCT-2 and asymmetric DST are listed in Table 1, such as the basis functions used in AV1. Table 1 Examples of primary transform basis functions

[0072] The availability of the hybrid transform kernel can be based on the transform block size and the prediction mode. The following Table 2 lists the example dependencies, where "→" and "↓" represent the horizontal and vertical dimensions, and "√" and "×" represent the availability of the kernel for the block size and prediction mode. IDTX (or IDT) represents the identity transform. Table 2 Availability of the hybrid transform kernel

[0073] For the chrominance component, the transform type selection can be performed implicitly. For the intra prediction residual, e.g., as specified in Table 3, the transform type can be selected according to the intra prediction mode. For the inter prediction residual, the transform type can be selected according to the transform type selection of the co-located luma block. Therefore, for the chrominance component, it may not be necessary to transmit the transform type in the bitstream. Table 3 Transform Type Selection for Intra Prediction Residual of Chrominance Intra Prediction Vertical Transformation Horizontal Transformation DC_PRED DCT DCT V_PRED ADST DCT H_PRED DCT ADST D45_PRED DCT DCT D135_PRED ADST ADST D113_PRED ADST DCT D157_PRED DCT ADST D203_PRED DCT ADST D67_PRED ADST DCT SMOOTH_PRED ADST ADST SMOOTH_V_PRED ADST DCT SMOOTH_H_PRED DCT ADST PAETH_PRED ADST ADST

[0074] Now, an example of encoding and decoding using a prediction block and a residual block is introduced. Figure 5A The calculation of a prediction block according to some embodiments is shown. In Figure 5A the example, intra prediction is performed on the current block 502 to generate a prediction block 504. In some embodiments, inter prediction is performed to generate a prediction block. The current block 502 includes a set of samples (e.g., a pixel block), and the prediction block 504 includes a set of predictions corresponding to the set of samples. Figure 5B The calculation of a residual block according to some embodiments is shown. As Figure 5B shown, the prediction block 504 is subtracted from the current block 502 to generate a residual block 506 including a set of residuals. For example, the corresponding difference between each sample and the corresponding prediction is calculated. Figure 5C The calculation of a reconstructed block according to some embodiments is shown. As Figure 5C shown, one or more transforms and quantizations are performed on the residual block 506 to generate a set of residual coefficients. The set of residual coefficients can be sent from the encoder component to the decoder component. The set of residual coefficients is dequantized and inverse-transformed to generate a reconstructed residual block 508. The reconstructed residual block 508 is combined with the prediction block 504 (e.g., the reconstructed residuals of the reconstructed residual block 508 are added to the predictions of the prediction block 504) to generate a reconstructed block 510 corresponding to the current block 502.

[0075] In some embodiments, a separable transform such as that shown in Table 1 is applied to the intra and inter prediction residual samples. In some embodiments, an Intra Secondary Transform (IST) scheme is customized for the video coding library (e.g., for transforming the intra residual block). Compared with the non-separable primary transform, the IST scheme can effectively capture the directional patterns in the intra residual samples with lower complexity. In the IST scheme, the nominal intra prediction angle can be used to classify the IST kernels.

[0076] Intra residual samples may present any directional texture pattern that can be more effectively acquired by non-separable transforms. However, due to its implementation complexity, the use of non-separable transforms for larger block sizes is limited. For non-separable secondary transform schemes that can capture most of the directionality but have lower complexity due to their application only to the low-frequency coefficients of separable primary transforms, they can be applied to larger block sizes with lower complexity.

[0077] In some embodiments, the IST scheme is incorporated into the intra prediction scheme. The IST scheme may include 12 sets of secondary transforms, each set having 3 kernels. Table 4 shows example secondary transform set selections and the corresponding indices for transform set selection. The left column indicates the intra prediction modes with available transform kernels, and the right column rows indicate the set indices. For example, in the encoder, for each mode, the best kernel is selected from the set based on RDO and signaled (4 symbols, excluding IST). In this example, in the decoder, the bitstream is parsed to obtain the kernel used. Table 4 Secondary Transform Set Selection Intra Prediction Set Index DC_PRED 0 V_PRED 1 H_PRED 2 D45_PRED 3 D135_PRED 4 D113_PRED 5 D157_PRED 6 D203_PRED 7 D67_PRED 8 SMOOTH_PRED 9 SMOOTH_V_PRED 10 SMOOTH_H_PRED 11

[0078] In some embodiments, the secondary transform set is derived based on the intra prediction direction, and the kernel type within the set is explicitly transmitted. In some embodiments, IST is enabled when DCT-2 or ADST is used as the horizontal and vertical primary transforms. In some embodiments, IST is enabled only for intra blocks of the luminance frame. For example, depending on the block size, a 4×4 non-separable transform or an 8×8 non-separable transform can be selected. If min(tx_width, tx_height) < 8, a 4×4 IST can be selected. For larger blocks where both tx_width and tx_height are greater than or equal to 8, an 8×8 IST can be used. Here, tx_width and tx_height correspond to the width and height of the transform block, respectively. The input to IST can be the low-frequency primary transform coefficients in zigzag scan order (which can be the default scan order). This helps to achieve more effective decorrelation of adjacent low-frequency coefficients.

[0079] In some embodiments, both intra-coded blocks and inter-coded blocks can be further divided into multiple transform units (e.g., with a division depth of 2 levels). In some embodiments, the application of IST is limited to the root (depth 0) of the transform division tree structure. By this limitation, a reduction in the overall coding time complexity (about 50%) can be achieved while having a minimal impact on the compression efficiency (about 0.25% loss). In some embodiments using the IST scheme, a square transform block size is used to derive context information, and thus context is derived for entropy coding of the kernel index. For rectangular transform blocks, the next smallest square size can be used.

[0080] In some embodiments, the IST scheme defines 14 secondary transform sets, each with 3 kernels. The IST set selection can depend on the intra prediction mode used for residual generation. Table 5 below describes the mapping between the intra prediction mode, the primary transform type, and the IST set index. The first column indicates the intra prediction modes with available kernels, the second column indicates the primary transform type, and the third column indicates the set index. Depending on the block size, 16-point or 64-point IST can be selected. If min(tx_width, tx_height) < 8, 16-point IST can be selected. For larger blocks where both tx_width and tx_height are greater than or equal to 8, 64-point IST can be used. Here, tx_width and tx_height correspond to the transform block width and height respectively. Transform coefficients outside the Region of Application (RoA) of the IST (only primary transform coefficients) can be zeroed out. Table 5 Secondary Transform Set Selection

[0081] Table 5 shows 14 secondary transform sets based on two primary transform types, DCT_DCT and ADST_ADST. Therefore, for one primary transform type, only 7 different sets need to be transmitted. In some embodiments, probability contexts for the selection of each set are derived from the intra prediction mode. In some embodiments, the encoder component implicitly selects the IST set based on a predefined mapping between the intra prediction mode and the IST set. In some embodiments, the encoder performs an additional search over all available IST sets (e.g., instead of only checking one IST set according to the intra prediction mode), such that the encoder can make a rate-distortion optimized decision on the selection of the IST set.

[0082] In some embodiments, IST is enabled for inter-coded blocks. For example, the IST kernels used for intra-coded blocks are also applied to inter-coded blocks. In some embodiments, the IST kernels used for intra-coded blocks are reused for inter-coded blocks without the need for change or addition. In some embodiments, the above description of IST for intra-coded blocks also applies to inter-coded blocks. For example, the encoder can select from multiple sets and multiple kernels within a set (e.g., 7 sets and 3 kernels within a set). The kernels and set indices can be transmitted explicitly.

[0083] In some embodiments, IST is enabled for inter-coded blocks using DCT_DCT or ADST_ADST as the primary transform. For example, for inter blocks, only the kernel index is transmitted, and the set used corresponds to DC_PRED or set index 0.

[0084] Figure 7A is a flowchart showing a method 600 for decoding video according to some embodiments. Method 600 may be executed on a computing system (e.g., server system 112, source device 102, or electronic device 120), which includes control circuitry and a memory storing instructions executed by the control circuitry. In some embodiments, method 600 is executed by executing instructions stored in the memory (e.g., memory 314) of the computing system.

[0085] The system receives (602) a video bitstream that includes a set of inter-prediction mode coded blocks and a corresponding set of transform coefficients. The system derives (604) a set of inter-prediction mode residual blocks from the set of transform coefficients. The system determines (606) whether to apply one or more non-separable transform kernels to the set of inter-prediction mode residual blocks based on the value of a first indicator (e.g., a flag) in the video bitstream. When the first indicator has a first value, the system applies (608) a first non-separable transform kernel to the set of inter-prediction mode residual blocks. When the indicator has a second value, the system abandons (610) applying the first non-separable transform kernel to the set of inter-prediction mode residual blocks. The system reconstructs (612) a set of video blocks using the set of inter-prediction mode residual blocks and a corresponding set of prediction blocks. For example, a flag may be received in the encoded bitstream, where the flag is associated with applying a non-separable transform kernel to an inter-prediction residual block. In some embodiments, the flag is binary and indicates whether to apply a non-separable transform kernel to an inter-prediction mode residual block.

[0086]

[0086] In some embodiments, the flag is transmitted in high-level syntax (e.g., sequence-level flag, picture-level flag, sub-picture-level flag, slice-level flag, tile-level flag) or block-level syntax (e.g., largest coding block-level flag, largest coding block row-level flag, coded block-level flag, or transform block-level flag).

[0087] In some embodiments, the flag may take N values, where N is an integer. For example, when the value of the flag is equal to the first value, the non-separable transform kernel is not applied to the inter-prediction residual block. Otherwise, the non-separable transform is applied to the inter-prediction mode residual block, and the non-separable transform kernel is determined by a kernel index specifying the received flag value.

[0088]

[0087] In some embodiments, when the flag value corresponds to a kernel index, a set index flag is further transmitted. The set index flag may take M values, where M is an integer. In one example, M may be the number of intra-prediction modes. In another example, M may be the number of sets of non-separable transforms (primary transform or secondary transform) applicable to intra-coded blocks. In some embodiments, M is less than (or greater than) the number of sets of non-separable transforms (primary transform or secondary transform) applicable to intra-coded blocks.

[0089] In some embodiments, the set index is binary - ized and encoded using binary symbols. In some embodiments, the flags are binary - ized and encoded using binary codewords. In some embodiments, one of the binary symbols in the binary codeword specifies whether the non - separable transform kernel is applied to the inter - frame prediction residual block.

[0090] In some embodiments, when the flag value corresponds to a kernel index, the set index is derived from the predicted samples of the inter - frame coded block. In some embodiments, all the predicted samples or a subset thereof are used to derive the set index.

[0091] In some embodiments, the method for deriving the set index includes performing a statistical analysis on the spatial direction pattern in the predicted samples. For example, the method may include simple techniques such as edge detection and classification, gradient analysis, and / or more complex techniques (e.g., Histogram of Oriented Gradients (HOG) or neural - network - based analysis).

[0092] In some embodiments, the corner samples of the prediction block (as shown in the shaded area in Figure 6 ) are used to analyze the gradients (denoted as dx, dy) in two dimensions, where the inverse tangent of dy / dx is mapped to the set index. In some embodiments, the boundary samples of the prediction block (e.g., the samples located at the top / bottom / left / right boundaries) are used to analyze the gradients (denoted as dx, dy) in two dimensions, where the inverse tangent of dy / dx is mapped to the set index.

[0093] In some embodiments, edge detection can be implemented by performing multiple filtering operations on the samples and the output, and the output of these filtering operations is used to determine the edge direction. Alternatively, known edge - detection methods can be applied, including but not limited to Canny edge detection.

[0094] In some embodiments, the prediction mode of adjacent blocks is used to derive the kernel index or the set index of the non - separable transform kernel. In some embodiments, the prediction mode includes the intra - frame / inter - frame prediction mode, the intra - frame prediction angle, and / or the segmentation angle of the geometric partition.

[0095] In some embodiments, when the flag value corresponds to a kernel index, a default set index is used. In some embodiments, the flags are transmitted through block - level syntax (e.g., the maximum coded block - level flag, the maximum coded block row - level flag, the coded block - level flag, the transform block - level flag).

[0096] In some embodiments, the non-separable transform kernels (or sets of non-separable transform kernels) used on intra-coded blocks and inter-coded blocks may overlap with each other. For example, some non-separable transform kernels used on intra-coded blocks are also applicable to inter-coded blocks, and some non-separable transform kernels used on intra- (or inter-) coded blocks are also applicable to inter- (or intra-) coded blocks.

[0097] In some embodiments, the non-separable transform kernels (or sets of non-separable transform kernels) used on inter-coded blocks are a subset of the non-separable transform kernels (or sets of non-separable transform kernels) used on intra-mode coded blocks.

[0098] In some embodiments, whether the non-separable transform kernels used in intra-coded blocks are also used for inter-mode coded blocks depends on the inter prediction mode (and / or other coding information).

[0099] Figure 7B FIG. 650 is a flowchart showing a method 650 for encoding a video according to some embodiments. Method 650 may be executed on a computing system (e.g., server system 112, source device 102, or electronic device 120). The computing system includes a control circuit and a memory storing instructions executed by the control circuit. In some embodiments, method 650 is executed by executing instructions stored in the memory (e.g., memory 314) of the computing system.

[0100] The system receives (652) video data including a set of video blocks. The system derives (654) a set of inter-mode residual blocks from the set of video blocks. The system determines (656) whether to apply one or more non-separable transform kernels to the set of inter-mode residual blocks. The system generates (658) a set of transform coefficients from the set of inter-mode residual blocks based on the determination of whether to apply one or more non-separable transform kernels to the set of inter-mode residual blocks. The system determines (660) the value of a first indicator based on whether to apply one or more non-separable transform kernels to the set of inter-mode residual blocks. The system transmits (662) the first indicator in the video bitstream. The system transmits (664) the set of transform coefficients in the video bitstream. As previously described, the encoding process may mirror the decoding process described herein (e.g., performing non-separable transforms). For the sake of brevity, these details are not repeated here.

[0101] Although Figure 7A and 7BA number of logical stages are shown in a particular order, but stages that are not order - dependent can be reordered, and other stages can be combined or split. Some reorderings or other groupings not specifically mentioned will be apparent to those of ordinary skill in the art, so the orderings and groupings presented herein are not exhaustive. Additionally, it should be recognized that these stages can be implemented in hardware, firmware, software, or any combination thereof.

[0102] Some example embodiments are described below:

[0103] (A1) In one aspect, some embodiments include a video decoding method (e.g., method 600). In some embodiments, the method is executed on a computing system (e.g., server system 112) including a memory and control circuitry. In some embodiments, the method is executed on an encoding module (e.g., codec module 320). In some embodiments, the method is executed on a source encoding component (e.g., source encoder 202), an encoding engine (e.g., encoding engine 212), and / or an entropy encoder (e.g., entropy encoder 214). The method includes (i) receiving a video bitstream that includes a set of inter - frame mode - encoded blocks and a corresponding set of transform coefficients; (ii) deriving a set of inter - frame mode residual blocks from the set of transform coefficients; (iii) determining whether to apply one or more non - separable transform kernels to the set of inter - frame mode residual blocks according to the value of a first indicator in the video bitstream; (iv) when the indicator has a first value, applying a first non - separable transform kernel to the set of inter - frame mode residual blocks; (v) when the indicator has a second value, forgoing applying the first non - separable transform kernel to the set of inter - frame mode residual blocks; and (vi) using the set of inter - frame mode residual blocks and a corresponding set of prediction blocks to reconstruct a set of video blocks. For example, one or more flags in an encoded bitstream can be received, where each flag is associated with applying a non - separable transform kernel on an inter - frame prediction residual block. Any of the above - mentioned sets can include one or more components. In some embodiments, the first non - separable transform kernel is applied to the set of residual blocks according to determining that the indicator has the first value, and the first non - separable transform kernel is not applied to the set of residual blocks according to determining that the indicator has the second value.

[0104] (A2) In some embodiments of A1, the first indicator is encoded in the video bitstream using binary coding. The method includes identifying the value of the first indicator by decoding the binary representation of the first indicator. For example, a binary codeword can be used to binary - encode and encode the first indicator (e.g., a flag).

[0105] (A3) In some embodiments of A2, the first indicator includes a binarized codeword, and the binary symbols of the binarized codeword indicate whether any non-separable transform kernel is applied to the set of inter-frame mode residual blocks. For example, one of the binary symbols in the binary codeword specifies whether the non-separable transform kernel is applied to the inter-frame prediction residual block.

[0106] (A4) In some embodiments of any one of A1 - A3, the value of the first indicator is selected from a set of N values, where N is a positive integer. The method includes using the value of the first indicator to identify a non-separable transform kernel from a set of non-separable transform kernels. For example, the flag can take N values. When the value of the flag is equal to the first value, the non-separable transform kernel is not applied to the inter-frame prediction residual block. Otherwise, the non-separable transform is applied to the inter-frame mode residual block, and the non-separable transform kernel is determined by the kernel index specifying the received flag value. For example, binary symbols can be used to binaryize and encode the set index.

[0107] (A5) In some embodiments of A4, the method includes identifying a set of non-separable transform kernels from a plurality of sets of non-separable transform kernels based on one or more prediction samples of the set of video blocks. For example, when the flag value corresponds to the kernel index, the set index is derived from the prediction samples of the inter-frame coded block. For example, all prediction samples or a subset thereof are used to derive the set index.

[0108] (A6) In some embodiments of A5, one or more prediction samples include corner samples from the corresponding set of prediction blocks. For example, the corner samples of the prediction block (as Figure 6 shown) are used to analyze the gradients (denoted as dx, dy) in two dimensions, where the inverse tangent of dy / dx is mapped to the set index.

[0109] (A7) In some embodiments of A5 or A6, one or more prediction samples include boundary samples from the corresponding set of prediction blocks. For example, the boundary samples of the prediction block (e.g., samples located at the top / bottom / left / right boundaries) are used to analyze the gradients (denoted as dx, dy) in two dimensions, where the inverse tangent of dy / dx is mapped to the set index.

[0110] (A8) In some embodiments of any one of A5 - A7, identifying a set of non-separable transform kernels based on one or more prediction samples includes using a statistical analysis of the spatial direction pattern in one or more prediction samples to identify the set of non-separable transform kernels. For example, the method for deriving the set index includes a statistical analysis of the spatial direction pattern in the prediction samples. As an example, such a method can include edge detection and classification, gradient analysis, and / or more complex techniques such as Histogram of Oriented Gradients (HOG) or neural network-based analysis.

[0111] (A9)In some embodiments of A8, the statistical analysis includes applying edge detection techniques. For example, edge detection can be achieved by performing multiple filtering operations on the samples and the output, and the output of these filtering operations is used to determine the edge direction. Alternatively, known edge detection methods can be applied, including but not limited to Canny edge detection.

[0112] (A10)In some embodiments of any one of A5 - A9, an inseparable transform kernel set is identified based on prediction mode information of adjacent blocks from a set of video blocks. As an example, the prediction mode of adjacent blocks can be used to derive the kernel index or set index of the inseparable transform kernel. For example, the prediction mode information can include intra / inter prediction mode, intra prediction angle, and / or segmentation angle of geometric partitioning.

[0113] (A11)In some embodiments of any one of A4 - A10, the method includes identifying an inseparable transform kernel set from a plurality of inseparable transform kernel sets according to the value of a second indicator in a video bitstream, where the value of the second indicator is selected from a set of M values, and M is a positive integer. For example, when the first indicator (e.g., flag value) corresponds to a kernel index, a set index flag is further transmitted. The set index flag can take M values.

[0114] (A12)In some embodiments of A11, M is based on the number of intra prediction modes of the video bitstream. For example, M can be, but is not limited to, the number of intra prediction modes.

[0115] (A13)In some embodiments of A11 or A12, M is based on the number of inseparable transform sets applicable to intra-mode coded blocks. For example, M can be, but is not limited to, the number of inseparable transform (primary transform or secondary transform) sets applicable to intra-coded blocks.

[0116] (A14)In some embodiments of A13, M is less than the number of inseparable transform sets applicable to intra-mode coded blocks. For example, M can be less than (or greater than) the number of inseparable transform (primary transform or secondary transform) sets applicable to intra-coded blocks.

[0117] (A15)In some embodiments of any one of A4 - A14, the method includes identifying an inseparable transform kernel set from a plurality of inseparable transform kernel sets according to a default set index. For example, when the flag value corresponds to a kernel index, the default set index is used.

[0118] (A16)In some embodiments of any one of A4 - A15, the non - separable transform kernel set includes one or more non - separable transform kernels for both inter - frame mode residual blocks and intra - frame mode residual blocks. For example, the non - separable transform kernel (or non - separable transform kernel set) used on intra - frame coded blocks and inter - frame coded blocks can overlap with each other. That is, some non - separable transform kernels used on intra - frame coded blocks are also applicable to inter - frame coded blocks, and some non - separable transform kernels used on intra - frame (or inter - frame) coded blocks are also applicable to inter - frame (or intra - frame) coded blocks. In some embodiments, the non - separable transform kernel (or non - separable transform kernel set) used on inter - frame coded blocks is a subset of the non - separable transform kernel (or non - separable transform kernel set) used on intra - frame mode coded blocks. In some embodiments, whether the non - separable transform kernel used in intra - frame coded blocks is also used for inter - frame mode coded blocks depends on the inter - frame prediction mode (and / or other coding information).

[0119] (A17)In some embodiments of any one of A1 - A16, the first indicator is a binary indicator indicating whether any non - separable transform kernel is applied to the set of inter - frame mode residual blocks. For example, the flag is binary and indicates whether a non - separable transform kernel is applied to the inter - frame mode residual blocks.

[0120] (A18)In some embodiments of any one of A1 - A17, the first indicator is transmitted in the high - level syntax of the video bitstream. For example, the first indicator (e.g., the flag) is transmitted in the high - level syntax (e.g., sequence - level flag, picture - level flag, sub - picture - level flag, slice - level flag, tile - level flag) or block - level syntax (e.g., maximum coded block - level flag, maximum coded block row - level flag, coded block - level flag, transform block - level flag). In some embodiments, the flag is transmitted in the block - level syntax (e.g., maximum coded block - level flag, maximum coded block row - level flag, coded block - level flag, or transform block - level flag).

[0121] (B1)On the other hand, some embodiments include a video coding method (e.g., method 650). In some embodiments, the method is executed on a computing system that includes a memory and one or more processors. The method includes: (i) receiving video data including a set of video blocks; (ii) deriving a set of inter - frame mode residual blocks from the set of video blocks; (iii) determining whether one or more non - separable transform kernels are applied to the set of inter - frame mode residual blocks; (iv) generating a set of transform coefficients from the set of inter - frame mode residual blocks according to whether one or more non - separable transform kernels are applied to the set of inter - frame mode residual blocks; (v) determining the value of a first indicator according to whether one or more non - separable transform kernels are applied to the set of inter - frame mode residual blocks; (vi) transmitting the first indicator in the video bitstream; and (vii) transmitting the set of transform coefficients in the video bitstream.

[0122] (B2) In some embodiments of B1, transmitting the first indicator includes transmitting a binary codeword.

[0123] (B3) In some embodiments of B1 or B2, the method includes determining a kernel index of one or more non-separable transform kernels and assigning a value to the first indicator based on the kernel index.

[0124] (B4) In some embodiments of any one of B1 - B3, the method includes determining a value of a second indicator based on a set index of a set of transform kernels, the set of transform kernels including one or more non-separable transform kernels, and transmitting the second indicator in a video bitstream.

[0125] (C1) In another aspect, some embodiments include a method for processing visual media data. In some embodiments, the method is performed on a computing system, the computing system including a memory and one or more processors. The method includes: (i) obtaining a source video sequence including a plurality of pictures; and (ii) performing a conversion between the source video sequence and a bitstream of visual media data, where the bitstream includes: (a) a plurality of encoded blocks corresponding to the plurality of pictures; (b) a set of transform coefficients corresponding to the plurality of encoded blocks; and (c) a first indicator that indicates whether one or more non-separable transform kernels are applied to a set of inter-frame mode residual blocks corresponding to the set of transform coefficients.

[0126] In another aspect, some embodiments include a computing system (e.g., server system 112), the computing system including control circuitry (e.g., control circuitry 302) and a memory (e.g., memory 314) coupled to the control circuitry, the memory storing one or more sets of instructions configured to be executed by the control circuitry, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 - A18, B1 - B4, and C1 above).

[0127] In yet another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more sets of instructions executed by control circuitry of a computing system, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 - A18, B1 - B4, and C1 above).

[0128] Unless otherwise specified, any syntax element described herein may be High-Level Syntax (HLS). As used herein, HLS is transmitted at a level higher than the block level. For example, HLS may correspond to the sequence level, frame level, slice level, or tile level. As another example, HLS elements may be transmitted in a Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Adaptation Parameter Set (APS), slice header, picture header, tile header, and / or CTU header.

[0129] It should be understood that although the terms "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to any and all possible combinations of one or more of the listed related items and includes any and all possible combinations of one or more of the listed related items. It should be further understood that when used in this specification, the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0130] As used herein, depending on the context, the term "when" may be interpreted to mean "if" or "after" or "in response to determining..." or "in accordance with determining..." or "in response to detecting" that the prerequisite is true. Similarly, depending on the context, the phrase "if it is determined (that the prerequisite is true)" or "if (the prerequisite is true)" or "when (the prerequisite is true)" may be interpreted to mean "after determining that the prerequisite is true" or "in response to determining that the prerequisite is true" or "in accordance with determining that the prerequisite is true" or "after detecting that the prerequisite is true" or "in response to detecting that the prerequisite is true". As used herein, N refers to a variable. Unless otherwise specified, different instances of N may refer to the same number (e.g., the same integer value, such as the number 2) or different numbers.

[0131] For purposes of explanation, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the principles of operation and practical application, thereby enabling others skilled in the art to implement.

Claims

1. A video decoding method performed on a computing system, the computing system including a memory and one or more processors, the method comprising: Receiving a video bitstream, the video bitstream including a set of inter-prediction mode coded blocks and a corresponding set of transform coefficients; Deriving a set of inter-prediction mode residual blocks from the set of transform coefficients; Determining whether to apply one or more non-separable transform kernels to the set of inter-prediction mode residual blocks according to the value of a first indicator in the video bitstream; When the indicator has a first value, applying a first non-separable transform kernel to the set of inter-prediction mode residual blocks; When the indicator has a second value, forgoing applying the first non-separable transform kernel to the set of inter-prediction mode residual blocks; And Using the set of inter-prediction mode residual blocks and a corresponding set of prediction blocks to reconstruct a set of video blocks.

2. The method according to claim 1, wherein, Encoding the first indicator in the video bitstream using binary coding; and The method further includes identifying the value of the first indicator by binary decoding for the first indicator.

3. The method according to claim 2, wherein The first indicator includes a binarized codeword, and the binary symbols of the binarized codeword indicate whether to apply any non-separable transform kernels to the set of inter-prediction mode residual blocks.

4. The method according to claim 1, wherein, The value of the first indicator is selected from a set of N values, where N is a positive integer; and The method further includes identifying a non-separable transform kernel from a set of non-separable transform kernels using the value of the first indicator.

5. The method according to claim 4, further comprising identifying the set of non-separable transform kernels from a plurality of sets of non-separable transform kernels based on one or more prediction samples of the set of video blocks.

6. The method according to claim 5, wherein, The one or more prediction samples include corner samples from the corresponding set of prediction blocks.

7. The method according to claim 5, wherein, The one or more prediction samples include boundary samples from the corresponding set of prediction blocks.

8. The method according to claim 5, wherein, Identifying the set of non-separable transform kernels based on the one or more prediction samples includes using a statistical analysis of the spatial direction patterns in the one or more prediction samples to identify the set of non-separable transform kernels.

9. The method according to claim 8, wherein The statistical analysis includes applying edge detection techniques.

10. The method according to claim 5, wherein, Identifying the set of non-separable transform kernels based on prediction mode information of adjacent blocks from the set of video blocks.

11. The method according to claim 4 further includes identifying the set of non-separable transform kernels from a plurality of sets of non-separable transform kernels according to the value of a second indicator in the video bitstream, wherein, The value of the second indicator is selected from a set of M values, where M is a positive integer.

12. The method according to claim 11, wherein, M is based on the number of intra-prediction modes of the video bitstream.

13. The method according to claim 11, wherein, M is based on the number of sets of non-separable transforms applicable to intra-mode coded blocks.

14. The method according to claim 13, wherein, M is less than the number of sets of non-separable transforms applicable to intra-mode coded blocks.

15. The method according to claim 4, further comprising identifying the set of non-separable transform kernels from a plurality of sets of non-separable transform kernels according to a default set index.

16. The method according to claim 4, wherein The set of non-separable transform kernels includes one or more non-separable transform kernels for both inter-prediction mode residual blocks and intra-prediction mode residual blocks.

17. The method according to claim 1, wherein The first indicator is a binary indicator indicating whether to apply any non-separable transform kernels to the set of inter-prediction mode residual blocks.

18. The method according to claim 1, wherein The first indicator is transmitted in the high-level syntax of the video bitstream.

19. A computing system, comprising: A control circuit; A memory; And One or more instruction sets stored in the memory and configured to be executed by the control circuit, the one or more instruction sets including instructions for the following operations: Receiving video data including a set of video blocks; Deriving a set of inter-mode residual blocks from the set of video blocks; Determining whether to apply one or more non-separable transform kernels to the set of inter-mode residual blocks; Generating a set of transform coefficients from the set of inter-mode residual blocks according to determining whether to apply the one or more non-separable transform kernels to the set of inter-mode residual blocks; Determining a value of a first indicator according to whether to apply the one or more non-separable transform kernels to the set of inter-mode residual blocks; Transmitting the first indicator in a video bitstream; And Transmitting the set of transform coefficients in the video bitstream.

20. A non-transitory computer-readable storage medium storing one or more instruction sets, the one or more instruction sets being configured to be executed by a computing device having a control circuit and a memory, the one or more instruction sets including instructions for the following operations: Obtain a source video sequence including multiple pictures; And Performing a conversion between the source video sequence and a bitstream of visual media data, wherein, The bitstream includes: A plurality of encoded blocks corresponding to the plurality of pictures; A set of transform coefficients corresponding to the plurality of encoded blocks; and A first indicator indicating whether to apply one or more non-separable transform kernels to a set of inter-mode residual blocks corresponding to the set of transform coefficients.