Systems and methods for partition-dependent quadratic transformations - Patents.com

JP2025509012A5Pending Publication Date: 2026-03-13TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-03-07
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently compressing video data while maintaining video quality, particularly due to limitations in bandwidth capacity and memory resources.

Method used

The implementation of advanced video coding techniques, specifically split-dependent quadratic conversion methods, which involve receiving a control flag from the video stream to determine if the interprediction mode is enabled, identifying transform units, applying a secondary transformation based on the relative position of the units, and reconstructing the video block accordingly.

Benefits of technology

This approach enhances video coding efficiency by allowing for adaptive secondary transformations within video blocks, leading to improved compression performance and reduced bitrate without significant degradation in video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Various embodiments described herein include a method and system for coding video. In one aspect, the method includes: determining whether a plurality of transform units are in a video block according to a determination that an inter-prediction mode is enabled; determining a transform unit of the plurality of transform units to apply a secondary transform based on a relative position of the transform unit in the video block according to a determination that the plurality of transform units are in the video block; applying the secondary transform to the transform unit; and reconstructing / processing the video block based at least on the secondary transform.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 318,262, entitled "Partition Dependent Secondary Transform," filed March 9, 2022, and is a continuation of and claims priority to U.S. Provisional Patent Application No. 18 / 117,218, entitled "Systems and Methods for Partition Dependent Secondary Transform," filed March 3, 2023, the entire contents of which are incorporated herein by reference.

[0002] TECHNICAL FIELD Embodiments of the present disclosure relate generally, but not exclusively, to video coding, including systems and methods for partition-dependent secondary transformations. [Background technology]

[0003] Digital video is supported by a variety of electronic devices, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. The electronic devices transmit, receive, or communicate digital video data over communication networks and / or store the digital video data in storage devices. Because communication networks have limited bandwidth capacity and storage devices have limited memory resources, video coding may be used to compress the video data according to one or more video coding standards before the video data is communicated or stored.

[0004] Several video codec standards have been developed. For example, video coding standards include AOMedia Video 1 (AV1), Versatile Video Coding (VVC), Joint Exploration test Model (JEM), High-Efficiency Video Coding (HEVC / H.265), Advanced Video Coding (AVC / H.264), and Moving Picture Expert Group (MPEG) coding. Video coding generally utilizes prediction methods (e.g., inter-prediction, intra-prediction, etc.) that exploit the redundancy inherent in video data. Video coding aims to compress video data into a format that uses a lower bitrate while avoiding or minimizing degradation of video quality.

[0005] HEVC, also known as H.265, is a video compression standard designed as part of the MPEG-H project. ITU-T and ISO / IEC published the HEVC / H.265 standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Versatile Video Coding (VVC), also known as H.266, is a video compression standard intended as a successor to HEVC. ITU-T and ISO / IEC published the VVC / H.266 standard in 2020 (version 1), and 2022 (version 2). AV1 is an open video coding format designed as a replacement for HEVC. The effective version 1.0.0 of this specification was released on January 8, 2019, containing Errata 1. Summary of the Invention [Means for solving the problem]

[0006] This disclosure describes advanced video coding techniques, and more specifically, partition-dependent secondary transformation methods.

[0007] According to some embodiments, a method of video coding is provided, which includes: receiving a first control flag from a video stream / data, where the first control flag indicates whether an inter-prediction mode is enabled for a video block of the video stream / data; determining whether a plurality of transform units are in the video block according to a determination that the inter-prediction mode is enabled; determining a transform unit of the plurality of transform units to apply a secondary transform based on a relative position of the transform unit in the video block according to a determination that the plurality of transform units are in the video block; applying the secondary transform to the transform unit; and reconstructing / processing the video block based at least on the secondary transform.

[0008] According to some embodiments, a computer system, such as a streaming system, a server system, a personal computer system, or other electronic device, is provided. The computer system includes control circuitry and a memory that stores one or more sets of instructions. The one or more sets of instructions include instructions for performing any of the methods described herein. In some embodiments, the computer system includes an encoder component and / or a decoder component.

[0009] According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more sets of instructions for execution by a computer system. The one or more sets of instructions include instructions for performing any of the methods described herein.

[0010] Accordingly, disclosed are devices and systems having methods for coding video, which may complement or replace conventional methods, devices, and systems for video coding.

[0011] The features and advantages described herein are not necessarily all-inclusive, and many additional features and advantages will be apparent to those of ordinary skill in the art, especially in view of the drawings, specification, and claims provided in this disclosure.Furthermore, it should be noted that the language used herein has been selected primarily for ease of reading and for explanatory purposes, and not to precisely describe or limit the subject matter described herein.

[0012] In order that the present disclosure may be more fully understood, a more detailed description may be made by reference to the features of various embodiments, some of which are illustrated in the accompanying drawings. However, the accompanying drawings merely illustrate the relevant features of the present disclosure, and therefore the description should not be considered necessarily limiting, since other useful features may be recognized as understood by those skilled in the art upon reading the present disclosure. [Brief description of the drawings]

[0013] [Figure 1] 1 is a block diagram illustrating an exemplary communication system, according to some embodiments. [Figure 2A] 2 is a block diagram illustrating exemplary elements of an encoder component according to some embodiments. [Figure 2B] 3 is a block diagram illustrating exemplary elements of a decoder component according to some embodiments. [Diagram 3] FIG. 1 is a block diagram illustrating an exemplary server system, according to some embodiments. [Figure 4] FIG. 2 illustrates an example of a linear transformation basis function, according to some embodiments. [Figure 5A] 1 is a diagram of a table of example dependencies of availability of various transform kernels based on transform block size and prediction mode, in accordance with some embodiments. [Figure 5B] 1 is a diagram of a table of example transform type selection based on intra-prediction mode, according to some embodiments. [Figure 6]FIG. 1 is a diagram of an exemplary use of intra-quadratic transform (IST) in the encoding and decoding process according to some embodiments. [Figure 7] FIG. 13 is a table of example mappings from intra-nominal modes and primary transform types to an IST set in accordance with some embodiments. [Figure 8] A diagram illustrating an exemplary squared error distribution within a 4x4 prediction unit (PU) in accordance with some embodiments. [Figure 9] A diagram showing an example split PU and error distribution within an upper-left transform unit (TU) in accordance with some embodiments. [Figure 10] FIG. 1 illustrates an example mapping from boundary types to transformations by using DST-VII, according to some embodiments. [Figure 11] FIG. 1 illustrates an example mapping from boundary types to transforms by using DCT-IV, according to some embodiments. [Figure 12] 10 illustrates an example transformation used for the TU of FIG. 9 in accordance with some embodiments. [Figure 13] A diagram showing an example of a TU in a coding block according to some embodiments. [Figure 14] 1 is an exemplary flow diagram illustrating a method for coding video according to some embodiments. [Figure 15] 1 is an exemplary flow diagram illustrating a method for coding video according to some embodiments. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0014] According to common practice, the various features illustrated in the drawings are not necessarily drawn to scale and like reference numerals are used throughout the specification and drawings to denote like features.

[0015] 1 is a block diagram illustrating a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 through 120-m) communicatively coupled to one another via one or more networks. In some embodiments, the communication system 100 is a streaming system for use in video-enabled applications, such as, for example, video conferencing applications, digital TV applications, and media storage and / or distribution applications.

[0016] The source device 102 includes a video source 104 (e.g., a camera component or a media storage device) and an encoder component 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to generate an uncompressed video sample stream). The encoder component 106 generates one or more encoded video bitstreams from the video stream. The video stream from the video source 104 may be of a high data volume compared to the encoded video bitstream 108 generated by the encoder component 106. Because the encoded video bitstream 108 has a low data volume (less data) compared to the video stream from the video source, the encoded video bitstream 108 requires less bandwidth to transmit and less storage space to store compared to the video stream from the video source 104. In some embodiments, the source device 102 does not include the encoder component 106 (e.g., configured to transmit uncompressed video data to the network 110).

[0017] The one or more networks 110 represent any number of networks that convey information between the source device 102, the server system 112, and / or the electronic device 120, including, for example, wireline and / or wireless communication networks. The one or more networks 110 may exchange data over circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0018] The one or more networks 110 include a server system 112 (e.g., a distributed / cloud computer system). In some embodiments, the server system 112 is or includes a streaming server (e.g., configured to store and / or distribute video content, such as an encoded video stream from the source device 102). The server system 112 includes a coder component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the coder component 114 includes an encoder component and / or a decoder component. In various embodiments, the coder component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, the coder component 114 is configured to decode the encoded video bitstream 108 and re-encode the video data using a different encoding standard and / or method to generate the encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings from the encoded video bitstream 108.

[0019] In some embodiments, server system 112 functions as a media aware network element (MANE). For example, server system 112 may be configured to prune encoded video bitstream 108 to tailor potentially different bitstreams to one or more of electronic devices 120. In some embodiments, a MANE is provided separate from server system 112.

[0020] Electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, decoder component 122 is configured to decode encoded video data 116 to generate an output video stream that may be rendered on a display or other type of rendering device. In some embodiments, one or more of electronic devices 120 do not include a display component (e.g., communicatively coupled to an external display device and / or includes a media storage device). In some embodiments, electronic device 120 is a streaming client. In some embodiments, electronic device 120 is configured to access server system 112 to obtain encoded video data 116.

[0021] The source device and / or the electronic devices 120 may also be referred to as “terminal devices,” or “user devices.” In some embodiments, the source device 102 and / or one or more of the electronic devices 120 are instances of a server system, a personal computer, a portable device (e.g., a smartphone, tablet, or laptop), a wearable device, a videoconferencing device, and / or other types of electronic devices.

[0022] In an exemplary operation of the communication system 100, the source device 102 transmits an encoded video bitstream 108 to the server system 112. For example, the source device 102 may code a stream of pictures captured by the source device. The server system 112 may receive the encoded video bitstream 108 and decode and / or encode the encoded video bitstream 108 using a coder component 114. For example, the server system 112 may apply an encoding to the video data that is optimal for network transmission and / or storage. The server system 112 may transmit the encoded video data 116 (e.g., one or more encoded video bitstreams) to one or more of the electronic devices 120. Each electronic device 120 may decode the encoded video data 116 to recover the video pictures and optionally display them.

[0023] In some embodiments, the above-mentioned transmission is a unidirectional data transmission. The unidirectional data transmission may be utilized in media serving applications, etc. In some embodiments, the above-mentioned transmission is a bidirectional data transmission. The bidirectional data transmission may be utilized in video conferencing applications, etc. In some embodiments, the encoded video bitstream 108 and / or the encoded video data 116 are encoded and / or decoded according to any of the video coding / compression standards described herein, such as HEVC, VVC, and / or AV1.

[0024] FIG. 2A is a block diagram illustrating exemplary elements of the encoder component 106, according to some embodiments. The encoder component 106 receives a source video sequence from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., a transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives a video sequence from a remote video source (e.g., a video source that is a component of a device different from the encoder component 106). The video source 104 may provide the source video sequence in the form of a digital video sample stream that may be of any suitable bit depth (e.g., 8-bit, 10-bit, or 12-bit), any color space (e.g., BT.601 Y CrCB, or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, the video source 104 is a storage device that stores previously captured / prepared video. In some embodiments, the video source 104 is a camera that captures local image information as a video sequence. Video data may be provided as a number of separate pictures which, when viewed in sequence, convey motion. The pictures themselves are organized as a spatial array of pixels, each of which may contain one or more samples depending on the sampling structure, color space, etc., in use. Those skilled in the art will readily appreciate the relationship between pixels and samples. The following description focuses on samples.

[0025] The encoder component 106 is configured to code and / or compress pictures of a source video sequence into a coded video sequence 216 in real time or under any other time constraint required by the application. Enforcing an appropriate coding rate is one function of the controller 204. In some embodiments, the controller 204 controls and is operatively coupled to other functional units described below. Parameters set by the controller 204 may include rate control related parameters (picture skip, quantizer, and / or lambda value for rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art may readily identify other functions of the controller 204 that may be associated with the encoder component 106 being optimized for a particular system design.

[0026] In some embodiments, the encoder component 106 is configured to operate in a coding loop. In a simplified example, the coding loop includes a source coder 202 (e.g., responsible for generating symbols, such as a symbol stream, based on an input picture to be coded and a reference picture) and a (local) decoder 210. The decoder 210 reconstructs the symbols to generate sample data, similar to a (remote) decoder (when the compression between the symbols and the coded video bitstream is lossless). The reconstructed sample stream (sample data) is input to a reference picture memory 208. Since the decoding of the symbol stream results in a bit-exact result regardless of the location of the decoder (local or remote), the contents of the reference picture memory 208 are also bit-exact between the local and remote encoders. In other words, the predictive part of the encoder reads as reference picture samples the same sample values ​​that the decoder would read when using prediction during decoding. This basic principle of reference picture synchrony (and the resulting drift if synchrony cannot be maintained, for example due to channel errors) is known to those skilled in the art.

[0027] The operation of the decoder 210 may be the same as that of a remote decoder, such as the decoder component 122, described in more detail below in connection with Figure 2B. However, with brief reference to Figure 2B, because symbols are available and the encoding / decoding of the symbols into a coded video sequence by the entropy coder 214 and parser 254 may be lossless, the entropy decoding portion of the decoder component 122, including the buffer memory 252 and the parser 254, may not be fully implemented in the local decoder 210.

[0028] At this point, it can be said that any decoder technology, except for parsing / entropy decoding, present in the decoder must necessarily also be present in the corresponding encoder in substantially the same functional form. For this reason, the subject matter of this disclosure focuses on the operation of the decoder. A description of the encoder technology may be omitted, since the encoder technology is the inverse of the decoder technology, which is described generically. Only in certain areas is a detailed description necessary, which is given below.

[0029] As part of its operation, the source coder 202 performs motion-compensated predictive coding, which predictively codes an input frame with reference to one or more previously coded frames from the video sequence, designated as "reference frames." In this method, the coding engine 212 codes differences between pixel blocks of the input frame and pixel blocks of the reference frames that may be selected as predictive references to the input frame. The controller 204 may manage the coding operations of the source coder 202, including, for example, setting parameters and subgroup parameters used to encode the video data.

[0030] The decoder 210 decodes the coded video data of a frame that may be designated as a reference frame based on the symbols generated by the source coder 202. The operation of the coding engine 212 may advantageously be a lossy process. When the coded video data is decoded in a video decoder (not shown in FIG. 2A), the reconstructed video sequence may be a replica of the source video sequence with some errors. The decoder 210 may reproduce the decoding process that may be performed by a remote video decoder on the reference frame and store the reconstructed reference frame in the reference picture memory 208. In this way, the encoder component 106 may locally store a copy of the reconstructed reference frame that has common content as the reconstructed reference frame that will be obtained by the remote video decoder (without transmission errors).

[0031] The predictor 206 may perform a predictive search for the coding engine 212. That is, for a new frame to be coded, the predictor 206 may search the reference picture memory 208 for sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture motion vectors, block shapes, etc., that may serve as suitable predictive references for the new picture. The predictor 206 may operate on sample blocks, pixel block by pixel block, to find suitable predictive references. In some cases, as determined by the search results obtained by the predictor 206, the input picture may have predictive references drawn from multiple reference pictures stored in the reference picture memory 208.

[0032] The output of all the aforementioned functional units may undergo entropy coding in entropy coder 214. Entropy coder 214 converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art (e.g., Huffman coding, variable length coding, arithmetic coding).

[0033] In some embodiments, the output of the entropy coder 214 is coupled to a transmitter. The transmitter is configured to buffer the coded video sequence generated by the entropy coder 214 to prepare it for transmission over a communication channel 218, which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter may be configured to merge the coded video data from the source coder 202 with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown). In some embodiments, the transmitter may transmit additional data along with the encoded video. The source coder 202 may include such data as part of the coded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, supplemental enhancement information (SEI) messages, visual usability information (VUI) parameter set fragments, etc.

[0034] The controller 204 may manage the operation of the encoder component 106. During coding, the controller 204 may assign a particular coding picture type to each coded picture, which may affect the coding technique that may be applied to the respective picture. For example, a picture may be assigned as an intra picture (I picture), a predictive picture (P picture), or a bidirectionally predictive picture (B picture). An intra picture may be a picture that may be coded and decoded without using any other frame in the sequence as a source of prediction. Some video codecs allow various types of intra pictures, including, for example, independent decoder refresh (IDR) pictures. Those skilled in the art are aware of those variations of I pictures, as well as their respective uses and characteristics, so they will not be repeated here. A predictive picture may be coded and decoded using intra prediction or inter prediction, which uses at most one motion vector and reference index to predict sample values ​​for each block. Bidirectionally predicted pictures may be coded and decoded using intra- or inter-prediction, which uses up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multi-predicted pictures may use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0035] A source picture may generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and coded block by block. A block may be predictively coded with reference to other (already coded) blocks as determined by the coding assignment applied to the block's respective picture. For example, a block of an I picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra prediction). A pixel block of a P picture may be non-predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. A block of a B picture may be non-predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.

[0036] Video may be captured in time sequence as multiple source pictures (video pictures). Intra-picture prediction (often abbreviated as intra-prediction) uses spatial correlation within a given picture, while inter-picture prediction uses correlation (temporal or other) between pictures. In one example, a particular picture of encoding / decoding, called the current picture, is divided into blocks. When a block in the current picture is similar to a reference block in a reference picture previously coded and still buffered in the video, the block in the current picture may be coded by a vector called a motion vector. The motion vector points to a reference block in the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0037] Encoder component 106 may perform coding operations in accordance with a given video coding technique or standard, such as any of those described herein. In its operations, encoder component 106 may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.

[0038] 2B is a block diagram illustrating example elements of the decoder component 122, according to some embodiments. The decoder component 122 of FIG. 2B is coupled to the channel 218 and the display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to the loop filter 256 and configured to transmit data to the display 124 (e.g., via a wired or wireless connection).

[0039] In some embodiments, the decoder component 122 includes a receiver coupled to the channel 218 and configured to receive data from the channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more coded video sequences to be decoded by the decoder component 122. In some embodiments, the decoding of each coded video sequence is independent of the other coded video sequences. The coded video sequences may be received from the channel 218, which may be a hardware / software link to a storage device that stores the encoded video data. The receiver receives the encoded video data along with other data, e.g., coded audio data and / or auxiliary data streams, which may be forwarded to their respective using entities (not shown). The receiver may separate the coded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the decoder component 122 to decode the data and / or to accurately reconstruct the original video data. The additional data may be in the form of, for example, a temporal, spatial or SNR enhancement layer, redundant slices, redundant pictures, forward error correction codes, etc.

[0040] According to some embodiments, the decoder component 122 includes a buffer memory 252, a parser 254 (sometimes referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra-picture prediction unit 262, a motion compensated prediction unit 260, an aggregator 268, a loop filter unit 256, a reference picture memory 266, and a current picture memory 264. In some embodiments, the decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuitry. In some embodiments, the decoder component 122 is implemented at least partially in software.

[0041] The buffer memory 252 is coupled between the channel 218 and the parser 254 (e.g., to combat network jitter). In some embodiments, the buffer memory 252 is separate from the decoder component 122. In some embodiments, a separate buffer memory is provided between the output of the channel 218 and the decoder component 122. In some embodiments, a separate buffer memory is provided external to the decoder component 122 (e.g., to combat network jitter) in addition to the buffer memory 252 in the decoder component 122 (e.g., which is configured to handle playout timing). When the receiver is receiving data from a storage / forwarding device of sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory 252 may be unnecessary or may be small. For use with best-effort packet networks such as the Internet, the buffer memory 252 may be used and may be relatively large, advantageously adaptively sized, and may be implemented at least in part in an operating system or similar element (not shown) external to the decoder component 122.

[0042] Parser 254 is configured to reconstruct symbols 270 from the coded video sequence. The symbols may include, for example, information used to manage the operation of decoder component 122 and / or information for controlling a rendering device such as display 124. The control information for the rendering device may be in the form of a Supplementary Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). Parser 254 parses (entropy decodes) the coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without contextual dependency, etc. Parser 254 may extract a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroups may include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. Parser 254 may also extract from the coded video sequence information such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0043] The reconstruction of symbols 270 may involve a number of different units, depending on the type of coded video picture or portion thereof (inter-picture and intra-picture, inter-block and intra-block, etc.), and other factors. Which units are involved and how they are involved may be controlled by subgroup control information parsed from the coded video sequence by parser 254. The flow of such subgroup control information between parser 254 and the following units is not depicted for convenience of explanation.

[0044] In addition to the functional blocks already mentioned, the decoder component 122 may be conceptually subdivided into several functional units, as described below. In an actual implementation operating under commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the conceptual subdivision into the following functional units is maintained.

[0045] The scalar / inverse transform unit 258 receives the quantized transform coefficients and control information (such as the transform to use, the block size, the quantization factor, and / or the quantization scaling matrix) from the parser 254 as symbols 270. The scalar / inverse transform unit 258 may output blocks containing sample values, which may be input to an aggregator 268.

[0046] In some cases, the output samples of the scaler / inverse transform unit 258 relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture, but may use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit 262. The intra-picture prediction unit 262 may generate a block of the same size and shape as the block being reconstructed using surrounding already reconstructed information fetched from the current (partially reconstructed) picture from the current picture memory 264. The aggregator 268 may append the prediction information generated by the intra-picture prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258 on a sample-by-sample basis.

[0047] In other cases, the output samples of the scalar / inverse transform unit 258 are associated with an inter-coded, potentially motion-compensated block. In such cases, the motion compensated prediction unit 260 may access the reference picture memory 266 to fetch samples used for prediction. After motion compensating the fetched samples according to the symbols 270 associated with the block, these samples may be added by the aggregator 268 to the output of the scalar / inverse transform unit 258 to generate output sample information (in this case referred to as residual samples or residual signals). The addresses in the reference picture memory 266 from which the motion compensated prediction unit 260 fetches the prediction samples may be controlled by a motion vector. The motion vector may be available to the motion compensated prediction unit 260 in the form of the symbols 270, which may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of fetched sample values ​​from the reference picture memory 266 when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, etc.

[0048] The output samples of aggregator 268 may be subjected to various loop filtering techniques in loop filter unit 256. Video compression techniques may include in-loop filter techniques controlled by parameters included in the coded video bitstream and made available to loop filter unit 256 as symbols 270 from parser 254, but may also be responsive to meta-information obtained during decoding of previous (in decoding order) portions of a coded picture or coded video sequence, or may be responsive to previously reconstructed and loop filtered sample values.

[0049] The output of the loop filter unit 256 may be a sample stream that may be output to a rendering device, such as the display 124, as well as stored in a reference picture memory 266 for use in future inter-picture prediction.

[0050] Once fully reconstructed, a particular coded picture may be used as a reference picture for future prediction. Once a coded picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by parser 254), the current reference picture may become part of reference picture memory 266, and a new current picture memory may be reallocated before beginning reconstruction of the next coded picture.

[0051] The decoder component 122 may perform decoding operations according to a given video compression technique, which may be documented in a standard, such as any of the standards described herein. The coded video sequence may conform to the syntax specified by the video compression technique or standard being used, in the sense of adhering to the syntax of the video compression technique or standard, as specified in the video compression technique document or standard, specifically the profile document therein. To conform to some video compression techniques or standards, the complexity of the coded video sequence may also be within a range prescribed by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may be further limited in some cases by the specification of a hypothetical reference decoder (HRD) and metadata for HRD buffer management signaled within the coded video sequence.

[0052] 3 is a block diagram illustrating a server system 112, according to some embodiments. The server system 112 includes a control circuit 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuit 302 includes one or more processors (e.g., a CPU, a GPU, and / or a DPU). In some embodiments, the control circuit includes one or more field programmable gate arrays (FPGAs), hardware accelerators, and / or one or more integrated circuits (e.g., application specific integrated circuits).

[0053] The network interface 304 may be configured to interface with one or more communication networks (e.g., wireless, wired, and / or optical networks). The communication networks may be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant, and the like. Examples of networks include local area networks such as Ethernet, cellular networks including WLAN, GSM, 3G, 4G, 5G, LTE, and the like, television wired or wireless wide area digital networks including cable, satellite, and terrestrial television, vehicular and industrial including CANBus, and the like. Such communications may be one-way, receive only (e.g., broadcast TV), one-way transmit only (e.g., CANbus to a particular CANbus device), or bidirectional (e.g., to other computer systems using local or wide area digital networks). Such communications may include communications to one or more cloud computer networks.

[0054] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input devices 310 may include one or more of a keyboard, a mouse, a trackpad, a touch screen, a data glove, a joystick, a microphone, a scanner, a camera, etc. The output devices 308 may include one or more of an audio output device (e.g., a speaker), a visual output device (e.g., a display or monitor), etc.

[0055] Memory 314 may include high-speed random access memory (such as DRAM, SRAM, DDR RAM, and / or other random access solid-state memory devices), and / or non-volatile memory (such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). Memory 314 optionally includes one or more storage devices located remotely from control circuitry 302. Memory 314, or a non-volatile solid-state memory device within memory 314, comprises a non-transitory computer-readable storage medium. In some embodiments, memory 314, or the non-transitory computer-readable storage medium of memory 314, an operating system 316 that includes instructions for handling various basic system services and for performing hardware-dependent tasks; A network communications module 318 used to connect the server system 112 to other computing devices via one or more network interfaces 304 (e.g., via wired and / or wireless connections); and a coding module 320 for performing various functions relating to encoding and / or decoding of data, such as video data; The coding module 320 may store any of the above programs, modules, instructions, and data structures, or a subset or superset thereof. In some embodiments, the coding module 320 is an instance of the coder component 114. The coding module 320 may include, but is not limited to, a decoding module 322 for performing various functions related to decoding of encoded data, such as those described above with respect to the decoder component 122; an encoding module 340 for performing various functions on the encoding data, such as those described above with respect to the encoder component 106; and and one or more of a picture memory 352 for storing pictures and picture data, e.g. for use with the coding module 320; In some embodiments, the picture memory 352 includes one or more of the reference picture memory 208, the buffer memory 252, the current picture memory 264, and the reference picture memory 266.

[0056] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions described above with respect to the parser 254), a transform module 326 (e.g., configured to perform various functions described above with respect to the scalar / inverse transform unit 258), a prediction module 328 (e.g., configured to perform various functions described above with respect to the motion compensated prediction unit 260 and / or the intra-picture prediction unit 262), and a filter module 330 (e.g., configured to perform various functions described above with respect to the loop filter 256).

[0057] In some embodiments, the encoding module 340 includes a code module 342 (e.g., configured to perform various functions described above with respect to the source coder 202 and / or the coding engine 212) and a prediction module 344 (e.g., configured to perform various functions described above with respect to the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 include a subset of the modules shown in FIG. 3. For example, a shared prediction module is used by both the decoding module 322 and the encoding module 340.

[0058] Each of the above identified modules stored in memory 314 corresponds to a set of instructions for performing functions described herein. The above identified modules (e.g., sets of instructions) need not be implemented as separate software programs, procedures, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, coding module 320 optionally does not include separate decoding and encoding modules, but rather uses the same set of modules to perform both sets of functions. In some embodiments, memory 314 stores a subset of the above identified modules and data structures. In some embodiments, memory 314 stores additional modules and data structures not described above, such as an audio processing module.

[0059] In some embodiments, the server system 112 includes a web or HyperText Transfer Protocol (HTTP) server, a File Transfer Protocol (FTP) server, as well as web pages and applications implemented using Common Gateway Interface (CGI) scripts, PHP HyperText Preprocessor (PHP), Active Server Pages (ASP), HyperText Markup Language (HTML), Extensible Markup Language (XML), Java, JavaScript, Asynchronous JavaScript and XML (AJAX), XHP, Javelin, Wireless Universal Resource Files (WURFL), and the like.

[0060] 3 illustrates a server system 112 according to some embodiments, however, FIG. 3 is not intended as a structural schematic of the embodiments described herein, but rather as a functional description of various features that may be present in one or more server systems. In practice, and as will be recognized by those skilled in the art, items shown separately may be combined and some items may be separated. For example, some items shown separately in FIG. 3 may be implemented on a single server, and single items may be implemented by one or more servers. The actual number of servers used to implement the server system 112, and how functionality is allocated among them, will vary from implementation to implementation and will, optionally, depend in part on the amount of data traffic the server system processes during peak and average usage periods.

[0061] An embodiment of a primary transform as used in AOMedia Video 1 (AV1) is described below. To support the extended coding block partitioning as described in this disclosure, as in AV1, multiple transform sizes (e.g., ranging from 4 points to 64 points per dimension) and transform shapes (e.g., square, rectangular shapes with width-to-height ratios of 2:1, 1:2, 4:1, or 1:4) may be used.

[0062] The 2D transform process may use a hybrid transform kernel that may include different 1D transforms for each dimension of the coded residual block. The linear 1D transform may include (a) 4-point, 8-point, 16-point, 32-point, 64-point DCT-2, (b) 4-point, 8-point, 16-point asymmetric DST (ADST) (e.g., DST-4, DST-7) and corresponding inverted versions (e.g., the inverted version of ADST or FlipADST may apply ADST in the reverse order), and / or (c) 4-point, 8-point, 16-point, 32-point identity transform (IDTX). Figure 4 shows an example of linear transform basis functions according to some embodiments. The linear transform basis functions in the example of Figure 4 include basis functions for DCT-2 and asymmetric DST (DST-4 and DST-7) with N-point input. The linear transform basis functions shown in Figure 4 may be used in AV1.

[0063] The availability of the hybrid transform kernels may depend on the transform block size and the prediction mode. Figure 5A shows a table of example dependencies of the availability of various transform kernels based on the transform block size and the prediction mode, according to some embodiments.

[0064] FIG. 5A illustrates an exemplary dependency of availability of various transform kernels (e.g., transform types shown in the first column and described in the second column) based on transform block size (e.g., sizes shown in the third column) and prediction mode (e.g., intra-prediction and inter-prediction shown in the third column). These exemplary hybrid transform kernels and their availability based on prediction mode and transform block size may be used in AV1. With reference to FIG. 5A, the symbols "→" and "↓" refer to the horizontal dimension (also called the horizontal direction) and the vertical dimension (also called the vertical direction), respectively. The symbols

number

number

[0065] In one example, the transform type (510) is represented as ADST_DCT, as shown in the first column of Figure 5A. The transform type (510) includes ADST in the vertical direction and DCT in the horizontal direction, as shown in the second column of Figure 5A. According to the third column of Figure 5A, when the block size is equal to or smaller than 16x16 (e.g., 16x16 samples, 16x16 luma samples), the transform type (510) is available for intra prediction and inter prediction.

[0066] In one example, the transform type (520) is represented as V_ADST, as shown in the first column of FIG. 5A. The transform type (520) includes ADST in the vertical direction and IDTX (i.e., identity matrix) in the horizontal direction, as shown in the second column of FIG. 5A. Thus, the transform type (520) (e.g., V_ADST) is performed in the vertical direction and not in the horizontal direction. According to the third column of FIG. 5A, the transform type (520) is not available for intra prediction regardless of the block size. The transform type (520) is available for inter prediction when the block size is less than 16×16 (e.g., 16×16 samples, 16×16 luma samples).

[0067] In one example, Figure 5A is applicable to the luma component. For the chroma components, the selection of the transform type (or transform kernel) may be performed implicitly. Figure 5B shows a table of example transform type selection based on intra-prediction mode according to some embodiments.

[0068] In one example, for intra prediction residuals, the transform type may be selected according to the intra prediction mode, as shown in FIG. 5B. In one example, the transform type selection shown in FIG. 5B is applicable to the chroma components. For inter prediction residuals, the transform type may be selected according to the transform type selection of the co-located luma block. Thus, in one example, the transform type for the chroma components is not signaled in the bitstream.

[0069] The Line Graph Transform (LGT) may be used in transforms such as the linear transform in AOMedia Video 2 (AV2), for example. AV2 may use an 8-bit / 10-bit transform core. In one example, the LGT includes various DCTs, discrete sine transforms (DSTs), as described below. The LGT may include 32-point and 64-point one-dimensional (1D) DSTs.

[0070] A graph is a general mathematical structure that includes a set of vertices and edges that can be used to model affinity relationships between objects of interest. A weighted graph, where a set of weights is assigned to the edges and, optionally, the vertices, can provide a sparse representation for robust modeling of signals / data. LGT can improve coding efficiency by providing good adaptability to diverse block statistics. A separable LGT can be designed and optimized by learning a line graph from the data to model the underlying row- and column-wise statistics of the block's residual signal, and the associated Generalized Graph Laplacian (GGL) matrix can be used to derive the LGT.

[0071] An embodiment of a secondary transform as used in AOMedia Video 2 (AV2) is described below. Figure 6 illustrates an exemplary use of an intra secondary transform (IST) in the encoding and decoding process according to some embodiments. For an original intra-predicted residual block 602 of a luma color component, a secondary transform 606 method, e.g., IST, may be applied to the primary transform 604 coefficient block before applying quantization 608 and entropy coding 610 in an encoder 624 for transmission of a bitstream 626. Thus, a secondary inverse transform 616 may be applied to the inverse quantized 614 transform coefficient block after bitstream parsing 612 before applying an inverse primary transform 618 to form a reconstructed residual block 620 in a decoder 622. An IST may not be applied to a chroma color component.

[0072] In one example, in IST, a non-separable transform process is applied. To apply a forward non-separable transform to a particular region of an input transform coefficient block of N samples, the N samples are stored in an N×1 vector (

number

number

number

[0073] In one example, to apply an inverse non-separable transform, an inverse quantized transform coefficient block is first given as an input, and then a particular region of the inverse quantized transform coefficient block is identified based on the transform block size.

number

number

number

number

[0074] In one example, the input to the forward IST is a coefficient vector consisting of low-frequency primary transform coefficients in a zigzag scan. Depending on the block size, either a 16-point or a 64-point non-separable secondary transform may be selected. When the minimum values ​​of the primary transform width and primary transform height are less than 8, a 16-point IST is used, and the low-frequency primary transform coefficients refer to the first 16 primary transform coefficients in the zigzag scan order. When both the width and height of the primary transform are 8 or greater, a 64-point IST is applied, and the low-frequency primary transform coefficients refer to the first 64 primary transform coefficients in the zigzag scan order. The 16-point non-separable transform uses an 8×16 transform kernel, and the 64-point non-separable transform uses a 32×64 transform kernel. Furthermore, when the IST is applied, high-frequency transform coefficients that are not processed by the secondary transform are zeroed out.

[0075] Figure 7 shows a table of an example mapping from intra nominal modes and primary transform types to IST sets according to some embodiments. A total of 12 secondary transform sets (or IST sets) are defined, each of which includes three secondary transform kernels. For each intra-coded transform block, a nominal intra prediction mode and a primary transform type are first identified, and then an IST set is selected based on Figure 7. For Paeth prediction mode and recursive intra prediction mode, IST is neither applied nor signaled.

[0076] In one example, given an IST set, there are four encoder choices: 1) no secondary transform; 2) a secondary transform using the first transform kernel in the given IST set; 3) a secondary transform kernel using the second transform kernel in the given IST set; and 4) a secondary transform kernel using the third transform kernel in the given IST set. The encoder signals the selection using the syntax element ist_idx. At the decoder, the value of the syntax element ist_idx is parsed first. Then, given the IST set and the value associated with ist_idx, the secondary transform kernel is identified. This syntax element ist_idx is signaled for each luma transform block after the signaling of the primary transform type. Signaling of ist_idx is performed when all of the following conditions are true: (1) the current block is an intra-coded luma transform block; (2) the primary transform type is DCT in both dimensions or ADST in both dimensions; (3) the intra prediction mode is neither Paeth prediction mode nor recursive intra prediction mode; (4) the transform split depth is 0; and (5) the end-of-block (EOB) position is within the low-frequency transform coefficient domain where a secondary transform can be applied.

[0077] In some examples, the entropy coding context for ist_idx is derived based on the transform block size.

[0078] An embodiment of boundary-dependent transformation is described below. Figure 8 shows an example of squared error distribution in a 4x4 prediction unit (PU) according to some embodiments. When using inter prediction for a PU, the prediction error is generally larger near the PU boundary than in the center of the PU. The figures in Figure 8 show an example of squared error distribution over a 4x4 PU.

[0079] FIG. 9 shows an example split PU and error distribution in an upper-left transform unit (TU) according to some embodiments. Similarly, when a PU is split into multiple transform units (TUs) as shown in FIG. 9, the prediction error is larger near the PU boundary than near the inner TU (non-PU) boundary. The numbers in FIG. 9 show the squared error distribution in TU0 at the upper-left corner of the PU. The error decreases from top to bottom and left to right. The reason for this effect may be due to different motion vectors (MVs) between two neighboring PUs. To handle this uneven error distribution, alternative transforms such as DST-VII and DCT-IV may be used. Equation (3) and Equation (4) respectively show the N-point DST-VII and DCT-IV of a signal f[n].

number

[0080] In some examples, based on the above observations, DST-VII and / or DCT-IV may be used instead of DCT-II when only one of the two TU boundaries is a PU boundary. FIG. 10 shows an example mapping from boundary type to transform by using DST-VII according to some embodiments. In some examples, FIG. 10 is applied to a 4-point (pt) transform. FIG. 11 shows an example mapping from boundary type to transform by using DCT-IV according to some embodiments. In some examples, FIG. 11 is applied to an 8pt and / or 16pt transform. DST-VII and DCT-IV are not used for 32pt transform because they have little gain and are significantly more complicated. In FIG. 10 and FIG. 11, the items "non-PU" and "PU" refer to non-PU and PU boundaries, respectively.

[0081] Figure 12 shows an example transformation used for the TU of Figure 9 according to some embodiments. F(DST-VII) in Figures 10 and 12 may mean flipping the DST matrix from left to right. When using F(DST-VII), it may be implemented to first flip the input data and then use DST-VII. The same is true for F(DCT-IV).

[0082] According to FIG. 10, the four TUs in FIG. 9 may perform the transformation as shown in FIG.

[0083] In some approaches, for transform units located at different relative positions within a coding block, the residuals exhibit different statistics, and different secondary transforms may be designed to achieve efficient transform coding of these residual blocks.

[0084] In some embodiments, the methods disclosed herein may be used separately or combined in any order. Furthermore, each of the methods (or embodiments), the encoder, and the decoder may be implemented by a processing circuit (e.g., one or more processors, or one or more integrated circuits). In one example, the one or more processors execute a program stored in a non-transitory computer-readable medium. Hereinafter, the term block may be interpreted as a prediction block, a coding block, or a coding unit, i.e., a CU. The term block here may also be used to refer to a transform block. In this specification, the size of a block may refer to a block width, a block height, a block aspect ratio, a block area size, or a minimum / maximum value between a block width and a block height.

[0085] In this disclosure, a transform block may also be referred to as a transform unit. A transform unit is located on a boundary of a coding block, or a transform unit is a boundary TU, which means that the coding block has multiple transform units, and the transform unit has at least one boundary that overlaps with one boundary of the coding block.

[0086] FIG. 13 illustrates an example of TUs in a coding block, according to some embodiments.

[0087] For example, in the left video block 1302 of Figure 13, the TUs with a gray background are border TUs, and TU5 / TU6 / TU9 / TU10 are not border TUs.

[0088] Alternatively, a transform unit is located on a boundary of a coding block, or is a boundary TU, means that the coding block has multiple transform units and that the transform unit has at least one boundary that overlaps with one boundary of the coding block but does not overlap with the coding block on the other side of the overlapping boundary. For example, in the right video block 1304 of Figure 13, TU0 and TU4 are boundary TUs, and both TU0 and TU4 have three boundaries that overlap with coding block boundaries. For TU0, the left boundary completely overlaps with the coding block boundary, while for TU4, the right boundary completely overlaps with the coding block boundary.

[0089] In some aspects / embodiments, when a coding block has multiple transform blocks, the secondary transform selection and / or signaling may differ depending on the relative position of the transform units within the coding block.

[0090] FIG. 14 is an example flow diagram illustrating a method 1400 of coding video, according to some embodiments. FIG. 15 is an example flow diagram illustrating a method 1500 of coding video, according to some embodiments. Methods 1400 and / or 1500 may be implemented in a computer system (e.g., server system 112, source device 102, or electronic device 120) having control circuitry and memory storing instructions for execution by the control circuitry. In some embodiments, methods 1400 and / or 1500 may be implemented by executing instructions stored in a memory (e.g., memory 314) of the computer system. Methods 1400 and / or 1500 may be implemented by an encoder (e.g., encoder 106) and / or a decoder (e.g., decoder 122).

[0091] Referring to FIG. 14, in one aspect, a video decoder (e.g., decoder 122 of FIG. 2B and / or decoder 622 of FIG. 6) receives a first control flag from a video stream (e.g., bitstream 626 of FIG. 6), where the first control flag indicates whether inter prediction mode is enabled for a video block of the video stream (e.g., predictive block (PU) of FIG. 9, video blocks 1302, 1304 of FIG. 13) (1410).

[0092] The video decoder determines (1420) whether multiple transform units (eg, TU0, TU1, TU2, and TU3 in FIG. 9, and TU0-TU15 in FIG. 13) are within the video block in accordance with the determination that the inter prediction mode is enabled.

[0093] In accordance with determining that the multiple transform units are within the video block, the video decoder determines a first transform unit of the multiple transform units to apply a first secondary transform (e.g., the inverse secondary transform 616 of FIG. 6) based on a first relative position of the first transform unit within the video block, applies the first secondary transform to the first transform unit, and reconstructs the video block based at least on the first secondary transform (1430).

[0094] Referring to FIG. 15, in one aspect, a video encoder (e.g., encoder 106 of FIG. 2B and / or encoder 624 of FIG. 6) receives a first control flag from video data (e.g., original residual block 602 of FIG. 6), the first control flag indicating whether inter prediction mode is enabled for a video block of the video data (e.g., predictive block (PU) of FIG. 9, video blocks 1302, 1304 of FIG. 13) (1510).

[0095] The video encoder determines (1520) whether multiple transform units (eg, TU0, TU1, TU2, and TU3 in FIG. 9, TU0-TU15 in FIG. 13) are within the video block in accordance with the determination that inter prediction mode is enabled.

[0096] In accordance with determining that multiple transform units are within the video block, the video encoder determines a first transform unit of the multiple transform units to apply a first secondary transform (e.g., secondary transform 606 of FIG. 6) based on a first relative position of the first transform unit within the video block, applies the first secondary transform to the first transform unit, and processes the video block based at least on the first secondary transform (1530).

[0097] In one embodiment and / or any combination of embodiments disclosed herein, the video decoder / encoder, in accordance with determining that the plurality of transform units are within the video block, further determines a second transform unit of the plurality of transform units to apply a second secondary transform or not apply a secondary transform based on a second relative position of the second transform unit within the video block, applies the second secondary transform or not apply the secondary transform to the second transform unit, and reconstructs / processes the video block at least further based on the second secondary transform. In some examples, the first relative position and the second relative position are different, and the first secondary transform and the second secondary transform are different.

[0098] In one embodiment and / or any combination of embodiments disclosed herein, the secondary transform is applied to selected transform units within an inter-coding block.

[0099] In one embodiment and / or any combination of embodiments disclosed herein, the selection of the transform unit that applies the secondary transform depends on the relative position of the transform unit within the coding block.

[0100] In one embodiment and / or any combination of embodiments disclosed herein, for inter-coding blocks (including coding blocks coded by combined intra-prediction modes), the secondary transform may be applied only to the boundary TU. In one embodiment, for inter-coding blocks using combined prediction, a different set of secondary transforms is applied to the boundary TU compared to the set of secondary transforms applied to inter-coding blocks using single reference prediction. For example, the first transform unit is placed at the boundary position of the video block.

[0101] In one embodiment and / or any combination of embodiments disclosed herein, the secondary transform set depends on the coding block boundary where the TU is located within the coding block. Different secondary transform sets may be applied for TUs located at various boundaries (top, left, right, bottom, top-left, top-right, bottom-left, bottom-right). For example, the method 1400 / 1500 further includes reconstructing / processing the video block by applying a third secondary transform to a third transform unit of the multiple transform units in the video block, where the first transform unit is located at a first boundary position of the video block, the third transform unit is located at a secondary boundary position of the video block, and the third secondary transform is different from the first secondary transform. In one embodiment, the secondary transform is applied only to TUs located at a corner (top-left, top-right, bottom-left, bottom-right) of the coding block. For example, the first transform unit is located at a corner of the video block.

[0102] In one embodiment and / or any combination of embodiments disclosed herein, when a coding block is inter-coded, the secondary transform is still applied to all TUs, but the context used to entropy code the secondary transform index / flag depends on the relative TU position within the coding block. For example, the method 1400 / 1500 further includes determining a secondary transform context for entropy coding, where the secondary transform context is determined based on a relative position of one of the transform units within the video block. In one example, the context used to entropy code the secondary transform index / flag depends on whether the TU is a boundary TU or not. For example, the method 1400 / 1500 further includes determining a first secondary transform context for entropy coding of a first transform unit and determining a second secondary transform context for entropy coding of a second transform unit, where the first transform unit is a boundary unit and the second transform unit is not a boundary unit, and the first secondary transform context is different from the second secondary transform context for entropy coding.

[0103] In one embodiment and / or any combination of embodiments disclosed herein, when a coding block is inter-coded, secondary transforms applied to TUs located at different relative positions in the coding block share the same elements / coefficients in the transform base. However, the elements / coefficients are arranged in different orders. For example, a first secondary transform and a second secondary transform share the same one or more coefficients in the transform base. In one embodiment, the elements are arranged in the basis vector according to an order that depends on the relative TU position in the coding block. For example, one or more coefficients in the transform base are arranged in the basis vector according to an order that depends on the relative position of one of the multiple transform units in the video block.

[0104] In one embodiment and / or any combination of embodiments disclosed herein, when a coding block is inter-coded, given a relative position of a TU in the coding block, an intra prediction mode is identified, and then a secondary transform set is identified and used to perform a secondary transform of the TU according to the intra prediction mode. For example, determining a first transform unit of the plurality of transform units to apply a first secondary transform based on a first relative position of the first transform unit in the video block includes identifying the first secondary transform according to a secondary transform used for the intra prediction mode of the first transform unit according to the first relative position of the first transform unit. In one embodiment, the relative position of the TU in the coding block is mapped to one of the following intra prediction modes: DC, SMOOTH, SMOOTH-H, SMOOTH-V, horizontal, vertical, diagonal (45 degrees), diagonal (135 degrees), diagonal (225 degrees). For example, a first transform unit is mapped to a first intra-prediction mode and a second transform unit is mapped to a second intra-prediction mode, where the first intra-prediction mode is different from the second intra-prediction mode.

[0105] In one embodiment and / or any combination of embodiments disclosed herein, for an inter-coding block, the signaling of the secondary transform of a transform unit also depends on the secondary transform index / flag signaled for the neighboring TU. For example, the first signaling flag of the first secondary transform is different from the second signaling flag of the second secondary transform. For example, the first signaling flag of the first secondary transform depends on the flag of the secondary transform signaled for one or more neighboring transform units of the first transform unit.

[0106] 14 and 15 show some logical stages in a particular order, stages that are not order dependent may be reordered and other stages may be combined or split. Some reordering or other groupings not specifically mentioned will be apparent to those of ordinary skill in the art, and thus the reordering and groupings presented herein are not exhaustive. Furthermore, it should be recognized that the stages may be implemented in hardware, firmware, software, or any combination thereof.

[0107] In another aspect, some embodiments include a computer system (e.g., server system 112) including a control circuit (e.g., control circuit 302) and a memory (e.g., memory 314) coupled to the control circuit, the memory storing one or more sets of instructions configured to be executed by the control circuit, the one or more sets of instructions including instructions for performing any of the methods described herein.

[0108] In yet another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more sets of instructions for execution by control circuitry of a computer system, the one or more sets of instructions including instructions for performing any of the methods described herein.

[0109] It will also be understood that although terms such as "first", "second", etc. are used herein to describe various elements, these elements are not intended to be limited by these terms, but rather are used only to distinguish one element from another.

[0110] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the scope of the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms "comprises" and / or "comprising", as used herein, specify the presence of stated features, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0111] As used herein, the term "if" may be interpreted to mean "when" or "upon" or "in response to determining" or "incorrespondence to determining" or "in response to detecting" that a stated precondition is true, depending on the context. Similarly, the phrases "if it is determined that [the precondition of the stated condition is true]" or "if [the precondition of the stated condition is true]" or "when [the precondition of the stated condition is true]" may be interpreted to mean "upon determining" or "in response to determining" or "in accordance with the determination" or "upon detecting" or "in response to detecting" that a stated precondition is true, depending on the context.

[0112] The foregoing description has been described with reference to specific embodiments for purposes of explanation. However, the above exemplary description is not intended to be exhaustive or to limit the scope of the claims to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments have been selected and described to best explain the principles of operation and practical application, thereby enabling others skilled in the art. [Explanation of symbols]

[0113] 100 Communication Systems 102 Source Device 104 Video Sources 106 Encoder Components 108 Encoded Video Bitstream 110 Network 112 Server System 114 Coda Constituents 116 Encoded Video Data 120-1 Electronic Devices 120-m Electronic Devices 122 Decoder Components 124 Display 202 Source Coder 204 Controller 206 Predictors 208 Reference Picture Memory 210 Local Decoder 212 Coding Engine 214 Entropy Coder 216 Coded Video Sequence 218 Communication Channels 252 Buffer Memory 254 Parser 256 Loop Filter Unit 258 Scaler / Descaler Unit 260 Motion Compensation Prediction Unit 262 Intra-picture Prediction Unit 264 Current Picture Memory 266 Reference Picture Memory 268 Aggregator 270 Symbols 302 Control circuit 304 Network Interface 306 User Interface 308 Output Device 310 Input Devices 312 Communication Bus 314 Memory 316 Operating Systems 318 Network Communication Module 320 Coding Module 322 Decoding Module 324 Syntax Analysis Module 326 Conversion Module 328 Prediction Module 330 Filter Module 340 Encoding Module 342 Code Module 344 Prediction Module 352 Picture Memory 510 Conversion Type 520 Conversion Type 602 Original intra prediction residual block 604 Primary Transformation 606 Secondary Transformation 608 Quantization 610 Entropy Coding 612 Bitstream Parsing 614 Inverse quantization 616 Secondary Inverse Transformation 618 Inverse Linear Transformation 620 Residual Block 622 Decoder 624 Encoder 626 bitstream 1302 Video Block 1304 Video Block

Claims

1. A method for processing a video stream performed in a computer system having memory and control circuits, wherein the method is A step of receiving a first control flag from the video stream, wherein the first control flag indicates whether or not an interprediction mode is enabled for the video block of the video stream. The steps include determining whether multiple conversion units are located within the video block, in accordance with the determination that the interpretation mode is enabled, In accordance with the determination that multiple conversion units are located within the video block, The steps include determining which of the plurality of conversion units is the first conversion unit to apply a first secondary conversion based on the first relative position of the first conversion unit within the video block, The steps include applying the first secondary conversion to the first conversion unit, The steps of reconstructing the video block based at least on the first secondary transformation and Methods that include...

2. The steps include determining which of the plurality of conversion units is the second conversion unit to apply a second secondary conversion based on the second relative position of the second conversion unit within the video block, or not to apply a secondary conversion; The steps of applying the second secondary conversion to the second conversion unit, or not applying the secondary conversion, The steps include: reconstructing the video block based at least on the second secondary transformation described above; It further includes, The first relative position and the second relative position are different, and the first quadratic transformation and the second quadratic transformation are different, The method according to claim 1.

3. The method according to claim 1, wherein the first conversion unit is positioned at the boundary of the video block.

4. A step of reconstructing the video block by applying a third secondary transformation to a third transformation unit among the plurality of transformation units in the video block, wherein the first transformation unit is located at a first boundary position of the video block, the third transformation unit is located at a secondary boundary position of the video block, and the third secondary transformation is different from the first secondary transformation. The method according to claim 3, further comprising:

5. The method according to claim 1, wherein the first conversion unit is positioned at the corner of the video block.

6. A step of determining a secondary transformation context for entropy coding, wherein the secondary transformation context is determined based on the first relative position of the first transformation unit of the plurality of transformation units in the video block. The method according to claim 1, further comprising:

7. A step of determining a first quadratic transformation context for entropy coding of the first transformation unit and a second quadratic transformation context for entropy coding of the second transformation unit, wherein the first transformation unit is a boundary unit, the second transformation unit is not a boundary unit, and the first quadratic transformation context is different from the second quadratic transformation context for entropy coding. The method according to claim 2, further comprising:

8. The method according to claim 2, wherein the first quadratic transformation and the second quadratic transformation share the same one or more coefficients in the transformation basis.

9. The method according to claim 8, wherein one or more coefficients in the transformation basis are arranged in the basis vector in an order that depends on the relative position of one of the plurality of transformation units in the video block.

10. The step of determining which of the plurality of conversion units is to apply the first secondary conversion based on the first relative position of the first conversion unit within the video block, Steps to identify the first secondary transformation according to the first relative position of the first transformation unit and according to the secondary transformation used for the intra-prediction mode of the first transformation unit. The method according to claim 2, including the method described in claim 2.

11. The method according to claim 10, wherein the first conversion unit is mapped to a first intra-prediction mode, the second conversion unit is mapped to a second intra-prediction mode, and the first intra-prediction mode is different from the second intra-prediction mode.

12. The method according to claim 2, wherein the first signaling flag of the first quadratic transformation is different from the second signaling flag of the second quadratic transformation.

13. The method according to claim 12, wherein the first signaling flag of the first quadratic transform depends on a flag of a quadratic transform that is signaled for one or more neighboring transform units of the first transform unit.

14. An electronic device configured to perform the method described in any one of claims 1 to 13.

15. A computer program for causing a processor to perform the method described in any one of claims 1 to 13.