Method and apparatus for video decoding

By decoding transform coefficients, determining residuals and reconstructing the encoding blocks of non-rectangular partitions in the processing circuit system, the transformation problem that the prior art cannot be applied to L-shaped partitions is solved, and efficient video encoding and decoding of non-rectangular partitions is realized.

CN114946183BActive Publication Date: 2025-05-27TENCENT AMERICA LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180006202.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-04-19
Filing Date
2021-08-05
Publication Date
2025-05-27
Estimated Expiration
2041-08-05

AI Technical Summary

Technical Problem

The prior art cannot be applied to the transformation scheme of L-shaped partitions, resulting in the inability to perform effective video encoding and decoding in the case of non-rectangular partitions.

Method used

By decoding the transform coefficients associated with the encoded blocks of the non-rectangular partitions in the processing circuitry, the residuals of the encoded blocks are determined and the samples of the encoded blocks are reconstructed based on the residuals. The specific method includes further dividing the non-rectangular partition into rectangular sub-blocks, or performing transformations on the entire non-rectangular portion and performing inverse transformations on the decoder side to determine the residual.

Benefits of technology

It realizes an applicable transformation solution for non-rectangular partitions, improves the flexibility and efficiency of video encoding and decoding, and can effectively process in non-rectangular situations such as L-shaped partitions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114946183B_ABST
    Figure CN114946183B_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure provide methods and apparatuses for video encoding / decoding. In some examples, an apparatus for video decoding includes processing circuitry. The processing circuitry decodes transform coefficients associated with an encoded block of a non-rectangular partition of a picture according to an encoded video bitstream. Additionally, the processing circuitry determines a residual of the encoded block based on the transform coefficients and reconstructs samples of the encoded block based on the residual of the encoded block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of priority of U.S. Patent Application No. 17 / 234,507, filed on Apr. 19, 2021, entitled “METHOD AND APPARATUS FOR VIDEO CODING”, which claims the benefit of priority of U.S. Provisional Application No. 63 / 087,042, filed on Oct. 2, 2020, entitled “TRANSFORMATION SCHEME FOR L - SHAPED PARTITION”. The entire disclosure of the prior application is incorporated herein by reference in its entirety. Technical Field

[0003] The present disclosure generally relates to video decoding techniques, and more particularly to a method and apparatus for video decoding. Background Art

[0004] The background art description provided herein is for the purpose of generally presenting the context of the present disclosure. To the extent that the work described in this background art section is not otherwise considered prior art as of the time of filing, neither the work of the currently named inventors nor aspects of the description that may not otherwise be considered prior art conditions are, either expressly or implicitly, admitted as prior art against the present disclosure.

[0005] Video coding and decoding can be performed using inter - picture prediction with motion compensation. Uncompressed digital video can include a series of pictures, each with spatial dimensions of, for example, 1920×1080 luminance samples and associated chrominance samples. This series of pictures can have a fixed or variable picture rate (also informally referred to as the frame rate), such as 60 pictures per second or 60 Hz. Uncompressed video has high bit - rate requirements. For example, 1080p60 4:2:0 video (1920×1080 luminance sample resolution at 60 Hz frame rate) at 8 bits per sample requires a bandwidth of nearly 1.5 Gbit / s. An hour of such video requires more than 600 gigabytes of storage space.

[0006] One purpose of video encoding and decoding can be to reduce the redundancy of an input video signal through compression. Compression can help reduce the above-mentioned bandwidth or storage space requirements, which can be reduced by two orders of magnitude or more in some cases. Both lossless compression and lossy compression, as well as combinations thereof, can be employed. Lossless compression refers to techniques by which an exact copy of the original signal can be reconstructed from the compressed original signal. When lossy compression is used, the reconstructed signal may not be the same as the original signal, but the distortion between the original signal and the reconstructed signal is small enough such that the reconstructed signal is useful for the intended application. In the case of video, lossy compression is widely adopted. The amount of distortion tolerated depends on the application; for example, users of certain consumer streaming applications can tolerate higher distortion than users of television distribution applications. The achievable compression ratio can reflect that a higher allowable / tolerable distortion can result in a higher compression ratio.

[0007] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transformation, quantization, and entropy coding.

[0008] Video codec techniques can include techniques referred to as intra coding. In intra coding, sample values are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, a picture is spatially subdivided into sample blocks. When all sample blocks are coded in the intra mode, the picture can be an intra picture. Intra pictures and their derived pictures (e.g., independent decoder refresh pictures) can be used to reset the decoder state and can thus be used as the first picture in an encoded video bitstream and a video session, or as a still image. A transformation can be applied to the samples of an intra block, and the transform coefficients can be quantized prior to entropy coding. Intra prediction can be a technique for minimizing sample values in the pre-transform domain. In some cases, the smaller the DC value and the AC coefficients after transformation, the fewer bits are required to represent the block after entropy coding for a given quantization step.

[0009] Traditional intra coding, such as intra coding known from, for example, MPEG-2 generation coding techniques, does not use intra prediction. However, some newer video compression techniques include techniques that attempt to utilize, for example, surrounding sample data and / or metadata obtained during the encoding / decoding of spatially adjacent and earlier in decoding order data blocks. Such techniques are hereafter referred to as "intra prediction" techniques. Note that in at least some cases, intra prediction uses only reference data from the current picture being reconstructed rather than reference data from reference pictures.

[0010] There can be many different forms of intra prediction. When more than one such technique can be used in a given video coding technology, the technique in use can be coded in an intra prediction mode. In some cases, the mode can have sub - modes and / or parameters, and these sub - modes and / or parameters can be coded separately or included in the mode codeword. Which codeword is used for a given mode / sub - mode / parameter combination can affect the coding efficiency obtained through intra prediction, and thus can affect the entropy coding technique used to convert the codeword into a bitstream.

[0011] Intra prediction for a specific mode was introduced in H.264, improved in H.265, and further improved in more recent coding technologies such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). Predictor blocks can be formed using neighboring sample values belonging to already available samples. The sample values of neighboring samples are copied into the predictor block according to a direction. The reference to the direction used can be coded in the bitstream or can itself be predicted.

[0012] Generally, a picture is partitioned into blocks, and the blocks can be units for various processes such as encoding, prediction, transformation, etc. In various block partitioning techniques, a picture is typically partitioned into rectangular or square blocks. Existing transformation schemes are designed for rectangular or square coding blocks. However, since the shape of an L - shaped partition is neither square nor rectangular, current transformation schemes cannot be applied to the case of L - shaped partitions.

[0013] Therefore, the problem to be solved by the present invention is to provide a transformation scheme applicable to L - shaped partitions. Summary of the Invention

[0014] Aspects of the present disclosure provide methods and apparatuses for video coding / decoding. In some examples, an apparatus for video decoding includes processing circuitry. The processing circuitry decodes transform coefficients associated with an encoded block of a non - rectangular partition of a picture according to an encoded video bitstream. In addition, the processing circuitry determines a residual of the encoded block based on the transform coefficients, and reconstructs samples of the encoded block based on the residual of the encoded block.

[0015] In some embodiments, the processing circuitry determines a first residual of a first rectangular sub - block based on a first transform coefficient among the transform coefficients. The first rectangular sub - block is a first partition of the non - rectangular partition. In addition, the processing circuitry determines a second residual of a second rectangular sub - block based on a second transform coefficient among the transform coefficients. In an example, the second rectangular sub - block is a second partition of the non - rectangular partition and has the same size as the first rectangular sub - block. In another example, the second rectangular sub - block is a second partition of the non - rectangular partition and has a different size from the first rectangular sub - block.

[0016] In some embodiments, the processing circuitry determines intermediate transform coefficients of an intermediate transform unit by performing an inverse transform on a transform unit in a first direction. The transform unit is formed of transform coefficients. Additionally, the processing circuitry determines a residual of an encoded block by performing an inverse transform on the intermediate transform unit in a second direction.

[0017] In some examples, the processing circuitry performs a first inverse transform operation on first plural columns of the transform unit respectively. The first inverse transform operations are inverse transforms of first numbers of points respectively. Additionally, the processing circuitry performs a second inverse transform operation on second plural columns of the transform unit respectively. The second inverse transform operations are inverse transforms of second numbers of points respectively. The processing circuitry performs a third inverse transform operation on first plural rows of the intermediate transform unit respectively. The third inverse transform operations are inverse transforms of third numbers of points respectively. Additionally, the processing circuitry performs a fourth inverse transform operation on second plural rows of the intermediate transform unit respectively. The fourth inverse transform operations are inverse transforms of fourth numbers of points respectively.

[0018] In some other examples, the processing circuitry performs a first inverse transform operation on first plural rows of the transform unit respectively. The first inverse transform operations are inverse transforms of first numbers of points respectively. The processing circuitry performs a second inverse transform operation on second plural rows of the transform unit respectively. The second inverse transform operations are inverse transforms of second numbers of points respectively. Additionally, the processing circuitry performs a third inverse transform operation on first plural columns of the intermediate transform unit respectively. The third inverse transform operations are inverse transforms of third numbers of points respectively. Then, the processing circuitry performs a fourth inverse transform operation on second plural columns of the intermediate transform unit respectively. The fourth inverse transform operations are inverse transforms of fourth numbers of points respectively.

[0019] According to an aspect of the present disclosure, the processing circuitry may determine a residual of an encoded block by performing an inverse Karhunen - Loève transform (KLT) on transform coefficients.

[0020] According to another aspect of the present disclosure, the processing circuitry determines a first residual of a rectangular unit including a non - rectangular partition by performing a two - dimensional inverse transform on transform coefficients, and selects a residual of the encoded block from the first residual of the rectangular block.

[0021] In some embodiments, the processing circuitry forms a transform unit for a non - rectangular partition by following a scan order of a rectangular unit including the non - rectangular partition and skipping scan positions outside the non - rectangular partition.

[0022] Aspects of the present disclosure also provide a method for video decoding, including: decoding transform coefficients associated with an encoded block of a non - rectangular partition of a picture according to an encoded video bitstream; determining a residual of the encoded block based on the transform coefficients; and reconstructing samples of the encoded block based on the residual of the encoded block.

[0023] Aspects of the present disclosure also provide an apparatus for video decoding, including: a decoding unit configured to decode transform coefficients associated with coded blocks of a non-rectangular partition of a picture according to an encoded video bitstream; a determination unit configured to determine a residual of the coded block based on the transform coefficients; and a reconstruction unit configured to reconstruct samples of the coded block based on the residual of the coded block.

[0024] Aspects of the present disclosure also provide a computer device, which includes a processor and a memory. The memory is used to store program codes and transmit the program codes to the processor; the processor is configured to execute according to instructions in the program codes: decode transform coefficients associated with coded blocks of a non-rectangular partition of a picture according to an encoded video bitstream; determine a residual of the coded block based on the transform coefficients; and reconstruct samples of the coded block based on the residual of the coded block.

[0025] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions, which, when executed by a computer for video decoding, cause the computer to perform a method for video decoding.

[0026] According to the method and apparatus for video decoding provided by the present disclosure, transform coefficients associated with coded blocks of a non-rectangular partition of a picture are decoded according to an encoded video bitstream, a residual of the coded block is determined based on the transform coefficients, and samples of the coded block are reconstructed based on the residual of the coded block. Through the solution of the present application, a transform scheme suitable for L-shaped partitions is provided, so that transforms can also be performed for non-rectangular partitions, thereby achieving more flexible and efficient video encoding and decoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Additional features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0028] Figure 1 is a schematic illustration of a simplified block diagram of a communication system according to an embodiment.

[0029] Figure 2 is a schematic illustration of a simplified block diagram of a communication system according to an embodiment.

[0030] Figure 3 is a schematic illustration of a simplified block diagram of a decoder according to an embodiment.

[0031] Figure 4 is a schematic illustration of a simplified block diagram of an encoder according to an embodiment.

[0032] Figure 5 Shows a block diagram of an encoder according to another embodiment.

[0033] Figure 6 Shows a block diagram of a decoder according to another embodiment.

[0034] Figure 7 Shows an example of a partitioning technique used in an example of a video coding format.

[0035] Figure 8 Shows an example of a partitioning technique used in another example of a video coding format.

[0036] Figure 9A and Figure 9B Shows an example of a partitioning technique used in another example of a video coding format.

[0037] Figure 10A and Figure 10B Shows examples of vertical center side three - way tree partitioning and horizontal center side three - way tree partitioning.

[0038] Figure 11 Shows an L - shaped block as one of the partitions in a partition from block partitioning.

[0039] Figure 12 Shows an example of an L - type partitioning tree according to some embodiments of the present disclosure.

[0040] Figure 13 Shows a table of size mapping from the transform size at the current depth to the transform size at the next depth.

[0041] Figure 14 Shows an example of transform partitioning for an intra - coded block.

[0042] Figure 15 Shows an example of transform partitioning for an inter - coded block.

[0043] Figure 16 Shows an example of transform partitioning for an L - shaped partition according to an embodiment of the present disclosure.

[0044] Figure 17 Shows an example of transform partitioning for an L - shaped partition according to an embodiment of the present disclosure.

[0045] Figure 18 Shows two examples of transform partitioning for an L - shaped partition according to some embodiments of the present disclosure.

[0046] Figure 19 Shows an example of transform partitioning for an L - shaped partition according to some embodiments of the present disclosure.

[0047] Figure 20 Shows an example of transform size adjustment according to an embodiment of the present disclosure.

[0048] Figure 21 Shows an example of transform size adjustment according to an embodiment of the present disclosure.

[0049] Figure 22 Shows an example of applying a two-dimensional transform in the order of horizontal transform and vertical transform.

[0050] Figure 23 Shows an example of applying a two-dimensional transform in the order of vertical transform and horizontal transform.

[0051] Figure 24 Shows an example of applying a two-dimensional transform in the order of horizontal transform and vertical transform.

[0052] Figure 25 Shows an example of applying a two-dimensional transform in the order of vertical transform and horizontal transform.

[0053] Figure 26 Shows an example of the scanning order of an L-shaped partition according to an embodiment of the present disclosure.

[0054] Figure 27 Shows a flowchart outlining an example of processing according to some embodiments of the present disclosure.

[0055] Figure 28 Is a schematic diagram of a computer system according to an embodiment. Detailed Description

[0056] Figure 1 Shows a simplified block diagram of a communication system (100) according to an embodiment of the present disclosure. The communication system (100) includes a plurality of terminal devices that can communicate with each other via, for example, a network (150). For example, the communication system (100) includes a first pair of terminal devices (110) and (120) interconnected via a network (150). In Figure 1 the example, the first pair of terminal devices (110) and (120) perform unidirectional transmission of data. For example, the terminal device (110) can encode video data (e.g., a video picture stream captured by the terminal device (110)) for transmission via the network (150) to another terminal device (120). The encoded video data can be transmitted in the form of one or more encoded video bitstreams. The terminal device (120) can receive the encoded video data from the network (150), decode the encoded video data to recover the video pictures, and display the video pictures according to the recovered video data. In media service applications and the like, unidirectional data transmission may be common.

[0057] In another example, a communication system (100) includes a second pair of terminal devices (130) and (140) that perform a two-way transmission of encoded video data, such as may occur during a video conference. For the two-way transmission of data, in the example, each of the terminal devices (130) and (140) may encode video data (e.g., a video picture stream captured by the terminal device) for transmission via a network (150) to the other of the terminal devices (130) and (140). Each of the terminal devices (130) and (140) may also receive the encoded video data transmitted by the other of the terminal devices (130) and (140), and may decode the encoded video data to recover the video pictures, and may display the video pictures on an accessible display device based on the recovered video data.

[0058] In Figure 1 the example, the terminal devices (110), (120), (130), and (140) may be shown as servers, personal computers, and smart phones, but the principles of the present disclosure may not be limited thereto. Embodiments of the present disclosure are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (150) represents any number of networks that convey encoded video data between the terminal devices (110), (120), (130), and (140), including, for example, wired connections (wired) and / or wireless communication networks. The communication network (150) may exchange data in circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, unless otherwise stated herein below, the architecture and topology of the network (150) may be immaterial to the operation of the present disclosure.

[0059] As an example of an application of the disclosed subject matter, Figure 2 the placement of a video encoder and a video decoder in a streaming environment is shown. The disclosed subject matter may be equivalently applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, and the like.

[0060] A streaming system can include a capture subsystem (213), which can include a video source (201), such as a digital imaging device, that creates a video picture stream (202), such as an uncompressed video picture stream. In an example, the video picture stream (202) includes samples taken by the digital imaging device. The video picture stream (202), depicted as a thick line to emphasize the high data volume when compared to the encoded video data (204) (or encoded video bitstream), can be processed by an electronic device (220) that includes a video encoder (203) coupled to the video source (201). The video encoder (203) can include hardware, software, or a combination thereof to implement or enforce aspects of the disclosed subject matter as described in more detail below. The encoded video data (204) (or encoded video bitstream (204)), depicted as a thin line to emphasize the lower data volume when compared to the video picture stream (202), can be stored on a streaming server (205) for future use. One or more streaming client subsystems, such as Figure 2 the client subsystems (206) and (208) in the example, can access the streaming server (205) to retrieve copies (207) and (209) of the encoded video data (204). The client subsystem (206) can include a video decoder (210) in an electronic device (230), for example. The video decoder (210) decodes an incoming copy (207) of the encoded video data and creates an outgoing video picture stream (211) that can be presented on a display (212), such as a display screen, or other rendering device (not depicted). In some streaming systems, the encoded video data (204), (207), and (209) (e.g., video bitstream) can be encoded according to a specific video coding / compression standard. Examples of such standards include ITU-T H.265 Recommendation. In an example, a video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter can be used in the context of VVC.

[0061] Note that the electronic devices (220) and (230) can include other components (not shown). For example, the electronic device (220) can include a video decoder (not shown), and the electronic device (230) can also include a video encoder (not shown).

[0062] Figure 3 A block diagram of a video decoder (310) according to an embodiment of the present disclosure is shown. The video decoder (310) can be included in an electronic device (330). The electronic device (330) can include a receiver (331) (e.g., receiving circuitry). The video decoder (310) can be used in place of Figure 2 the video decoder (210) in the example.

[0063] A receiver (331) may receive one or more encoded video sequences to be decoded by a video decoder (310); in the same or another embodiment, one encoded video sequence is received at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequences may be received from a channel (301), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (331) may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective consuming entities (not depicted). The receiver (331) may separate the encoded video sequences from the other data. To prevent network jitter, a buffer memory (315) may be coupled between the receiver (331) and an entropy decoder / parser (320) (hereinafter referred to as "parser (320)"). In some applications, the buffer memory (315) is part of the video decoder (310). In other applications, the buffer memory (315) may be external to the video decoder (310) (not depicted). In still other applications, a buffer memory (not depicted) may exist external to the video decoder (310) to, for example, prevent network jitter, and additionally a further buffer memory (315) may exist internal to the video decoder (310) to, for example, handle playout timing. When the receiver (331) is receiving data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (315) may not be required, or the buffer memory (315) may be small. For use on a best effort packet network such as the Internet, a buffer memory (315) may be required, which may be relatively large and may advantageously have an adaptive size, and may be implemented at least in part in an operating system or similar element (not depicted) external to the video decoder (310).

[0064] The video decoder (310) may include a parser (320) to reconstruct symbols (321) from the encoded video sequences. The categories of these symbols include: information for managing the operation of the video decoder (310); and potential information for controlling a rendering device such as a rendering device (312) (e.g., a display screen), which is not part of the electronic device (330) but may be coupled to the electronic device (330), as Figure 3As shown. The control information for presenting the device can be in the form of supplementary enhancement information (SEI message) or a video usability information (VUI) parameter set segment (not depicted). The parser (320) can parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence can be performed according to a video coding technology or standard and can follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (320) can extract subgroup parameter sets of at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroup can include a group of pictures (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (320) can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.

[0065] The parser (320) can perform an entropy decoding / parsing operation on the video sequence received from the buffer memory (315) to create symbols (321).

[0066] The reconstruction of the symbols (321) may involve multiple different units, depending on the type of the encoded video picture or a part thereof (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors. Which units are involved and how they are involved can be controlled by subgroup control information parsed by the parser (320) from the encoded video sequence. For clarity, the flow of such subgroup control information between the parser (320) and multiple units below is not depicted.

[0067] In addition to the functional blocks already mentioned, the video decoder (310) can be conceptually subdivided into multiple functional units as described below. In an actual implementation operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the following functional units.

[0068] The first unit is the scaler / inverse transform unit (351). The scaler / inverse transform unit (351) receives the quantized transform coefficients as symbols (321) from the parser (320) and control information, including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (351) can output blocks including sample values that can be input into the aggregator (355).

[0069] In some cases, the output samples of the scaler / inverse transform unit (351) can belong to an intra-coded block; that is, a block that does not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed part of the current picture. Such prediction information can be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) uses the reconstructed surrounding information extracted from the current picture buffer (358) to generate a block having the same size and shape as the block being reconstructed. The current picture buffer (358) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (355) adds the prediction information that the intra-picture prediction unit (352) has generated to the output sample information provided by the scaler / inverse transform unit (351) on a per-sample basis.

[0070] In other cases, the output samples of the scaler / inverse transform unit (351) can belong to an inter-coded block and a potential motion compensation block. In such a case, the motion compensation prediction unit (353) can access the reference picture memory (357) to extract samples for prediction. After motion compensation of the extracted samples according to the sign (321) belonging to the block, these samples can be added by the aggregator (355) to the output of the scaler / inverse transform unit (351) (referred to as residual samples or residual signals in this case), thereby generating output sample information. The address from which the motion compensation prediction unit (353) in the reference picture memory (357) extracts prediction samples can be controlled by a motion vector, and the motion vector can be available to the motion compensation prediction unit (353) in the form of a sign (321), and the sign (321) can have, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of sample values extracted from the reference picture memory (357) when sub-sample accurate motion vectors are in use, a motion vector prediction mechanism, and the like.

[0071] The output samples of the aggregator (355) can undergo various loop filtering techniques in the loop filter unit (356). The video compression technique can include in-loop filter techniques that are controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream) and available to the loop filter unit (356) as signs (321) from the parser (320), but the video compression technique can also respond to meta-information obtained during decoding of a previous (in decoding order) part of the encoded picture or the encoded video sequence, and to previously reconstructed and loop-filtered sample values.

[0072] The output of the loop filter unit (356) can be a sample stream, which can be output to a rendering device (312) and stored in a reference picture memory (357) for future inter-picture prediction.

[0073] Once fully reconstructed, some of the encoded pictures can be used as reference pictures for future prediction. For example, once the encoded picture corresponding to the current picture has been fully reconstructed and the encoded picture has been identified as a reference picture (by, for example, a parser (320)), the current picture buffer (358) can become part of the reference picture memory (357), and a new current picture buffer can be reallocated before starting to reconstruct subsequent encoded pictures.

[0074] The video decoder (310) can perform decoding operations according to a predetermined video compression technique in a standard such as ITU-T Recommendation H.265. The encoded video sequence can conform to the syntax specified by the video compression technique or standard in the sense that the encoded video sequence follows both the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard. Specifically, the profile can select certain tools from all the tools available in the video compression technique or standard as tools that are only available under the profile. For compliance, it can also be required that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by an hypothetical reference decoder (HRD) specification and metadata for signaling HRD buffer management in the encoded video sequence.

[0075] In an embodiment, the receiver (331) can receive additional (redundant) data and the encoded video. The additional data can be included as part of the encoded video sequence. The additional data can be used by the video decoder (310) to correctly decode the data and / or more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0076] Figure 4 A block diagram of a video encoder (403) according to an embodiment of the present disclosure is shown. The video encoder (403) is included in an electronic device (420). The electronic device (420) includes a transmitter (440) (e.g., transmission circuitry). The video encoder (403) can be used instead of Figure 2 the video encoder (203) in the example.

[0077] A video encoder (403) may receive video samples from a video source (401) (which, in Figure 4 the example, is not part of the electronic device (420)), and the video source (401) may capture video images to be encoded by the video encoder (403). In another example, the video source (401) is part of the electronic device (420).

[0078] The video source (401) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (403), and the digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, …), any color space (e.g., BT.601 Y CrCb, RGB, …), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (401) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (401) may be a camera device that captures local image information as a video sequence. Video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. The pictures themselves may be organized as a spatial pixel array, where each pixel may include one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0079] According to an embodiment, the video encoder (403) may encode pictures of the source video sequence in real time or under any other time constraints required by the application and compress them into an encoded video sequence (443). Implementing an appropriate encoding speed is a function of the controller (450). In some embodiments, the controller (450) controls other functional units as described below and is functionally coupled to other functional units. For clarity, the couplings are not depicted. The parameters set by the controller (450) may include rate control related parameters (picture skip, quantizer, λ value of rate distortion optimization techniques, …), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (450) may be configured to have other suitable functions that belong to the video encoder (403) optimized for certain system designs.

[0080] In some embodiments, the video encoder (403) is configured to operate in an encoding loop. As an overly simplified description, in an example, the encoding loop may include: a source encoder (430) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and reference pictures); and a (local) decoder (433) embedded in the video encoder (403). The decoder (433) reconstructs the symbols to create sample data in a manner similar to the way a (remote) decoder will also create it (since in the video compression techniques contemplated in the disclosed subject matter, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input to the reference picture memory (434). Since the decoding of the symbol stream produces bit-exact results independent of the decoder location (local or remote), the content in the reference picture memory (434) is also bit-exact between the local encoder and the remote encoder. In other words, the sample values of the reference picture samples that the prediction part of the encoder "sees" are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This reference picture synchronization principle (and the drift that occurs in cases where synchronization cannot be maintained, e.g., due to channel errors) is also used in some related fields.

[0081] The operation of the "local" decoder (433) may be the same as the operation of the "remote" decoder, such as the video decoder (310), already described in detail above. However, also briefly referring to Figure 3 , when symbols are available and the entropy encoder (445) and the parser (320) can encode / decode the symbols losslessly into an encoded video sequence, the entropy decoding part of the video decoder (310), including the buffer memory (315) and the parser (320), may not be fully implementable in the local decoder (433). Figure 3

[0082] At this point, it can be observed that any decoder technology other than the parsing / entropy decoding present in the decoder must also necessarily exist in the corresponding encoder in substantially the same functional form. Therefore, the disclosed subject matter focuses on decoder operations. The description of encoder technologies can be simplified since they are the opposite of the fully described decoder technologies. More detailed descriptions are only needed in certain areas and are provided below.

[0083] During operation, in some examples, the source encoder (430) may perform motion-compensated predictive coding that predicts an input picture by referring to one or more previously encoded pictures in the video sequence designated as "reference pictures". In this way, the encoding engine (432) encodes the difference between a pixel block of the input picture and a pixel block of a reference picture that can be selected as a prediction reference for the input picture.​

[0084] The local video decoder (433) can decode the encoded video data of a picture that can be designated as a reference picture based on the symbols created by the source encoder (430). The operation of the encoding engine (432) can advantageously be a lossy process. When the encoded video data can be decoded at a video decoder ( Figure 4 not shown), the reconstructed video sequence can generally be a copy of the source video sequence with some errors. The local video decoder (433) replicates the decoding process that can be performed by the video decoder on the reference picture, and can cause the reconstructed reference picture to be stored in the reference picture cache (434). In this way, the video encoder (403) can locally store a copy of the reconstructed reference picture, which has common content (no transmission errors) with the reconstructed reference picture that will be obtained by the remote video decoder.

[0085] The predictor (435) can perform a prediction search for the encoding engine (432). That is, for a new picture to be encoded, the predictor (435) can search the reference picture memory (434) for sample data (as a candidate reference pixel block) or some metadata that can be used as an appropriate prediction reference for the new picture, such as a reference picture motion vector, block shape, etc. The predictor (435) can operate on a per-pixel block basis of the sample blocks to find an appropriate prediction reference. In some cases, as determined by the search results obtained by the predictor (435), the input picture can have prediction references extracted from multiple reference pictures stored in the reference picture memory (434).

[0086] The controller (450) can manage the encoding operations of the source encoder (430), including, for example, setting parameters and subgroup parameters for encoding the video data.

[0087] The outputs of all the above-mentioned functional units can be subjected to entropy encoding in the entropy encoder (445). The entropy encoder (445) can convert these symbols into an encoded video sequence by losslessly compressing the symbols generated by various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0088] The transmitter (440) can buffer the encoded video sequence created by the entropy encoder (445) so as to prepare for transmission via a communication channel (460), which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter (440) can merge the encoded video data from the video encoder (403) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).

[0089] The controller (450) may manage the operation of the video encoder (403). During encoding, the controller (450) may specify a particular encoded picture type for each encoded picture, which may affect the encoding techniques that can be applied to the corresponding picture. For example, a picture may typically be specified as one of the following picture types:

[0090] An intra picture (I picture) may be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including for example instant decoder refresh (“IDR”) pictures. Those skilled in the art are aware of these variants of I pictures and their corresponding applications and characteristics.

[0091] A predictive picture (P picture) may be a picture that can be encoded and decoded using intra prediction or inter prediction that predicts the sample values of each block with at most one motion vector and a reference index.

[0092] A bi-predictive picture (B picture) may be a picture that can be encoded and decoded using intra prediction or inter prediction that predicts the sample values of each block with at most two motion vectors and reference indices. Similarly, a multi-predictive picture may use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0093] A source picture may typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples respectively), and encoded block by block. These blocks may be predictively encoded with reference to other (already encoded) blocks, which are determined by the encoding assignment applied to the corresponding picture of the block. For example, blocks of an I picture may be non-predictively encoded, or these blocks may be predictively encoded with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture may be predictively encoded with reference to one previously encoded reference picture via spatial prediction or via temporal prediction. Blocks of a B picture may be predictively encoded with reference to one or two previously encoded reference pictures via spatial prediction or via temporal prediction.

[0094] The video encoder (403) may perform encoding operations according to a predetermined video coding technique or standard such as the ITU-T H.265 recommendation. In the operation of the video encoder (403), the video encoder (403) may perform various compression operations, including predictive coding operations that exploit the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video coding technique or standard used.

[0095] In an embodiment, the transmitter (440) may transmit additional data and the encoded video. The source encoder (430) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set segments, etc.

[0096] Video may be captured as a plurality of source pictures (video pictures) in a time series. Intra-picture prediction (usually abbreviated as intra prediction) exploits the spatial correlation within a given picture, while inter-picture prediction exploits the (temporal or other) correlation between pictures. In an example, a particular picture being encoded / decoded, referred to as the current picture, is partitioned into blocks. When a block in the current picture is similar to a reference block in a reference picture that was previously encoded in the video and is still buffered, the block in the current picture may be encoded by a vector referred to as a motion vector. The motion vector points to the reference block in the reference picture, and in the case where multiple reference pictures are in use, the motion vector may have a third dimension that identifies the reference picture.

[0097] In some embodiments, dual-prediction techniques may be used in inter-picture prediction. According to the dual-prediction technique, two reference pictures are used, for example, a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be past and future respectively in display order). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be predicted by a combination of the first reference block and the second reference block.

[0098] In addition, the merge mode technique may be used in inter-picture prediction to improve the encoding efficiency.

[0099] According to some embodiments of the present disclosure, predictions such as inter - picture prediction and intra - picture prediction are performed on a block - by - block basis. For example, according to the HEVC standard, pictures in a video picture sequence are segmented into coding tree units (CTUs) for compression. CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luminance CTB and two chrominance CTBs. Each CTU can be recursively split into one or more coding units (CUs) in a quadtree. For example, a 64×64 - pixel CTU can be split into a 64×64 - pixel CU, or four 32×32 - pixel CUs, or sixteen 16×16 - pixel CUs. In an example, each CU is analyzed to determine the prediction type for the CU, e.g., an inter - prediction type or an intra - prediction type. Depending on the temporal and / or spatial predictability, the CU is split into one or more prediction units (PUs). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, the prediction operation in decoding (encoding / decoding) is performed on a prediction - block basis. Using a luminance prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values (e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0100] Figure 5 FIG. shows a video encoder (503) according to another embodiment of the present disclosure. The video encoder (503) is configured to receive a processing block (e.g., a prediction block) of sample values within a current video picture in a video picture sequence and encode the processing block into an encoded picture that is part of an encoded video sequence. In an example, the video encoder (503) is used instead of Figure 2 the video encoder (203) in the example.

[0101] In the HEVC example, a video encoder (503) receives a matrix of sample values for processing blocks such as prediction blocks of 8×8 samples. The video encoder (503) uses, for example, rate-distortion optimization to determine whether to use an intra mode, an inter mode, or a bi-prediction mode to optimally encode the processing block. When encoding the processing block in the intra mode, the video encoder (503) may use intra prediction techniques to encode the processing block into the encoded picture; and when encoding the processing block in the inter mode or the bi-prediction mode, the video encoder (503) may use inter prediction techniques or bi-prediction techniques respectively to encode the processing block into the encoded picture. In some video coding techniques, the merge mode may be an inter-picture prediction sub-mode, where a motion vector is derived from one or more motion vector predictors without resorting to encoded motion vector components external to the predictor. In some other video coding techniques, there may be motion vector components applicable to the subject block. In the example, the video encoder (503) includes other components such as a mode decision module (not shown) for determining the mode of the processing block.

[0102] In Figure 5 the example, the video encoder (503) includes as Figure 5 shown an inter-frame encoder (530), an intra-frame encoder (522), a residual calculator (523), a switch (526), a residual encoder (524), a general controller (521), and an entropy encoder (525) coupled together.

[0103] The inter-frame encoder (530) is configured to: receive samples of a current block (e.g., the processing block); compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a subsequent picture); generate inter-frame prediction information (e.g., a motion vector, merge mode information, a description of redundant information according to inter-frame coding techniques); and calculate an inter-frame prediction result (e.g., a predicted block) based on the inter-frame prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture decoded based on the encoded video information.

[0104] The intra-frame encoder (522) is configured to: receive samples of a current block (e.g., the processing block); compare the block with already encoded blocks in the same picture in some cases; generate quantization coefficients after transformation; and also generate intra-frame prediction information (e.g., intra-frame prediction direction information according to one or more intra-frame coding techniques) in some cases. In the example, the intra-frame encoder (522) also calculates an intra-frame prediction result (e.g., a predicted block) based on the intra-frame prediction information and reference blocks in the same picture.

[0105] The general controller (521) is configured to determine general control data and control other components of the video encoder (503) based on the general control data. In an example, the general controller (521) determines the mode of a block and provides a control signal to the switch (526) based on the mode. For example, when the mode is an intra mode, the general controller (521) controls the switch (526) to select the intra mode result for use by the residual calculator (523) and controls the entropy encoder (525) to select the intra prediction information and include the intra prediction information in the bitstream; and when the mode is an inter mode, the general controller (521) controls the switch (526) to select the inter prediction result for use by the residual calculator (523) and controls the entropy encoder (525) to select the inter prediction information and include the inter prediction information in the bitstream.

[0106] The residual calculator (523) is configured to calculate the difference (residual data) between a received block and a prediction result selected from the intra encoder (522) or the inter encoder (530). The residual encoder (524) is configured to operate on the residual data to encode the residual data to generate transform coefficients. In an example, the residual encoder (524) is configured to transform the residual data from the spatial domain to the frequency domain and generate transform coefficients. The transform coefficients are then subjected to a quantization process to obtain quantized transform coefficients. In various embodiments, the video encoder (503) further includes a residual decoder (528). The residual decoder (528) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be suitably used by the intra encoder (522) and the inter encoder (530). For example, the inter encoder (530) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (522) can generate a decoded block based on the decoded residual data and the intra prediction information. In some examples, the decoded blocks are suitably processed to generate a decoded picture, and the decoded picture can be buffered in a memory circuit (not shown) and used as a reference picture.

[0107] The entropy encoder (525) is configured to format the bitstream to include the encoded blocks. The entropy encoder (525) is configured to include various information according to a suitable standard such as the HEVC standard. In an example, the entropy encoder (525) is configured to include the general control data, the selected prediction information (e.g., intra prediction information or inter prediction information), the residual information, and other suitable information in the bitstream. Note that according to the disclosed subject matter, there is no residual information when encoding a block in an inter mode or the merge submode of the bi-prediction mode.

[0108] Figure 6FIG. showing a video decoder (610) according to another embodiment of the present disclosure. The video decoder (610) is configured to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In an example, the video decoder (610) is used in place of Figure 2 the video decoder (210) in the example.

[0109] In Figure 6 the example, the video decoder (610) includes an entropy decoder (671), an inter-frame decoder (680), a residual decoder (673), a reconstruction module (674), and an intra-frame decoder (672) coupled together as Figure 6 shown.

[0110] The entropy decoder (671) may be configured to reconstruct certain symbols from the encoded picture, where such symbols represent syntax elements that make up the encoded picture. Such symbols may include, for example, the mode in which a block is encoded (such as, for example, an intra mode, an inter mode, a bi-prediction mode, a merge sub-mode of the latter two, or another sub-mode); prediction information that can respectively identify certain samples or metadata used for prediction by the intra-frame decoder (672) or the inter-frame decoder (680) (such as, for example, intra prediction information or inter prediction information); residual information in the form of, for example, quantized transform coefficients, etc. In an example, when the prediction mode is an inter mode or a bi-prediction mode, the inter prediction information is provided to the inter-frame decoder (680); and when the prediction type is an intra prediction type, the intra prediction information is provided to the intra-frame decoder (672). The residual information may be inverse quantized and provided to the residual decoder (673).

[0111] The inter-frame decoder (680) is configured to: receive inter prediction information and generate an inter prediction result based on the inter prediction information.

[0112] The intra-frame decoder (672) is configured to: receive intra prediction information and generate a prediction result based on the intra prediction information.

[0113] The residual decoder (673) is configured to: perform inverse quantization to extract dequantized transform coefficients and process the dequantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (673) may also require certain control information (to include quantization parameter (QP)), and such information may be provided by the entropy decoder (671) (the data path is not depicted as this may be only low-volume control information).

[0114] The reconstruction module (674) is configured to: combine, in the spatial domain, the residual output by the residual decoder (673) with the prediction result (output by the inter-frame prediction module or the intra-frame prediction module as the case may be) to form a reconstructed block, which can be part of a reconstructed picture, and the reconstructed picture can in turn be part of a reconstructed video. Note that other suitable operations such as deblocking operations can be performed to improve the visual quality.

[0115] Note that any suitable technology can be used to implement the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610). In an embodiment, one or more integrated circuits can be used to implement the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610). In another embodiment, one or more processors executing software instructions can be used to implement the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610).

[0116] Aspects of the present disclosure provide transform schemes for non-rectangular partitions such as L-shaped partitions and the like.

[0117] Generally, a picture is divided into blocks, and the blocks can be units for various processes such as encoding, prediction, transformation, etc. Various block partitioning techniques can be used.

[0118] Figure 7 An example of a partitioning technique used by the Alliance for Open Media (AOMedia) in the video coding format VP9 is shown. For example, a picture (710) is divided into multiple blocks (720) of size 64×64 (e.g., 64 samples × 64 samples). Additionally, a 4-way partitioning tree can start from the 64×64 level and go down to smaller blocks, and the lowest level can be the 4×4 level (e.g., block size of 4 samples × 4 samples). In some examples, some additional restrictions can apply to blocks of 8×8 and below. In Figure 7 the example, one of the first path (721), the second path (722), the third path (723), and the fourth path (724) can be used to divide the 64×64 block (720) into smaller blocks. Note that the partition designated as R (shown in the fourth path (724)) is called a recursive partition because the same partitioning tree can be repeated at lower levels until the lowest 4×4 level.

[0119] Figure 8Shows an example of the partitioning technique used in the video coding format AOMediaVideo 1 (AV1) designed for video transmission over the Internet. AV1 was developed as the successor to VP9. For example, a picture (810) is divided into multiple blocks (820) of size 128×128 (e.g., 128 samples × 128 samples). Additionally, a 10-way partitioning structure can start from 128×128 and go down to smaller blocks. In Figure 8 the example, one of the ten ways (821) to (830) can be used to divide the 128×128 block into smaller blocks. AV1 not only extends the partitioning tree to a 10-way structure but also increases the maximum size (referred to as a superblock in VP9 / AV1 terminology) to start from 128×128. Note that the 10-way structure includes 4:1 / 1:4 rectangular partitions that do not exist in VP9. In the example, none of the rectangular partitions can be further subdivided. Additionally, AV1 adds greater flexibility for using partitions at levels lower than 8×8, in the sense that 2×2 chroma inter-frame prediction can now be performed in some cases.

[0120] In some examples, the block partitioning structure is referred to as an encoding tree. In an example (e.g., HEVC), the encoding tree can have a quadtree structure where each split divides a larger square block into four smaller square blocks. In some examples, a picture is split into coding tree units (CTUs), and then the CTUs are split into smaller blocks using a quadtree structure. According to the quadtree structure, the coding tree unit (CTU) is split into coding units (CUs) to adapt to various local features. A decision can be made at the CU level on whether to use inter-picture (temporal) prediction or intra-picture (spatial) prediction to encode a picture region. Each CU can be further split into one, two, or four prediction units (PUs) according to the PU split type. Within a PU, the same prediction process is applied, and relevant information is transmitted to the decoder based on the PU.

[0121] After obtaining the residual block by applying the prediction process based on the PU split type, the CU can be divided into transform units (TUs) according to another quadtree structure. In the example of HEVC, there are multiple partitioning concepts including CUs, PUs, and TUs. In some embodiments, the CU or TU can only be square-shaped, while the PU can be square or rectangular-shaped. In some embodiments, an encoding block can be further split into four square sub-blocks, and a transform is performed on each sub-block, i.e., the TU. Each TU can be further recursively split into smaller TUs using a quadtree structure called the residual quadtree (RQT).

[0122] In some examples (e.g., HEVC), at the picture boundary, implicit quadtree splitting can be adopted such that the blocks will maintain quadtree splitting until the size fits the picture boundary.

[0123] In some examples (e.g., VVC), the block partitioning structure can use a quadtree plus binary tree (QTBT) block partitioning structure. The QTBT structure can remove the concepts of multiple partitioning types (CU, PU, and TU concepts) and support greater flexibility in the CU partitioning shape. In the QTBT block partitioning structure, a CU can have a square or rectangular shape.

[0124] Figure 9A Shows a CTU (910) segmented by using Figure 9B the shown QTBT block partitioning structure (920). The CTU (910) is first segmented by a quadtree structure. The quadtree leaf nodes are further segmented by a binary tree structure or a quadtree structure. There can be two types of splits in the binary tree splitting, symmetric horizontal split (e.g., labeled as "0" in the QTBT block partitioning structure (920)) and symmetric vertical split (e.g., labeled as "1" in the QTBT block partitioning structure (920)). The leaf nodes that do not require further splitting are called CUs, and the CUs can be used for prediction and transform processing without any further splitting. Therefore, the CUs, PUs, and TUs have the same block size in the QTBT block partitioning structure.

[0125] In some examples (e.g., JEM), a CU can include coding blocks (CBs) of different color components. For example, in the case of P slices and B slices in 4:2:0 chroma format, one CU contains one luma CB and two chroma CBs. A CU can include CBs of a single color component. For example, in the case of I slices, one CU contains only one luma CB or only two chroma CBs.

[0126] In some embodiments, the following parameters are defined for the QTBT block partitioning scheme:

[0127] – CTU size: The size of the root node of the quadtree, e.g., the same concept as in HEVC.

[0128] – MinQTSize: The minimum allowed quadtree leaf node size.

[0129] – MaxBTSize: The maximum allowed binary tree root node size.

[0130] – MaxBTDepth: The maximum allowed binary tree depth.

[0131] – MinBTSize: The minimum allowed binary tree leaf node size.

[0132] In an example of the QTBT block partitioning structure, the CTU size is set to 128×128 luma samples and two corresponding 64×64 chroma samples, the MinQTSize is set to 16×16, the MaxBTSize is set to 64×64, the MinBTSize (for both width and height) is set to 4×4, and the MaxBTDepth is set to 4. First, quadtree partitioning is applied to the CTU to generate quadtree leaf nodes. The size of the quadtree leaf nodes can range from 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). If the leaf quadtree node is 128×128, then since the size exceeds MaxBTSize (i.e., 64×64), the leaf quadtree node will not be further split by the binary tree. Otherwise, the leaf quadtree node can be further split by the binary tree. Thus, the quadtree leaf node is also the root node of the binary tree and the quadtree leaf node has a binary tree depth of 0.

[0133] When the binary tree depth reaches MaxBTDepth (i.e., 4), further splitting is no longer considered. When the width of the binary tree node equals MinBTSize (i.e., 4), further horizontal splitting is no longer considered. Similarly, when the height of the binary tree node equals MinBTSize, further vertical splitting is no longer considered. The leaf nodes of the binary tree are further processed by prediction and transform processing without any further partitioning. In an embodiment, the maximum CTU size is 256×256 luma samples.

[0134] In Figure 9A and Figure 9B the solid lines indicate quadtree splitting and the dashed lines indicate binary tree splitting. In each split (i.e., non-leaf) node of the binary tree, a flag indicating which split type (i.e., horizontal or vertical) is used is signaled. For example, 0 indicates horizontal splitting and 1 indicates vertical splitting. For quadtree splitting, no indication of the split type is needed because quadtree splitting can split the block both horizontally and vertically to produce 4 sub-blocks of equal size.

[0135] In some embodiments, the QTBT block partitioning scheme supports the flexibility of having separate QTBT block partitioning structures for luma and chroma. For example, for P slices and B slices, the luma blocks and chroma blocks in a CTU share the same QTBT block partitioning structure. However, for I slices, the luma CTB is split into CUs by the QTBT block partitioning structure and the chroma blocks are split into chroma CUs by another QTBT block partitioning structure. Thus, the CUs in I slices consist of coded blocks of the luma component or coded blocks of the two chroma components, and the CUs in P slices or B slices consist of coded blocks of all three color components.

[0136] In some examples (e.g., HEVC), inter - prediction of small blocks is restricted to reduce memory access for motion compensation. For example, 4×8 and 8×4 blocks do not support dual prediction, and 4×4 blocks do not support inter - prediction.

[0137] In addition, in some examples (e.g., VCC), a multi - type tree (MTT) block - partitioning structure is used. The MTT block - partitioning structure is a more flexible tree structure than the QTBT block - partitioning structure. In the MTT, in addition to quadtree partitioning and binary - tree partitioning, horizontal center - side ternary - tree partitioning and vertical center - side ternary - tree partitioning can be used.

[0138] Figure 10A An example of vertical center - side ternary - tree partitioning is shown, and Figure 10B An example of horizontal center - side ternary - tree partitioning is shown. Ternary - tree partitioning can complement quadtree partitioning and binary - tree partitioning. For example, ternary - tree partitioning can capture an object located at the center of a block, while quadtree and binary - tree can split along the block center. The width and height of the partition by the ternary - tree are powers of 2, such that no additional transform partitioning is required.

[0139] The design of block partitioning is mainly for reducing complexity. Theoretically, the complexity of traversing the tree is T D , where T represents the number of split types, and D represents the depth of the tree.

[0140] According to some aspects of the present disclosure, non - rectangular shape partitioning, such as L - shaped partitioning, etc., can be used in block partitioning. For example, a coding unit or a coding block can have an L - shape.

[0141] Figure 11 An L - shaped block (1100) is shown as one of the partitions from block partitioning. For example, instead of using rectangular block partitioning, L - type partitioning can split a block into one or more L - shaped partitions and one or more rectangular partitions. As Figure 11 shown, an L - shaped (or L - type) partition can be defined by four parameters, which are called width, shorter width, height, and shorter height. The L - shaped partition (1100) can be appropriately rotated or flipped.

[0142] Figure 12 Examples of L - type partition trees (1210), (1220), (1230), and (1240) according to some embodiments of the present disclosure are shown. In the examples of (1210), (1220), (1230), and (1240), a rectangular block can be split into two partitions, including an L - shaped partition (shown by "1") and a rectangular partition (shown by "0").

[0143] In some examples (e.g., AV1), transform block partitioning may be performed to generate transform blocks. For both intra-coded blocks and inter-coded blocks, the coded blocks may be further divided into multiple transform units, where the partitioning depth is up to two levels (two depths). In an example, a 1-1 size mapping from the current depth to the next depth is used.

[0144] Figure 13 A table showing the size mapping from the transform size at the current depth to the transform size at the next depth is shown. Note that the numbers in the table refer to the number of samples. For example, the last row indicates that when the transform unit size at the current depth is 64 samples by 16 samples and the transform unit is further split, the transform unit size at the next depth is 32 samples by 16 samples.

[0145] In some examples, different transform partitioning techniques may be used for intra-coded blocks and inter-coded blocks.

[0146] Figure 14 An example of transform partitioning for intra-coded blocks is shown. For intra-coded blocks, transform partitioning is performed in such a way that all transform blocks have the same size, and the transform blocks are coded in raster scan order. For example, an intra-coded block (1410) may be divided into four transform blocks of the same size (e.g., using quadtree partitioning), as shown by (1420). In another example, an intra-coded block (1410) may be divided into 16 blocks of the same size (e.g., using quadtree partitioning with two depths), as shown by (1430). The transform blocks may be coded in the raster scan order shown by the arrow lines in Figure 14 as shown.

[0147] For inter-coded blocks, transform unit partitioning may be performed recursively, where the partitioning depth is up to two levels (2 depths). The transform partitioning in inter-coded blocks may support different shapes, e.g., transform unit sizes of 1:1 (square shape), 1:2 / 2:1, and 1:4 / 4:1, ranging from 4×4 to 64×64. In an example, when the coded block is less than or equal to 64×64, the transform block partitioning may be applied only to the luminance component; then for the chrominance blocks, the transform block size is the same as the coded block size. Otherwise, when the coded block width or height is greater than 64, both the luminance coded block and the chrominance coded block may be implicitly split into multiples of min(W,64)×min(H,64) transform blocks and min(W,32)×min(H,32) transform blocks, respectively.

[0148] Figure 15Shows an example of a transform partition for an inter-coded block. For example, at the first transform partition depth, an intra-coded block (1510) is divided into four transform blocks A, B, C, and D of the same size (e.g., using quadtree partitioning). Then, at the second transform partition depth, transform block B is divided into four transform blocks B1, B2, B3, and B4 (e.g., using quadtree partitioning). The transform blocks can be encoded in raster scan order, for example, in the order of A, B1, B2, B3, B4, C, and D, as Figure 15 shown by the arrow lines in

[0149] Note that when the coded block is an L-shaped partition, and after prediction, the residual can be L-shaped and the residual block can be referred to as an L-shaped residual block. In the following description of the transform scheme, L-shaped partitioning can be used to define the geometric properties of the processing unit, and the processing unit can be any one of a residual block, a transform coefficient block, an intermediate transform coefficient block, etc. Aspects of the present disclosure provide a transform scheme for non-rectangular shaped residual blocks such as L-shaped residual blocks, etc.

[0150] It should also be noted that in the following description, although the transform scheme is described for L-shaped partitioning, the transform scheme can be suitably applied to other non-rectangular blocks.

[0151] According to an aspect of the present disclosure, for an L-shaped partition (e.g., an L-shaped coded block), a transform can be applied to the entire L-shaped residual block. In some embodiments, a separable two-dimensional transform can be applied to the entire L-shaped residual block. In some embodiments, an inseparable transform can be applied to the entire L-shaped residual block. In some embodiments, the L-shaped residual block is suitably adjusted to form a rectangular residual block, and a transform can be applied to the rectangular residual block.

[0152] According to another aspect of the present disclosure, an L-shaped partition (e.g., an L-shaped residual block, an L-shaped transform coefficient block) can be further split into multiple transform units, such as multiple transform units of rectangular shape. Then, the transform can be performed on the transform units, for example, multiple transform units of rectangular shape, respectively.

[0153] In some embodiments, it can be signaled in the bitstream whether the L-shaped partition is split. In an example, a sequence-level flag is used to indicate whether further splitting is allowed for L-shaped partitions in the current sequence. When the sequence-level flag indicates that further splitting of the current sequence is not allowed, then the transform is applied to each of the entire L-shaped residual blocks in the sequence. When the sequence-level flag indicates that further splitting of the current sequence is allowed, a coding-block-level flag can be used to signal whether further splitting is applied to the current L-shaped coded block.

[0154] According to aspects of the present disclosure, an L-shaped coding block can be divided into two or more rectangular sub-blocks, such as at least two square or non-square rectangular sub-blocks. The rectangular sub-blocks can be used as transform units. Then, a transform scheme for rectangular blocks can be applied to the residuals of each rectangular sub-block.

[0155] In some embodiments, an L-shaped coding block can be divided into three rectangular transform units of the same size. The rectangular transform units can be square or non-square rectangles. Additionally, in some examples, each transform unit can be split or not split individually.

[0156] Figure 16 An example of transform partitioning for L-shaped partitioning according to an embodiment of the present disclosure is shown. In Figure 16 , the L-shaped block (1600) (e.g., coding unit) is divided into three rectangular blocks (1601) to (1603) (e.g., transform units). The partitioning (also referred to as transform splitting) is indicated by the dashed lines. In the example, the three rectangular blocks (1601) to (1603) can be three transform units.

[0157] In an embodiment, for an intra-coded block of L-shaped partitioning, all transform units within the L-shaped partitioning have the same size. Using Figure 16 as an example, the L-shaped block (1600) is an intra-coded block, and the rectangular blocks (1601) to (1603) have a square shape and the same size.

[0158] In another embodiment, for an inter-coded block of L-shaped partitioning, the transform units can be further split individually into smaller transform units.

[0159] Figure 17 An example of transform partitioning for L-shaped partitioning according to an embodiment of the present disclosure is shown. In Figure 17 , the L-shaped block (1700) is divided into three rectangular blocks (1701) to (1703). The partitioning (also referred to as transform splitting) is indicated by the dashed lines. Additionally, the rectangular blocks (1701) and (1703) are not further split, but the rectangular block (1702) is further split into four rectangular blocks A to D. Then, in the example, the blocks (1701), (1703), and A to D can be transform units. In the example, the blocks (1701), (1703), and A to D can have a square shape.

[0160] In some embodiments, an L-shaped coding block can be divided into multiple transform units having different sizes and / or aspect ratios. In an example, an L-shaped coding block can be divided into two transform units, such as a square transform unit and a non-square rectangular transform unit. Additionally, further splitting of each transform unit can be determined individually.

[0161] Figure 18 show two examples of transformed partitions for an L-shaped partition according to some embodiments of the present disclosure. In Figure 18 , the L-shaped block (1810) is divided into two rectangular blocks (1811) and (1812) by a vertical split represented by a dashed line. In the example, at least one of the blocks (1811) and (1812), e.g., (1811), has a square shape. Further, in Figure 18 , the L-shaped block (1820) is divided into two rectangular blocks (1821) and (1822) by a horizontal split represented by a dashed line. In the example, at least one of the blocks (1821) and (1822), e.g., (1821), has a square shape.

[0162] In an embodiment, when the transform unit has a square shape, further splitting of the transform unit can be determined separately, and further splitting cannot be applied to a non-square rectangular transform unit.

[0163] In another embodiment, further splitting of both the square transform unit and the non-square rectangular transform unit can be determined separately.

[0164] In some embodiments, an L-shaped coding block can be divided into a square transform unit and a non-square rectangular transform unit only when the L-shaped coding unit is an inter-coded block.

[0165] In some embodiments, an L-shaped coding block can be divided into three square or non-square rectangular transform units. The sizes of the three transform units can be the same or can be different. Further, further splitting of each transform unit can be determined separately.

[0166] Figure 19 show an example of transformed partition for an L-shaped partition according to some embodiments of the present disclosure. In Figure 19 , the L-shaped block (1910) is divided into three rectangular blocks (1911), (1912), and (1913) by a vertical split and a horizontal split represented by dashed lines.

[0167] Note that the residual of the L-shaped coding block can have the same L shape as the coding block and can be referred to as an L-shaped residual block. In the following description, the L-shaped partition can be used to refer to the L-shaped residual block. Some aspects of the present disclosure also provide the following technique: using a separable transform, e.g., a separable two-dimensional transform including a horizontal transform and a vertical transform, for the L-shaped residual block without further splitting the L-shaped residual block into rectangular blocks. When performing a separable two-dimensional transform on the L-shaped partition, the transform size can be different for different rows / columns of the L-shaped residual block.

[0168] In an embodiment, when performing a horizontal transform on the residual samples (also referred to as input samples) of each row of the L-shaped partition, the transform size is adjusted according to the number of residual samples (input samples) in the row. For example, the transform size is the same as the number of residual samples in the corresponding row.

[0169] Figure 20 Shows an example of transform size adjustment according to an embodiment of the present disclosure. In Figure 20 , the L-shaped transform unit (2000) includes 8 rows respectively referred to as ROW1 to ROW8. Each of ROW1 to ROW4 includes 8 samples, and each of ROW5 to ROW8 includes 4 samples. In the example, an 8-point transform can be performed on each of the first four rows (ROW1 to ROW4), and a 4-point transform can be performed on each of the last four rows ROW5 to ROW8.

[0170] In another embodiment, when performing a vertical transform on the residual samples (also referred to as input samples) of each column of the L-partition, the transform size is adjusted according to the number of residual samples in the column. For example, the transform size is the same as the number of residual samples in the corresponding column.

[0171] Figure 21 Shows an example of transform size adjustment according to an embodiment of the present disclosure. In Figure 21 , the L-shaped transform unit (2100) includes 8 columns respectively referred to as COL1 to COL8. Each of COL1 to COL4 includes 8 samples, and each of COL5 to COL8 includes 4 samples. In the example, an 8-point transform is performed on each of the first four columns COL1 to COL4, and a 4-point transform is performed on each of the last four columns COL5 to COL8.

[0172] According to an aspect of the present disclosure, a two-dimensional transform for the L-shaped partition can be performed by sequentially performing a horizontal transform and a vertical transform, for example, in the order of the horizontal transform and the vertical transform, or in the order of the vertical transform and the horizontal transform.

[0173] In some embodiments, the two-dimensional transform is performed in the order of the horizontal transform and the vertical transform. When applying the horizontal transform to the residual block of the L-shape, for a row of K samples (K is a positive integer), K transform coefficients in the row are determined. Note that for different rows, the number of transform coefficients can be different. After applying the horizontal transform, the transform coefficients of the row are properly aligned into a transform coefficient block, and then the vertical transform can be applied to the transform coefficient block. In an embodiment, the transform coefficients of the row are aligned to the left, and the transform coefficient block has the same shape as the residual block.

[0174] Figure 22An example of applying a two-dimensional transform in the order of a horizontal transform and a vertical transform is shown. In Figure 22 the example, a horizontal transform is applied to an L-shaped residual block (2210). Specifically, the L-shaped residual block (2210) includes 8 rows respectively called ROW1 to ROW8. Each of ROW1 to ROW4 includes 8 residual samples, and each of ROW5 to ROW8 includes 4 residual samples. In the example, an 8-point transform can be performed on each of the first four rows ROW1 to ROW4 to determine 8 transform coefficients in each of the first four rows ROW1 to ROW4; and an 4-point transform can be performed on each of the last four rows ROW5 to ROW8 to determine 4 transform coefficients in each of the last four rows ROW5 to ROW8.

[0175] Then, the transform coefficients of ROW1 to ROW8 are aligned to the left to form a transform coefficient block (2220) having the same L-shape as the residual block (2210). Then, a vertical transform is applied to the L-shaped transform coefficient block (2220). Specifically, the L-shaped transform coefficient block (2220) includes 8 columns respectively called COL1 to COL8. Each of COL1 to COL4 includes 8 transform coefficient samples, and each of COL5 to COL8 includes 4 transform coefficient samples. In the example, an 8-point transform can be performed on each of the first four columns COL1 to COL4 to determine 8 final transform coefficients in each of the first four columns COL1 to COL4; and an 4-point transform can be performed on each of the last four columns COL5 to COL8 to determine 4 final transform coefficients in each of the last four columns COL5 to COL8.

[0176] In some embodiments, the two-dimensional transform is performed in the order of a vertical transform and a horizontal transform. When a vertical transform is applied to an L-shaped residual block, for a column of K samples (K is a positive integer), K transform coefficients in the column are determined. Note that for different columns, the number of transform coefficients of the column can be different. After applying the vertical transform, the transform coefficients of the column can be appropriately aligned into a transform coefficient block, and then a horizontal transform can be applied to the transform coefficient block. In the embodiment, the transform coefficients of the column are aligned with the top, and the transform coefficient block has the same shape as the residual block.

[0177] Figure 23 An example of applying a two-dimensional transform in the order of a vertical transform and a horizontal transform is shown. In Figure 23In the example, a vertical transform is applied to an L-shaped residual block (2310). Specifically, the L-shaped residual block (2310) includes eight columns respectively referred to as COL1 to COL8. Each of COL1 to COL4 includes eight residual samples, and each of COL5 to COL8 includes four residual samples. In the example, an 8-point transform can be performed on each of the first four columns COL1 to COL4 to determine eight transform coefficients in each of the first four columns COL1 to COL4; and a 4-point transform can be performed on each of the last four columns COL5 to COL8 to determine four transform coefficients in each of the last four columns COL5 to COL8.

[0178] Then, the transform coefficients of COL1 to COL8 are aligned at the top to form a transform coefficient block (2320) having the same L-shape as the residual block (2310). Then, a horizontal transform is applied to the L-shaped transform coefficient block (2320). Specifically, the L-shaped transform coefficient block (2320) includes eight rows respectively referred to as ROW1 to ROW8. Each of ROW1 to ROW4 includes eight transform coefficient samples, and each of ROW5 to ROW8 includes four transform coefficient samples. In the example, an 8-point transform can be performed on each of the first four rows ROW1 to ROW4 to determine eight final transform coefficients in each of the first four rows ROW1 to ROW4; and a 4-point transform can be performed on each of the last four rows ROW5 to ROW8 to determine four final transform coefficients in each of the last four rows ROW5 to ROW8.

[0179] According to aspects of the present disclosure, for a two-dimensional transform, after the first transform, the transform coefficients can be aligned in other suitable ways to form a transform coefficient block having a shape different from that of the residual block, and then a second transform can be applied to the transform coefficient block to determine the final transform coefficients.

[0180] In the example, the two-dimensional transform is performed in the order of horizontal transform and vertical transform. When a horizontal transform is applied to an L-shaped residual block, for a row of K samples (K is a positive integer), K transform coefficients in the row are determined. Note that the number of transform coefficients can be different for different rows. After applying the horizontal transform, the transform coefficients of the row are suitably aligned (e.g., based on energy curve similarity for better coding efficiency) into the transform coefficient block, and then a vertical transform can be applied to the transform coefficient block. In the example, the transform coefficients of the row with a smaller number of transform coefficients are aligned with the selected positions (e.g., odd positions or even positions) of the transform coefficients of the row with a larger number of transform coefficients.

[0181] Figure 24 An example of applying the two-dimensional transform in the order of horizontal transform and vertical transform is shown. In Figure 24In the example, a horizontal transform is applied to an L-shaped residual block (2410). Specifically, the L-shaped residual block (2410) includes 8 rows respectively referred to as ROW1 to ROW8. Each of ROW1 to ROW4 includes 8 residual samples, and each of ROW5 to ROW8 includes 4 residual samples. In the example, an 8-point transform can be performed on each of the first four rows ROW1 to ROW4 to determine 8 transform coefficients in each of the first four rows ROW1 to ROW4; and a 4-point transform can be performed on each of the last four rows ROW5 to ROW8 to determine 4 transform coefficients in each of the last four rows ROW5 to ROW8.

[0182] Then, the transform coefficients of ROW5 to ROW8 are aligned to the odd-column positions of ROW1 to ROW4 to form a transform coefficient block (2420) having a different shape from the residual block (2410). Then, a vertical transform is applied to the non-rectangular-shaped transform coefficient block (2420). Specifically, the transform coefficient block (2420) includes 8 columns respectively referred to as COL1 to COL8. Each of the odd columns COL1, COL3, COL5, and COL7 includes 8 transform coefficient samples, and each of the even columns COL2, COL4, COL6, and COL8 includes 4 transform coefficient samples. In the example, an 8-point transform can be performed on each of the odd columns COL1, COL3, COL5, and COL7 to determine 8 final transform coefficients in each of the odd columns COL1, COL3, COL5, and COL7; and a 4-point transform can be performed on each of the even columns COL2, COL4, COL6, and COL8 to determine 4 final transform coefficients in each of the even columns COL2, COL4, COL6, and COL8.

[0183] In some embodiments, the two-dimensional transform is performed in the order of the vertical transform and the horizontal transform. When the vertical transform is applied to an L-shaped residual block, for a column of K samples (K being a positive integer), K transform coefficients in the column are determined. Note that for different columns, the number of transform coefficients of the column can be different. After applying the vertical transform, the transform coefficients of the column are appropriately aligned (e.g., based on energy) into the transform coefficient block, and then the horizontal transform can be applied to the transform coefficient block. In the example, the transform coefficients of the column with a smaller number of transform coefficients are aligned to the selected positions (e.g., odd positions or even positions) of the transform coefficients of the column with a larger number of transform coefficients.

[0184] Figure 25 An example of applying the two-dimensional transform in the order of the vertical transform and the horizontal transform is shown. In Figure 25In the example, a vertical transform is applied to an L-shaped residual block (2510). Specifically, the L-shaped residual block (2510) includes eight columns respectively referred to as COL1 to COL8. Each of COL1 to COL4 includes eight residual samples, and each of COL5 to COL8 includes four residual samples. In the example, an 8-point transform can be performed on each of the first four columns COL1 to COL4 to determine eight transform coefficients in each of the first four columns COL1 to COL4; and a 4-point transform can be performed on each of the last four columns COL5 to COL8 to determine four transform coefficients in each of the last four columns COL5 to COL8.

[0185] Then, the transform coefficients of COL5 to COL8 are aligned to the odd-row positions of COL1 to COL4 to form a transform coefficient block (2520) having a different shape from the residual block (2510). Then, a horizontal transform is applied to the L-shaped transform coefficient block (2520). Specifically, the L-shaped transform coefficient block (2520) includes eight rows respectively referred to as ROW1 to ROW8. Each of the odd rows ROW1, ROW3, ROW5, and ROW7 includes eight transform coefficient samples, and each of the even rows ROW2, ROW4, ROW6, and ROW8 includes four transform coefficient samples. In the example, an 8-point transform can be performed on each of the odd rows ROW1, ROW3, ROW5, and ROW7 to determine eight final transform coefficients in each of the odd rows ROW1, ROW3, ROW5, and ROW7; and a 4-point transform can be performed on each of the even rows ROW2, ROW4, ROW6, and ROW8 to determine four final transform coefficients in each of the even rows ROW2, ROW4, ROW6, and ROW8.

[0186] In some embodiments, the scanning technique for rectangular-shaped transform coefficients can be appropriately adjusted for scanning non-rectangular-shaped transform coefficients during entropy coding. For example, when performing entropy coding on the transform coefficients of an L-partition having a maximum of W samples in the horizontal direction and a maximum of H samples in the vertical direction, first obtain the scanning order of a W×H rectangular block. Then, along the scanning order, if the scanning position points to a sample position outside the L-partition, the scanning position can be skipped, and the scanning continues to the next scanning position. The coefficient scanning process ends only when all samples of the L-partition have been scanned.

[0187] Figure 26 An example of the scanning order of an L-shaped partition (2600) according to an embodiment of the present disclosure is shown. In Figure 26 the example, first obtain the zigzag scanning order (2610) for an 8×8 block, and then obtain the scanning order of the L-shaped partition by skipping samples outside the L-shaped partition. InFigure 26 In the example, the actual scan positions in the L-shaped partition are shown by solid lines, and the skipped positions outside the L-shaped partition are shown by dashed lines.

[0188] In some embodiments, a non-separable transform is used in the two-dimensional transform on the L-partition. In an embodiment, the non-separable transform can be the KLT. Using Figure 26 the L-shaped partition in as an example, the L-shaped partition (2600) includes 48 samples, and the 48 samples can form a vector of 48 elements, and then a KLT transform can be performed on the vector to generate, for example, 48 transform coefficients.

[0189] In some embodiments, when performing a two-dimensional transform on the L-shaped partition, padding can be performed to add additional samples at positions outside the L-shaped partition, and then the added samples and the samples in the L-shaped partition can form a virtual rectangular block. In addition, 2D transform techniques for rectangular blocks can be applied to the virtual rectangular block to obtain the transform coefficients for the L-partition.

[0190] In an embodiment, the existing samples in the L-shaped partition are used to obtain the additional samples.

[0191] In another embodiment, predefined values are used to obtain the additional samples.

[0192] Figure 27 A flowchart outlining a process (2700) according to an embodiment of the present disclosure is shown. The process (2700) can be used for the reconstruction of a block to generate a predicted block for the block being reconstructed. In various embodiments, the process (2700) is performed by a processing circuitry such as, for example, the processing circuitry in the terminal devices (110), (120), (130), and (140), the processing circuitry performing the functions of the video encoder (203), the processing circuitry performing the functions of the video decoder (210), the processing circuitry performing the functions of the video decoder (310), the processing circuitry performing the functions of the video encoder (403), etc. In some embodiments, the process (2700) is implemented as software instructions, and thus when the processing circuitry executes the software instructions, the processing circuitry performs the process (2700). The process starts at (S2701) and proceeds to (S2710).

[0193] At (S2710), the transform coefficients associated with the encoded block of the non-rectangular partition of the picture are decoded according to the encoded video bitstream.

[0194] In some embodiments, on the encoder side, the transform coefficients in the transform unit of the non-rectangular partition can be scanned by following the scan order of the rectangular units including the non-rectangular partition and skipping the positions outside the non-rectangular partition. Accordingly, on the decoder side, the transform unit for the non-rectangular partition can be formed by following the scan order of the rectangular units including the non-rectangular partition and skipping the positions outside the non-rectangular partition.

[0195] At (S2720), the residual of the coded block is determined based on the transform coefficients.

[0196] According to some aspects of the present disclosure, on the encoder side, the non-rectangular partition is further divided into two or more rectangular partitions to generate the transform coefficients, and then on the decoder side, the residuals for the respective rectangular partitions can be determined based on the transform coefficients. In some embodiments, on the decoder side, the first residual of the first rectangular sub-block is determined based on the first transform coefficient among the transform coefficients. The first rectangular sub-block is the first partition of the non-rectangular partition. In addition, the second residual of the second rectangular sub-block is determined based on the second transform coefficient among the transform coefficients. The second rectangular sub-block is the second partition of the non-rectangular partition. The second rectangular sub-block may have the same size as the first rectangular sub-block, or may have a different size from the first rectangular sub-block. In some embodiments, the residuals of more than two rectangular sub-blocks can be determined according to the transform coefficients. In some examples, the rectangular sub-blocks may have a square shape.

[0197] According to some other aspects of the present disclosure, on the encoder side, the transform is performed on the entire non-rectangular portion to generate the transform coefficients, and then on the decoder side, the inverse transform is appropriately applied to the transform coefficients to determine the residual.

[0198] In some embodiments, a separable two-dimensional transform is applied on the encoder side. In some examples, on the decoder side, the intermediate transform coefficients of the intermediate transform unit are determined by performing an inverse transform on the transform unit in a first direction. The transform unit is formed by the transform coefficients. Then, the residual of the coded block is determined by performing an inverse transform on the intermediate transform unit in a second direction. In some examples, the intermediate transform unit and the transform unit may have the same shape and size as the non-rectangular partition. In some other examples, alignment may be performed on the encoder side to achieve better coding efficiency, and thus, on the decoder side, realignment may be performed before the inverse transform.

[0199] In some examples, the first direction is the vertical direction and the second direction is the horizontal direction. For example, a first inverse transform operation is performed on the first plurality of columns of the transform unit respectively. The first inverse transform operations are inverse transforms of the first number of points, such as the 8-point inverse transform in the example. A second inverse transform operation is performed on the second plurality of columns of the transform unit respectively. The second inverse transforms are inverse transforms of the second number of points, such as the 4-point inverse transform in the example. In addition, a third inverse transform operation is performed on the first plurality of rows of the intermediate transform unit respectively. The third inverse transform operations are inverse transforms of the third number of points, such as the 8-point inverse transform in the example. Then a fourth inverse transform operation is performed on the second plurality of rows of the intermediate transform unit respectively. The fourth inverse transform operations are inverse transforms of the fourth number of points, such as the 4-point inverse transform in the example.

[0200] In some examples, the first direction is the horizontal direction and the second direction is the vertical direction. For example, a first inverse transform operation is performed on the first plurality of rows of the transform unit respectively. The first inverse transform operations are inverse transforms of the first number of points, such as the 4-point inverse transform in the example. A second inverse transform operation is performed on the second plurality of rows of the transform unit respectively. The second inverse transforms are inverse transforms of the second number of points, such as the 8-point inverse transform in the example. In addition, a third inverse transform operation is performed on the first plurality of columns of the intermediate transform unit respectively. The third inverse transform operations are inverse transforms of the third number of points, such as the 4-point inverse transform in the example. Then a fourth inverse transform operation is performed on the second plurality of columns of the intermediate transform unit respectively. The fourth inverse transform operations are inverse transforms of the fourth number of points, such as the 8-point inverse transform in the example.

[0201] According to another aspect of the present disclosure, on the encoder side, a Karhunen-Loeve transform (KLT) is performed on the entire non-rectangular portion to generate transform coefficients, and then on the decoder side, the inverse Karhunen-Loeve transform (KLT) can be appropriately applied to the transform coefficients to determine the residual.

[0202] According to another aspect of the present disclosure, on the encoder side, padding can be performed to add additional samples to the coding block to form a rectangular unit including a non-rectangular partition, and then a transform can be performed on the rectangular unit. Thus, in some embodiments, on the decoder side, a two-dimensional inverse transform can be performed on the transform coefficients to determine a first residual of the rectangular unit including the non-rectangular partition, and then the residual of the coding block can be selected from the first residual of the rectangular unit.

[0203] At (S2730), the samples of the coding block are reconstructed based on the residual. For example, appropriate prediction of the coding block can be performed, and then the residual can be appropriately added to the prediction. Then, the process proceeds to S2799 and terminates.

[0204] Embodiments of the present disclosure also provide a device for video decoding, including: a decoding unit configured to decode transform coefficients associated with an encoded block of a non-rectangular partition of a picture according to an encoded video bitstream; a determining unit configured to determine a residual of the encoded block based on the transform coefficients; and a reconstruction unit configured to reconstruct samples of the encoded block based on the residual of the encoded block.

[0205] In some examples, the determining unit is further configured to: determine a first residual of a first rectangular sub-block based on a first transform coefficient among the transform coefficients, where the first rectangular sub-block is a first partition of the non-rectangular partition.

[0206] In some examples, the determining unit is further configured to: determine a second residual of a second rectangular sub-block based on a second transform coefficient among the transform coefficients, where the second rectangular sub-block is a second partition of the non-rectangular partition and has the same size as the first rectangular sub-block.

[0207] In some examples, the determining unit is further configured to: determine a second residual of a second rectangular sub-block based on a second transform coefficient among the transform coefficients, where the second rectangular sub-block is a second partition of the non-rectangular partition and has a different size from the first rectangular sub-block.

[0208] In some examples, the determining unit is further configured to: determine intermediate transform coefficients of an intermediate transform unit by performing an inverse transform on a transform unit in a first direction, where the transform unit is formed by the transform coefficients; and determine the residual of the encoded block by performing an inverse transform on the intermediate transform unit in a second direction.

[0209] In some examples, the device further includes an inverse transform unit configured to: respectively perform a first inverse transform operation on a first plurality of columns of the transform unit, where the first inverse transform operations are inverse transforms of a first number of points; respectively perform a second inverse transform operation on a second plurality of columns of the transform unit, where the second inverse transform operations are inverse transforms of a second number of points; respectively perform a third inverse transform operation on a first plurality of rows of the intermediate transform unit, where the third inverse transform operations are inverse transforms of a third number of points; and respectively perform a fourth inverse transform operation on a second plurality of rows of the intermediate transform unit, where the fourth inverse transform operations are inverse transforms of a fourth number of points.

[0210] In some examples, the inverse transformation unit is further configured to: perform a first inverse transformation operation on a first plurality of rows of the transformation unit respectively, where the first inverse transformation operations are inverse transformations of a first number of points; perform a second inverse transformation operation on a second plurality of rows of the transformation unit respectively, where the second inverse transformation operations are inverse transformations of a second number of points; perform a third inverse transformation operation on a first plurality of columns of the intermediate transformation unit respectively, where the third inverse transformation operations are inverse transformations of a third number of points; and perform a fourth inverse transformation operation on a second plurality of columns of the intermediate transformation unit respectively, where the fourth inverse transformation operations are inverse transformations of a fourth number of points.

[0211] In some examples, the determination unit is further configured to: determine the residual of the coding block by performing an inverse Karhunen - Loève transform (KLT) on the transform coefficients.

[0212] In some examples, the determination unit is further configured to: determine a first residual of a rectangular unit including the non - rectangular partition by performing a two - dimensional inverse transformation on the transform coefficients; and select the residual of the coding block from the first residual of the rectangular unit.

[0213] In some examples, the apparatus further includes a formation unit, which is configured to: form a transformation unit for the non - rectangular partition by following the scanning order of a rectangular unit including the non - rectangular partition and skipping scanning positions outside the non - rectangular partition.

[0214] The techniques described above can be implemented as computer software using computer - readable instructions and physically stored in one or more computer - readable media. For example, Figure 28 FIG. shows a computer system (2800) suitable for implementing certain embodiments of the disclosed subject matter.

[0215] The computer software can be encoded using any suitable machine code or computer language, which may be subject to mechanisms such as assembly, compilation, linking, etc. to create code including instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc. or executed through interpretation, microcode execution, etc.

[0216] The instructions can be executed on various types of computers or their components, including for example personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0217] Figure 28The components shown for the computer system (2800) are exemplary in nature and are not intended to impose any limitations on the scope of use or functionality of the computer software implementing the present disclosure. Nor should the configuration of the components be construed as having any dependencies or requirements related to any one component or combination of components shown in the exemplary embodiments of the computer system (2800).

[0218] The computer system (2800) may include certain human - machine interface input devices. Such human - machine interface input devices can respond to inputs made by one or more human users through, for example, tactile inputs (such as keystrokes, swipes, data - glove movements), audio inputs (such as speech, clapping), visual inputs (such as gestures), and olfactory inputs (not depicted). The human - machine interface devices can also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (such as speech, music, ambient sounds), images (such as scanned images, photographic images obtained from a still - image camera), and video (such as two - dimensional video, three - dimensional video including stereoscopic video).

[0219] The input human - machine interface devices may include one or more of the following (only one of each depicted): keyboard (2801), mouse (2802), touchpad (2803), touchscreen (2810), data glove (not shown), joystick (2805), microphone (2806), scanner (2807), camera device (2808).

[0220] The computer system (2800) may also include certain human - machine interface output devices. Such human - machine interface output devices can stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human - machine interface output devices may include tactile output devices (such as tactile feedback through the touchscreen (2810), data glove (not shown), or joystick (2805), but there may also be tactile feedback devices that do not serve as input devices), audio output devices (such as speakers (2809), headphones (not depicted)), visual output devices (such as screens (2810), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input capabilities, each with or without tactile feedback capabilities - some of which may be capable of outputting two - dimensional visual output or more than three - dimensional output in ways such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays, and smokeboxes (not depicted)), and printers (not depicted).

[0221] The computer system (2800) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (2820) with media such as CD / DVD (2821), thumb drives (2822), removable hard disk drives or solid state drives (2823), legacy magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD-based devices such as security dongles (not depicted), etc.

[0222] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0223] The computer system (2800) may also include an interface to one or more communication networks. The network may be, for example, wireless, wired-connected, optical. The network may also be local, wide-area, metro-area, vehicular and industrial, real-time, delay-tolerant, etc. Examples of networks include: local area networks such as Ethernet; wireless LAN; cellular networks including GSM, 3G, 4G, 5G, LTE, etc.; television wired connections or wireless wide-area digital networks including cable television, satellite television, and terrestrial broadcast television; vehicular and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter attached to certain common data ports or peripheral buses (2849) (such as, for example, the USB port of the computer system (2800)); others are typically integrated into the core of the computer system (2800) by attaching to the system bus as described below (for example, an Ethernet interface in a PC computer system or a cellular network interface in a smart phone computer system). Using any of these networks, the computer system (2800) can communicate with other entities. Such communication can be one-way, receive-only (e.g., broadcast television), one-way transmit-only (e.g., CANBus to certain CANBus devices), or two-way (e.g., using a local digital network or a wide-area digital network to other computer systems). Certain protocols and protocol stacks can be used on each of these networks and network interfaces as described above.

[0224] The above-described human-machine interface devices, human-accessible storage devices, and network interfaces may be attached to the core (2840) of the computer system (2800).

[0225] The core (2840) may include one or more central processing units (CPUs) (2841), a graphics processing unit (GPU) (2842), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (2843), a hardware accelerator (2844) for certain tasks, etc. These devices, together with a read-only memory (ROM) (2845), a random access memory (2846), and an internal mass storage device such as an internal non-user-accessible hard disk drive, SSD, etc. (2847), may be connected via a system bus (2848). In some computer systems, the system bus (2848) may be accessible in the form of one or more physical plugs to enable expansion by attaching additional CPUs, GPUs, etc. Peripheral devices may be attached directly or via a peripheral bus (2849) to the system bus (2848) of the core. The architecture of the peripheral bus includes PCI, USB, etc.

[0226] The CPU (2841), GPU (2842), FPGA (2843), and accelerator (2844) may execute certain instructions that, when combined, may form the aforementioned computer code. The computer code may be stored in the ROM (2845) or the RAM (2846). Interim data may also be stored in the RAM (2846), while permanent data may be stored in, for example, the internal mass storage device (2847). Fast storage and retrieval of any of the memory devices in the memory devices may be achieved by using a cache memory that may be closely associated with one or more CPUs (2841), GPUs (2842), mass storage devices (2847), ROM (2845), RAM (2846), etc.

[0227] Computer-readable media may have computer code for performing various computer-implemented operations. The media and the computer code may be media and computer code specially designed and constructed for the purposes of this disclosure, or the media and the computer code may be of the type well-known and available to those of ordinary skill in the computer software art.

[0228] By way of example and not limitation, having Figure 28A computer system (2800) of the architecture shown - particularly a core (2840) - can provide functionality due to software executed by processors (including CPUs, GPUs, FPGAs, accelerators, etc.) embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with a user-accessible mass storage device as introduced above, as well as certain storage devices of the core (2840) having a non-transitory nature, such as an on-core mass storage device (2847) or a ROM (2845). The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (2840). Depending on specific needs, the computer-readable media can include one or more memory devices or chips. The software can cause the core (2840) - particularly the processors (including CPUs, GPUs, FPGAs, etc.) therein - to perform specific processing or specific portions of specific processing described herein, including defining data structures stored in a RAM (2846) and modifying such data structures according to the processing defined by the software. Additionally or alternatively, the computer system can provide functionality due to being logically hardwired or otherwise embodied in a circuit (e.g., an accelerator (2844)), which can operate in place of or in conjunction with the software to perform specific processing or specific portions of specific processing described herein. In appropriate cases, the software mentioned can include logic, and conversely, the logic mentioned can also include software. In appropriate cases, the computer-readable media mentioned can include a circuit (e.g., an integrated circuit (IC)) storing software for execution, a circuit embodying logic for execution, or both of the above. The present disclosure encompasses any suitable combination of hardware and software.

[0229] Appendix A: Abbreviations

[0230] JEM: Joint Exploration Model

[0231] VVC: Versatile Video Coding

[0232] BMS: Benchmark Set

[0233] MV: Motion Vector

[0234] HEVC: High Efficiency Video Coding

[0235] SEI: Supplemental Enhancement Information

[0236] VUI: Video Usability Information

[0237] GOP: Group of Pictures

[0238] TU: Transform Unit

[0239] PU: Prediction Unit

[0240] CTU: Coding Tree Unit

[0241] CTB: Coding Tree Block

[0242] PB: Prediction Block

[0243] HRD: Hypothetical Reference Decoder

[0244] SNR: Signal-to-Noise Ratio

[0245] CPU: Central Processing Unit

[0246] GPU: Graphics Processing Unit

[0247] CRT: Cathode Ray Tube

[0248] LCD: Liquid Crystal Display

[0249] OLED: Organic Light-Emitting Diode

[0250] CD: Compact Disc

[0251] DVD: Digital Versatile Disc

[0252] ROM: Read-Only Memory

[0253] RAM: Random Access Memory

[0254] ASIC: Application-Specific Integrated Circuit

[0255] PLD: Programmable Logic Device

[0256] LAN: Local Area Network

[0257] GSM: Global System for Mobile Communications

[0258] LTE: Long-Term Evolution

[0259] CANBus: Controller Area Network Bus

[0260] USB: Universal Serial Bus

[0261] PCI: Peripheral Component Interconnect

[0262] FPGA: Field-Programmable Gate Array

[0263] SSD: Solid State Drive

[0264] IC: Integrated Circuit

[0265] CU: Coding Unit

[0266] TSM: Transform Skip Mode

[0267] IBC: Intra Block Copy

[0268] DPCM: Differential Pulse Code Modulation

[0269] BDPCM: Block-based DPCM

[0270] Although the present disclosure has described several exemplary embodiments, there are variations, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it should be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.

Claims

1. A method for video decoding, comprising: decoding transform coefficients associated with coded blocks of an L-shaped partition of a picture according to an encoded video bitstream; determining intermediate transform coefficients of an intermediate transform unit by performing an inverse transform on a transform unit in a first direction, the transform unit being formed by the transform coefficients; and determining a residual of the coded block by performing an inverse transform on the intermediate transform unit in a second direction; reconstructing samples of the coded block based on the residual of the coded block; when the first direction is a vertical direction and the second direction is a horizontal direction, performing a first inverse transform operation on each column of a first plurality of columns of the transform unit to determine a first number of transform coefficients for each column of the first plurality of columns; the first inverse transform operation is respectively an inverse transform of a first number, the first number being the number of residual samples included in each column of the first plurality of columns of the transform unit; performing a second inverse transform operation on each column of a second plurality of columns of the transform unit to determine a second number of transform coefficients for each column of the second plurality of columns, the second inverse transform operation being respectively an inverse transform of a second number, the second number being the number of residual samples included in each column of the second plurality of columns of the transform unit, the sum of the first plurality of columns and the second plurality of columns being the total number of columns of the transform unit, the first plurality of columns being a continuous part of the transform unit and each column including a first number of residual samples, the second plurality of columns being a continuous part of the transform unit and each column including a second number of residual samples; performing a third inverse transform operation on each row of a first plurality of rows of the intermediate transform unit to determine a third number of final transform coefficients for each row of the first plurality of rows; the third inverse transform operation is respectively an inverse transform of a third number, the third number being the number of transform coefficient samples included in each row of the first plurality of rows of the transform unit; and performing a fourth inverse transform operation on each row of a second plurality of rows of the intermediate transform unit to determine a fourth number of final transform coefficients for each row of the second plurality of rows, the fourth inverse transform operation being respectively an inverse transform of a fourth number, the fourth number being the number of transform coefficient samples included in each row of the second plurality of rows of the transform unit, the sum of the first plurality of rows and the second plurality of rows being the total number of rows of the intermediate transform unit, the first plurality of rows being a continuous part of the intermediate transform unit and each row including a third number of transform coefficient samples, the second plurality of rows being a continuous part of the intermediate transform unit and each row including a fourth number of transform coefficient samples.

2. The method according to claim 1, further comprising: When the first direction is the horizontal direction and the second direction is the vertical direction, perform a first inverse transform operation on each row of the first plurality of rows of the transform unit respectively to determine the first number of transform coefficients of each row of the first plurality of rows; the first inverse transform operations are respectively inverse transforms of the first number, and the first number is the number of samples included in each row of the first plurality of rows in the transform unit; Perform a second inverse transform operation on each row of the second plurality of rows of the transform unit respectively to determine the second number of transform coefficients in each row of the second plurality of rows, the second inverse transform operations are respectively inverse transforms of the second number, the second number is the number of samples included in each row of the second plurality of rows in the transform unit, and the sum of the first plurality of rows and the second plurality of rows is the total number of rows of the transform unit; Perform a third inverse transform operation on each column of the first plurality of columns of the intermediate transform unit respectively to determine the third number of final transform coefficients in each column of the first plurality of columns; the third inverse transform operations are respectively inverse transforms of the third number, the third number is the number of samples included in each column of the first plurality of columns in the transform unit; and Perform a fourth inverse transform operation on each column of the second plurality of columns of the intermediate transform unit respectively to determine the fourth number of final transform coefficients in each column of the second plurality of columns, the fourth inverse transform operations are respectively inverse transforms of the fourth number, the fourth number is the number of samples included in each column of the second plurality of columns in the transform unit, and the sum of the first plurality of columns and the second plurality of columns is the total number of columns of the intermediate transform unit.

3. The method according to claim 1, further comprising: Determine the residual of the coding block by performing an inverse Karhunen - Loève transform (KLT) on the transform coefficients.

4. An apparatus for video decoding, comprising: A decoding unit configured to decode transform coefficients associated with a coding block of an L - shaped partition of a picture according to an encoded video bitstream; A determination unit configured to: Determine intermediate transform coefficients of an intermediate transform unit by performing an inverse transform on the transform unit in a first direction, the transform unit being formed by the transform coefficients; and Determine the residual of the coding block by performing an inverse transform on the intermediate transform unit in a second direction; A reconstruction unit configured to reconstruct samples of the coding block based on the residual of the coding block; An inverse transform unit configured to: When the first direction is the vertical direction and the second direction is the horizontal direction, perform a first inverse transform operation on each row of the first plurality of rows of the transform unit respectively to determine the first number of transform coefficients of each row of the first plurality of rows; the first inverse transform operations are respectively inverse transforms of the first number, and the first number is the number of samples included in each row of the first plurality of rows in the transform unit; Perform a second inverse transform operation on each row of the second plurality of rows of the transform unit to determine a second number of transform coefficients in each row of the second plurality of rows, where the second inverse transform operations are respectively inverse transforms of a second number of points, and the second number of points is the number of samples included in each row of the second plurality of rows of the transform unit, and the sum of the first plurality of rows and the second plurality of rows is the total number of rows of the transform unit; Perform a third inverse transform operation on each column of the first plurality of columns of the intermediate transform unit to determine a third number of final transform coefficients in each column of the first plurality of columns; the third inverse transform operations are respectively inverse transforms of a third number of points, and the third number of points is the number of samples included in each column of the first plurality of columns of the transform unit; and Perform a fourth inverse transform operation on each column of the second plurality of columns of the intermediate transform unit to determine a fourth number of final transform coefficients in each column of the second plurality of columns, where the fourth inverse transform operations are respectively inverse transforms of a fourth number of points, and the fourth number of points is the number of samples included in each column of the second plurality of columns of the transform unit, and the sum of the first plurality of columns and the second plurality of columns is the total number of columns of the intermediate transform unit.

5. The apparatus according to claim 4, wherein, the inverse transform unit is further configured to: when the first direction is the horizontal direction and the second direction is the vertical direction, perform a first inverse transform operation on each row of the first plurality of rows of the transform unit to determine a first number of transform coefficients in each row of the first plurality of rows; the first inverse transform operations are respectively inverse transforms of a first number of points, and the first number of points is the number of samples included in each row of the first plurality of rows of the transform unit; perform a second inverse transform operation on each row of the second plurality of rows of the transform unit to determine a second number of transform coefficients in each row of the second plurality of rows, where the second inverse transform operations are respectively inverse transforms of a second number of points, and the second number of points is the number of samples included in each row of the second plurality of rows of the transform unit, and the sum of the first plurality of rows and the second plurality of rows is the total number of rows of the transform unit; perform a third inverse transform operation on each column of the first plurality of columns of the intermediate transform unit to determine a third number of final transform coefficients in each column of the first plurality of columns; the third inverse transform operations are respectively inverse transforms of a third number of points, and the third number of points is the number of samples included in each column of the first plurality of columns of the transform unit; and perform a fourth inverse transform operation on each column of the second plurality of columns of the intermediate transform unit to determine a fourth number of final transform coefficients in each column of the second plurality of columns, where the fourth inverse transform operations are respectively inverse transforms of a fourth number of points, and the fourth number of points is the number of samples included in each column of the second plurality of columns of the transform unit, and the sum of the first plurality of columns and the second plurality of columns is the total number of columns of the intermediate transform unit.

6. The apparatus according to claim 4, wherein, the determination unit is further configured to: The residual of the coded block is determined by performing an inverse Karhunen-Loève transform (KLT) on the transform coefficients.

7. A computer device, characterized in that the device includes a processor and a memory, the memory is used to store program code and transmit the program code to the processor; the processor is used to execute the video decoding method according to any one of claims 1-3 according to the instructions in the program code.

8. A method for storing or transmitting a video bitstream, characterized in that the video bitstream can be decoded based on the decoding method according to any one of claims 1 to 3.

9. A computer storage medium, characterized in that a bitstream formed by a computer program is stored thereon, and when the computer program is executed by a computer, the computer executes the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Signaling of simplified depth coding (SDC) for depth intra-and inter-prediction modes in 3D video coding

    CN105934948A

  • Image encoder, image decoder, image encoding method, and image decoding method

    CN111034197A

  • Method and Apparatus of Flexible Block Partition for Video Coding

    US20170244964A1