Scanning order of secondary conversion coefficients

The method improves video encoding and decoding efficiency by extracting data blocks, performing irreducible transforms, and optimizing entropy encoding, addressing the challenges of redundancy reduction and compression ratio enhancement in existing video coding technologies.

JP7683894B2Active Publication Date: 2025-05-27TENCENT AMERICA LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023532416
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-01-28
Filing Date
2022-02-04
Publication Date
2025-05-27
Estimated Expiration
2042-02-04

AI Technical Summary

Technical Problem

Current video coding technologies face challenges in efficiently encoding and decoding video data, particularly in reducing redundancy and achieving high compression ratios while maintaining acceptable distortion levels.

Method used

The proposed method involves extracting data blocks from video data, scanning them according to specific orders, performing irreducible transforms, and replacing portions of the data blocks with transformed sequences, thereby optimizing entropy encoding and decoding processes.

Benefits of technology

This approach enhances video encoding and decoding efficiency by reducing redundancy and improving compression ratios, while maintaining acceptable distortion levels, thus addressing the limitations of existing video coding technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007683894000012
    Figure 0007683894000012
  • Figure 0007683894000013
    Figure 0007683894000013
  • Figure 0007683894000014
    Figure 0007683894000014
Patent Text Reader

Abstract

A method, apparatus, and computer-readable storage medium for processing video data includes extracting a data block from the video data, scanning a first number of data items in the data block according to a first scanning order to generate a first data sequence, performing a non-separable transform on the first data sequence to obtain a second data sequence having a second number of data items, and replacing at least some of the first number of data items in the data block with some or all of the second data sequence according to a second scanning order.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 238,646, filed Aug. 30, 2021, and U.S. Non - Provisional Patent Application No. 17 / 587,164, filed Jan. 28, 2022, the entireties of both applications are incorporated herein by reference.

[0002] This disclosure describes a collection of the latest video coding techniques. More specifically, the disclosed techniques include the implementation of inseparable transforms of data blocks in video encoding and decoding.

Background Art

[0003] The description of the background art provided herein is for the purpose of generally presenting the context of the present disclosure. The inventors' research is not admitted as prior art to the present disclosure, either expressly or implicitly, to the extent that the research is described in this background art section and to the extent that aspects of the description that may not be recognized as prior art at the time of filing of the present application other than those described in this background art section.

[0004] Video coding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video can include a series of pictures, each picture having a spatial dimension of, for example, 1920×1080 luminance samples and associated full-sampled or subsampled chrominance samples. The series of pictures can have a fixed or variable picture rate (or frame rate, also called) of, for example, 60 pictures per second or 60 frames per second. Uncompressed video has specific bitrate requirements for streaming or data processing. For example, video having a pixel resolution of 1920×1080, a frame rate of 60 frames / second, and 4:2:0 chroma subsampling with 8 bits per pixel per color channel requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires storage space exceeding 600 GByte.

[0005] One purpose of video coding and decoding can be the reduction of redundancy of an uncompressed input video signal by compression. Compression can help reduce the aforementioned bandwidth and / or storage space requirements, in some cases by more than two orders of magnitude. Both reversible compression and irreversible compression, and combinations thereof, can be used. Reversible compression refers to a technique in which an exact copy of the original signal can be reconstructed from the original signal compressed by a decoding process. Irreversible compression refers to a coding / decoding process in which the original video information is not fully retained during coding and cannot be fully restored during decoding. When using irreversible compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough to be useful for the intended application, even with some information loss. In the case of video, irreversible compression is widely adopted in many applications. The amount of acceptable distortion depends on the application. For example, users of certain consumer video streaming applications may tolerate higher distortion than users of movie or television broadcast applications. The compression ratio achievable by a particular coding algorithm can be selected or adjusted to reflect various distortion tolerances. That is, generally, the higher the distortion tolerance, the more possible it is to have a coding algorithm that results in high loss and a high compression ratio.

[0006] Video encoders and video decoders can utilize techniques from several broad categories and steps, including, for example, motion compensation, Fourier transform, quantization, and entropy coding.

[0007] Video codec technology may include a technology known as intra coding. In intra coding, sample values are represented without referring to samples or other data from previously reconstructed reference pictures. In some video codecs, a picture is spatially subdivided into blocks of samples. If all blocks of samples are coded in an intra mode, that picture can be called an intra picture. Those derived pictures such as intra pictures and independent decoder refresh pictures can be used to reset the decoder state and thus can be used as the first picture in a coded video bitstream and video session or as a still image. Next, the samples of the block after intra prediction can be transformed into the frequency domain, and the transformation coefficients thus generated can be quantized before entropy coding. Intra prediction represents a technique for minimizing sample values in the pre-transformation region. In some cases, the smaller the DC value after transformation and the smaller the AC coefficients, the fewer bits are required with a given quantization step size to represent the block after entropy coding.

[0008] For example, conventional intra coding, such as known from MPEG-2 production coding technology, does not use intra prediction. However, some newer video compression technologies include techniques that attempt to code / decode a block based on surrounding sample data and / or metadata that are obtained during spatial adjacent encoding and / or decoding and that precede in decoding order a block of data that is intra-coded or intra-decoded. Such techniques will hereafter be referred to as "intra prediction" techniques. Note that in at least some cases, intra prediction uses reference data only from the current picture being reconstructed and does not use reference data from other reference pictures.

[0009] Intra prediction can have many different forms. When two or more of such techniques are available in a given video coding technique, the technique used can be referred to as an intra prediction mode. One or more intra prediction modes can be provided in a particular codec. In certain cases, a mode can have sub - modes and / or can be associated with various parameters. The mode / sub - mode information and the intra - coding parameters of the video block can be coded individually or can be included together in the codeword of the mode. Which codeword to use for a given combination of mode, sub - mode, and / or parameters can affect the coding efficiency improvement through intra prediction, and thus can also affect the entropy coding technique used to convert the codeword into the bitstream.

[0010] A particular mode of intra prediction was introduced in H.264, improved in H.265, and further improved in more recent coding techniques such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). Generally, in intra prediction, a predictor block can be formed using the available adjacent sample values. For example, the available values of a particular set of adjacent samples along a particular direction and / or line can be copied into the predictor block. The reference to the direction used can be coded within the bitstream or can itself be predicted.

[0011] Referring to FIG. 1A, shown in the lower right is a subset of nine predictor directions specified in 33 possible intra predictor directions of H.265 (corresponding to 33 of the 35 intra modes specified in H.265). The point (101) where the arrows converge represents the sample being predicted. The arrows represent the directions in which adjacent samples are used to predict sample 101 therefrom. For example, arrow (102) indicates that sample (101) is predicted from one or more adjacent samples at an angle of 45 degrees from the horizontal direction towards the upper right. Similarly, arrow (103) indicates that sample (101) is predicted from one or more adjacent samples at an angle of 22.5 degrees from the horizontal direction towards the lower left of sample (101).

[0012] Referring further to FIG. 1A, depicted in the upper left is a square block (104) of 4×4 samples (indicated by the thick dashed line). Square block (104) contains 16 samples, each labeled with its "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample from the top in the Y dimension and the first sample from the left in the X dimension. Similarly, sample S44 is the fourth sample in both the Y and X dimensions within block (104). Since the block size is 4×4 samples, S44 is in the lower right. Further examples of reference samples following a similar numbering scheme are shown. The reference samples are labeled with an "R" for their Y position (e.g., row index) and X position (column index) relative to block (104). In both H.264 and H.265, predicted samples adjacent to the block being reconstructed are used.

[0013] The intra-picture prediction of block 104 may start by copying the reference sample value from adjacent samples according to the signaled prediction direction. For example, the coded video bitstream may include signaling indicating the prediction direction of arrow (102) for this block 104, i.e., it is assumed that the samples are predicted from one or more prediction samples at an angle of 45 degrees from the horizontal direction towards the upper right. In such a case, samples S41, S32, S23, S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.

[0014] In certain cases, especially when the direction is not evenly divisible by 45 degrees, the values of multiple reference samples may be combined, for example, by interpolation, to calculate the reference sample.

[0015] The number of possible directions has been increasing as video coding technology continues to evolve. In H.264 (2003), for example, 9 different directions are available for intra prediction. This has increased to 33 in H.265 (2013), and JEM / VVC / BMS can support up to 65 directions at the time of this disclosure. Experimental studies have been conducted to help identify the most appropriate intra prediction direction, and using certain techniques of entropy coding, the specific bit penalty for the direction can be accepted to encode those most appropriate directions with a small number of bits. Further, the direction itself may sometimes be predicted from the adjacent directions used in the intra prediction of the decoded adjacent blocks.

[0016] FIG. 1B shows a schematic diagram (180) showing 65 intra prediction directions by JEM to illustrate the increasing number of prediction directions in various coding technologies that have evolved over time.

[0017] Methods for mapping bits representing an intra prediction direction in a coded video bitstream to a prediction direction may vary depending on the video coding technology and may range from a simple direct mapping of prediction direction to intra prediction mode to complex adaptive schemes including codewords, most probable modes, and similar techniques. However, in all cases, there may be certain directions of intra prediction that are statistically less likely to occur in the video content than other specific directions. Since the purpose of video compression is to reduce redundancy, in well-designed video coding techniques, those less likely directions may be represented with more bits than the more likely directions.

[0018] Inter-picture prediction, or inter prediction, may be based on motion compensation. In motion compensation, sample data from a previously reconstructed picture or a portion thereof (reference picture) is spatially shifted in a direction indicated by a motion vector (hereinafter MV) and then used for prediction of a newly reconstructed picture or picture portion (e.g., block). In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions X and Y, or three dimensions, where the third dimension is an indication of the reference picture used (similar to the temporal dimension).

[0019] In some video compression techniques, the current MV applicable to a particular area of sample data can be predicted from other MVs, for example, from other MVs related to other areas of sample data that are spatially adjacent to the area being reconstructed and that precede the current MV in the decoding order. By doing so, the overall data amount required to code the MVs can be significantly reduced by relying on the removal of redundancy of the correlated MVs, thereby increasing the compression efficiency. MV prediction can function effectively, for example, when coding an input video signal derived from a camera (known as natural video), because areas larger than the area to which a single MV is applicable have a statistical likelihood of moving in the same direction in the video sequence and thus, in some cases, can be predicted using a similar motion vector derived from the MVs of adjacent areas. As a result, the actual MV of a given area becomes similar or identical to the predicted MV from the surrounding MVs. Such an MV can further be represented in fewer bits than the number of bits that would be used if the MV were coded directly rather than being predicted from one or more adjacent MVs after entropy coding. In some cases, MV prediction can be regarded as an example of lossless compression of a signal (i.e., an MV) derived from the original signal (i.e., the sample stream). In other cases, for example, due to rounding errors when calculating predictors from several surrounding MVs, MV prediction itself can be lossy.

[0020] H.265 / HEVC (ITU-T Rec. H.265, “High Efficiency Video Coding”, December 2016) describes various MV prediction mechanisms. Among the many MV prediction mechanisms specified by H.265, the one described below is a technique hereafter referred to as “spatial merge”.

[0021] Specifically, referring to FIG. 2, the current block (201) contains samples detected by the encoder as being predictable from a previous block of the same size that has been spatially shifted during motion search. Instead of directly coding the MV, the MV can be derived from metadata associated with one or more reference pictures using an MV associated with any one of five surrounding samples represented by A0, A1, and B0, B1, B2 (202 to 206 respectively), for example, from the last reference picture (in decoding order). In H.265, MV prediction can use predictors from the same reference picture that the neighboring blocks are using. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0022] Aspects of the present disclosure provide methods and apparatuses for video encoding and video decoding.

[0023] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that cause a computer to execute a method for video decoding and / or encoding when executed by the computer for video decoding and / or encoding.

[0024] According to one aspect, embodiments of the present disclosure provide a method for processing video data. The method includes extracting a data block from the video data, scanning a first number of data items within the data block according to a first scan order to generate a first data sequence, performing an irreducible transform on the first data sequence to obtain a second data sequence having a second number of data items, and replacing at least a portion of the first number of data items within the data block with a portion or all of the second data sequence according to a second scan order.

[0025] According to another aspect, one embodiment of the present disclosure provides a method for entropy encoding a conversion coefficient associated with video data. The method includes scanning the conversion coefficient when performing entropy encoding of the conversion coefficient using a first scanning order that is one of a horizontal scanning order or a vertical scanning order in response to the conversion associated with the conversion coefficient being non-separable, and scanning the conversion coefficient in a second scanning order different from the first scanning order when performing entropy encoding of the conversion coefficient in response to the conversion associated with the conversion coefficient being separable.

[0026] According to another aspect, an embodiment of the present disclosure provides a method for processing video data. The method includes receiving video data, determining whether a non-separable conversion is applied to the video data as a secondary conversion, scanning a first number of primary conversion coefficients in response to the non-separable conversion being applied to the video data as a secondary conversion, the primary conversion coefficients following a first scanning order, performing a non-separable conversion using the first number of primary conversion coefficients as an input to obtain a second number of secondary conversion coefficients as an output, the secondary conversion coefficients following a second scanning order, replacing at least the second number of primary conversion coefficients with the secondary conversion coefficients following the second scanning order, performing an inverse secondary conversion corresponding to the non-separable conversion using the second number of secondary conversion coefficients as an input to obtain the first number of primary conversion coefficients as an output, and replacing at least the first number of secondary conversion coefficients with the primary conversion coefficients following the first scanning order.

[0027] According to another aspect, one embodiment of the present disclosure provides an apparatus for video encoding and / or decoding. The apparatus includes a memory that stores instructions and a processor that communicates with the memory. When the processor executes the instructions, the processor is configured to cause the apparatus to perform the above-described method for video decoding and / or encoding.

[0028] According to yet another aspect, one embodiment of the present disclosure provides a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding and / or encoding, cause the computer to execute the above method for video decoding and / or encoding.

[0029] The above and other aspects and their implementations will be described in more detail in the drawings, the specification, and the claims.

[0030] Further features, properties, and various advantages of the subject matter of the present disclosure will become more apparent from the following detailed description and the accompanying drawings.

Brief Description of the Drawings

[0031]

Figure 1A

Figure 1B

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Mode for Carrying Out the Invention

[0032] Figure 3 shows a simplified block diagram of a communication system (300) according to an embodiment of the present disclosure. The communication system (300) includes, for example, a plurality of terminal devices that can communicate with each other via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via a network (350). In the example of Figure 3, the first pair of terminal devices (310) and (320) can perform unidirectional transmission of data. For example, the terminal device (310) can code video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to the other terminal device (320) via the network (350). The encoded video data can be transmitted in the form of one or more encoded video bitstreams. The terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to restore the video pictures, and display the video pictures according to the restored video data. Unidirectional data transmission can be performed, for example, in media serving applications.

[0033] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data, which can be performed, for example, during video conferencing applications. For bidirectional data transmission, in one example, each of the terminal devices (330) and (340) can code video data (e.g., a stream of video pictures captured by that terminal device) for transmission to the other of the terminal devices (330) and (340) via the network (350). Each of the terminal devices (330) and (340) can also receive the encoded video data transmitted by the other of the terminal devices (330) and (340), decode the encoded video data to restore the video pictures, and display the video pictures on an accessible display device according to the restored video data.

[0034] In the example of FIG. 3, the terminal devices (310), (320), (330), and (340) can be implemented as servers, personal computers, and smartphones, but the applicability of the principles underlying the present disclosure is not so limited. Embodiments of the present disclosure can be implemented in desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated video conferencing devices, and the like. The network (350) represents any number of networks and any type of network that transmit coded video data between the terminal devices (310), (320), (330), and (340), including, for example, wired (wired) and / or wireless communication networks. The communication network (350)9 may exchange data over a circuit-switched channel, a packet-switched channel, and / or other types of channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of the network (350) may not be important for the operation of the present disclosure, unless explicitly described herein.

[0035] FIG. 4 shows the arrangement of a video encoder and a video decoder in a video streaming environment as an example of the use of the subject matter of the present disclosure. The subject matter of the present disclosure can be equally applied to other video-related applications, including, for example, video conferencing, digital television broadcasting, gaming, virtual reality, storage of compressed video on digital media including CDs, DVDs, memory sticks, and the like.

[0036] A video streaming system may include a video source (401), such as a digital camera, for creating a stream (402) of uncompressed video pictures or images, which may include a video capture subsystem (413). In one example, the stream (402) of video pictures includes samples recorded by the digital camera of the video source 401. The stream (402) of video pictures is shown in bold lines to emphasize the high data volume when compared to the encoded video data (404) (or encoded video bitstream), and may be processed by an electronic device (420) including a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof to enable or implement aspects of the subject matter of this disclosure, as will be described in more detail below. The encoded video data (404) (or encoded video bitstream (404)) is shown in thin lines to emphasize the low data volume when compared to the stream (402) of uncompressed video pictures, and may be stored in the streaming server (405) for future use, or directly in a downstream video device (not shown). One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, may access the streaming server (405) to obtain copies (407) and (409) of the encoded video data (404). The client subsystem (406) may include, for example, a video decoder (410) within an electronic device (430). The video decoder (410) decodes an input copy (407) of the encoded video data and creates an output stream (411) of video pictures that are not compressed and can be rendered on a display (412) (e.g., a display screen) or other rendering device (not shown). The video decoder 410 may be configured to implement some or all of the various functions described in this disclosure.In some streaming systems, encoded video data (404), (407), and (409) (e.g., video bitstreams) can be encoded according to a particular video coding / compression standard. Examples of such standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC). The subject matter of the present disclosure can be used in the context of VVC and other video coding standards.

[0037] Note that electronic devices (420) and (430) can include other components (not shown). For example, electronic device (420) can include a video decoder (not shown), and electronic device (430) can also include a video encoder (not shown).

[0038] FIG. 5 shows a block diagram of a video decoder (510) according to any embodiment of the present disclosure below. The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., receiving circuitry). The video decoder (510) can be used in place of the video decoder (410) in the example of FIG. 4.

[0039] The receiver (531) may receive one or more coded video sequences to be decoded by the video decoder (510). In the same or another embodiment, one coded video sequence may be decoded at a time, and the decoding of each coded video sequence is independent of other coded video sequences. Each video sequence may be associated with a plurality of video frames or video images. The coded video sequence may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data or a streaming source transmitting the encoded video data. The receiver (531) may receive the encoded video data together with other data such as coded audio data and / or auxiliary data streams that may be transferred to respective processing circuits (not shown). The receiver (531) may separate the coded video sequence from other data. To counter network jitter, a buffer memory (515) may be disposed between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). For certain applications, the buffer memory (515) may be implemented as part of the video decoder (510). For other applications, the buffer memory (515) may be external and separated from the video decoder (510) (not shown). For still other applications, for example, a buffer memory (not shown) may be external to the video decoder (510) to counter network jitter, and another buffer memory (515) may be inside the video decoder (510) to process playback timing. When the receiver (531) is receiving data from a storage / transfer device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (515) may be unnecessary or can be made small. For use in a best-effort packet network such as the Internet, a buffer memory (515) of sufficient size may be required, and its size may be relatively large.Such a buffer memory may be implemented in an adaptive size and may be implemented at least partially in an operating system external to the video decoder (510) or a similar element (not shown).

[0040] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the coded video sequence. The categories of those symbols include information used to manage the operation of the video decoder (510) and, potentially, information for controlling a rendering device such as a display (512) (e.g., a display screen) that can be coupled to the electronic device (530), whether or not it is an essential part of the electronic device (530), as shown in FIG. 5. The control information for the (one or more) rendering devices may be in the form of supplementary enhancement information (SEI message) or a video user capability information (VUI) parameter set fragment (not shown). The parser (520) may parse / entropy-decode the coded video sequence received by the parser (520). The entropy coding of the coded video sequence can be in accordance with a video coding technology or standard and can follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context dependency, etc. The parser (520) may extract a set of subgroup parameters of at least one subgroup of pixels in the video decoder based on at least one parameter corresponding to a subgroup from the coded video sequence. Subgroups can include group of pictures (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (520) may also extract information such as transform coefficients (e.g., Fourier transform coefficients), quantization parameter values, motion vectors, etc. from the coded video sequence.

[0041] The parser (520) can perform entropy decoding / syntax analysis operations on the video sequence received from the buffer memory (515) to create symbols (521).

[0042] The reconstruction of the symbol (521) can include multiple different processing units or functional units depending on the type of the coded video picture or a part thereof (such as inter-picture and intra-picture, inter-block and intra-block, etc.), and other factors. The units included and how the units are included can be controlled by subgroup control information parsed from the video sequence coded by the parser (520). Such a flow of subgroup control information between the parser (520) and the following multiple processing units or functional units is not illustrated for simplicity.

[0043] In addition to the functional blocks already described, the video decoder (510) can be conceptually subdivided into several functional units as described below. In an actual implementation operating under commercial constraints, many of these functional units interact closely with each other and can be at least partially integrated with each other. However, for the purpose of clearly explaining the various functions of the subject matter of the present disclosure, a conceptual subdivision into functional units is adopted in the following disclosure.

[0044] The first unit is the scaler / inverse transform unit (551). The scaler / inverse transform unit (551) can receive control information including quantization transform coefficients, information indicating which type of inverse transform to use, block size, quantization coefficient / parameter, quantization scaling matrix, etc. from the parser (520) as one or more symbols (521). The scaler / inverse transform unit (551) can output a block comprising sample values that can be input to the aggregator (555).

[0045] In some cases, the output samples of the scaler / inverse transform (551) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information can be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) may generate a block of the same size and shape as the block being reconstructed, using information of surrounding blocks that have already been reconstructed and stored in the current picture buffer (558). The current picture buffer (558) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. The aggregator (555) may, in some implementations, add, sample by sample, the prediction information generated by the intra prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).

[0046] In other cases, the output samples of the scaler / inverse transform unit (551) may be related to inter-coded and potentially motion-compensated blocks. In such cases, the motion compensation prediction unit (553) can access the reference picture memory (557) to fetch the samples used for inter-picture prediction. After motion-compensating the samples fetched according to the symbols (521) related to the block, these samples can be added to the output of the scaler / inverse transform unit (551) by the aggregator (555) to generate output sample information (the output of unit 551 may be referred to as residual samples or a residual signal). The address in the reference picture memory (557) from which the motion compensation prediction unit (553) fetches the prediction samples can be controlled by a motion vector in the form of a symbol (521) that can have, for example, an X component, a Y component (shift), and a reference picture component (time) available to the motion compensation prediction unit (553). Motion compensation may also include interpolation of the sample values fetched from the reference picture memory (557) when an exact motion vector of sub-samples is used, and may be associated with a motion vector prediction mechanism and the like.

[0047] The output samples of the aggregator (555) can be subjected to various loop filtering techniques in the loop filter unit (556). The video compression technology is controlled by the parameters included in the coded video sequence (also referred to as the coded video bitstream), and can include in-loop filter techniques available to the loop filter unit (556) as symbols (521) from the parser (520), but can also respond to meta-information obtained during the decoding of the previous part of the coded picture or coded video sequence (in decoding order), and can also respond to previously reconstructed and loop-filtered sample values. As will be described in more detail below, several types of loop filters can be included as part of the loop filter unit 556 in various orders.

[0048] The output of the loop filter unit (556) can be output to the rendering device (512) and can also be a sample stream that can be stored in the reference picture memory (557) for use in future inter-picture prediction.

[0049] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future inter-picture prediction. For example, when the coded picture corresponding to the current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and the unused current picture buffer can be reallocated before starting the reconstruction of the next coded picture.

[0050] The video decoder (510) may perform a decoding operation according to a predetermined video compression technique adopted in a standard such as ITU-T Rec.H.265. The coded video sequence may conform to the syntax specified by the video compression technique or standard being used in the sense that the coded video sequence is faithful to both the syntax of the video compression technique or standard and the profile documented in the video compression technique or standard. Specifically, the profile can select specific tools from all the tools available in the video compression technique or standard as tools that are only used under that profile. To conform to the standard, the complexity of the coded video sequence may be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, the maximum frame rate, the maximum reconstruction sample rate (e.g., measured in megasamples per second), the maximum reference picture size, etc. The limits set by the level may in some cases be further restricted by the virtual reference decoder (HRD) specification and the metadata for HRD buffer management signaled in the coded video sequence.

[0051] In some exemplary embodiments, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the (one or more) coded video sequences. The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0052] FIG. 6 shows a block diagram of a video encoder (603) according to an exemplary embodiment of the present disclosure. The video encoder (603) may be included in an electronic device (620). The electronic device (620) may further include a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used in place of the video encoder (403) in the example of FIG. 4.

[0053] The video encoder (603) may receive video samples from a video source (601) (not part of the electronic device (620) in the example of FIG. 6) that may capture one or more video images to be coded by the video encoder (603). In another example, the video source (601) may be implemented as a part of the electronic device (620).

[0054] The video source (601) may provide the source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream having any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit,...), any color space (e.g., BT.601 YCrCb, RGB, XYZ,...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (601) may be a storage device capable of storing previously prepared videos. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures or images that give motion when viewed in sequence. The picture itself may be organized as a spatial array of pixels, and each pixel may include one or more samples depending on the sampling structure, color space, etc. being used. One skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0055] According to some exemplary embodiments, a video encoder (603) may code and compress pictures of a source video sequence into a coded video sequence (643) in real time or under any other arbitrary time constraints required by the application. Enforcing an appropriate coding speed constitutes one function of the controller (650). In some embodiments, the controller (650) may be functionally coupled to and control other functional units, as described below. For the sake of brevity, the couplings are not shown. Parameters set by the controller (650) may include rate control related parameters (such as picture skip, quantizer, lambda value of the rate distortion optimization method), picture size, group of pictures (GOP) layout, maximum motion vector search range, and the like. The controller (650) may be configured to have other appropriate functions related to the video encoder (603) optimized for a particular system design.

[0056] In some exemplary embodiments, the video encoder (603) may be configured to operate in a coding loop. As an overly simplified explanation, in one example, the coding loop may include a source coder (630) (e.g., which is responsible for creating symbols such as a symbol stream based on an input picture to be coded and one or more reference pictures), and a (local) decoder (633) incorporated in the video encoder (603). The decoder (633) can reconstruct symbols and create sample data in the same way as a (remote) decoder would, even if the incorporated decoder 633 processes a video stream coded by the source coder 630 without entropy coding (in the video compression techniques contemplated by the subject matter of the present disclosure, any compression between symbols and the coded video bitstream can be reversible). The reconstructed sample stream (sample data) is input to the reference picture memory (634). Since decoding of the symbol stream leads to bit-exact results regardless of the location of the decoder (local or remote), the content in the reference picture memory (634) is also bit-exact between a local encoder and a remote encoder. In other words, the prediction part of the encoder "sees" the same sample values as reference picture samples that the decoder will "see" when using prediction during decoding. This basic principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained, e.g., due to channel errors) is used to improve coding quality.

[0057] The operation of the "local" decoder (633) can be the same as that of a "remote" decoder such as the video decoder (510), which has already been described in detail above with reference to FIG. 5. Referring briefly to FIG. 5, however, since symbols are available and the encoding / decoding of symbols to the coded video sequence by the entropy encoder (645) and the parser (520) can be reversible, the entropy decoding portion of the video decoder (510) including the buffer memory (515) and the parser (520) may not be fully implemented in the local decoder (633) within the encoder.

[0058] What can be said at this point is that any decoder technology, except for syntax analysis / entropy decoding that can only exist within the decoder, may also necessarily exist in substantially the same functional form in the corresponding encoder. For this reason, the subject matter of the present disclosure may focus on decoder operations, which are similar to the decoding portion of the encoder. Thus, the description of encoder technology can be omitted since it is the reverse of the decoder technology that is comprehensively described. A more detailed description of the encoder is presented below only in certain areas or aspects.

[0059] During operation, in some exemplary implementations, the source coder (630) may perform motion-compensated predictive coding that predictively codes an input picture by referring to one or more previously coded pictures from the video sequence designated as "reference pictures". In this way, the coding engine (632) codes the difference (or residual) in color channels between a pixel block of the input picture and a pixel block of the (one or more) reference pictures that can be selected as the (one or more) prediction references to the input picture. The term "residual" and its adjectival form "residual" may be used interchangeably.

[0060] The local video decoder (633) can decode the coded video data of a picture that can be specified as a reference picture based on the symbols created by the source coder (630). The operation of the coding engine (632) may advantageously be an irreversible process. When the coded video data can be decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence can usually be a replica of the source video sequence with some errors. The local video decoder (633) can replicate the decoding process that can be performed by the video decoder for the reference picture and store the reconstructed reference picture in the reference picture cache (634). In this way, the video encoder (603) can locally store a copy of the reconstructed reference picture having the same content as the reconstructed reference picture obtained by the remote video decoder (without transmission errors).

[0061] The predictor (635) can perform a prediction search for the coding engine (632). That is, for a new picture to be coded, the predictor (635) can search the reference picture memory (634) for specific metadata such as sample data (as candidate reference pixel blocks) or reference picture motion vectors, block shapes, etc. that can serve as an appropriate prediction reference for the new picture. The predictor (635) can operate on the sample blocks for each pixel block to find an appropriate prediction reference. In some cases, the input picture can have a prediction reference drawn from a plurality of reference pictures stored in the reference picture memory (634) as determined by the search result obtained by the predictor (635).

[0062] The controller (650) can manage the coding operation of the source coder (630), including, for example, setting the parameters and subgroup parameters used to encode the video data.

[0063] The outputs of all the aforementioned functional units can be entropy-coded within an entropy coder (645). The entropy coder (645) converts the symbols generated by the various functional units into a coded video sequence by means of reversible compression of the symbols according to techniques such as Huffman coding, variable-length coding, arithmetic coding, etc.

[0064] The transmitter (640) can buffer the coded video sequence created by the entropy coder (645) in preparation for transmission via a communication channel (660), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (640) can merge the coded video data from the video coder (603) with other data to be transmitted, such as coded audio data and / or an auxiliary data stream (source not shown).

[0065] The controller (650) can manage the operation of the video encoder (603). During coding, the controller (650) can assign a specific coded picture type to each coded picture, which may affect the coding technique applicable to each picture. For example, a picture may often be assigned as one of the following picture types.

[0066] An intra picture (I picture) may be coded and decoded without using any other picture in the sequence as a source of prediction. Some video codecs allow for different types of intra pictures, for example, independent decoder refresh (「IDR」) pictures. Those skilled in the art are aware of those variations of I pictures as well as their respective uses and characteristics.

[0067] A predicted picture (P picture) can be coded and decoded using intra prediction or inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.

[0068] A bi-directionally predicted picture (B picture) can be coded and decoded using intra prediction or inter prediction that uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predicted pictures can use three or more reference pictures and associated metadata for the reconstruction of a single block.

[0069] A source picture is generally spatially subdivided into a plurality of sample coding blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and can be coded block by block. The blocks can be coded predictively by referring to other (already coded) blocks as determined by the coding assignment applied to each block of the picture. For example, blocks of an I picture can be coded non-predictively or predictively (spatial prediction or intra prediction) by referring to already coded blocks of the same picture. Pixel blocks of a P picture can be coded predictively via spatial prediction or via temporal prediction by referring to one previously coded reference picture. Blocks of a B picture can be coded predictively by spatial prediction or by temporal prediction by referring to one or two previously coded reference pictures. The source picture or an intermediate processed picture may be subdivided into other types of blocks for other purposes. The division of coding blocks and other types of blocks may or may not follow the same method as will be described in more detail below.

[0070] The video encoder (603) can perform a coding operation according to a predetermined video coding technology or standard such as ITU-T Rec.H.265. In that operation, the video encoder (603) can perform various compression operations including a predictive coding operation that utilizes the temporal and spatial redundancies in the input video sequence. Thus, the coded video data can conform to the syntax specified by the video coding technology or standard being used.

[0071] In some exemplary embodiments, the transmitter (640) can transmit additional data along with the encoded video. The source coder (630) can include such data as part of the encoded video sequence. The additional data can include, for example, temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures or slices, SEI messages, VUI parameter set fragments, and the like.

[0072] Video may be captured as a plurality of source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often abbreviated as intra prediction) utilizes the spatial correlation within a given picture, and inter-picture prediction utilizes the temporal or other correlations between pictures. For example, a particular picture being encoded / decoded, called the current picture, can be divided into blocks. A block within the current picture can be coded by a vector called a motion vector if it is similar to a reference block within a reference picture that was previously coded and is subsequently buffered within the video. The motion vector points to the reference block within the reference picture and can have a third dimension that identifies the reference picture if multiple reference pictures are being used.

[0073] In some exemplary embodiments, a dual prediction technique can be used for inter-picture prediction. According to such a dual prediction technique, two reference pictures, such as a first reference picture and a second reference picture, are used, both of which advance the current picture in the video in decoding order (however, in display order, they can be in the past or future respectively). A block in the current picture can be coded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block can be predicted cooperatively by a combination of the first reference block and the second reference block.

[0074] Furthermore, a merge mode technique may be used to improve coding efficiency in inter-picture prediction.

[0075] According to some exemplary embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed on a block-by-block basis. For example, pictures in a sequence of video pictures are divided into coding tree units (CTUs) for compression, and the CTUs within a picture can have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU can include three parallel coding tree blocks (CTBs), namely, one luma CTB and two chroma CTBs. Each CTU can be recursively quad-tree divided into one or more coding units (CUs). For example, a 64×64 pixel CTU can be divided into one 64×64 pixel CU or four 32×32 pixel CUs. Each of one or more of the 32×32 blocks can be further divided into four 16×16 pixel CUs. In some exemplary embodiments, each CU can be analyzed during encoding to determine the prediction type of that CU from various prediction types, such as an inter prediction type or an intra prediction type. A CU can be divided into one or more prediction units (PUs) according to temporal and / or spatial predictability. Generally, each PU includes one luma prediction block (PB) and two chroma PBs. In one embodiment, the prediction operation in coding (encoding / decoding) is performed on a prediction block-by-block basis. The division of a CU into PUs (or PBs of different color channels) can be performed in various spatial patterns. A luma PB or a chroma PB can include a matrix of sample values (e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0076] FIG. 7 shows a diagram of a video encoder (703) according to another exemplary embodiment of the present disclosure. The video encoder (703) receives a processing block (e.g., a prediction block) of sample values within a current video picture in a sequence of video pictures and is configured to encode the processing block into a coded picture that is part of a coded video sequence. An exemplary video encoder (703) may be used in place of the video encoder (403) of the example of FIG. 4.

[0077] For example, the video encoder (703) receives a matrix of sample values of a processing block such as an 8×8 sample prediction block. The video encoder (703) then determines, for example using rate distortion optimization (RDO), whether the processing block is best coded using it in an intra mode, an inter mode, or a bi-prediction mode. If it is determined that the processing block is to be coded in the intra mode, the video encoder (703) encodes the processing block into the coded picture using intra prediction techniques, and if it is determined that the processing block is to be coded in the inter mode or the bi-prediction mode, the video encoder (703) may encode the processing block into the coded picture using inter prediction techniques or bi-prediction techniques, respectively. In some exemplary embodiments, as a sub-mode of inter-picture prediction, a merge mode derived from one or more motion vector predictors without the benefit of coded motion vector components outside the predictor may be used. In some other exemplary embodiments, there may be motion vector components applicable to the target block. Thus, the video encoder (703) may include components not explicitly shown in FIG. 7, such as a mode decision module, to determine the prediction mode of the processing block.

[0078] In the example of FIG. 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general-purpose controller (721), and an entropy encoder (725) that are coupled to each other as shown in the exemplary configuration of FIG. 7.

[0079] The inter-encoder (730) receives samples of the current block (e.g., a processing block), compares the block with one or more reference blocks (e.g., blocks in the previous and subsequent pictures in display order) in a reference picture, generates inter-prediction information (e.g., description of redundant information, motion vectors, merge mode information by inter-encoding techniques), and is configured to calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture that is decoded based on video information encoded using a decoding unit 633 incorporated in the exemplary encoder 620 of FIG. 6 (shown as the residual decoder 728 of FIG. 7 and described in more detail below).

[0080] The intra-encoder (722) receives samples of the current block (e.g., a processing block), compares the block with already-coded blocks in the same picture, generates quantized coefficients after transformation, and optionally also generates intra-prediction information (e.g., intra-prediction direction information by one or more intra-encoding techniques). The intra-encoder (722) can calculate an intra-prediction result (e.g., a predicted block) based on the intra-prediction information and reference blocks in the same picture.

[0081] The general-purpose controller (721) may be configured to determine general-purpose control data and control other components of the video encoder (703) based on the general-purpose control data. In one example, the general-purpose controller (721) determines the prediction mode of a block and provides a control signal to the switch (726) based on the prediction mode. For example, when the prediction mode is the intra mode, the general-purpose controller (721) controls the switch (726) to select the intra mode result for use by the residual calculator (723), controls the entropy encoder (725) to select the intra prediction information and include the intra prediction information in the bitstream, and when the description mode of the block is the inter mode, the general-purpose controller (721) controls the switch (726) to select the inter prediction result for use by the residual calculator (723), controls the entropy encoder (725) to select the inter prediction information and include the inter prediction information in the bitstream.

[0082] The residual calculator (723) may be configured to calculate the difference (residual data) between the received block and the prediction result for the block selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) may be configured to encode the residual data to generate a transformation coefficient. For example, the residual encoder (724) may be configured to convert the residual data from the spatial domain to the frequency domain to generate a transformation coefficient. Next, the transformation coefficient is subjected to quantization processing to obtain a quantized transformation coefficient. In various exemplary embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transformation and generate decoded residual data. The decoded residual data can be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and the intra prediction information. The decoded block is appropriately processed to generate a decoded picture, and the decoded picture can be buffered in a memory circuit (not shown) and used as a reference picture.

[0083] The entropy encoder (725) is configured to format the bitstream to include the encoded block and perform entropy coding. The entropy encoder (725) is configured to include various information in the bitstream. For example, the entropy encoder (725) may be configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other appropriate information in the bitstream. When coding a block in either the inter mode or the merge submode of the bi-prediction mode, there may be no residual information.

[0084] FIG. 8 shows a diagram of an exemplary video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive a coded picture that is part of a coded video sequence and decode the coded picture to generate a reconstructed picture. In one example, the video decoder (810) can be used in place of the video decoder (410) of the example of FIG. 4.

[0085] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter-decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-decoder (872) coupled to each other as shown in the exemplary configuration of FIG. 8.

[0086] The entropy decoder (871) can be configured to reconstruct from the coded picture specific symbols representing the syntax elements that the coded picture is composed of. Such symbols can include, for example, the mode in which a block is coded (e.g., intra mode, inter mode, bi-prediction mode, merge sub-mode or another sub-mode), prediction information (e.g., intra prediction information or inter prediction information) that can identify specific samples or metadata used for prediction by the intra-decoder (872) or the inter-decoder (880), and residual information such as in the form of quantized transform coefficients. In one example, when the prediction mode is inter mode or bi-prediction mode, the inter prediction information is provided to the inter-decoder (880), and when the prediction type is intra prediction type, the intra prediction information is provided to the intra-decoder (872). The residual information can be inverse quantized and provided to the residual decoder (873).

[0087] The inter-decoder (880) can be configured to receive the inter prediction information and generate an inter prediction result based on the inter prediction information.

[0088] The intra decoder (872) may be configured to receive intra prediction information and generate a prediction result based on the intra prediction information.

[0089] The residual decoder (873) may be configured to perform inverse quantization to extract inverse quantization transform coefficients, and process the inverse quantization transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also use certain control information (for including quantization parameter (QP)), and that information may be provided by the entropy decoder (871) (since this may be only a small amount of control information, the data path is not shown).

[0090] The reconstruction module (874) may be configured to combine, in the spatial domain, the residual as the output by the residual decoder (873) and the prediction result (optionally, as the output by the inter prediction module or the intra prediction module) to form a reconstructed block that forms a part of the reconstructed picture as a part of the reconstructed video. Note that other appropriate operations such as deblocking operations may be performed to improve visual quality.

[0091] Note that the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) may be implemented using any appropriate technique. In some exemplary embodiments, the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (603), and the video decoders (410), (510), and (810) can be implemented using one or more processors that execute software instructions.

[0092] When returning to intra prediction processing, samples within a block (e.g., a luma or chroma prediction block, or a coding block if not further divided into prediction blocks) are predicted by samples in adjacent lines, the next adjacent lines, or one or more other lines, or combinations thereof, to generate a prediction block. Thereafter, the residual between the actual block and the prediction block during coding may be processed by transformation after quantization. Various intra prediction modes can be made available, and parameters related to the selection of the intra mode and other parameters can be signaled in the bitstream. The various intra prediction modes can relate to, for example, one or more line positions for predicting samples, the direction in which predicted samples are selected from predicting one or more lines, and other special intra prediction modes.

[0093] For example, a set of intra prediction modes (also referred to interchangeably as "intra modes") can include a predetermined number of directional intra prediction modes. As described above with respect to the exemplary implementation of FIG. 1, these intra prediction modes may correspond to a predetermined number of directions to proceed when selecting samples outside the block as the prediction destination of samples being predicted within a particular block. In another particular exemplary implementation, eight main direction modes corresponding to angles from 45 degrees to 207 degrees with respect to the horizontal axis can be supported and predefined.

[0094] In some other implementations of intra prediction, to further utilize more diverse spatial redundancy in the direction texture, the directional intra mode can be further extended to an angle set with finer granularity. For example, as shown in FIG. 9, the above eight-angle implementation may be configured to provide eight nominal angles (referred to as V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED), and for each nominal angle, a predetermined number (e.g., seven) of finer angles may be added. By expanding in this way, the total number of direction angles increases (e.g., 56 in this example), and the total number of these direction angles can be used for intra prediction, which corresponds to the same number of predetermined directional intra modes. The prediction angles may be represented by the nominal intra angle and the angle difference associated therewith. In the above specific example having seven finer angle directions for each nominal angle, the angle difference may be -3 to 3 times the step size of 3 degrees.

[0095] In some implementations, instead of or in addition to the above directional intra mode, a predetermined number of non-directional intra prediction modes may also be predefined and made available. For example, five non-directional intra modes, referred to as smooth intra prediction modes, may be specified. These non-directional intra mode prediction modes may be particularly referred to as DC intra mode, PAETH intra mode, SMOOTH intra mode, SMOOTH_V intra mode, and SMOOTH_H intra mode. An example of predicting a sample of a specific block using these non-directional modes is shown in FIG. 10. For example, FIG. 10 shows how a 4×4 block 1002 is predicted by samples obtained from the upper adjacent line and / or the left adjacent line. A specific sample 1010 within block 1002 may correspond to the sample 1004 directly above sample 1010 in the upper adjacent line of block 1002, the sample 1006 in the upper left of sample 1010 as the intersection of the upper adjacent line and the left adjacent line, and the sample 1008 directly to the left of sample 1010 in the left adjacent line of block 1002. In an example of the DC intra prediction mode, the average value of the left adjacent sample 1008 and the upper adjacent sample 1004 may be used as the predictor for sample 1010. In an example of the PAETH intra prediction mode, the upper, left, and upper left reference samples 1004, 1008, and 1006 may be obtained, and then any value among these three reference samples that is closest to (upper + left - upper left) may be set as the predictor for sample 1010. In an example of the SMOOTH_V intra prediction mode, sample 1010 may be predicted by vertical quadratic interpolation of the upper left adjacent sample 1006 and the left adjacent sample 1008. In an example of the SMOOTH_H intra prediction mode, sample 1010 may be predicted by horizontal quadratic interpolation of the upper left adjacent sample 1006 and the upper adjacent sample 1004. In an example of the SMOOTH intra prediction mode, sample 1010 may be predicted by the average of vertical and horizontal quadratic interpolations. The above examples of implementing the non-directional intra mode are shown only as non-limiting examples.It is also possible to combine other adjacent lines, and other non-directional selections of samples, and prediction samples for predicting specific samples within the prediction block.

[0096] The selection of a specific intra prediction mode by an encoder from the above-mentioned directional mode or non-directional mode at various coding levels (picture, slice, block, unit, etc.) can be signaled in the bitstream. In some exemplary implementations, first, eight exemplary nominal directional modes may be signaled together with five smooth modes without using angles (a total of 13 options). Thereafter, if the signaled mode is one of the eight nominal angular intra modes, an index indicating the selected angular difference with respect to the corresponding signaled nominal angle is further signaled. In some other exemplary implementations, all intra prediction modes may be indexed together for signaling (e.g., adding five non-directional modes to 56 directional modes to generate 61 intra prediction modes).

[0097] In some exemplary implementations, the exemplary 56 or other number of directional intra prediction modes can be implemented using a unified directional predictor that projects each sample of the block onto a reference subsample position and interpolates the reference samples by a 2-tap bilinear filter.

[0098] In some implementations, an additional filter mode called FILTER INTRA mode can be designed to capture the decaying spatial correlation with references on the edge. In this mode, in addition to samples outside the block, samples predicted within the block may be used as intra prediction reference samples for some patches within the block. For example, the mode may be a default mode, or the mode may be made available for at least intra prediction of luma blocks (or intra prediction used only for luma blocks). A default number (e.g., 5) of filter intra modes may be designed in advance. For example, each filter intra mode may be represented by a set of n-tap filters (e.g., 7-tap filters) that reflect the correlation between samples within a 4×2 patch and n adjacent neighbors. In other words, the weight coefficients of the n-tap filters may depend on the position. As shown in FIG. 11, as an example, when using an 8×8 block, a 4×2 patch, and 7-tap filtering, the 8×8 block 1102 may be divided into 8 4×2 patches. In FIG. 11, these patches are denoted as B0, B1, B1, B3, B4, B5, B6, and B7. For each patch, the 7 neighbors of the patch (denoted as R0 to R7 in FIG. 11) may be used to predict the samples within the target patch. For patch B0, all neighbors may already be reconstructed. On the other hand, for other patches, since some of the neighbors are within the target block, they may not be reconstructed. In that case, the predicted values of the directly adjacent ones are used as references. For example, since all neighbors of patch B7 as shown in FIG. 11 are not reconstructed, the predicted samples of the neighbors of patch B7 are used instead.

[0099] In some implementations of intra prediction, one color component can be predicted using one or more other color components. The color component can be any one of the components such as the YCrCb color space, the RGB color space, the XYZ color space, etc. For example, a prediction of predicting a chroma component (e.g., a chroma block) from a luma component (e.g., a luma reference sample), luma to chroma, i.e., CfL (Chroma from Luma) can be performed. In some exemplary implementations, much of the cross-color prediction is only allowed from luma to chroma. For example, the chroma samples within a chroma block can be modeled as a linear function of the corresponding reconstructed luma samples. CfL prediction can be performed as follows. CfL(α)=α×L AC +DC (1)

[0100] Here, L AC represents the AC contribution of the luma component, α represents the parameter of the linear model, and DC represents the DC contribution of the chroma component. For example, while the AC component is obtained for each sample of the block, the DC component is obtained for the entire block. Further, subsampling may be performed on the reconstructed luma samples to obtain the chroma resolution, and then the average luma value (luma DC) may be subtracted from each luma value to generate the AC contribution of the luma. Then, the AC contribution of the luma is used in the linear mode of Equation (1) to predict the AC value of the chroma component. Instead of requiring the decoder to calculate a scaling parameter to obtain or predict an approximation of the chroma AC component from the luma AC contribution, in an implementation example of CfL, the parameter α can be determined based on the original chroma samples and signaled in the bitstream. This reduces the complexity of the decoder and obtains a more accurate prediction. Regarding the DC contribution of the chroma component, in some exemplary implementations, it can be calculated using the intra DC mode within the chroma component.

[0101] Subsequently, following the quantization of the transform coefficients, the transform of the residual of either an intra-predicted block or an inter-predicted block may be performed. To perform the transform, the intra-coded block and the inter-coded block may be further divided into a plurality of transform blocks (when the term "unit" is used in its normal usage to represent a set of three color channels (for example, when a "coding unit" includes one luma coding block and a plurality of chroma coding blocks), it may instead be used as a "transform unit") before the transform. In some implementations, the maximum division depth of the coded block (i.e., the prediction block) may be specified (the term "coded block" may be used instead of "coding block"). For example, the division may be at a level of two steps or less. When dividing the prediction block into transform blocks, different processes may be performed for the intra-predicted block and the inter-predicted block. On the other hand, in some implementations, the same process may be performed for the intra-predicted block and the inter-predicted block during the division.

[0102] In some exemplary implementations, for an intra-coded block, the transform division may be performed such that all transform blocks have the same size, and the transform blocks are coded in raster scan order. An example of such transform block division of an intra-coded block is shown in FIG. 12. Specifically, FIG. 12 shows the coded block 1202 being divided into 16 transform blocks of the same block size, indicated by 1206 via an intermediate-level quadtree division 1204. FIG. 12 shows an example of the raster scan order of coding by arrows arranged in sequence.

[0103] In some exemplary implementations, for inter-coded blocks, recursive transform unit splitting may be performed using a split depth up to a default number of levels (e.g., the level of the second stage). As shown in FIG. 13, the splitting may be aborted or the recursive splitting may be continued at any level for some subdivision. Specifically, FIG. 13 shows an example in which block 1302 is split into four quadtree sub-blocks 1304, and one of the sub-blocks is further split into four transform blocks at the level of the second stage, while the splitting of the other sub-blocks is aborted after the level of the first stage, resulting in a total of seven transform blocks of two different sizes. FIG. 13 is further shown by arrows arranged in order with an example of the raster scan order of coding. FIG. 13 shows an exemplary implementation of the quadtree splitting of square transform blocks up to the level of the second stage, but in some implementations regarding generation, the splitting for the transform may support transform block shapes of 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1, and sizes ranging from 4×4 to 64×64. In some exemplary implementations, when the coding block is 64×64 or less, the splitting of the transform block may be applied only to the luma component (in other words, in this state, the chroma transform block becomes the same as the coding block). Different from the above, when the width or height of the coding block exceeds 64, the luma coding block and the chroma coding block may be implicitly split into transform blocks in multiples of min(W,64)×min(H,64) and min(W,32)×min(H,32), respectively.

[0104] Thereafter, each of the above transform blocks may be subjected to a primary transform. By the primary transform, the residual of the transform block is substantially moved from the spatial domain to the frequency domain. In some implementations of the actual primary transform, in order to support the above example of the extended coding block splitting, multiple transform sizes (ranging from 4 points to 64 points for each dimension of the two dimensions) and transform shapes (square, rectangle with a width / height ratio of 2:1 / 1:2 and 4:1 / 1:4) may be allowed.

[0105] Focusing on the actual primary transformation, in some exemplary implementations, the use of a composite transformation kernel (e.g., which may be composed of different one-dimensional transformations for each dimension of the coded residual transformation block) may be required for the two-dimensional transformation process. Examples of one-dimensional transformation kernels may include, but are not limited to: a) DCT-2 of 4 points, 8 points, 16 points, 32 points, 64 points; b) Asymmetric DST types (DST-4, DST-7) of 4 points, 8 points, 16 points and their inverted types; c) Identity transformation of 4 points, 8 points, 16 points, 32 points. The selection of the transformation kernel used for each dimension may be based on rate-distortion (RD) criteria. For example, the basis functions of DCT-2 and asymmetric DST that can be implemented are listed in Table 1.

[0106] [Table 1]

[0107] In some exemplary implementations, the effectiveness of the composite transformation kernel of a particular implementation of the primary transformation may be based on the size of the transformation block and the prediction mode. An example of the dependency is shown in Table 2. For the chroma component, the selection of the transformation type may be implicitly performed. For example, for the intra prediction residual, as described in Table 3, the transformation type may be selected according to the intra prediction mode. For the inter prediction residual, the transformation type of the chroma block may be selected according to the selection of the transformation type of the luma block at the same position. Therefore, in the case of the chroma component, there is no signaling of the transformation type in the bitstream.

[0108] [Table 2]

[0109] [Table 3]

[0110] In some implementations, a secondary transformation may be performed on the primary transformation coefficients. For example, as shown in FIG. 14, in order to further decorrelate the primary transformation coefficients, a LFNST (Low Frequency Non-Separable Transform), known as a reduced secondary transformation, can be applied between the forward primary transformation and quantization (in the encoder) and between inverse quantization and the inverse primary transformation (on the decoder side). Essentially, the LFNST can take a portion of the primary transformation coefficients, for example, the low-frequency portion (thus, a “reduced” portion from the complete set of primary transformation coefficients of the transformation block), in order to proceed with the secondary transformation. In an example of LFNST, a 4×4 non-separable transform or an 8×8 non-separable transform may be applied according to the transformation block size. For example, a 4×4 LFNST may be applied to a small transformation block (e.g., min(width, height) < 8), and an 8×8 LFNST may be applied to a large transformation block (e.g., min(width, height) > 8). For example, when an 8×8 transformation block is affected by a 4×4 LFNST, only the low-frequency 4×4 portion of the 8×8 primary transformation coefficients is further subjected to the secondary transformation.

[0111] As specifically shown in FIG. 14, the transformation block may be 8×8 (or 16×16). Thus, the forward primary transformation 1402 of the transformation block generates an 8×8 (or 16×16) primary transformation coefficient matrix 1404, and each square unit represents a 2×2 (or 4×4) portion. The input to the forward LFNST may not be, for example, the entire 8×8 (or 16×16) primary transformation coefficients. For example, a 4×4 (or 8×8) LFNST may be used for the secondary transformation. Thus, as shown in the hatched portion (upper left) 1406, only the 4×4 (or 8×8) low-frequency primary transformation coefficients of the primary transformation coefficient matrix 1404 can be used as the input to the LFNST. The remaining portion of the primary transformation coefficient matrix may not be subjected to the secondary transformation. In this way, after the secondary transformation, the portion of the primary transformation coefficients affected by the LFNST becomes the secondary transformation coefficients, but the remaining portion not affected by the LFNST (e.g., the non-shadowed portion of the matrix 1404) maintains the corresponding primary transformation coefficients. In some exemplary implementations, the remaining portion not subject to the secondary transformation may all be set to 0 coefficients.

[0112] An example of the application of the inseparable transform used in the LFNST will be described below. To apply an example 4×4 LFNST, a 4×4 input block X (for example, representing the 4×4 low-frequency part of the primary transform coefficient block such as the hatched part 1406 of the primary transform matrix 1404 in FIG. 14) can be represented as follows.

Number

[0113] This two-dimensional input matrix is first linearized or scanned into a vector

Number

Number

[0114] Next, the inseparable transform of the 4×4 LFNST can be calculated as

Number

Number

Number

[0115] The LFNST of the above example is based on a direct matrix multiplication technique for applying non-separable transforms, and as a result, it is implemented in a single pass without multiple iterations. In some further exemplary implementations, the dimension of the non-separable transform matrix (T) of the example 4×4 LFNST can be further reduced in order to minimize the computational complexity and memory space requirements for storing the transform coefficients. Such an implementation may be referred to as a reduced non-separable transform (RST). More specifically, the main concept of RST is to map an N (where N is 4×4 = 16 in the above example, but may be equal to 64 for an 8×8 block) -dimensional vector to an R -dimensional vector in a different space, and N / R (R < N) represents the dimension reduction factor. Therefore, instead of an N×N transform matrix, the RST matrix becomes an R×N matrix as follows, [Number]

[0116] where the R rows of the transform matrix are the reduced R basis of the N -dimensional space. Therefore, the transform converts the input vector or N dimensions to an output vector of the reduced R dimensions. Thus, as shown in FIG. 14, the secondary transform coefficients 1408 transformed from the primary coefficients 1406 are reduced in dimension by a factor of the coefficient or N / R. The three squares around 1408 in FIG. 14 may be zero - padded.

[0117] The inverse transformation matrix of the RTS can be the transpose of its forward transformation. In the case of an 8×8 LFNST for example (contrasted with the above 4×4 LFNST for more diverse explanations), a reduction factor of 4 can be applied. Thus, the 64×64 non-separable transformation matrix is correspondingly reduced to a 16×64 direct matrix. Further, in some implementations, not all but some of the input primary coefficients may be linearized into the input vector of the LFNST. For example, only a part of the 8×8 input primary transformation coefficients may be linearized into the above X vector. In a specific example, among the four 4×4 quadrants of the 8×8 primary transformation coefficient matrix, the lower right (high-frequency coefficients) can be excluded, and only the other three quadrants are linearized into a 64×1 vector using a predefined scanning order instead of a 48×1 vector. In such an implementation, the non-separable transformation matrix may be further reduced from 16×64 to 16×48.

[0118] Therefore, a reduced 48×16 inverse RST matrix can be used on the decoder side to generate the upper left, upper right, and lower left 4×4 quadrants of the 8×8 core (primary) transformation coefficients. Specifically, when a further reduced 16×48 RST matrix is applied instead of the 16×64 RST with the same transformation set configuration, the non-separable secondary transformation takes as input the 48 matrix elements vectorized from three 4×4 quadrant blocks of the 8×8 primary coefficient block excluding the lower right 4×4 block. In such an implementation, the omitted lower right 4×4 primary transformation coefficients are ignored by the secondary transformation. This further reduced transformation converts a 48×1 vector into a 16×1 output vector, which is scanned in reverse into a 4×4 matrix to satisfy 1408 in FIG. 14. The three squares of the secondary transformation coefficients surrounding 1408 may be zero-padded.

[0119] With the help of such a reduction in the dimensions of the RST, the memory usage for storing all LFNST matrices is reduced. In the above example, for instance, the memory usage can be reduced from 10KB to 8KB with a reasonably small performance degradation compared to an implementation without dimensional reduction.

[0120] In some implementations, to reduce complexity, LFNST may be further restricted to be applicable only when all coefficients outside the primary transform coefficient part targeted by LFNST (e.g., outside the 1406 part of 1404 in FIG. 14) are insignificant. Thus, when LFNST is applied, all primary-only transform coefficients (e.g., the unshaded part of the primary coefficient matrix 1404 in FIG. 4) can be close to 0. Such a restriction enables adjustment of the LFNST index signal transmission at the final significant position, and thus avoids additional coefficient scans that may be required to check for significant coefficients at specific positions when this restriction is not applied. In some implementations, the worst-case processing of LFNST (with respect to multiplications per pixel) can limit the inseparable transforms of 4×4 and 8×8 blocks to 8×16 and 8×48 transforms, respectively. In such cases, when LFNST is applied, for other sizes less than 16, the final significant scan position must be less than 8. For blocks having a shape of 4×N and N×4 and N>8, the above restriction means that LFNST is applied only once to the upper left 4×4 region. When LFNST is applied, since all primary-only coefficients are 0, in such cases, the number of operations required for the primary transform is reduced. From the perspective of the encoder, coefficient quantization can be simplified when the LFNST transform is tested. Rate-distortion optimized quantization (RDO) may have to be performed at most for the first 16 coefficients (in scan order), and the remaining coefficients may be forced to be 0.

[0121] In some exemplary implementations, the available RST kernels may be specified as a number of transform sets, where each transform set includes a number of non-separable transform matrices. For example, there may be a total of four transform sets and two non-separable transform matrices (kernels) for each transform set used in LFNST. These kernels may be pre-trained offline and are thus data-driven. The offline-trained transform kernels may be stored in memory for use during encoding / decoding processing, or may be hard-coded into the encoding device or the decoding device. The selection of the transform set during encoding or decoding processing may be determined by the intra prediction mode. The mapping from the intra prediction mode to the transform set may be pre-defined. An example of such a pre-defined mapping is shown in Table 4. For example, as shown in Table 4, if one of the three cross-component linear model (CCLM) modes (INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the current block (i.e., 81 <= predModeIntra <= 83), transform set 0 may be selected for the current chroma block. For each transform set, the selected non-separable secondary transform candidate may be further specified by an explicitly signaled LFNST index. For example, the index may be signaled in the bitstream once per intra CU, after the transform coefficients.

[0122] [Table 4]

[0123] In the above exemplary implementation, since LFNST is restricted to be applicable only when all coefficients outside the first coefficient subgroup or part are not significant, the LFNST index coding depends on the position of the final significant coefficient. In addition, the LFNST index may be context-coded, but it does not depend on the intra prediction mode, and only the first bin may be context-coded. Furthermore, LFNST can be applied to both intra CUs in both intra and inter slices, as well as to both luma and chroma. When the dual tree is enabled, the LFNST indexes for the luma component and the chroma component can be signaled separately. In the case of inter slice (when the dual tree is disabled), a single LFNST index is signaled and used for both luma and chroma.

[0124] In some exemplary implementations, when the intra subpartitioning (ISP) mode is selected, since there may be a high limit to the performance improvement even if RST is applied to all executable partition blocks, LFNST may be disabled and the RST index may not be signaled. Furthermore, disabling the RST of the ISP prediction residual can reduce the encoding complexity. In some further implementations, when the multiple linear regression intra prediction (MIP) mode is selected, LFNST may also be disabled and the RST index may not need to be signaled.

[0125] Considering that due to the existing maximum transform size limit (e.g., 64×64), large CUs (or any other predefined size representing the maximum transform block size) exceeding 64×64 are implicitly divided (e.g., TU tiling), the LFNST index search can quadruple the data buffering for a specific number of decoding pipeline stages. Therefore, in some implementations, the maximum size allowed for LFNST can be limited to, for example, 64×64. In some implementations, LFNST can be enabled only with DCT2 as the primary transform.

[0126] In some other implementations, for example, by using three kernels within each set to define, for example, 12 sets of secondary transforms, an intra secondary transform (IST) is provided for the luma component. An intra mode-dependent index may be used for transform set selection. The kernel selection within a set may be based on the signaling syntax elements. The IST may be enabled when either DCT2 or ADST is used as both the horizontal and vertical primary transforms. In some implementations, a non-separable 4×4 transform or an 8×8 non-separable transform can be selected according to the block size. When min(tx_width, tx_height) < 8, a 4×4 IST can be selected. For larger blocks, an 8×8 IST can be used. Here, tx_width and tx_height correspond to the width and height of the transform block, respectively. The input to the IST may be the low-frequency primary transform coefficients in zigzag scan order.

[0127] In various transformations in video coding or decoding processes, such as either a primary transformation of samples within a residual block or a secondary transformation of a block of primary transformation coefficient processes, when only separable transformation methods are used, it may not always be efficient when capturing directional texture patterns such as edges in a 45-degree direction (e.g., a direction substantially away from the horizontal or vertical direction). As described above, in some exemplary implementations, for the secondary transformation of primary transformation coefficients, one or more non-separable transformation designs can be used. As further described below, such non-separable transformation methods can also be used for the primary transformation of the residual sample block to generate primary transformation coefficients. In some exemplary implementations, the primary transformation coefficients generated via either separable or non-separable transformation can be directly subjected to quantization followed by entropy coding, or can be subjected to either separable or non-separable secondary transformation before quantization and entropy coding. In implementations where the primary transformation coefficients obtained via non-separable transformation are further subjected to non-separable secondary transformation, two non-separable transformations have been utilized in a cascaded manner.

[0128] The following disclosure further describes some exemplary implementations of non-separable transformation methods applicable to both primary and secondary transformations. These non-separable transformation designs described below are particularly aimed at improving the coding efficiency of directional image patterns. In particular, non-separable primary and / or secondary transformation methods dependent on intra mode are disclosed. These exemplary embodiments focus on the data scanning order within a block that is coded / decoded during either the forward transformation process or the inverse transformation process. The block being processed may sometimes generally be referred to as a data block, and in the case of forward transformation, includes a residual transformation / coding / prediction block or samples of primary transformation coefficients, and in the case of inverse transformation, may include secondary transformation coefficients that are inverse-transformed from primary transformation coefficients, or primary transformation coefficients that are inverse-transformed from residual samples.

[0129] Optimize the energy compression achieved by non-separable primary or secondary conversion and effectively quantize the primary or secondary conversion coefficients (to discard high-frequency coefficients), as will be described in more detail below. The input and output scanning processes of non-separable primary or secondary conversion can be performed considering the intra prediction mode, block size, and / or primary conversion type.

[0130] In some exemplary implementations, the set of conversions may refer to a group of one or more conversion kernels as candidates or options for the encoder to select data blocks during the coding process.

[0131] In some implementations, the primary conversion may be performed using a non-separable conversion or may be achieved by performing a series of one-dimensional conversions. For example, in the case of a DCT_DCT combination, the DCT is applied to the block horizontally and vertically. In another example, in the case of an ADST_ADST combination, the one-dimensional ADST is applied to the block horizontally and vertically. In some implementations, different conversion types can be used horizontally and vertically. Such conversions may sometimes be referred to as composite primary conversions.

[0132] FIG. 15 shows an exemplary data flow 1500 for non-separable conversions in the forward and reverse directions. Various data scanning processes are identified as S1 and S2, and the non-separable conversions in the forward and reverse directions are identified as b 1502 and 1504 in FIG. 15. The operation of the exemplary data flow 1500 applies to both primary and secondary non-separable conversion processes and will be described in more detail below.

[0133] Data Scanning in Non-Separable Secondary Conversion In this exemplary embodiment, the forward and inverse secondary conversions may be inseparable conversions, as shown by 1502 and 1504 in FIG. 15. From the encoding side, the input of the forward secondary conversion 1502 may be N primary conversion coefficients 1506 of the primary coefficient data block 1508 scanned in a data sequence 1510 according to the first scan order S1 in the scan process 1512, and the output of the forward secondary conversion 1502 may be a secondary conversion coefficient sequence 1514 of K data items that replace K input primary conversion coefficients 1516 of the primary conversion coefficients to generate a modified primary conversion coefficient data block 1507 according to the data scan 1518 using the second scan order S2. N and K are positive integers.

[0134] From the decoding side, the input of the inverse secondary conversion 1504 is a secondary conversion coefficient sequence 1520 of K data items obtained by reversely scanning the K input coefficients 1517 of the modified primary coefficient data block 1509 in the reverse S2 scan order in the scan process 1522, and the output of the inverse secondary conversion 1504 may be a primary conversion coefficient sequence 1524 of N coefficients that replace the N primary conversion coefficients 1526 of 1509 to generate a primary coefficient data block 1528 using the reverse scan process 1530 according to the reverse S1 scan order.

[0135] In some exemplary implementations, N may be greater than K. In other words, the inseparable secondary conversion 1502 may be a reduced inseparable secondary conversion. In some exemplary embodiments, the N primary conversion coefficients 1506 may be the first N coefficients of the primary coefficient data block 1506, for example, the upper left corner (or low frequency portion) of the primary coefficient data block 1506.

[0136] In one exemplary implementation, the scan order S1 may depend on at least one of the intra prediction mode associated with the input data block 1508, the primary conversion type associated with the data block, or the block size of the data block.

[0137] In one exemplary implementation, the scanning order S2 may depend on at least one of the intra prediction mode associated with the input data block 1508, the primary conversion type associated with the data block, or the block size of the data block.

[0138] In one exemplary implementation, the scanning orders of S1 and S2 may include at least one of a zigzag scanning order, a diagonal scanning order, or a row and column scanning order.

[0139] In one implementation, when the same set of secondary conversions (having one or more second conversion kernels) is used as secondary conversion candidates for multiple intra prediction modes, the scanning orders of S1 and S2 may depend on the intra prediction mode associated with the data block.

[0140] In some exemplary implementations, the scanning orders of S1 and S2 may be different.

[0141] In some exemplary implementations, N and K include, but are not limited to, any integer between 0 and 127, inclusive.

[0142] In some exemplary implementations, when the data block is a square block, the scanning order of S1 and / or S2 may include a zigzag order. When the block is a non-square block, the scanning order of S1 and / or S2 may include a diagonal scanning order.

[0143] FIG. 15 further shows quantization and entropy encoding 1540, as well as corresponding entropy decoding and inverse quantization 1550.

[0144] Data Scanning in Inseparable Primary Conversion as Primary Conversion In this exemplary embodiment, the irreducible transformation processes 1502 and 1504 may be performed with respect to the primary transformation. Thus, the input to the data flow of 1500(1508) is, for example, a residual sample data block rather than a primary transformation coefficient, and the output 1507 is a modified residual sample data block including K primary transformation coefficients 1516 from the irreducible transformation process 1502. Specifically, from the encoder side, the input to the forward primary transformation may be the first N residual samples according to the first scanning order S1, and the output of the forward primary transformation is K irreducible transformation coefficients that replace K input residual samples in the second scanning order S2. N and K are positive integers. From the decoder side, the input to the inverse irreducible transformation is K irreducible transformation coefficients according to the S2 scanning order, and the output of the inverse irreducible transformation is N residual samples that replace N irreducible transformation coefficients according to the S1 scanning order, where N is an integer.

[0145] In one exemplary implementation, the scanning order S1 may depend on at least one of the intra prediction mode associated with the data block, the primary transformation type associated with the data block, or the block size of the data block.

[0146] In one exemplary implementation, the scanning order S2 may depend on at least one of the intra prediction mode associated with the data block, the primary transformation type associated with the data block, or the block size of the data block.

[0147] In one exemplary implementation, the scanning orders of S1 and S2 may include at least one of a zigzag scanning order, a diagonal scanning order, or a row and column scanning order.

[0148] In one exemplary implementation, when the same set of transformations (having one or more second transformation kernels) is used as a secondary transformation candidate for multiple intra prediction modes, the scanning orders of S1 and S2 may depend on the intra prediction mode associated with the data block.

[0149] In some exemplary implementations, the scanning orders of S1 and S2 may be different.

[0150] In one exemplary implementation, N and K include, but are not limited to, any integer between 0 and 127 including both ends.

[0151] In one exemplary implementation, when the data block is a square block, the scanning order of S1 and / or S2 may include a zigzag scanning order. When the block is a non-square block, the scanning order of S1 and / or S2 may include a diagonal scanning order. In other words, for either a square or non-square data block, a diagonal scanning order may be enforced.

[0152] In one exemplary implementation, the S1 scanning order may include a horizontal (row) scanning order or a vertical (column) scanning order.

[0153] In these implementations, the quantization / entropy processing 1540 and the entropy decoding / inverse quantization processing 1550 may be performed directly on 1507 or 1509 without additional secondary transformation. In some other alternative implementations, additional secondary transformation may be performed prior to the encoding-side processing 1540 and following the decoding-side processing 1550. The additional secondary transformation may be separable or non-separable, for example, following the implementation described above for non-separable secondary transformation.

[0154] Data Scanning in Non-Separable Primary Transformation as Primary Transformation In some embodiments, the type of transformation (i.e., separable or non-transcribable) may affect the scanning order when performing transform coefficient coding (e.g., entropy coding of transform coefficients). In one implementation, when non-separable transformation is applied (to the primary transformation and / or secondary transformation), a horizontal (row) or vertical (column) scanning order may be used for transform coefficient coding.

[0155] FIG. 16 shows an exemplary method 1600 for processing video data. The method 1600 may include some or all of the following steps: step 1610 of extracting a data block from the video data; step 1620 of scanning a first number of data items in the data block according to a first scanning order to generate a first data sequence; step 1630 of performing an irreducible transform on the first data sequence to obtain a second data sequence having a second number of data items; and step 1640 of replacing at least a part of the first number of data items in the data block with part or all of a second data sequence according to a second scanning order.

[0156] In embodiments of the present disclosure, any steps and / or operations may be combined or arranged in any quantity or order as necessary. Two or more of the steps and / or operations may be performed in parallel.

[0157] Embodiments of the present disclosure may be used separately or combined in any order. Further, each of the methods (or embodiments), encoders, and decoders may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium. Embodiments of the present disclosure may be applied to luma blocks or chroma blocks.

[0158] The techniques described above may be implemented as computer software using computer-readable instructions and may be physically stored on one or more computer-readable media. For example, FIG. 17 shows a computer system (1700) suitable for implementing a particular embodiment of the subject matter of the present disclosure.

[0159] Computer software can be coded using any suitable machine code or computer language that can be subjected to assembly, compilation, linking, or similar mechanisms to generate code containing instructions that can be executed directly by one or more computer central processing units (CPUs) and graphics processing units (GPUs), or through interpretation and execution of microcode.

[0160] The instructions may be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.

[0161] The components shown in FIG. 17 with respect to the computer system (1700) are essentially illustrative and are not intended to imply any limitation regarding the use or functionality scope of the computer software implementing the embodiments of the present disclosure. The configuration of the components should not be construed as having any dependency or requirement regarding any one or combination of the components shown in the exemplary embodiment of the computer system (1700).

[0162] The computer system (1700) may include specific human interface input devices. Such human interface input devices may respond to input by one or more human users via, for example, tactile input (e.g., keystrokes, swipes, movement of a data glove), audio input (e.g., voice, clapping), visual input (e.g., gesture), olfactory input (not shown). The human interface device may also be used to capture specific media that is not necessarily directly related to conscious input by humans, such as audio (voice, music, ambient sound, etc.), images (scanned images, photographic images obtained from a still image camera, etc.), video (2D video, 3D video including stereoscopic video, etc.).

[0163] The input human interface device may include one or more (only one of each) of a keyboard (1701), a mouse (1702), a trackpad (1703), a touch screen (1710), a data glove (not shown), a joystick (1705), a microphone (1706), a scanner (1707), and a camera (1708).

[0164] The computer system (1700) may also include certain human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, by tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., a touch screen (1710), a data glove (not shown), or a joystick (1705) may include tactile feedback, but there may also be a tactile feedback device that does not function as an input device), audio output devices (such as speakers (1709), headphones (not shown), etc.), visual output devices (regardless of whether each has a touch screen input function and regardless of whether each has a tactile feedback function, such as a screen (1710) including a CRT screen, an LCD screen, a plasma screen, an OLED screen, etc., some of which can output two-dimensional vision or three-dimensional or higher-dimensional output through means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and a printer (not shown).

[0165] The computer system (1700) can also include human-accessible storage devices and their associated media, such as an optical medium including a CD / DVD ROM / RW (1720) having a medium (1721) such as a CD / DVD, a thumb drive (1722), a removable hard drive or solid state drive (1723), a legacy magnetic medium such as a tape or floppy disk (not shown), a dedicated ROM / ASIC / PLD-based device such as a security dongle (not shown), etc.

[0166] One of ordinary skill in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include a transmission medium, a carrier wave, or other transient signals.

[0167] The computer system (1700) can also include an interface (1754) to one or more communication networks (1755). The network can be, for example, wireless, wired, optical. The network can further be local, wide area, metropolitan, vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicle and industrial ones including CAN bus, etc. A particular network typically requires an external network interface adapter attached to a particular general-purpose data port or peripheral bus (1749) (such as a USB port of the computer system (1700)), and other networks are typically integrated into the core of the computer system (1700) by attaching to the system bus described below (such as an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (1700) can communicate with other entities. Such communication can be only unidirectional reception (such as broadcast TV), only unidirectional transmission (such as from a CANbus to a particular CANbus device), or bidirectional, for example, communication with other computer systems using a local area digital network or a wide area digital network. Specific protocols and protocol stacks can be used for each of those networks and network interfaces as described above.

[0168] The aforementioned human interface device, human-accessible storage device, and network interface can be attached to the core (1740) of the computer system (1700).

[0169] The core (1740) can include one or more central processing units (CPUs) (1741), a graphics processing unit (GPU) (1742), a dedicated programmable processing device in the form of a field programmable gate array (FPGA) (1743), a hardware accelerator for specific tasks (1744), a graphics adapter (1750), etc. These devices can be connected via a system bus (1748) together with a read-only memory (ROM) (1745), a random access memory (1746), an internal hard drive that cannot be accessed by the user, an internal mass storage device such as an SSD (1747). In some computer systems, the system bus (1748) can be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the system bus (1748) of the core or via a peripheral bus (1749). In one example, a screen (1710) can be connected to a graphics adapter (1750). The architecture of the peripheral bus includes PCI, USB, etc.

[0170] The CPU (1741), GPU (1742), FPGA (1743), and accelerator (1744) can execute specific instructions that can be combined to form the above computer code. The computer code can be stored in the ROM (1745) or RAM (1746). Also, transfer data can be stored in the RAM (1746), and persistent data can be stored, for example, in the internal mass storage device (1747). The use of cache memory that can be closely associated with one or more CPUs (1741), GPUs (1742), mass storage devices (1747), ROM (1745), RAM (1746), etc. can enable fast storage and retrieval to any of the memory devices.

[0171] A computer-readable medium can have computer code for performing various computer-implemented operations. The medium and the computer code can be specially designed and constructed for the purposes of this disclosure, or they can be of the kind well-known and available to persons having skill in the art of computer software.

[0172] As a non-limiting example, a computer system (1700) having an architecture, and in particular a core (1740), can provide functionality as a result of one or more processors (including, e.g., a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be associated with media other than the mass storage devices accessible to the user as introduced above, and can also be associated with specific storage devices of the core (1740) of a non-transitory nature, such as the core internal mass storage device (1747) or ROM (1745). Software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core (1740). The computer-readable media can include one or more memory devices or chips, depending on the specific requirements. The software can cause the core (1740), and in particular the processors (including, e.g., a CPU, GPU, FPGA, etc.) within the core (1740), to perform specific processes or specific portions of specific processes, including determining data structures stored in the RAM (1746) as described herein and modifying such data structures according to processes defined by the software. Additionally, or alternatively, the computer system can provide functionality as a result of circuitry (e.g., an accelerator (1744)) wired or otherwise embodied to operate instead of or in conjunction with software to perform the specific processes or specific portions of specific processes described herein. References to software can, where appropriate, include logic, and vice versa. Where necessary, references to computer-readable media can include circuitry (such as an integrated circuit (IC)) storing software for execution, circuitry embodying logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0173] Although the present disclosure has described several exemplary embodiments, there are modifications, substitutions, and various alternative equivalents within the scope of the present disclosure. Accordingly, it will be understood by those skilled in the art that many systems and methods can be devised that embody the principles of the present disclosure and are thus within its spirit and scope, even though not explicitly shown or described herein.

[0174] Appendix A: Acronyms JEM: Joint Exploration Model VVC: Versatile Video Coding BMS: Benchmark Set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplementary Enhancement Information VUI: Video User Interface Information GOP: Group of Pictures TU: Transform Unit PU: Prediction Unit CTU: Coding Tree Unit CTB: Coding Tree Block PB: Prediction Block HRD: Hypothetical Reference Decoder SNR: Signal-to-Noise Ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: Cathode Ray Tube LCD: Liquid Crystal Display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile Communications LTE: Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Array SSD: Solid State Drive IC: Integrated Circuit HDR: High Dynamic Range SDR: Standard Dynamic Range JVET: Joint Video Exploration Team MPM: Most Probable Mode WAIP: Wide Angle Intra Prediction CU: Coding Unit PU: Prediction Unit TU: Transform Unit CTU: Coding Tree Unit PDPC: Position Dependent Prediction Combination ISP: Intra Sub - partition SPS: Sequence Parameter Set PPS: Picture Parameter Set APS: Adaptive Parameter Set VPS: Video Parameter Set DPS: Decoding Parameter Set ALF: Adaptive Loop Filter SAO: Sample Adaptive Offset CC - ALF: Cross - Component Adaptive Loop Filter CDEF: Constrained Directional Enhancement Filter CCSO: Cross - Component Sample Offset LSO: Local Sample Offset LR: Loop Restoration Filter AV1: AOMedia Video 1 AV2: AOMedia Video 2

Explanation of Symbols

[0175] 101 Samples 102 Arrow 103 Arrow 104 Square blocks 201 Blocks 300 Communication system 310 Terminal device 320 Terminal device 330 Terminal device 340 Terminal device 350 Communication network 400 Communication system 401 Video source 402 Video picture or image stream 403 Video encoder 404 Encoded video data, encoded video bitstream 405 Streaming server 406 Client subsystem 407 Input copy 408 Client subsystem 409 Copy 410 Video decoder 411 Output stream of video picture 412 Display 413 Video capture subsystem 420 Electronic device 430 Electronic device 501 Channel 510 Video decoder 512 Display, rendering device 515 Buffer memory 520 Entropy decoder / parser 521 Symbol 530 Electronic device 531 Receiver 551 Scaler / inverse transform unit 552 Intra-picture prediction unit 553 Motion compensation prediction unit 555 Aggregator 556 Loop filter unit 557 Reference picture memory 558 Picture buffer 601 Video source 603 Video encoder, video coder 620 Electronic device, encoder 630 Source coder 632 Coding engine 633 Local video decoder, decoding unit 634 Reference picture memory, reference picture cache 635 Predictor 640 Transmitter 643 Coded video sequence 645 Entropy coder 650 Controller 660 Communication channel 703 Video encoder 721 General-purpose controller 722 Intra encoder 723 Residual calculator 724 Residual encoder 725 Entropy encoder 726 Switch 728 Residual decoder 730 Inter encoder 810 Video decoder 871 Entropy decoder 872 Intra decoder 873 Residual decoder 874 Reconstruction module 880 Inter decoder 1002 4×4 block 1004 Reference sample 1006 Adjacent sample, reference sample 1008 Adjacent sample, reference sample 1010 Sample 1102 8×8 block 1202 Coded block 1204 Intermediate-level quadtree partitioning 1302 Block 1304 Quarter-tree sub-block 1402 Forward first transformation 1404 First transformation coefficient matrix 1406 Diagonal part, first coefficient 1408 Second transformation coefficient 1500 Data flow 1502 Forward second transformation, inseparable second transformation, inseparable transformation processing 1504 Inverse second transformation 1506 First transformation coefficient data block 1507 Modified first transformation coefficient data block, output 1508 First coefficient data block 1509 Modified first coefficient data block 1510 Data sequence 1512 Scanning process 1514 Second transformation coefficient sequence 1516 Input first transformation coefficient 1517 Input coefficient 1518 Data scanning 1520 Second transformation coefficient sequence 1522 Scanning process 1524 First transformation coefficient sequence 1526 First transformation coefficient 1528 First coefficient data block 1530 Reverse scanning process 1540 Quantization / entropy processing, quantization and entropy encoding 1550 Entropy decoding / inverse quantization processing, entropy decoding and inverse quantization 1700 Computer system 1701 Keyboard 1702 Mouse 1703 Track pad 1705 Joystick 1706 Microphone 1707 Scanner 1708 Camera 1709 Audio output device speaker 1710 Touch screen 1720 CD / DVD ROM / RW 1721 Media such as CD / DVD 1722 Thumb drive 1723 Removable hard drive or solid state drive 1740 Core 1741 Central Processing Unit (CPU) 1742 Graphics Processing Unit (GPU) 1743 Field Programmable Gate Array (FPGA) 1744 Hardware accelerator 1745 Read Only Memory (ROM) 1746 Random Access Memory 1747 Large-capacity internal storage device of the core 1748 System bus 1749 Peripheral bus 1750 Graphics adapter 1754 Interface 1755 Communication network S1 First scanning order S2 Second scanning order

Claims

Claim 1 A method for processing video data, comprising: extracting a data block from the video data; scanning a first number of data items within the data block according to a first scanning order to generate a first data sequence; performing a non-separable transform with a non-separable second-order transform that is not separable in the forward direction on the first data sequence to obtain a second data sequence having a second number of data items; replacing at least a part of the first number of data items within the data block with a part or all of the second data sequence according to a second scanning order different from the first scanning order; performing quantization / entropy processing on the data block in which at least a part of the first number of data items has been replaced with a part or all of the second data sequence according to the second scanning order; wherein the first scanning order is determined based on at least one of an intra prediction mode associated with the data block, a type of the non-separable second-order transform in the forward direction, or a size of the data block; the second scanning order is determined based on at least one of an intra prediction mode associated with the data block, a type of the non-separable second-order transform in the forward direction, or a size of the data block; and when the data block is a non-square block, the first scanning order and / or the second scanning order is a diagonal scanning order. Claim 2 wherein the data block has primary transform coefficients, and the step of replacing at least a part of the first number of data items within the data block includes replacing a second number of data items within the first number of data items within the data block with the second data sequence according to the second scanning order. The method according to claim 1. Claim 3 wherein the first scanning order or the second scanning order includes a zigzag scanning order, a diagonal scanning order, or a row and column scanning order. Claim 4 When the same set of transform kernels is shared by a plurality of intra prediction modes with respect to the video data, the first scanning order is determined based on a first intra prediction mode among the plurality of intra prediction modes associated with the data block, and a second scanning order different from the first scanning order is determined based on a second intra prediction mode different from the first intra prediction mode associated with the data block. The method according to claim 2.

5. The method according to claim 2, wherein the first number and the second number are integers between 0 and 127 inclusive.

6. The first scanning order or the second scanning order is a zigzag scanning order when the data block is a square block, or a diagonal scanning order when the data block is a non-square block, and the second scanning order is different from the first scanning order. The method according to claim 2.

7. The data block comprises residual samples, The step of replacing at least a part of the data items of the first number in the data block includes replacing the data items of the second number in the data items of the first number in the data block with the second data sequence according to the second scanning order. The method according to claim 1.

8. The first scanning order or the second scanning order is a zigzag scanning order, a diagonal scanning order, or comprises one of a row and column scanning order. The method according to claim 6.

9. When the same set of transform kernels is shared by a plurality of intra prediction modes with respect to the video data, the first scanning order is determined based on a first intra prediction mode among the plurality of intra prediction modes associated with the data block, and a second scanning order different from the first scanning order is determined based on a second intra prediction mode different from the first intra prediction mode associated with the data block. The method according to claim 6.

10. The method according to claim 6, wherein the first number and the second number are integers between 0 and 127 inclusive.

11. The first scanning order or the second scanning order is a zigzag scanning order when the data block is a square block, or The method according to claim 6, wherein the diagonal scanning order is when the data block is a non-square block, and the second scanning order is different from the first scanning order.

12. The method according to claim 6, wherein the first scanning order is a horizontal or vertical scanning order.

13. A method for entropy encoding a transform coefficient associated with video data, comprising: In response to the transform associated with the transform coefficient being a non-separable transform comprising a non-separable forward quadratic transform, Scanning the transform coefficient using a first scanning order that is one of a horizontal scanning order or a vertical scanning order; and In response to the transform associated with the transform coefficient being separable, scanning the transform coefficient in a second scanning order different from the first scanning order when performing entropy encoding of the transform coefficient. comprising wherein the first scanning order is determined based on at least one of an intra prediction mode associated with a data block included in the video data, the type of the non-separable forward quadratic transform, or the size of the data block; wherein the second scanning order is determined based on at least one of an intra prediction mode associated with the data block, the type of the non-separable forward quadratic transform, or the size of the data block; wherein when the data block is a non-square block, the first scanning order and / or the second scanning order is a diagonal scanning order.

14. The method according to claim 13, wherein the transform comprises one of a primary transform or a secondary transform.

15. A method for processing video data, comprising: receiving the video data; determining whether a non-separable transform comprising a non-separable forward quadratic transform is applied to the video data as a secondary transform; in response to the non-separable transform being applied to the video data as the secondary transform, scanning a first number of primary transform coefficients, wherein the primary transform coefficients follow a first scanning order. Performing a non-separable transform using the first-order transform coefficients of the first number as input to obtain second-order transform coefficients of a second number as output, wherein the second-order transform coefficients follow a second scanning order different from the first scanning order; Replacing at least the first-order transform coefficients of the second number with the second-order transform coefficients following the second scanning order; Performing an inverse second-order transform corresponding to the non-separable transform using the second-order transform coefficients of the second number as input to obtain first-order transform coefficients of the first number as output; Replacing at least the second-order transform coefficients of the first number with the first-order transform coefficients following the first scanning order; comprising; wherein the first scanning order is determined based on at least one of an intra prediction mode associated with a data block included in the video data, the type of the forward non-separable second-order transform, or the size of the data block; wherein the second scanning order is determined based on at least one of an intra prediction mode associated with the data block, the type of the forward non-separable second-order transform, or the size of the data block; wherein when the data block is a non-square block, the first scanning order and / or the second scanning order is a diagonal scanning order.

16. A device comprising a circuit configured to perform the method according to any one of claims 1 to 15.

17. A computer program comprising computer code which, when executed by one or more processors, causes the one or more processors to perform the method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Context adaptive entropy coding for non-square blocks in video coding

    US20130064294A1

  • One-dimensional transform modes and coefficient scan order

    US20170280163A1

  • Non-separable secondary transform for video coding

    US20200092583A1

  • Method and apparatus for video coding

    US20210160519A1