Orthogonal transform generation with subspace constraints

Transform kernel sharing in video encoding and decoding optimizes intra-prediction and motion compensation by sharing high-frequency basis vectors across kernels, enhancing compression efficiency and reducing redundancy in video coding.

JP7761348B2Active Publication Date: 2025-10-28TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024034075
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-01-05
Filing Date
2024-03-06
Publication Date
2025-10-28
Estimated Expiration
2042-01-28

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently reducing redundancy and achieving high compression ratios while maintaining acceptable video quality, particularly in intra-prediction and motion compensation processes.

Method used

Implementing transform kernel sharing methods in video encoding and decoding, where multiple transform kernels share high-frequency basis vectors while individualizing low-frequency basis vectors, tailored for specific intra-picture prediction modes.

Benefits of technology

Enhances compression efficiency by reducing bit usage for transform coefficients, thereby improving video coding performance and reducing storage and bandwidth requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007761348000012
    Figure 0007761348000012
  • Figure 0007761348000013
    Figure 0007761348000013
  • Figure 0007761348000014
    Figure 0007761348000014
Patent Text Reader

Abstract

To disclose a transform kernel sharing in video encoding and decoding.SOLUTION: This disclosure relates to a transform kernel sharing in video encoding and decoding. For example, a method is disclosed for such transform kernel sharing. The method may include a step of identifying a plurality of transform kernels, wherein each of the plurality of transform kernels comprises a set of basis vectors from low to high frequencies; N high-frequency basis vectors of two or more of the plurality of transform kernels are shared, N being a positive integer; and low-frequency basis vectors of the two or more of the plurality of the transform kernels other than the N high-frequency basis vectors are individualized. The method may further include: extracting a data block from a video bitstream; selecting a transform kernel from the plurality of transform kernels on the basis of information associated with the data block; and applying the transform kernel to at least a portion of the data block to generate a transformed block.SELECTED DRAWING: Figure 15
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is based on and claims the benefit of priority to U.S. Non-provisional Patent Application No. 17 / 568,871, filed January 5, 2022, which claims priority to U.S. Provisional Patent Application No. 63 / 172,060, entitled "ORTHOGONAL TRANSFORM GENERATION WITH SUBSPACE CONSTRAINT," filed April 7, 2021. Both applications are incorporated herein by reference in their entireties.

[0002] This disclosure describes a set of advanced video coding techniques. More specifically, the disclosed techniques include transform kernel sharing methods in video encoding and video decoding. [Background technology]

[0003] The discussion of the background art provided herein is intended to generally present the context for the present disclosure. The inventors' work is not admitted expressly or implicitly as prior art to the present disclosure to the extent that that work is described in this background section, along with aspects of the description that may not otherwise be admitted as prior art at the time of filing of this application.

[0004] Video coding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video can include a series of pictures, each with spatial dimensions of, for example, 1920 x 1080 luma samples and associated fully sampled or subsampled chroma samples. The series of pictures can have a fixed or variable picture rate (also called frame rate), for example, 60 pictures per second or 60 frames per second. Uncompressed video has specific bitrate requirements for streaming or data processing. For example, a video with a pixel resolution of 1920 x 1080, a frame rate of 60 frames per second, and 4:2:0 chroma subsampling with 8 bits per pixel per color channel requires a bandwidth approaching 1.5 Gbit / s. One hour of such video requires more than 600 Gbytes of storage space.

[0005] One goal of video coding and video decoding is to reduce redundancy in an uncompressed input video signal through compression. Compression can help reduce the aforementioned bandwidth and / or storage space requirements by more than two orders of magnitude, in some cases. Both lossless and lossy compression, as well as combinations thereof, can be used. Lossless compression refers to techniques in which an exact copy of the original signal can be reconstructed from the compressed original signal through the decoding process. Lossy compression refers to a coding / decoding process in which the original video information is not fully preserved during coding and cannot be fully recovered during decoding. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signal is small enough to make the reconstructed signal useful for its intended purpose, even with some information loss. For video, lossy compression is widely adopted in many applications. The amount of tolerable distortion depends on the application. For example, users of certain consumer video streaming applications may tolerate higher distortion than users of film or television broadcast applications. The compression ratio achievable by a particular coding algorithm can be selected or adjusted to reflect different distortion tolerances. That is, generally, higher distortion tolerance allows for coding algorithms that result in higher losses and higher compression ratios.

[0006] Video encoders and decoders can utilize techniques from several broad categories and steps, including, for example, motion compensation, Fourier transforms, quantization, and entropy coding.

[0007] Video codec technology can include a technique known as intra-coding. In intra-coding, sample values ​​are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, pictures are spatially subdivided into blocks of samples. If all blocks of samples are coded in intra mode, the picture can be called an intra-picture. Intra-pictures and their derivatives, such as independent decoder refresh pictures, can be used to reset the decoder state and can therefore be used as the first picture in a coded video bitstream and video session, or as still images. The samples of the intra-predicted block can then be transformed into the frequency domain, and the transform coefficients thus generated can be quantized before entropy coding. Intra-prediction refers to a technique that minimizes sample values ​​in the pre-transform domain. In some cases, the smaller the DC value and the smaller the AC coefficients after the transform, the fewer bits are required at a given quantization step size to represent the block after entropy coding.

[0008] For example, traditional intra-coding, such as that known from MPEG-2 generation coding techniques, does not use intra-prediction. However, some newer video compression techniques include techniques that attempt to code / decode blocks based on surrounding sample data and / or metadata that precede the block of data being intra-coded or intra-decoded in decoding order, for example, obtained during the encoding and / or decoding of spatial neighbors. Such techniques are hereinafter referred to as "intra-prediction" techniques. It should be noted that, at least in some cases, intra-prediction uses reference data only from the current picture being reconstructed, and not from other reference pictures.

[0009] Intra-prediction may take many different forms. When two or more such techniques are available in a given video coding technology, the technique used may be referred to as an intra-prediction mode. One or more intra-prediction modes may be provided in a particular codec. In certain cases, a mode may have sub-modes and / or may be associated with various parameters, and the mode / sub-mode information and the intra-coding parameters of a block of video may be coded separately or collectively included in the mode's codeword. The codeword used for a given mode, sub-mode, and / or parameter combination may affect coding efficiency gains via intra-prediction, and therefore may also affect the entropy coding technique used to convert the codeword into a bitstream.

[0010] Certain modes of intra prediction were introduced in H.264, refined in H.265, and further refined in newer coding techniques such as joint search model (JEM), versatile video coding (VVC), and benchmark set (BMS). In general, intra prediction can use available neighboring sample values ​​to form a predictor block. For example, available values ​​of a particular neighboring sample set along a particular direction and / or line can be copied into the predictor block. A reference to the direction in use can be coded in the bitstream or can itself be predicted.

[0011] Referring to FIG. 1A, shown at the bottom right is a subset of nine predictor directions specified by the 33 possible predictor directions of H.265 (corresponding to the 33 angle modes of the 35 intra modes specified in H.265). The point where the arrows converge (101) represents the sample being predicted. The arrows represent the direction from which neighboring samples are used to predict sample 101. For example, arrow (102) indicates that sample 101 is predicted from one or more neighboring samples to the upper right, at an angle of 45 degrees from horizontal. Similarly, arrow (103) indicates that sample 101 is predicted from one or more neighboring samples to the lower left of sample 101, at an angle of 22.5 degrees from horizontal.

[0012] 1A, a square block (104) of 4x4 samples (indicated by a thick dashed line) is depicted in the upper left. The square block (104) contains 16 samples, each labeled with "S," its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample (from the top) in the Y dimension and the first sample (from the left) in the X dimension. Similarly, sample S44 is the fourth sample in both the Y and X dimensions within the block (104). Because the size of the block is 4x4 samples, S44 is located in the lower right. Also shown are examples of reference samples that follow a similar numbering scheme. The reference samples are labeled R, their Y position (e.g., row number) and X position (column number) relative to the block (104). Both H.264 and H.265 use predicted samples that neighbor the block being reconstructed.

[0013] Intra-picture prediction of block 104 may begin by copying reference sample values ​​from neighboring samples according to a signaled prediction direction. For example, assume that the coded video bitstream includes signaling for this block 104 indicating the prediction direction of the arrow (102), i.e., the sample is predicted from one or more prediction samples to the upper right, at a 45-degree angle from the horizontal. In such a case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.

[0014] In certain cases, the values ​​of multiple reference samples may be combined, for example by interpolation, to calculate a reference sample, especially when the direction is not evenly divisible by 45 degrees.

[0015] The number of possible directions has increased as video coding technology continues to develop. In H.264 (2003), for example, nine different directions are available for intra prediction. This increased to 33 in H.265 (2013), and JEM / VVC / BMS can support up to 65 directions as of the time of this disclosure. Experimental studies have been conducted to help identify the most appropriate intra prediction directions, and these most appropriate directions can be encoded with fewer bits using specific techniques of entropy coding, accepting a specific bit penalty for the direction. Furthermore, the direction itself may be predictable from neighboring directions used in the intra prediction of neighboring decoded blocks.

[0016] FIG. 1B shows a schematic diagram (180) showing 65 intra-prediction directions according to JEM to illustrate the increasing number of prediction directions in various encoding techniques that have evolved over time.

[0017] The mapping of bits representing intra-prediction directions to prediction directions in a coded video bitstream can vary between video coding techniques, ranging from simple direct mappings of prediction directions to intra-prediction modes to complex adaptation schemes involving codewords, most-likely modes, and similar techniques. In all cases, however, there may be certain directions of intra-prediction that are statistically less likely to occur in the video content than certain other directions. Because the goal of video compression is to reduce redundancy, in a well-designed video coding technique, less-likely directions are represented by more bits than more-likely directions.

[0018] Inter-picture prediction, or inter-prediction, may be based on motion compensation. In motion compensation, sample data from a previously reconstructed picture or part thereof (reference picture) may be used to predict a newly reconstructed picture or picture part (e.g., block) after being spatially shifted in a direction indicated by a motion vector (hereafter MV). In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions, X and Y, or three dimensions, with the third dimension being an indication of the reference picture to be used (similar to the temporal dimension).

[0019] In some video compression techniques, a current MV applicable to a particular area of ​​sample data can be predicted from other MVs, e.g., from other MVs related to other areas of sample data that are spatially adjacent to the area being reconstructed and precede the current MV in decoding order. Doing so can significantly reduce the overall amount of data required to code the MV by relying on the removal of redundancy in correlated MVs, thereby increasing compression efficiency. MV prediction can work effectively because, for example, when coding an input video signal derived from a camera (known as natural video), areas larger than the area to which a single MV is applicable have a statistical likelihood of moving in a similar direction in the video sequence and therefore, in some cases, can be predicted using similar motion vectors derived from MVs in neighboring areas. As a result, the actual MV of a given area is similar or identical to the MV predicted from the surrounding MVs. Such MVs can further be represented, after entropy coding, with fewer bits than would be used if the MV were coded directly rather than predicted from one or more neighboring MVs. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., an MV) derived from an original signal (i.e., a sample stream). In other cases, the MV prediction itself may be lossy, for example due to rounding errors when computing the predictor from several surrounding MVs.

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). Among the many MV prediction mechanisms specified by H.265, the one described below is a technique hereafter referred to as "spatial merging".

[0021] Specifically, referring to Figure 2, a current block (201) contains samples that the encoder detected during the motion search process as being predictable from a spatially shifted previous block of the same size. Instead of directly coding its MV, the MV can be derived from metadata associated with one or more reference pictures, e.g., from the last reference picture (in decoding order), using the MV associated with any one of five surrounding samples represented by A0, A1, and B0, B1, B2 (202 to 206, respectively). In H.265, MV prediction can use predictors from the same reference picture as neighboring blocks. Summary of the Invention [Means for solving the problem]

[0022] Aspects of the present disclosure provide methods and apparatus for transform kernel sharing in video encoding and video decoding.

[0023] In some implementations, a method for such transform kernel sharing is disclosed. The method may include identifying a plurality of transform kernels, each of the plurality of transform kernels including a set of basis vectors ranging from low frequency to high frequency, wherein N high-frequency basis vectors of two or more of the plurality of transform kernels are shared, where N is a positive integer, and two or more low-frequency basis vectors of the plurality of transform kernels other than the N high-frequency basis vectors are individualized. The method may further include extracting a data block from a video bitstream, selecting a transform kernel from the plurality of transform kernels based on information associated with the data block, and applying the transform kernel to at least a portion of the data block to generate a transform block.

[0024] In the above implementations, the multiple transformation kernels may be pre-trained offline. In some further implementations, the multiple transformation kernels may be jointly trained offline to determine the N shared high-frequency basis vectors.

[0025] In any of the above implementations, the N high-frequency basis vectors may include basis vectors having frequencies higher than a predetermined threshold frequency. In some further implementations, the multiple transformation kernels and the predetermined threshold frequency may be pre-trained offline.

[0026] In any of the above implementations, the multiple transform kernels can include secondary transform kernels applicable to transform primary transform coefficients, and the data block includes an array of primary transform coefficients.

[0027] In any of the above implementations, two or more of the multiple transform kernels that share the N high-frequency basis vectors correspond to the same one of the multiple intra-picture prediction modes, and all of the multiple transform kernels assigned to the same one of the multiple intra-picture prediction modes share the N high-frequency basis vectors while the other low-frequency basis vectors are individualized.

[0028] In any of the above implementations, two or more of the multiple transform kernels that share the N high-frequency basis vectors may be assigned to two or more different intra-picture prediction modes.

[0029] In any of the above implementations, the multiple transform kernels are secondary transform kernels, and two or more of the multiple transform kernels that share the N high-frequency basis vectors are configured to transform primary transform coefficients generated from a primary transform having the same transform type.

[0030] In any of the above implementations, N is an integer power of 2, and / or N is less than a predefined upper bound, and / or N depends on the transform size.

[0031] In any of the above implementations, the multiple transform kernels may be configured for an intra-quadratic transform (IST). In some implementations, the multiple transform kernels may be quadratic low-frequency non-separable transform (LFNST) kernels. In some implementations, the multiple transform kernels are configured for a line graph transform (LGT).

[0032] In some other implementations, a device for processing video information is disclosed. The device can include circuitry configured to perform any one of the above method implementations.

[0033] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform a method for video decoding and / or video encoding.

[0034] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Brief explanation of the drawings]

[0035] [Figure 1A] FIG. 10 is a schematic diagram of an example subset of intra-prediction directional modes. [Figure 1B] FIG. 2 illustrates exemplary intra-prediction directions. [Figure 2] FIG. 1 is a schematic diagram illustrating a current block and its surrounding spatial merge candidates for motion vector prediction in one example. [Figure 3] FIG. 3 is a schematic diagram illustrating a simplified block diagram of a communication system (300) according to an example embodiment. [Figure 4]FIG. 4 is a schematic diagram illustrating a simplified block diagram of a communication system (400) according to an example embodiment. [Figure 5] FIG. 2 is a schematic diagram illustrating a simplified block diagram of a video decoder according to an example embodiment. [Figure 6] FIG. 1 is a schematic diagram illustrating a simplified block diagram of a video encoder according to an example embodiment. [Figure 7] FIG. 10 is a block diagram illustrating a video encoder according to another example embodiment. [Figure 8] FIG. 10 is a block diagram illustrating a video decoder according to another example embodiment. [Figure 9] FIG. 10 illustrates directional intra-prediction modes according to an example embodiment of the present disclosure. [Figure 10] FIG. 10 illustrates a non-directional intra-prediction mode according to an example embodiment of the present disclosure. [Figure 11] FIG. 10 illustrates a recursive intra-prediction mode according to an example embodiment of the present disclosure. [Figure 12] FIG. 10 illustrates transform block partitioning and scanning of intra-predicted blocks according to an example embodiment of this disclosure. [Figure 13] FIG. 10 illustrates transform block partitioning and scanning of inter-predicted blocks according to an example embodiment of this disclosure. [Figure 14] FIG. 1 illustrates a low frequency non-separable transformation process according to an example embodiment of the present disclosure. [Figure 15] FIG. 10 illustrates sharing of high frequency basis vectors between different transform kernels according to an example embodiment of the present disclosure. [Figure 16] FIG. 1 illustrates a flowchart according to an example embodiment of the present disclosure. [Figure 17] FIG. 1 is a schematic diagram of a computer system according to an example embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0036] Figure 3 shows a simplified block diagram of a communication system (300) according to one embodiment of the present disclosure. The communication system (300) includes, for example, multiple terminal devices that can communicate with each other via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). In the example of Figure 3, the first pair of terminal devices (310) and (320) may perform unidirectional transmission of data. For example, the terminal device (310) may code video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to the other terminal device (320) via the network (350). The encoded video data may be transmitted in the form of one or more coded video bitstreams. The terminal device (320) may receive the coded video data from the network (350), decode the coded video data to reconstruct the video pictures, and display the video pictures according to the reconstructed video data. One-way data transmission may be implemented for media serving applications, etc.

[0037] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) performing bidirectional transmission of coded video data, such as may be performed during video conferencing applications. For the bidirectional transmission of data, in one example, each of the terminal devices (330) and (340) may code video data (e.g., a stream of video pictures captured by that terminal device) for transmission to the other of the terminal devices (330) and (340) over the network (350). Each of the terminal devices (330) and (340) may also receive coded video data transmitted by the other of the terminal devices (330) and (340), decode the coded video data to recover the video pictures, and display the video pictures on an accessible display device according to the recovered video data.

[0038] In the example of FIG. 3 , the terminal devices 310, 320, 330, and 340 may be embodied as a server, a personal computer, and a smartphone, although the applicability of the principles underlying the present disclosure is not so limited. Embodiments of the present disclosure may be implemented in desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated videoconferencing equipment, and the like. The network 350 represents any number or type of network that conveys coded video data between the terminal devices 310, 320, 330, and 340, including, for example, wired (wired) and / or wireless communication networks. The communication network 350 may exchange data over circuit-switched channels, packet-switched channels, and / or other types of channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this discussion, the architecture and topology of the network 350 may not be important to the operation of the present disclosure unless explicitly described herein.

[0039] 4 illustrates the arrangement of a video encoder and a video decoder in a video streaming environment as an example of an application of the disclosed subject matter. The disclosed subject matter may be equally applied to other video-enabled applications including, for example, video conferencing, digital television broadcasting, gaming, virtual reality, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0040] A video streaming system may include a video source (401), such as a video capture subsystem (413), which may include a digital camera, for creating a stream of uncompressed video pictures or images (402). In one example, the stream of video pictures (402) includes samples recorded by the digital camera of the video source 401. The stream of video pictures (402), depicted in bold to emphasize its high data volume compared to the encoded video data (404) (or coded video bitstream), may be processed by an electronic device (420) including a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (404) (or encoded video bitstream (404)), depicted with thin lines to emphasize its low data volume compared to the uncompressed stream of video pictures (402), can be stored directly on the streaming server (405) or on a downstream video device (not shown) for future use. One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to obtain copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include a video decoder (410), for example, within the electronic device (430). The video decoder (410) decodes the input copy of the encoded video data (407) and creates an output stream of video pictures (411) that is uncompressed and can be rendered on a display (412) (e.g., a display screen) or other rendering device (not shown). Video decoder 410 may be configured to perform some or all of the various functions described in this disclosure.In some streaming systems, the encoded video data (404), (407), and (409) (e.g., video bitstreams) may be encoded according to a particular video coding / compression standard. Examples of such standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC, as well as other video coding standards.

[0041] It should be noted that the electronic devices (420) and (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).

[0042] 5 shows a block diagram of a video decoder (510) according to any of the following embodiments of the present disclosure. The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., receiving circuitry). The video decoder (510) can be used in place of the video decoder (410) in the example of FIG. 4.

[0043] The receiver (531) may receive one or more coded video sequences to be decoded by the video decoder (510). In the same or another embodiment, one coded video sequence may be decoded at a time, with the decoding of each coded video sequence being independent of other coded video sequences. Each video sequence may be associated with multiple video frames or images. The coded video sequences are received from a channel (501), which may be a hardware / software link to a storage device that stores the encoded video data or a streaming source that transmits the encoded video data. The receiver (531) may receive the encoded video data along with other data, such as coded audio data and / or auxiliary data streams, which may be forwarded to respective processing circuits (not shown). The receiver (531) may separate the coded video sequences from other data. To combat network jitter, a buffer memory (515) may be located between the receiver (531) and the entropy decoder / parser (520) (hereinafter "parser (520)"). In certain applications, the buffer memory (515) may be implemented as part of the video decoder (510). In other applications, the buffer memory (515) may be separate and external to the video decoder (510) (not shown). In still other applications, there may be a buffer memory (not shown) external to the video decoder (510), for example, to combat network jitter, or there may be another additional buffer memory (515) internal to the video decoder (510), for example, to handle playback timing. When the receiver (531) is receiving data from a storage / forwarding device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory (515) may be unnecessary or may be small.For use with best-effort packet networks such as the Internet, a buffer memory (515) of sufficient size may be required, which may be relatively large. Such a buffer memory may be implemented with an adaptive size and may be implemented at least in part in an operating system or similar element (not shown) external to the video decoder (510).

[0044] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the coded video sequence. These symbol categories include information used to manage the operation of the video decoder (510) and, in some cases, information for controlling a rendering device, such as a display (512) (e.g., a display screen), which may or may not be an integral part of the electronic device (530) but may be coupled to the electronic device (530), as shown in FIG. 5. The control information for the rendering device may be in the form of a supplemental enhancement information (SEI) message or a video usability information (VUI) parameter set fragment (not shown). The parser (520) may parse / entropy decode the coded video sequence received by the parser (520). The entropy coding of the coded video sequence may be in accordance with a video coding technique or standard and may be in accordance with various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) may extract, from the coded video sequence, a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the subgroup. The subgroups may include Groups of Pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (520) may also extract information from the coded video sequence, such as transform coefficients (e.g., Fourier transform coefficients), quantization parameter values, motion vectors, etc.

[0045] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to produce symbols (521).

[0046] The reconstruction of the symbols (521) may involve several different processing or functional units, depending on the type of coded video picture or portion thereof (inter-picture and intra-picture, inter-block and intra-block, etc.) and other factors. The units included and how they are included may be controlled by subgroup control information parsed from the coded video sequence by the parser (520). The flow of such subgroup control information between the parser (520) and the following several processing or functional units is not shown for simplicity.

[0047] In addition to the functional blocks already mentioned, the video decoder (510) may be conceptually subdivided into several functional units, as described below. In an actual implementation operating under commercial constraints, many of these functional units will interact closely with each other and may be, at least partially, integrated with each other. However, to clearly describe the various functions of the disclosed subject matter, a conceptual subdivision into functional units will be adopted in the following disclosure.

[0048] The first unit is a scalar / inverse transform unit (551), which may receive quantized transform coefficients as well as control information from the parser (520) as symbols (521), including information indicating which type of inverse transform to use, block size, quantization coefficients / parameters, quantization scaling matrices, etc. The scalar / inverse transform unit (551) may output blocks having sample values ​​that can be input to an aggregator (555).

[0049] In some cases, the output samples of the scaler / inverse transform (551) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) may generate blocks of the same size and shape as the block being reconstructed using information from surrounding blocks that have already been reconstructed and stored in the current picture buffer (558). The current picture buffer (558), for example, buffers the partially reconstructed and / or fully reconstructed current picture. In some implementations, the aggregator (555) may add, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (552) to the output sample information as provided by the scaler / inverse transform unit (551).

[0050] In other cases, the output samples of the scalar / inverse transform unit (551) may relate to an inter-coded and possibly motion-compensated block. In such cases, the motion-compensated prediction unit (553) may access a reference picture memory (557) to fetch samples used for inter-picture prediction. After motion-compensating the fetched samples according to the symbols (521) associated with the block, these samples may be added by an aggregator (555) to the output of the scalar / inverse transform unit (551) (the output of unit 551 may be referred to as a residual sample or residual signal) to generate output sample information. The address in the reference picture memory (557) from which the motion-compensated prediction unit (553) fetches prediction samples may be controlled by a motion vector available to the motion-compensated prediction unit (553), e.g., in the form of a symbol (521) that may have an X component, a Y component (shift), and a reference picture component (time). Motion compensation may also include interpolation of sample values ​​fetched from the reference picture memory (557) when sub-sample accurate motion vectors are used, and may be associated with a motion vector prediction mechanism, etc.

[0051] The output samples of the aggregator (555) may undergo various loop filtering techniques in a loop filter unit (556). Video compression techniques may include in-loop filter techniques controlled by parameters included in the coded video sequence (also called the coded video bitstream) and made available to the loop filter unit (556) as symbols (521) from the parser (520), but may also respond to previously reconstructed, loop-filtered sample values ​​as well as meta-information obtained during decoding of a coded picture or previous portion (in decoding order) of the coded video sequence. As described in more detail below, several types of loop filters may be included as part of the loop filter unit 556, in various orders.

[0052] The output of the loop filter unit (556) may be a sample stream that can be output to a rendering device (512) and also stored in a reference picture memory (557) for use in future inter-picture prediction.

[0053] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future inter-picture prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and any unused current picture buffer can be reallocated before beginning reconstruction of the next coded picture.

[0054] The video decoder (510) may perform decoding operations according to a predetermined video compression technique adopted in a standard such as ITU-T Rec. H.265. A coded video sequence may conform to the syntax specified by the video compression technique or standard being used, in the sense that the coded video sequence adheres to both the syntax of the video compression technique or video compression standard and a profile documented in the video compression technique. Specifically, a profile may select specific tools from among all tools available in the video compression technique or video compression standard as tools available for use only under that profile. To conform to a standard, the complexity of the coded video sequence may be within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further constrained by a hypothetical reference decoder (HRD) specification and HRD buffer management metadata signaled in the coded video sequence.

[0055] In some example embodiments, the receiver (531) may receive additional (redundant) data along with the coded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0056] 6 shows a block diagram of a video encoder (603) according to an example embodiment of the present disclosure. The video encoder (603) may be included in an electronic device (620). The electronic device (620) may further include a transmitter (640) (e.g., a transmitting circuit). The video encoder (603) may be used in place of the video encoder (403) in the example of FIG. 4.

[0057] The video encoder (603) may receive video samples from a video source (601) (which in the example of FIG. 6 is not part of the electronic device (620)) that can capture video images to be coded by the video encoder (603). In another example, the video source (601) may be implemented as part of the electronic device (620).

[0058] The video source (601) may provide a source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream, which may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 YCrCb, RGB, XYZ, ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (601) may be a storage device capable of storing previously prepared video. In a video conferencing system, the video source (601) may be a camera capturing local image information as a video sequence. The video data may be provided as multiple individual pictures or images that impart motion when viewed sequentially. The picture itself may be organized as a spatial array of pixels, each of which may contain one or more samples, depending on the sampling structure, color space, etc., used. Those skilled in the art will readily understand the relationship between pixels and samples. The following discussion focuses on samples.

[0059] According to some example embodiments, the video encoder (603) may code and compress pictures of a source video sequence into a coded video sequence (643) in real time or under any other time constraints required by the application. Enforcing an appropriate coding rate constitutes one function of the controller (650). In some embodiments, the controller (650) may be functionally coupled to and control other functional units as described below. For simplicity, coupling is not shown. Parameters set by the controller (650) may include rate control-related parameters (e.g., picture skip, quantizer, lambda value for rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) may be configured with other appropriate functions associated with the video encoder (603) optimized for a particular system design.

[0060] In some example embodiments, the video encoder (603) may be configured to operate in a coding loop. As an overly simplified explanation, in one example, the coding loop may include a source coder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be coded and a reference picture) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs symbols to create sample data in a manner similar to that which a (remote) decoder would create, even though the embedded decoder 633 processes the video stream coded by the source coder 630 without entropy coding (because in the video compression techniques contemplated in the disclosed subject matter, any compression between the symbols and the coded video bitstream may be lossless). The reconstructed sample stream (sample data) is input to a reference picture memory (634). Because decoding the symbol stream leads to bit-accurate results regardless of the decoder's location (local or remote), the contents in the reference picture memory (634) are also bit-accurate between the local and remote encoders. In other words, the predictive portion of the encoder "sees" the exact same sample values ​​for reference picture samples that the decoder will "see" when using prediction during decoding. This fundamental principle of reference picture synchrony (and the resulting drift if synchrony cannot be maintained, for example, due to channel error) is used to improve coding quality.

[0061] The operation of the "local" decoder (633) may be the same as the operation of a "remote" decoder, such as the video decoder (510), already described in detail above in conjunction with Figure 5. Referring also briefly to Figure 5, however, because symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder (645) and parser (520) may be lossless, the entropy decoding portion of the video decoder (510), including the buffer memory (515) and parser (520), may not be fully implemented in the local decoder (633) within the encoder.

[0062] At this point, it can be said that any decoder technology, except for parsing / entropy decoding, which may only exist in the decoder, may also necessarily need to exist in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter may focus on decoder operation, which is similar to the decoding part of the encoder. Therefore, a description of the encoder technology can be omitted, since it is the opposite of the decoder technology described comprehensively. Only in certain areas or aspects will a more detailed description of the encoder be provided below.

[0063] In operation, in some example implementations, the source coder (630) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as "reference pictures." In this manner, the coding engine (632) codes color channel differences (or residuals) between pixel blocks of the input picture and pixel blocks of reference pictures that may be selected as predictive references for the input picture. The terms "residual" and its adjective form "residual" may be used interchangeably.

[0064] The local video decoder (633) may decode coded video data of pictures that may be designated as reference pictures based on symbols created by the source coder (630). The operation of the coding engine (632) may advantageously be a lossy process. When the coded video data is decoded in a video decoder (not shown in FIG. 6), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (633) may replicate the decoding process that may be performed by the video decoder on the reference pictures and store the reconstructed reference pictures in a reference picture cache (634). In this way, the video encoder (603) may locally store copies of reconstructed reference pictures that have content in common with reconstructed reference pictures obtained (without transmission errors) by a far-end (remote) video decoder.

[0065] The predictor (635) may perform the prediction search for the coding engine (632). That is, for a new picture to be coded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture motion vectors, block shapes, etc., that can serve as appropriate prediction references for the new picture. The predictor (635) may operate on sample blocks, pixel blocks at a time, to find an appropriate prediction reference. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (634).

[0066] The controller (650) may manage the coding operations of the source coder (630), including, for example, setting the parameters and subgroup parameters used to encode the video data.

[0067] The output of all the aforementioned functional units can undergo entropy coding in an entropy coder (645), which converts the symbols produced by the various functional units into a coded video sequence by lossless compression of the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0068] The transmitter (640) may buffer the coded video sequence created by the entropy coder (645) and prepare it for transmission over a communication channel (660), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (640) may merge the coded video data from the video coder (603) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0069] The controller (650) may manage the operation of the video encoder (603). During coding, the controller (650) may assign a particular coded picture type to each coded picture, which may affect the coding technique that may be applied to the respective picture. For example, pictures may often be assigned as any of the following picture types:

[0070] An intra-picture (I-picture) may be a picture that can be coded and decoded without relying on any other picture in a sequence for prediction. Some video codecs allow different types of intra-pictures, including, for example, independent decoder refresh ("IDR") pictures. Those skilled in the art are aware of these variations of I-pictures and their respective uses and characteristics.

[0071] A predictive picture (P picture) may be a picture that can be coded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict the sample values ​​of each block.

[0072] A bidirectionally predicted picture (B picture) may be a picture that can be coded and decoded using intra- or inter-prediction, which uses up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, a multi-predicted picture may use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0073] A source picture is generally spatially subdivided into multiple sample coding blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks as determined by the coding assignment applied to the block's respective picture. For example, blocks of an I-picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P-picture may be predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be predictively coded by spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures. Source pictures or intermediate processed pictures may also be subdivided into other types of blocks for other purposes. The division of coding blocks and other types of blocks may or may not follow the same method, as described in more detail below.

[0074] The video encoder (603) may perform coding operations in accordance with a predetermined video coding technique or standard, such as ITU-T Rec. H.265. In its operations, the video encoder (603) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancies in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.

[0075] In some example embodiments, the transmitter (640) may transmit additional data along with the encoded video. The source coder (630) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures or slices, SEI messages, VUI parameter set fragments, etc.

[0076] Video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits temporal or other correlation between pictures. For example, a particular picture being encoded / decoded, called the current picture, may be partitioned into blocks. If a block in the current picture resembles a reference block in a previously coded and still buffered reference picture in the video, it may be coded by a vector called a motion vector. A motion vector points to a reference block within the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0077] In some exemplary embodiments, bi-prediction techniques can be used for inter-picture prediction. Such bi-prediction techniques use two reference pictures, such as a first reference picture and a second reference picture, both of which follow the current picture in the video in decoding order (but may be past or future, respectively, in display order). A block in the current picture can be coded with a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. A block can be jointly predicted by a combination of the first reference block and the second reference block.

[0078] Additionally, merge mode techniques may be used to improve coding efficiency in inter-picture prediction.

[0079] According to some example embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed on a block-by-block basis. For example, pictures in a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and the CTUs within a picture may have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. In general, a CTU may include three parallel coding tree blocks (CTBs), i.e., one luma CTB and two chroma CTBs. Each CTU may be recursively quadtree-decomposed into one or more coding units (CUs). For example, a 64x64 pixel CTU may be partitioned into one 64x64 pixel CU or four 32x32 pixel CUs. One or more of the 32x32 blocks may each be further partitioned into four 16x16 pixel CUs. In some example embodiments, each CU may be analyzed during encoding to determine the CU's prediction type from various prediction types, such as inter-prediction and intra-prediction. A CU may be divided into one or more prediction units (PUs) according to temporal predictability and / or spatial predictability. Generally, each PU includes one luma prediction block (PB) and two chroma PBs. In one embodiment, prediction operations in coding (encoding / decoding) are performed on a prediction block basis. The division of a CU into PUs (or PBs of different color channels) may be performed in various spatial patterns. A luma PB or a chroma PB may include a matrix of sample values ​​(e.g., luma values), such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.

[0080] 7 shows a diagram of a video encoder (703) according to another example embodiment of this disclosure. The video encoder (703) is configured to receive a processed block (e.g., a predictive block) of sample values ​​in a current video picture in a sequence of video pictures and encode the processed block into a coded picture that is part of a coded video sequence. The example video encoder (703) may be used in place of the example video encoder (403) of FIG. 4.

[0081] For example, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as a predictive block of 8x8 samples. The video encoder (703) then determines, using, for example, rate-distortion optimization (RDO), whether the processing block is best coded using intra-mode, inter-mode, or bi-predictive mode. If it is determined that the processing block is coded in intra-mode, the video encoder (703) may encode the processing block into a coded picture using intra-prediction techniques, and if it is determined that the processing block is coded in inter-mode or bi-predictive mode, the video encoder (703) may encode the processing block into a coded picture using inter-prediction techniques or bi-prediction techniques, respectively. In some example embodiments, a merge mode may be used as a sub-mode of inter-picture prediction, in which motion vectors are derived from one or more motion vector predictors without the benefit of coded motion vector components outside the predictors. In some other example embodiments, there may be motion vector components applicable to the current block. Thus, the video encoder (703) may include components not explicitly shown in FIG. 7, such as a mode decision module, to determine the prediction mode of a processing block.

[0082] In the example of Figure 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general-purpose controller (721), and an entropy encoder (725), coupled together as shown in the example configuration of Figure 7.

[0083] The inter-encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block to one or more reference blocks in a reference picture (e.g., blocks in previous and subsequent pictures in display order), generate inter-prediction information (e.g., a description of redundancy information, motion vectors, merge mode information according to the inter-encoding technique), and calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a reference picture decoded based on video information encoded using a decoding unit 633 incorporated in the example encoder 620 of FIG. 6 (shown as residual decoder 728 of FIG. 7, as described in further detail below).

[0084] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), compare the block with previously coded blocks in the same picture, generate transformed quantized coefficients, and possibly also generate intra prediction information (e.g., intra prediction direction information according to one or more intra encoding techniques). The intra encoder (722) may calculate intra prediction results (e.g., predicted blocks) based on the intra prediction information and reference blocks in the same picture.

[0085] The general-purpose controller (721) may be configured to determine general-purpose control data and control other components of the video encoder (703) based on the general-purpose control data. In one example, the general-purpose controller (721) determines a prediction mode for a block and provides a control signal to the switch (726) based on the prediction mode. For example, if the prediction mode is intra-mode, the general-purpose controller (721) controls the switch (726) to select an intra-mode result for use by the residual calculator (723) and to control the entropy encoder (725) to select intra-prediction information and include the intra-prediction information in the bitstream. If the predicate mode of the block is inter-mode, the general-purpose controller (721) controls the switch (726) to select an inter-prediction result for use by the residual calculator (723) and to control the entropy encoder (725) to select inter-prediction information and include the inter-prediction information in the bitstream.

[0086] The residual calculator (723) may be configured to calculate the difference (residual data) between a received block and a prediction result for a block selected from the intra-encoder (722) or the inter-encoder (730). The residual encoder (724) may be configured to encode the residual data to generate transform coefficients. For example, the residual encoder (724) may be configured to transform the residual data from the spatial domain to the frequency domain to generate the transform coefficients. The transform coefficients then undergo a quantization process to obtain quantized transform coefficients. In various example embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be used appropriately by the intra-encoder (722) and the inter-encoder (730). For example, the inter-encoder (730) can generate decoded blocks based on the decoded residual data and inter-prediction information, and the intra-encoder (722) can generate decoded blocks based on the decoded residual data and intra-prediction information. The decoded blocks are processed appropriately to generate decoded pictures, which can be buffered in a memory circuit (not shown) and used as reference pictures.

[0087] The entropy encoder (725) is configured to format a bitstream to include the encoded blocks and to perform entropy coding. The entropy encoder (725) is configured to include various information in the bitstream. For example, the entropy encoder (725) may be configured to include general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other appropriate information in the bitstream. Residual information may not be present when coding a block in a merged sub-mode of either an inter mode or a bi-prediction mode.

[0088] 8 shows a diagram of an example video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive coded pictures that are part of a coded video sequence and decode the coded pictures to generate reconstructed pictures. In one example, the video decoder (810) may be used in place of the example video decoder (410) of FIG. 4.

[0089] In the example of Figure 8, the video decoder (810) includes an entropy decoder (871), an inter-decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-decoder (872) coupled to each other as shown in the example configuration of Figure 8.

[0090] The entropy decoder (871) can be configured to reconstruct, from a coded picture, certain symbols that represent the syntax elements of which the coded picture is composed. Such symbols can include, for example, prediction information (e.g., intra- or inter-prediction information) that can identify the mode in which the block is coded (e.g., intra-mode, inter-mode, bi-prediction mode, merged submode, or another submode), certain samples or metadata used for prediction by the intra decoder (872) or inter decoder (880), residual information, for example, in the form of quantized transform coefficients, etc. In one example, if the prediction mode is an inter-mode or bi-prediction mode, the inter-prediction information is provided to the inter decoder (880), and if the prediction type is an intra-prediction type, the intra-prediction information is provided to the intra decoder (872). The residual information can undergo inverse quantization and be provided to the residual decoder (873).

[0091] The inter decoder (880) may be configured to receive the inter prediction information and generate inter prediction results based on the inter prediction information.

[0092] The intra decoder (872) may be configured to receive intra prediction information and generate a prediction result based on the intra prediction information.

[0093] The residual decoder (873) may be configured to perform inverse quantization to extract inverse quantized transform coefficients and process the inverse quantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) may also utilize certain control information (to include quantization parameters (QPs)), which may be provided by the entropy decoder (871) (datapath not shown as this may only be a small amount of control information).

[0094] The reconstruction module (874) may be configured to combine, in the spatial domain, the residual as output by the residual decoder (873) and the prediction result (possibly as output by an inter-prediction module or an intra-prediction module) to form a reconstructed block that forms part of a reconstructed picture as part of the reconstructed video. Note that other appropriate operations, such as deblocking operations, may also be performed to improve visual quality.

[0095] It should be noted that the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using any suitable technique. In some example embodiments, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (603) and the video decoders (410), (510), and (810) may be implemented using one or more processors executing software instructions.

[0096] Returning to the intra-prediction process, samples within a block (e.g., a luma or chroma prediction block, or a coding block if not further divided into prediction blocks) are predicted by samples of adjacent, next-adjacent, or one or more other lines, or a combination thereof, to generate a prediction block. The residual between the actual block being coded and the prediction block may be processed by a transform followed by quantization. Various intra-prediction modes may be made available, and parameters related to intra-prediction mode selection and other parameters may be signaled in the bitstream. The various intra-prediction modes may relate, for example, to one or more line positions for predicting samples, the direction in which prediction samples are selected from predicting one or more lines, and other special intra-prediction modes.

[0097] For example, a set of intra-prediction modes (interchangeably referred to as "intra modes") may include a predefined number of directional intra-prediction modes. As described above in connection with the example implementation of FIG. 1, these intra-prediction modes may correspond to a predefined number of directions in which an out-of-block sample is selected as a prediction for a sample predicted within a particular block. In another particular example implementation, eight main directional modes may be supported and predefined, corresponding to angles from 45 degrees to 207 degrees relative to the horizontal axis.

[0098] In some other implementations of intra prediction, directional intra modes may be further expanded to a set of angles with finer granularity to further exploit the more diverse spatial redundancy in directional textures. For example, the eight-angle implementation described above may be configured to provide eight nominal angles designated V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED, as shown in FIG. 9, with a predefined number (e.g., seven) of finer angles added for each nominal angle. Such expansion allows a larger total number (e.g., 56 in this example) of directional angles to be utilized for intra prediction, corresponding to the same number of predefined directional intra modes. The prediction angle may be represented by the nominal intra angle plus an angle delta. In the specific example described above with seven finer angle directions for each nominal angle, the angle delta may be -3 to 3 times the 3-degree step size.

[0099] In some implementations, instead of or in addition to the above-described directional intra modes, a predefined number of omnidirectional intra prediction modes may also be predefined and made available. For example, five omnidirectional intra modes called smooth intra prediction modes may be specified. These omnidirectional intra prediction modes may be specifically called DC, PAETH, SMOOTH, SMOOTH_V, and SMOOTH_H intra modes. Prediction of samples of a particular block under these example omnidirectional modes is shown in FIG. 10. As an example, FIG. 10 shows a 4×4 block 1002 predicted by samples from the top-most and / or left-most neighboring lines. A particular sample 1010 in the block 1002 may correspond to a sample 1004 immediately above the sample 1010 in the top-most neighboring line of the block 1002, a sample 1006 above and to the left of the sample 1010 as the intersection of the top-most and left-most neighboring lines, and a sample 1008 immediately to the left of the sample 1010 in the left-most neighboring line of the block 1002. In an example DC intra prediction mode, the average of the left and above neighboring samples 1008 and 1004 may be used as a predictor for sample 1010. In an example PAETH intra prediction mode, the above, left, and above-left reference samples 1004, 1008, and 1006 may be fetched, and then any value among these three reference samples that is closest to (above + left - above-left) may be set as the predictor for sample 1010. In an example SMOOTH_V intra prediction mode, sample 1010 may be predicted by quadratic interpolation in the vertical direction of the above-left neighboring sample 1006 and the left neighboring sample 1008. For an example SMOOTH_H intra prediction mode, sample 1010 may be predicted by quadratic interpolation in the horizontal direction of the above-left neighboring sample 1006 and the above neighboring sample 1004. For an example SMOOTH intra prediction mode, sample 1010 may be predicted by an average of quadratic interpolation in the vertical and horizontal directions. The above omni-directional intra mode implementations are provided merely as non-limiting examples.Other adjacent lines, and other non-directional selection of samples, and methods of combining predicted samples to predict a particular sample within a prediction block are also possible.

[0100] The encoder's selection of a particular intra-prediction mode from the above directional or omnidirectional modes at various coding levels (picture, slice, block, unit, etc.) may be signaled in the bitstream. In some example implementations, the eight exemplary nominal directional modes may be signaled first, along with the five non-angle smooth modes (for a total of 13 options). Then, if the signaled mode is one of the eight nominal angle intra-modes, an index is further signaled to indicate the selected angle delta to the corresponding signaled nominal angle. In some other example implementations, all intra-prediction modes may be indexed together for signaling (e.g., 56 directional modes plus 5 omnidirectional modes to generate 61 intra-prediction modes).

[0101] In some example implementations, the example 56 or other number of directional intra-prediction modes may be implemented using a unified directional predictor that projects each sample of a block to a reference sub-sample position and interpolates the reference samples with a 2-tap bilinear filter.

[0102] In some implementations, additional filter modes, called FILTER INTRA modes, can be designed to account for decaying spatial correlation with edge references. In these modes, predicted samples within a block may be used as intra-prediction reference samples for some patches within the block, in addition to out-of-block samples. These modes may be predefined and made available for intra-prediction of at least the luma block (or only the luma block). A predefined number (e.g., five) of filter intra modes may be predesigned, each represented by a set of n-tap filters (e.g., 7-tap filters) that reflect the correlation between a sample within a 4x2 patch and its n neighbors. In other words, the weight coefficients of the n-tap filters may depend on the position. Taking an 8x8 block, 4x2 patch, and 7-tap filtering as an example, the 8x8 block 1102 may be divided into eight 4x2 patches, as shown in Figure 11. These patches are shown as B0, B1, B1, B3, B4, B5, B6, and B7 in Figure 11. For each patch, its seven neighboring patches, denoted by R0 through R7 in FIG. 11, can be used to predict samples within the current patch. For patch B0, all neighboring elements may already be reconstructed. However, for other patches, not all neighboring elements have been reconstructed, so the predicted values ​​of the nearest neighboring elements are used as reference values. For example, as pointed out in FIG. 11, all neighboring elements of patch B7 have not been reconstructed, so the predicted samples of the neighboring elements are used instead.

[0103] In some implementations of intra prediction, one color component may be predicted using one or more other color components. The color components may be in any of the YCrCb, RGB, XYZ color spaces, etc. For example, prediction of a chroma component (e.g., a chroma block) from a luma component (e.g., a luma reference sample) (called chroma from luma, or CfL) may be performed. In some example implementations, most cross-color predictions are only allowed from luma to chroma. For example, chroma samples within a chroma block may be modeled as a linear function of the corresponding reconstructed luma sample. CfL prediction may be performed as follows: CfL(α)=α×L AC +DC (1)

[0104] In the formula, L AC where α represents the AC contribution of the luma component, α represents a parameter of a linear model, and DC represents the DC contribution of the chroma component. For example, the AC components are obtained for each sample of the block, while the DC component is obtained for the entire block. Specifically, the reconstructed luma samples may be subsampled to the chroma resolution, and then the average luma value (the luma DC) may be subtracted from each luma value to form the AC contribution in luma. The luma AC contribution is then used in the linear mode of Equation (1) to predict the AC value of the chroma component. Instead of requiring the decoder to calculate scaling parameters to approximate or predict the chroma AC components from the luma AC contribution, an example CfL implementation can determine the parameter α based on the original chroma samples and signal them in the bitstream. This reduces decoder complexity and results in more accurate prediction. Regarding the DC contribution of the chroma components, in some example implementations, it may be calculated using an intra-DC mode within the chroma component.

[0105] A transform of the residual of either the intra-predicted block or the inter-predicted block may then be performed, followed by quantization of the transform coefficients. For the purpose of performing the transform, both intra- and inter-coding blocks may be further partitioned into multiple transform blocks (the term "unit" is typically used to refer to a collection of three color channels, but may also be used interchangeably as "transform unit"; e.g., a "coding unit" includes a luma coding block and a chroma coding block) before the transform. In some implementations, a maximum partitioning depth of a coded block (or predictive block) may be specified (the term "coded block" may be used interchangeably with "coding block"). For example, such partitioning may not exceed two levels. The partitioning of a predictive block into transform blocks may be handled differently for intra- and inter-predicted blocks. However, in some implementations, such partitioning may be similar between intra- and inter-predicted blocks.

[0106] In some example implementations, and for intra-coding blocks, transform partitioning may be performed such that all transform blocks have the same size and the transform blocks are coded in raster scan order. An example of such partitioning of transform blocks for intra-coding blocks is shown in Figure 12. Specifically, Figure 12 shows that a coded block 1202 is partitioned into 16 transform blocks of the same block size via mid-level quadtree partitioning 1204, as indicated by 1206. An example raster scan order for coding is indicated by the ordered arrows in Figure 12.

[0107] In some example implementations, and for inter-coding blocks, transform unit partitioning may be performed recursively with a partitioning depth up to a predefined number of levels (e.g., two levels). The partitioning can stop or continue recursively for any sub-partition and at any level, as shown in FIG. 13. In particular, FIG. 13 shows an example in which a block 1302 is partitioned into four quadtree sub-blocks 1304, one of the sub-blocks is further partitioned into four second-level transform blocks, while the partitioning of the other sub-block stops after the first level, resulting in a total of seven transform blocks of two different sizes. An example raster scan order for coding is indicated by the ordered arrows in FIG. 13. While FIG. 13 shows an example implementation of quadtree partitioning of square transform blocks with up to two levels, in some production implementations, the transform partitioning may support 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1 transform block shapes and sizes ranging from 4x4 to 64x64. In some example implementations, if the coding block is 64x64 or smaller, transform block partitioning may be applied only to the luma component (in other words, the chroma transform block is the same as the coding block under that condition). Otherwise, if the width or height of the coding block is greater than 64, both the luma coding block and the chroma coding block may be implicitly divided into multiples of min(W,64)xmin(H,64) and min(W,32)xmin(H,32) transform blocks, respectively.

[0108] Each of the above transform blocks may then undergo a linear transform, which essentially moves the residual within the transform block from the spatial domain to the frequency domain. Some implementations of the actual linear transform may allow for multiple transform sizes (ranging from 4 points to 64 points for each of the two dimensions) and transform shapes (square, rectangular with width / height ratios of 2:1 / 1:2 and 4:1 / 1:4) to support the example extended coding block partition above.

[0109] Turning to the actual linear transform, in some example implementations, the 2D transform process can include the use of hybrid transform kernels (which may, for example, be composed of different 1D transforms for each dimension of the coding residual transform block). Example 1D transform kernels may include, but are not limited to, the following: a) 4-point, 8-point, 16-point, 32-point, and 64-point DCT-2; b) 4-point, 8-point, and 16-point asymmetric DST (DST-4, DST-7) and their flipped versions; and c) 4-point, 8-point, 16-point, and 32-point equivalent transforms. The selection of the transform kernel to be used for each dimension can be based on a rate-distortion (RD) criterion. For example, the basis functions of the DCT-2 and asymmetric DST that may be implemented are listed in Table 1.

[0110] [Table 1]

[0111] In some example implementations, the availability of hybrid transform kernels for a particular primary transform implementation can be based on the transform block size and prediction mode. Example dependencies are listed in Table 2. For chroma components, the selection of the transform type may be performed in an implicit manner. For example, for intra-prediction residuals, the transform type may be selected according to the intra-prediction mode, as specified in Table 3. For inter-prediction residuals, the transform type of the chroma block may be selected according to the transform type selection of the co-located luma block. Therefore, for chroma components, there is no transform type signaling in the bitstream.

[0112] [Table 2]

[0113] [Table 3]

[0114] In some implementations, a secondary transform may be performed on the primary transform coefficients. For example, as shown in FIG. 14, a LFNST (low-frequency separable transform), also known as a reduced secondary transform, may be applied between the forward primary transform and quantization (at the encoder) and between the inverse quantization and the inverse primary transform (at the decoder side) to further decorrelate the primary transform coefficients. Essentially, the LFNST may take a portion of the primary transform coefficients, e.g., a low-frequency portion (thus a "reduced" portion from the full set of primary transform coefficients of the transform block), to proceed to the secondary transform. In an example of LFNST, a 4x4 non-separable transform or an 8x8 non-separable transform may be applied according to the transform block size. For example, a 4x4 LFNST may be applied to a small transform block (e.g., min(width, height)<8), and an 8x8 LFNST may be applied to a large transform block (e.g., min(width, height)>8). For example, if an 8x8 transform block is subjected to a 4x4 LFNST, only the low frequency 4x4 portion of the 8x8 primary transform coefficients undergoes a further secondary transform.

[0115] As specifically shown in FIG. 14, a transform block may be 8×8 (or 16×16). Thus, a forward primary transform 1402 of the transform block generates an 8×8 (or 16×16) primary transform coefficient matrix 1404, with each square unit representing a 2×2 (or 4×4) portion. The input to the forward LFNST need not be, for example, the entire 8×8 (or 16×16) primary transform coefficients. For example, a 4×4 (or 8×8) LFNST may be used for the secondary transform. Thus, as shown in the shaded area (top left) 1406, only the 4×4 (or 8×8) low-frequency primary transform coefficients of the primary transform coefficient matrix 1404 may be used as input to the LFNST. The remaining portion of the primary transform coefficient matrix may not undergo a secondary transform. In this way, after the secondary transform, the portions of the primary transform co-effects affected by the LFNST become secondary transform coefficients, while the remaining portions not affected by the LFNST (e.g., the unshaded portions of matrix 1404) retain their corresponding primary transform coefficients. In some example implementations, the remaining portions not subject to the secondary transform may be set to all zero coefficients.

[0116] An example of the application of a non-separable transform used in LFNST is described below. To apply an example 4×4 LFNST, a 4×4 input block X (representing, e.g., a 4×4 low-frequency portion of a block of linear transform coefficients such as the shaded portion 1406 of the linear transform matrix 1404 in FIG. 14) can be expressed as follows:

number

[0117] This 2D input matrix is ​​first linearized or converted to a vector in the order

number

number

[0118] Next, the inseparable transform of the 4×4 LFNST can be calculated as

Number

Number

Number

[0119] The LFNST of the above example is based on a direct matrix multiplication method for applying separable transforms and, as a result, is performed in a single pass without multiple iterations. In some further example implementations, in order to minimize the computational complexity and memory space requirements for storing transform coefficients, the dimension of the inseparable transform matrix (T) of, for example, the 4×4 LFNST can be further reduced. Such an implementation may be referred to as a reduced inseparable transform (RST). More specifically, the main concept of RST is to map an N (where N is 4×4 = 16 in the above example, but may also be equal to 64 for an 8×8 block) - dimensional vector to an R - dimensional vector in a different space, and N / R (R < N) represents the dimension reduction coefficient. Therefore, instead of an N×N transform matrix, the RST matrix becomes an R×N matrix as follows,

Number

[0120] Here, the R rows of the transformation matrix are a reduced R basis of the N-dimensional space. Thus, the transformation converts an input vector or N dimensions into an output vector of reduced R dimensions. Thus, as shown in Figure 14, the secondary transformation coefficients 1408 transformed from the primary coefficients 1406 are reduced in dimension by a factor or N / R. The three squares around 1408 in Figure 14 may be padded with zeros.

[0121] The inverse transform matrix of the RTS may be the transpose of its forward transform. For an example 8×8 LFNST (contrasted with the 4×4 LFNST described above for more detailed explanation), an example reduction factor of 4 may be applied, and thus the 64×64 direct non-separable transform matrix is ​​correspondingly reduced to a 16×64 direct matrix. Furthermore, in some implementations, some, but not all, of the input primary coefficients may be linearized into the input vectors of the LFNST. For example, only some of the example 8×8 input primary transform coefficients may be linearized into the X vector described above. In a particular example, of the four 4×4 quadrants of the 8×8 primary transform coefficient matrix, the lower right (high-frequency coefficients) may be excluded, and only the other three quadrants are linearized into a 64×1 vector using a predefined scan order rather than a 48×1 vector. In such an implementation, the non-separable transform matrix may be further reduced from 16×64 to 16×48.

[0122] Therefore, an example reduced 48x16 inverse RST matrix can be used at the decoder side to generate the upper-left, upper-right, and lower-left 4x4 quadrants of the 8x8 core (primary) transform coefficients. Specifically, if a further reduced 16x48 RST matrix is ​​applied instead of a 16x64 RST matrix with the same transform set configuration, the non-separable secondary transform takes as input 48 matrix elements vectorized from the three 4x4 quadrant blocks of the 8x8 primary coefficient block, excluding the lower-right 4x4 block. In such an implementation, the omitted lower-right 4x4 primary transform coefficient is ignored in the secondary transform. This further reduced transform converts the 48x1 vector into a 16x1 output vector, which is scanned back to the 4x4 matrix to fill 1408 in Figure 14. The three squares of secondary transform coefficients surrounding 1408 may be padded with zeros.

[0123] With the help of this reduction in the dimensions of the RST, the memory usage for storing all LFNST matrices is reduced. In the above example, for example, the memory usage can be reduced from 10 KB to 8 KB with a reasonably small performance degradation compared to an implementation without dimensionality reduction.

[0124] In some implementations, to reduce complexity, the LFNST may be further restricted so that it is only applicable if all coefficients outside the portion of primary transform coefficients targeted by the LFNST (e.g., outside 1406 portion of 1404 in Figure 14) are insignificant. Thus, when the LFNST is applied, all primary-only transform coefficients (e.g., the unshaded portion of the primary coefficient matrix 1404 in Figure 4) may be near zero. Such a restriction allows for adjustment of the LFNST index signaling at the least significant positions, thus avoiding additional coefficient scans that may be required to check for significant coefficients at certain positions when this restriction does not apply. In some implementations, the worst-case processing (in terms of multiplications per pixel) of the LFNST may limit non-separable transforms of 4x4 and 8x8 blocks to 8x16 and 8x48 transforms, respectively. In such cases, for other sizes less than 16, the final significant scan position must be less than 8 when the LFNST is applied. For blocks with shapes of 4xN, Nx4, and N>8, the above restriction means that LFNST is applied only once to the top-left 4x4 region. When LFNST is applied, all first-order-only coefficients are zero, so in such cases the number of operations required for the first-order transform is reduced. From the encoder's perspective, quantization of coefficients can be simplified when the LFNST transform is tested. Rate-distortion optimized quantization (RDO) must be performed on a maximum of the first 16 coefficients (in scan order), and the remaining coefficients may be forced to be zero.

[0125] In some example implementations, the available RST kernels may be specified as several transform sets, with each transform set containing several non-separable transform matrices. For example, there may be a total of four transform sets and two non-separable transform matrices (kernels) for each transform set used in LFNST. These kernels may be pre-trained offline and thus data-driven. The offline trained transform kernels may be stored in memory or hard-coded into the encoding or decoding device for use during the encoding / decoding process. The selection of a transform set during the encoding or decoding process may be determined by the intra-prediction mode. The mapping from intra-prediction mode to transform sets may be pre-defined. An example of such a pre-defined mapping is shown in Table 4. For example, as shown in Table 4, if one of three cross-component linear model (CCLM) modes (INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the current block (i.e., 81<=predModeIntra<=83), transform set 0 may be selected for the current chroma block. For each transform set, the selected non-separable secondary transform candidate may be further specified by an explicitly signaled LFNST index, e.g., the index may be signaled in the bitstream once per intra CU after the transform coefficients.

[0126] [Table 4]

[0127] In the above example implementation, LFNST is restricted to be applicable only when all coefficients outside the first coefficient subgroup or portion are insignificant, so that LFNST index coding depends on the position of the last significant coefficient. Additionally, the LFNST index may be context coded, but it may be independent of the intra prediction mode, and only the first bin may be context coded. Furthermore, LFNST may be applied to intra CUs in both intra and inter slices, and to both luma and chroma. When dual tree is enabled, LFNST indices for the luma and chroma components can be signaled separately. For inter slices (when dual tree is disabled), a single LFNST index is signaled and used for both luma and chroma.

[0128] In some example implementations, when an intra subpartitioning (ISP) mode is selected, LFNST may be disabled and the RST index may not be signaled because performance gains are likely to be limited even if RST is applied to all feasible partition blocks. Furthermore, disabling RST for ISP prediction residuals can reduce encoding complexity. In some further implementations, when a multiple linear regression intra prediction (MIP) mode is selected, LFNST may also be disabled and the RST index may not be signaled.

[0129] Considering that due to the existing maximum transform size limitation (e.g., 64x64), large CUs larger than 64x64 (or any other predefined size representing the maximum transform block size) are implicitly split (e.g., TU tiling), LFNST index lookup can increase data buffering by four times for a certain number of decoding pipeline stages. Therefore, in some implementations, the maximum size allowed for LFNST may be limited to, for example, 64x64. In some implementations, LFNST may be enabled with DCT2 only as the primary transform.

[0130] In some other implementations, an intra-secondary transform (IST) is provided for the luma component by defining, for example, 12 sets of secondary transforms, with, for example, three kernels in each set. An intra-mode dependent index may be used for transform set selection. Kernel selection within a set may be based on a signaled syntax element. IST may be enabled when either DCT2 or ADST is used as both the horizontal and vertical primary transforms. In some implementations, a 4x4 non-separable transform or an 8x8 non-separable transform may be selected according to the block size. If min(tx_width, tx_height)<8, a 4x4 IST may be selected. For larger blocks, an 8x8 IST may be used, where tx_width and tx_height correspond to the width and height of the transform block, respectively. The input to the IST may be low-frequency primary transform coefficients in zigzag scan order.

[0131] The use of data-driven or offline-trained transform kernels, particularly non-separable transform kernels, for either primary or secondary transforms can facilitate improved coding performance compared to using only fixed transforms such as DCT / DST. However, applying such data-driven transform kernels can increase the complexity of a video codec. For example, data-driven transform kernels may require a significant amount of memory to store all possible kernels, particularly for non-separable transforms. Specifically, for one of the example implementations described above, there may be 12 possible kernel sets. Each set may include, for example, three kernels for different LFNSTs (e.g., 4x4 and 8x8). Furthermore, the LFNST kernels may be different for different primary transform types (e.g., DCT or ADST primary types for which LFNST may be permitted). In other words, if memory is used to store offline-trained LFNST kernels, a significant amount of memory is required. As described in further detail below, various implementations described below for using data-driven transforms can be designed to reduce the memory footprint of the stored kernels and the complexity of the codec.

[0132] These various implementations or embodiments are merely examples. They may be used separately or combined in any order or manner. Furthermore, each implementation may be embodied as an encoder or decoder including a method and processing circuitry (e.g., one or more processors or one or more integrated circuits) for performing the method. In one example, the processing circuitry may include one or more processors executing a program stored on a non-transitory computer-readable medium. In another example, the processing circuitry may be hard-coded to perform the method. In yet another example, the processing circuitry may be a mixture of processor components and hard-coded circuitry for executing computer-readable instructions.

[0133] In the following example implementation, the term "block" may be interpreted as a prediction block or a coding block, etc. The term "block" here may also be used to refer to a transform block, which may be part of a coding block as described above. The term "size" of a block is used to refer to any of the block's width, height, block aspect ratio, block region size, or minimum / maximum value between width and height.

[0134] In various example implementations below, one or more indices for identifying one or more separable or non-separable transforms (or transform kernels) among a set of separable or non-separable secondary transforms used to decode a current transform block may be signaled in a video bitstream. Each such index may be represented as stIdx. These implementations may apply to either a primary transform, a secondary transform, or an additional transform applied after a secondary transform. Furthermore, the underlying principles may be applied to either separable or non-separable transforms.

[0135] In the following example implementation, the term "basis vector" is used to refer to the spatial frequency components of a transform kernel for image transformation from the spatial domain to the frequency domain. Terms such as "high-frequency basis vector" are used to refer to basis vectors that can be used to generate transform coefficients that are scanned from low-frequency components to high-frequency components after N coefficients, and can also refer to basis vectors located after the Nth row (or column) of a non-separable transform kernel (and therefore higher in frequency than N). Example values ​​of N include, but are not limited to, any integer value between 1 and 128 (inclusive). The delineation of high and low frequencies (number N) can be determined via RDO.

[0136] In some common example implementations, multiple transform kernels may be applied, and when offered as selectable options during the encoding and decoding process, some of the basis vectors of the multiple transform kernels may be shared. For example, high-frequency basis vectors between some or all of the multiple transform kernels may be shared, and low-frequency basis vectors may be personalized among the multiple transform kernels. In such problems, such offline-trained and data-driven transform kernels may require less memory overall. In particular, shared basis vectors do not need to be replicated; one copy may need to be kept in memory between the shared transform kernels. The manner in which the shared and personalized basis vectors are stored in memory may be predefined.

[0137] Such an implementation is illustrated in FIG. 15, which shows an example of how basis vectors are shared between two LFNST transform matrices or kernels used in a forward quadratic transform. Two blocks 1502 and 1504 show two transform matrices, namely, transform matrix A (left) and B (right), where each row of the transform matrix represents one basis vector, and the bottom row (shaded) refers to the basis vector shared between A and B. The bottom basis vectors may be of higher frequency. The higher frequency basis vectors may be shared in memory between these example kernels. For the inverse transform, the shared basis vectors of multiple transform matrices may be the right N columns of the transform matrix as a result of the transposition for the inverse transform.

[0138] The example transform kernel in Figure 15 is shown for a quadratic transform kernel applied to a linearization vector of a linear transform coefficient. Therefore, the frequency axis of a transform kernel as shown in Figure 15 is one-dimensional (top to bottom). This applies to all 1D kernels. For a transform kernel that is 2D (e.g., when a data-driven 2D kernel is used for the linear transform), the lower frequency components are in the upper left part of the 2D transform kernel and the higher frequency components are in the lower right part of the 2D transform kernel. In that case, the lower right part may be shared between different 2D kernels.

[0139] In some example implementations, for a set of multiple transform kernels applied to a particular intra-prediction mode, the high-frequency basis vectors may be identically constructed (shared), and the low-frequency basis vectors may be different (individualized). In other words, a set of kernels corresponding to a particular intra-prediction mode (e.g., the various directions, omnidirectional, and other intra-prediction modes described above) may have similar dimensions, and thus, for example, the high-frequency basis vectors may be shared with individualized low-frequency basis vectors. A predefined threshold may be used for delineating the low and high frequencies. In some implementations, the delineation of the high and low frequencies may be data-driven and thus pre-trained. Thus, the frequency delineation threshold may be different for different intra-prediction modes.

[0140] In some example implementations, when multiple secondary transform kernels are applied with different primary transforms, some or all of the multiple secondary transform kernels can be shared among different primary transform types. In other words, different types of primary transforms may correspond to different kernels for the secondary transforms. These different secondary transform kernels may include individualized low-frequency basis vectors, yet be constructed with shared high-frequency basis vectors. That is, when the secondary transform kernel selection is applicable to multiple primary transform types, it may be independent of the primary transform type (e.g., independent of the DCT or ADST primary transform type).

[0141] In some example implementations, when multiple secondary transform kernels are applied with different primary transforms, the high-frequency basis vectors in the multiple secondary transform kernels may be shared between different primary transform types, and the low-frequency basis vectors may be different or individualized. A predefined threshold can be used to distinguish between low and high frequencies. Again, the distinction between high and low frequencies (or number N) may be data-driven and therefore pre-trained. Thus, the frequency distinction threshold may be different for different intra-prediction modes.

[0142] In some example implementations, for multiple secondary transform kernels applied with a particular type of linear transform, the high-frequency basis vectors may be the same, but the low-frequency basis vectors may be different. In some further implementations, both the high-frequency and low-frequency basis vectors of kernels corresponding to different linear transform types may be independent.

[0143] In some example implementations, the number N of high frequency basis vectors shared among various transform kernels may be an integer power of two, e.g., 1, 2, 4, 8, 16, 32, 64, 128. In particular for reduced quadratic transform matrices, the number of N high frequency rows may be a power of two value, e.g., 1, 2, 4, 8, 16, 32, 64, 128.

[0144] In some example implementations, the value of N may depend on the transform size. For example, the ratio of N to the total number of basis vectors of the transform kernel may be specified.

[0145] In some examples, each and any combination of the above implementations can be applied to the LFNST quadratic transform. Thus, a set of kernels that share high-frequency basis vectors but include individualized low-frequency basis vectors can comprise quadratic LFNST transform kernels.

[0146] In some examples, each and any combination of the above implementations may be applied to an intra-quadratic transform (IST).

[0147] In some examples, each and any combination of the above implementations may be applied to a line graph transformation (LGT).

[0148] The various kernels may be predetermined. They may be pre-trained offline. In the case of kernels that share high-frequency basis vectors, they may be trained together. The threshold frequencies representing the shared basis vectors and the individualized basis vectors may be trained offline.

[0149] During the encoding or decoding process, the various kernels described above may reside in memory space for use by the encoder or decoder, and only one copy of the shared basis vectors needs to be stored in memory space. When the various kernels are used during the encoding or decoding process, these shared basis vectors can be accessed using pointers to the memory locations for the shared basis vectors.

[0150] FIG. 16 shows a flowchart 1600 of an example method according to the principles underlying the above implementations. The example method flow starts at 1601. At S1610, a plurality of transform kernels are identified, each of the plurality of transform kernels including a set of basis vectors ranging from low frequency to high frequency, wherein N high-frequency basis vectors of two or more of the plurality of transform kernels are shared, where N is a positive integer, and low-frequency basis vectors of two or more of the plurality of transform kernels other than the N high-frequency basis vectors are individualized. At S1620, a data block is extracted from a video bitstream. At S1630, a transform kernel is selected from the plurality of transform kernels based on information associated with the data block. At S1640, the transform kernel is applied to at least a portion of the data block to generate a transform block. The example method flow ends at S1699.

[0151] The embodiments of the present disclosure may be used separately or combined in any order. Furthermore, each of the methods (or embodiments), the encoder, and the decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium. The embodiments of the present disclosure may be applied to luma blocks or chroma blocks.

[0152] The techniques described above may be implemented as computer software using computer-readable instructions physically stored on one or more computer-readable media. For example, Figure 17 illustrates a computer system (1700) suitable for implementing certain embodiments of the disclosed subject matter.

[0153] Computer software may be coded using any suitable machine code or computer language that can be subjected to assembly, compilation, linking, or similar mechanisms to create code containing instructions that can be executed by one or more computer central processing units (CPUs) and graphics processing units (GPUs), etc., directly, or through interpretation and execution of microcode, etc.

[0154] The instructions may be executed by various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming consoles, Internet of Things devices, and the like.

[0155] 17 for computer system (1700) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing embodiments of the present disclosure. The arrangement of components should not be interpreted as having a dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of computer system (1700).

[0156] The computer system (1700) may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users, for example, via tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), or olfactory input (not shown). The human interface devices may also be used to capture certain media that do not necessarily involve direct conscious human input, such as audio (speech, music, ambient sounds, etc.), images (scanned images, photographic images obtained from a still image camera, etc.), and video (two-dimensional video, three-dimensional video, including stereoscopic video, etc.).

[0157] The input human interface devices may include one or more (only one of each) of a keyboard (1701), a mouse (1702), a trackpad (1703), a touch screen (1710), a data glove (not shown), a joystick (1705), a microphone (1706), a scanner (1707), and a camera (1708).

[0158] The computer system (1700) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen (1710), data gloves (not shown), or joystick (1705), although some haptic feedback devices may not function as input devices), audio output devices (such as speakers (1709), headphones (not shown)), visual output devices (such as screens (1710), including CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touchscreen input capability, each with or without haptic feedback capability, some of which may be capable of two-dimensional visual output or three-dimensional or higher-dimensional output via means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0159] The computer system (1700) may also include human-accessible storage devices and associated media such as optical media including CD / DVD ROM / RW (1720) with media such as CD / DVD (1721), thumb drives (1722), removable hard drives or solid state drives (1723), legacy magnetic media such as tape or floppy disks (not shown), and dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0160] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.

[0161] The computer system (1700) also includes an interface (1754) to one or more communication networks (1755). The network may be, for example, wireless, wired, or optical. The network may further be local, wide-area, metropolitan, vehicular, industrial, real-time, delay-tolerant, or the like. Examples of networks include local area networks such as Ethernet and wireless LAN; cellular networks including GSM, 3G, 4G, 5G, LTE, and the like; television wired or wireless wide-area digital networks including cable television, satellite television, and terrestrial broadcast television; and vehicular and industrial networks including CANbus. Certain networks typically require an external network interface adapter attached to a particular general-purpose data port (e.g., a USB port on the computer system (1700)) or peripheral bus (1749), while other networks are typically integrated into the core of the computer system (1700) by attaching to a system bus (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system), as described below. Using any of these networks, the computer system (1700) can communicate with other entities. Such communication may be unidirectional, receive only (e.g., broadcast TV), unidirectional transmit only (e.g., from the CANbus to a particular CANbus device), or bidirectional, e.g., communication with other computer systems using local area digital networks or wide area digital networks. Specific protocols and protocol stacks may be used with each of these networks and network interfaces, as described above.

[0162] The aforementioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the core (1740) of the computer system (1700).

[0163] The core (1740) may include one or more central processing units (CPUs) (1741), graphics processing units (GPUs) (1742), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (1743), task-specific hardware accelerators (1744), graphics adapters (1750), etc. These devices may be connected via a system bus (1748), along with read-only memory (ROM) (1745), random access memory (1746), and internal mass storage devices (1747) such as internal hard drives or SSDs that are not user-accessible. In some computer systems, the system bus (1748) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (1748) or via a peripheral bus (1749). In one example, a screen (1710) may be connected to the graphics adapter (1750). Peripheral bus architectures include PCI, USB, and the like.

[0164] The CPU (1741), GPU (1742), FPGA (1743), and accelerator (1744) can execute several instructions that can combine to form the above-mentioned computer code. The computer code can be stored in ROM (1745) or RAM (1746). Transient data can also be stored in RAM (1746), while permanent data can be stored, for example, in internal mass storage (1747). Fast storage and retrieval from any of the memory devices can be enabled through the use of cache memory, which can be closely associated with one or more of the CPU (1741), GPU (1742), mass storage (1747), ROM (1745), RAM (1746), etc.

[0165] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.

[0166] As a non-limiting example, a computer system (1700) having the architecture, and in particular a core (1740), can provide functionality as a result of processor(s) (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be user-accessible mass storage devices, as introduced above, as well as media associated with specific storage devices of the core (1740) that are non-transitory in nature, such as the core's internal mass storage device (1747) or ROM (1745). Software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core (1740). The computer-readable media can include one or more memory devices or chips, depending on particular needs. The software can cause the core (1740), and in particular the processor (including a CPU, GPU, FPGA, etc.) within the core (1740), to perform particular processes or particular portions of particular processes, including determining data structures stored in RAM (1746) and modifying such data structures according to processes defined by the software, as described herein. Additionally, or alternatively, a computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator (1744)) that can operate in place of or together with software to perform particular processes or portions of particular processes described herein. References to software may encompass logic, and vice versa, where appropriate. Where appropriate, references to computer-readable media may encompass circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both. The present disclosure encompasses any appropriate combination of hardware and software.

[0167] While this disclosure has described several exemplary embodiments, there are modifications, permutations, and various substitute equivalents that fall within the scope of this disclosure. In the implementations and embodiments described above, any operations of the processes can be combined or configured in any quantity or order as needed. Also, two or more of the operations of each process described above may be performed in parallel. Thus, it will be appreciated that those skilled in the art can devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and thus are within its spirit and scope. Appendix A: Acronyms JEM: Joint Exploration Model VVC: Versatile Video Coding BMS: Benchmark Set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplemental Enhancement Information VUI: Video Usability Information GOP: Group of Pictures TU: Conversion unit PU: Prediction Unit CTU: Coding Tree Unit CTB: coding tree block PB: Predicted Block HRD: Hypothetical Reference Decoder SNR: Signal-to-Noise Ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: cathode ray tube LCD: Liquid crystal display OLED: Organic Light Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile Communications LTE: Long Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Area SSD: Solid State Drive IC: Integrated Circuit HDR: High Dynamic Range SDR: Standard Dynamic Range JVET: Joint Video Exploration Team MPM: Most Probable Mode WAIP: Wide-angle Intra Prediction CU: Coding Unit PU: Prediction Unit TU: Conversion unit CTU: Coding Tree Unit PDPC: Position-dependent prediction combination ISP: Intra-subpartition SPS: Sequence parameter settings PPS: Picture Parameter Set APS: Adaptive Parameter Set VPS: Video Parameter Set DPS: Decoding Parameter Set ALF: Adaptive Loop Filter SAO: Sample Adaptive Offset CC-ALF: Cross-component adaptive loop filter CDEF: Constrained Directivity Enhancement Filter CCSO: Cross-component sample offset LSO: Local Sample Offset LR: Loop Recovery Filter AV1:AOMedia Video 1 AV2:AOMedia Video 2 [Explanation of symbols]

[0168] 101 Samples 102 Arrow 103 Arrow 104 blocks 180 Schematic showing intra-predicted aroma 201 Current Block 202 Surrounding Samples 203 Surrounding Samples 204 Surrounding Samples 205 Surrounding Samples 206 Surrounding Samples 300 Communication Systems 310 Terminal Equipment 320 Terminal Equipment 330 Terminal Equipment 340 Terminal Equipment 350 Network 400 Communication Systems 401 Video Source 402 Video Picture Stream 403 Video Encoder 404 encoded video bitstream 405 Streaming Server 406 Client Subsystem 408 Client Subsystem 407 Copy Video Data 409 Copy Video Data 410 Video Decoder 411 Output Stream 412 Display 413 Video Capture Subsystem 420 Electronic equipment 430 Electronic equipment 501 Channel 510 Video Decoder 512 display 515 buffer memory 520 Parser 521 Symbol 530 Electronic equipment 531 Receiver 551 Scaler / Descaler Unit 552 Intra Prediction Units 553 Motion Compensation Prediction Unit 555 Aggregator 556 Loop Filter 557 Reference Picture Memory 558 Current Picture Buffer 601 Video Sources 603 Video Encoder 620 Electronic equipment 630 Source Coder 632 Coding Engine 633 Decoder, Decoding Unit 634 Reference Picture Memory 635 Predictor 640 Transmitter 643 Video Sequences 645 Entropy Coder 650 Controller 660 channels 703 Video Encoder 721 General-purpose controller 722 Intra Encoder 723 Residual Calculator 724 Residual Encoder 725 Entropy Encoder 726 Switch 728 Residual Decoder 730 Interencoder 810 Video Encoder 871 Entropy Decoder 872 Intra Decoder 873 Residual Decoder 874 Reconstruction Module 880 Interdecoder 1002 blocks 1010 samples 1004 adjacent samples 1006 adjacent samples 1008 adjacent samples 1202 coded blocks 1204 Quadtree splitting 1206 Conversion Block Block 1302 1304 Quadtree Subblocks 1402 Forward Linear Transform 1404 Linear transformation coefficient matrix 1406 Shaded area of ​​linear transformation coefficient matrix 1408 Quadratic Conversion Factors 1502 blocks 1504 blocks 1600 Method Flowchart 1700 Computer Systems 1701 Keyboard 1702 Mouse 1703 Trackpad 1705 Joystick 1706 Microphone 1707 Scanner 1708 Camera 1709 Audio Output Device 1710 Touch Screen 1720 Optical media 1721 CD / DVD and other media 1722 thumb drive 1723 Removable Hard Drives and Solid State Drives 1740 cores 1741 Central Processing Unit 1742 Graphics Processing Unit 1743 Field Programmable Gate Area 1744 Hardware Accelerator 1745 Read-Only Memory 1746 Random Access Memory 1747 Internal mass storage 1748 System Bus 1749 Peripheral bus 1750 graphics adapter 1754 Network Interface 1755 Communication Network

Claims

Claim 1: A method for encoding a video bitstream, comprising: generating a data block and a plurality of transformation kernels from the video information; a transformation kernel of the plurality of transformation kernels is associated with the data block; each of the plurality of transformation kernels includes a set of basis vectors ranging from low frequency to high frequency; N high frequency basis vectors of two or more of the plurality of transformation kernels are shared, where N is a positive integer; low-frequency basis vectors of the two or more transformation kernels among the plurality of transformation kernels other than the N high-frequency basis vectors are individualized; generating a data block and a plurality of transformation kernels; generating a bitstream including the data block and the plurality of transform kernels; 1. A method for encoding a video bitstream, comprising:

2. The method of claim 1 , wherein the plurality of transformation kernels are pre-trained offline.

3. The method of claim 2 , wherein the multiple transformation kernels are jointly trained offline to determine the N shared high-frequency basis vectors.

4. The method of claim 1 , wherein the N high frequency basis vectors include basis vectors having frequencies higher than a predetermined threshold frequency.

5. The method of claim 4 , wherein the plurality of transformation kernels and the predetermined threshold frequency are pre-trained offline.

6. the plurality of transform kernels includes a secondary transform kernel applicable to transform primary transform coefficients; the data block comprises an array of linear transform coefficients; The method of claim 1.

7. the two or more transform kernels among the plurality of transform kernels that share the N high-frequency basis vectors correspond to the same one of a plurality of intra-picture prediction modes; all transform kernels of the plurality of transform kernels assigned to the same one of the plurality of intra-picture prediction modes share the N high-frequency basis vectors while other low-frequency basis vectors are individualized; 7. The method according to any one of claims 1 to 6.

8. The method of claim 1 , wherein the two or more transform kernels among the plurality of transform kernels that share the N high-frequency basis vectors are assigned to two or more different intra-picture prediction modes.

9. the plurality of transformation kernels are quadratic transformation kernels; the two or more transform kernels of the plurality of transform kernels that share the N high frequency basis vectors are configured to transform linear transform coefficients generated from linear transforms having the same transform type.

7. The method according to any one of claims 1 to 6.

10. 7. The method of claim 1, wherein N is an integer power of two.

11. The method of claim 10 , wherein N is less than a predefined upper limit.

12. The method of claim 1 , wherein N depends on the transform size.

13. The method of claim 1 , wherein the plurality of transform kernels are second-order low-frequency non-separable transform (LFNST) kernels.

14. The method of claim 1 , wherein the plurality of transform kernels are configured for an intra-quadratic transform (IST).

15. The method of claim 1 , wherein the plurality of transformation kernels are configured for line graph transformation (LGT).

16. A device for encoding a video bitstream, comprising:

16. A device for encoding a video bitstream, comprising circuitry configured to cause the device to perform the method of any one of claims 1 to 15.

17. A computer program comprising computer instructions which, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 15.

Citation Information

Patent Citations

  • JPP7451772B