Video encoding and decoding method and device and storage medium

By optimizing intra-frame prediction mode and motion compensation video coding methods, and utilizing quadratic transform indexing and entropy decoding techniques, the problem of insufficient redundancy reduction efficiency in high-resolution and high-frame-rate videos is solved, achieving more efficient video compression and decoding.

CN120935359APending Publication Date: 2025-11-11TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511128952.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2021-06-28
Filing Date
2021-06-29
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies suffer from insufficient efficiency in reducing redundancy in intra-frame prediction and motion compensation, especially in high-resolution and high-frame-rate video compression, resulting in excessive bandwidth and storage requirements.

Method used

A novel video coding method is adopted, which determines the context of the secondary transform index by determining the intra-frame prediction mode, block size, and main transform type. The entropy decoding technique is used to optimize the encoding and decoding process of video blocks, including the use of transform methods such as discrete cosine transform and line graph transform, and motion vector prediction is combined to reduce redundancy.

Benefits of technology

It improves the compression efficiency of video encoding and decoding, reduces bandwidth and storage requirements, and enhances video quality and encoding/decoding efficiency, making it suitable for efficient compression of high-resolution and high-frame-rate videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120935359A_ABST
    Figure CN120935359A_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure provide a method, apparatus, and storage medium including processing circuitry for video decoding. A processing circuit decodes encoding information of a transform block (TB) from an encoded video bitstream. The encoding information indicates one of intra prediction mode information of the TB, a size of the TB, and a main transform type of the TB, and the intra prediction mode information indicates an intra prediction mode of the TB. Processing circuitry determines a context for entropy decoding a secondary transform index based on one of intra prediction mode information of the TB, a size of the TB, and a primary transform type of the TB. The secondary transform index indicates a secondary transform to be performed on the TB in a set of secondary transforms. Processing circuitry entropy decodes the secondary transform index based on the context and performs the secondary transform.
Need to check novelty before this filing date? Find Prior Art

Description

References merged

[0001] This application claims priority to U.S. Patent Application No. 17 / 360,431, filed June 28, 2021, entitled “Video Encoding / Decoding Method and Apparatus”, and U.S. Provisional Application No. 63 / 112,529, filed November 11, 2020, entitled “Context Design for Entropy Encoding / Decoding of Quadratic Transform Indexes”. The entire disclosure of the earlier applications is incorporated herein by reference. Technical Field

[0002] This disclosure describes embodiments that generally relate to video encoding and decoding. Background Technology

[0003] The background description provided herein is intended to provide a general overview of the context of this disclosure. The extent to which the work of the currently attributed inventors is described in this background section, and aspects of the description that may not constitute prior art at the time of filing, are neither explicitly nor implicitly considered to be prior art of this disclosure.

[0004] Video encoding and decoding can be performed using inter-frame prediction with motion compensation. Uncompressed digital video can comprise a series of pictures, each with spatial dimensions, for example, 1920×1080 luma samples and associated chroma samples. The series of pictures can have a fixed or variable picture rate (also informally referred to as the frame rate), such as 60 pictures per second or 60Hz. Uncompressed video has specific bitrate requirements. For example, 1080p60 4:2:0 video with 8 bits per sample (1920×1080 luma sample resolution at a frame rate of 60Hz) requires close to 1.5 Gbit / s of bandwidth. One hour of such video would require more than 600 GB of storage space.

[0005] One objective of video encoding and decoding is to reduce redundancy in the input video signal through compression. Compression can help reduce the aforementioned bandwidth and / or storage requirements, in some cases by two orders of magnitude or more. Lossless compression and lossy compression, as well as combinations thereof, can be employed. Lossless compression refers to a technique that can reconstruct an exact copy of the original signal from the compressed original signal. When using lossy compression, the reconstructed signal may differ from the original signal, but the distortion between the original and reconstructed signals is small enough that the reconstructed signal can be used for the intended application. In the case of video, lossy compression is widely used. The amount of distortion that is tolerated depends on the application; for example, users of some consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio can be reflected in the fact that higher permissible / acceptable distortion can result in a higher compression ratio.

[0006] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transform, quantization, and entropy encoding / decoding.

[0007] Video codec techniques can include techniques called intra-frame coding and decoding. In intra-frame coding and decoding, sample values ​​are represented without reference to samples or other data from previously reconstructed reference images. In some video codecs, an image is spatially subdivided into sample blocks. When all sample blocks are encoded in intra-frame mode, the image can be an intra-frame image. Intra-frame images and their derived images (such as independent decoder refresh images) can be used to reset the decoder state and therefore can be used as the first image in an encoded video stream and video session, or as a still image. Samples of an intra-frame block can be exposed to a transform, and the transform coefficients can be quantized before entropy coding and decoding. Intra-frame prediction can be a technique that minimizes the sample values ​​in the pre-transform domain. In some cases, the smaller the transformed DC value and the smaller the AC coefficients, the fewer bits are needed to represent the block after entropy coding and decoding at a given quantization step size.

[0008] Traditional intra-frame encoding and decoding techniques, such as those known from MPEG-2 generation codecs, do not use intra-frame prediction. However, some newer video compression techniques include techniques that attempt to extract data blocks from, for example, surrounding sample data and / or metadata, which are obtained during the encoding and / or decoding of spatially adjacent data blocks and precede the data blocks in the decoding order. Such techniques are hereinafter referred to as "intra-frame prediction" techniques. Note that in at least some cases, intra-frame prediction uses only reference data from the current picture being reconstructed, and not reference data from a reference picture.

[0009] There can be many different forms of intra-frame prediction. When more than one such technique can be used in a given video codec, the techniques used can be encoded and decoded in an intra-frame prediction mode. In some cases, a mode can have sub-modes and / or parameters, and these sub-modes and / or parameters can be encoded and decoded individually or included in the mode codeword. Such a codeword used for a given combination of mode, sub-mode, and / or parameters can affect the codec efficiency gain through intra-frame prediction, and therefore can affect the entropy codec technique used to convert the codeword into a bitstream.

[0010] A certain mode of intra-frame prediction was introduced with H.264, improved in H.265, and further refined in newer codec techniques such as the Joint Exploratory Model (JEM), Universal Video Coding (VVC), and Baseline Set (BMS). Predictor blocks can be formed using neighboring sample values ​​belonging to already available samples. The sample values ​​of neighboring samples are copied into the predictor block according to the orientation. The reference to the orientation can be encoded in the bitstream or can be predicted itself.

[0011] Referring to Figure 1A, a subset of nine predictor directions known from the 33 possible predictor directions of H.265 (corresponding to the 33 angular modes of the 35 intra-frame modes) is depicted in the lower right. The point (101) where the arrows converge represents the sample being predicted. The arrow indicates the direction in which the sample is being predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples in the upper right at a 45-degree angle to the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples in the lower left of sample (101) at a 22.5-degree angle to the horizontal.

[0012] Referring again to Figure 1A, a square block (104) of 4×4 samples is depicted in the upper left (indicated by the dashed bold line). The square block (104) comprises 16 samples, each labeled with “S”, its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (from the top) and the first sample in the X dimension (from the left). Similarly, sample S44 is the fourth sample in both the Y and X dimensions of the block (104). Since the block size is 4×4 samples, S44 is in the lower right. Reference samples following a similar numbering scheme are further shown. Reference samples are labeled with R, their Y position relative to the block (104) (e.g., row index), and X position (column index). In H.264 and H.265, the blocks in the reconstruction are predicted from the nearest neighbor; therefore, negative values ​​are not required.

[0013] Intra-frame image prediction can work by copying reference sample values ​​from neighboring samples appropriate to the prediction direction, as indicated by a signal. For example, suppose an encoded video stream includes signaling that, for a block, indicates a prediction direction consistent with arrow (102)—that is, predicting samples from one or more prediction samples at a 45-degree angle to the horizontal. In this case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Sample S44 is then predicted from reference sample R08.

[0014] In some cases, the values ​​of multiple reference samples can be combined, for example, by interpolation, in order to calculate the reference sample; especially when the direction is not divisible by 45 degrees.

[0015] With the development of video codec technology, the number of possible directions has increased. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013), and JEM / VVC / BMS, when publicly released, could support up to 65 directions. Experiments have been conducted to identify the most likely directions, and certain techniques from entropy coding have been used to represent those possible directions with a small number of bits, accepting some penalty for less likely directions. Furthermore, sometimes the direction itself can be predicted based on adjacent directions used in adjacent decoded blocks.

[0016] Figure 1B shows a schematic diagram (180) depicting 65 intra-frame prediction directions according to JEM, illustrating the increase in the number of prediction directions over time.

[0017] This represents the mapping of intra-prediction direction bits in the encoded video bitstream, which can have directions different from the video codec technique, to the video codec technique; and it can be, for example, a simple direct mapping from prediction direction to intra-prediction mode, to codeword, to complex adaptive schemes involving the most probable mode, and similar techniques. However, in all cases, there may be certain directions in the video content that are statistically less likely to occur than some other directions. Since the goal of video compression is to reduce redundancy, in well-functioning video codec techniques, those less likely directions will be represented by a larger number of bits than the more likely directions.

[0018] Motion compensation can be a lossy compression technique and can involve techniques where sample data blocks from a previously reconstructed image or a portion thereof (the reference image) are used to predict a newly reconstructed image or image portion after being spatially shifted to a direction indicated by a motion vector (hereinafter referred to as MV). In some cases, the reference image can be the same as the image currently being reconstructed. The MV can have two dimensions, X and Y, or three dimensions, with the third dimension being an indication of the reference image in use (the latter could indirectly be a temporal dimension).

[0019] In some video compression techniques, a motion vector (MV) applicable to a region of sample data can be predicted from other MVs, such as those related to another region of sample data in spatially adjacent reconstructed regions and preceding that MV in the decoding order. This substantially reduces the amount of data required to encode and decode MVs, thereby eliminating redundancy and increasing compression. MV prediction can work efficiently, for example, because when encoding and decoding an input video signal derived from a camera (called natural video), there is a statistical likelihood that a region larger than the area to which a single MV can be applied moves in a similar direction, and therefore, in some cases, similar motion vectors derived from MVs of neighboring regions can be used for prediction. This results in an MV found for a given region being similar to or the same as an MV predicted from surrounding MVs, and after entropy encoding and decoding, this can be represented with fewer bits than would be used if the MV were directly encoded and decoded. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., the MV) derived from the original signal (i.e., the sample stream). In other cases, MV prediction itself may be lossy, for example, due to rounding errors when calculating predictions from several surrounding MVs.

[0020] H.265 / HEVC (ITU-T H.265 Recommendation, “High Efficiency Video Coding”, December 2016) describes various MV prediction mechanisms. Among the various MV prediction mechanisms provided by H.265, this application describes the technique hereinafter referred to as “spatial combining”.

[0021] Referring to Figure 2, the current block (201) includes samples discovered by the encoder during the motion search process, which can be predicted based on previous blocks of the same size that have generated spatial offsets. Alternatively, the MV can be derived from metadata associated with one or more reference images, rather than being directly encoded. For example, using the MV associated with any of the five surrounding samples A0, A1 and B0, B1, B2 (corresponding to 202 to 206 respectively), the MV is derived from the metadata of the nearest reference image (in decoding order). In H.265, MV prediction can use predictions from the same reference image that is also being used in adjacent blocks. Summary of the Invention

[0022] This disclosure provides methods and apparatus for video encoding and / or decoding. In some examples, an apparatus for video decoding includes processing circuitry. The processing circuitry can decode encoded information of a transform block (TB) from an encoded video bitstream. The encoded information may indicate one of intra-frame prediction mode information of the TB, the size of the TB, and the primary transform type of the TB. The intra-frame prediction mode information of the TB may indicate an intra-frame prediction mode of the TB. The processing circuitry may determine a context for entropy decoding of a secondary transform index based on one of the intra-frame prediction mode information of the TB, the size of the TB, and the primary transform type of the TB. The secondary transform index may indicate a secondary transform to be performed on the TB from a set of secondary transforms. The processing circuitry may perform entropy decoding of the secondary transform index based on the context and perform the secondary transform indicated by the secondary transform index on the TB.

[0023] In an embodiment, the size of the TB can be indicated by one of the intra-frame prediction mode information, the size of the TB, and the primary transform type of the TB. The processing circuitry can determine the context for entropy decoding of the secondary transform index based on the size of the TB. In the example, the size of the TB indicates the width W and height H of the TB, the minimum of the width W and height H being L, and the processing circuitry can determine the context based on L or L×L.

[0024] In an embodiment, one of the intra-frame prediction mode information of the TB, the size of the TB, and the main transform type of the TB can indicate the intra-frame prediction mode information of the TB. The processing circuitry can determine the context for entropy decoding of the secondary transform index based on the intra-frame prediction mode information of the TB.

[0025] In the example, the intra-frame prediction mode information of the TB indicates the nominal mode index, the TB is predicted by a directional prediction mode determined based on the nominal mode index and the angle offset, and the processing circuitry can determine the context for entropy decoding of the secondary transform index based on the nominal mode index.

[0026] In the example, the intra-frame prediction mode information of the TB indicates the nominal mode index. The TB is predicted using a directional prediction mode determined based on the nominal mode index and an angle offset. The processing circuitry can determine the context for entropy decoding of the secondary transform index based on the index value associated with the nominal mode index.

[0027] In the example, the intra-frame prediction mode information of the TB indicates a non-directional prediction mode index. The TB makes predictions using the non-directional prediction mode indicated by the non-directional prediction mode index. The processing circuitry can determine the context for entropy decoding of the quadratic transform index based on the non-directional prediction mode index.

[0028] In the example, the intra-frame prediction mode information of the TB indicates the recursive filtering mode used to predict the TB. The processing circuitry can determine the nominal mode index based on the recursive filtering mode. The nominal mode index can indicate the nominal mode. The processing circuitry can determine the context used for entropy decoding of the secondary transform index based on the nominal mode index.

[0029] In an embodiment, one of the intra-frame prediction mode information of the TB, the size of the TB, and the main transform type of the TB can indicate the main transform type of the TB. The processing circuitry can determine the context for entropy decoding of the secondary transform index based on the main transform type of the TB.

[0030] The primary transform indicated by the primary transform type can include a horizontal transform indicated by a horizontal primary transform type and a vertical transform indicated by a vertical primary transform type. In the example, the processing circuitry determines the context for entropy decoding of the quadratic transform index based on whether both the horizontal and vertical primary transform types are Discrete Cosine Transform (DCT) or Asymmetric Discrete Sine Transform (ADST).

[0031] In the example, the processing circuitry determines the context for entropy decoding of the quadratic transform index based on whether both the horizontal and vertical main transform types are discrete cosine transforms (DCTs) or line graph transforms (LGTs).

[0032] In the example, the processing circuit determines the context for entropy decoding of the quadratic transform index based on whether the horizontal principal transform type and the vertical principal transform type are (i) both discrete cosine transform (DCT), (ii) both line graph transform (LGT), (iii) DCT and LGT respectively, or (iv) LGT and DCT respectively.

[0033] In the example, the processing circuit determines the context for entropy decoding of the quadratic transform index based on whether the horizontal principal transform type and the vertical principal transform type are (i) both discrete cosine transform (DCT), (ii) both line graph transform (LGT), (iii) DCT and identity transform (IDTX) respectively, or (iv) IDTX and DCT respectively.

[0034] Various aspects of this disclosure also provide a non-volatile computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to perform methods for video decoding and / or encoding. Attached Figure Description

[0035] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:

[0036] Figure 1A is a schematic diagram of an exemplary subset of intra-prediction modes.

[0037] Figure 1B is an illustration of an exemplary intra-frame prediction direction.

[0038] Figure 2 is a schematic diagram of the current block and its surrounding space merge candidates in an example.

[0039] Figure 3 This is a simplified block diagram of a communication system (300) according to an embodiment.

[0040] Figure 4 This is a simplified block diagram of a communication system (400) according to an embodiment.

[0041] Figure 5 This is a simplified block diagram of the decoder according to an embodiment.

[0042] Figure 6 This is a schematic diagram of a simplified block diagram of an encoder according to an embodiment.

[0043] Figure 7 A block diagram of an encoder according to another embodiment is shown.

[0044] Figure 8 A block diagram of a decoder according to another embodiment is shown.

[0045] Figure 9 An example of the nominal pattern of an encoded block according to an embodiment of the present disclosure is shown.

[0046] Figure 10 Examples of non-directional smoothing intra-frame prediction according to various aspects of this disclosure are shown.

[0047] Figure 11 An example of an intra-frame predictor based on recursive filtering according to an embodiment of the present disclosure is shown.

[0048] Figure 12 An example of multiple reference lines for an encoded block according to an embodiment of the present disclosure is shown.

[0049] Figure 13 An example of transform block partitioning on a block according to an embodiment of this disclosure is shown.

[0050] Figure 14 An example of transform block partitioning on a block according to an embodiment of this disclosure is shown.

[0051] Figure 15 An example of the principal transformation basis function according to an embodiment of the present disclosure is shown.

[0052] Figure 16A Exemplary dependencies on the availability of various transform kernels based on transform block size and prediction mode according to embodiments of this disclosure are shown.

[0053] Figure 16B An exemplary transform type selection based on an intra-prediction mode according to an embodiment of the present disclosure is shown.

[0054] Figure 16C An example of a general line graph transformation (LGT) characterized by self-loop weights and edge weights according to an embodiment of the present disclosure is shown.

[0055] Figure 16D An exemplary generalized graph Laplace (GGL) matrix according to an embodiment of this disclosure is shown.

[0056] Figures 17 to 18 Examples of two transform encoding / decoding processes (1700) and (1800) using 16×64 transform and 16×48 transform respectively, according to embodiments of the present disclosure, are shown.

[0057] Figure 19 A flowchart outlining a process (1900) according to an embodiment of the present disclosure is shown.

[0058] Figure 20 This is a schematic diagram of a computer system according to an embodiment. Detailed Implementation

[0059] Figure 3 This is a simplified block diagram of a communication system (300) according to an embodiment disclosed in this application. The communication system (300) includes a plurality of terminal devices that can communicate with each other via, for example, a network (350). For example, the communication system (300) includes a first terminal device (310) and a second terminal device (320) interconnected via a network (350). Figure 3 In this embodiment, the first terminal device (310) and the second terminal device (320) perform unidirectional data transmission. For example, the first terminal device (310) may encode video data (e.g., a video image stream captured by the terminal device (310)) for transmission over a network (350) to the second terminal device (320). The encoded video data is transmitted in the form of one or more encoded video streams. The second terminal device (320) may receive the encoded video data from the network (350), decode the encoded video data to recover the video data, and display video images based on the recovered video data. Unidirectional data transmission is common in applications such as media services.

[0060] In another embodiment, the communication system (300) includes a third terminal device (330) and a fourth terminal device (340) that perform bidirectional transmission of encoded video data, which may occur, for example, during a video conference. For bidirectional data transmission, each of the third terminal device (330) and the fourth terminal device (340) may encode video data (e.g., a stream of video images captured by the terminal device) for transmission over a network (350) to the other terminal device. Each of the third terminal device (330) and the fourth terminal device (340) may also receive encoded video data transmitted by the other terminal device and may decode the encoded video data to recover the video data, and may display the video images on an accessible display device based on the recovered video data.

[0061] exist Figure 3 In the embodiments disclosed herein, the terminal devices (310), (320), (330), and (340) may be servers, personal computers, and smartphones, but the principles disclosed herein are not limited thereto. The embodiments disclosed herein are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network (350) refers to any number of networks that transmit encoded video data between the terminal devices (310), (320), (330), and (340), including, for example, wired (connected) and / or wireless communication networks. The communication network (350) may exchange data in circuit-switched and / or packet-switched channels. This network may include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this application, unless explained below, the architecture and topology of the network (350) may be irrelevant to the operation of this application.

[0062] As an example, Figure 4 The diagram illustrates the placement of a video encoder and a video decoder in a streaming environment. The subject matter disclosed in this application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0063] The streaming system may include an acquisition subsystem (413) that may include a video source (401) such as a digital camera, which creates an uncompressed video image stream (402). In an embodiment, the video image stream (402) includes samples captured by a digital camera. The video image stream (402) is depicted as a thick line to emphasize the high data volume of the video image stream compared to encoded video data (404) (or encoded video bitstream). The video image stream (402) may be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination of hardware and software to implement or carry out aspects of the disclosed subject matter as described in more detail below. Compared to the video image stream (402), the encoded video data (404) (or the encoded video bitstream (404)) is depicted as a thin line to emphasize the lower data volume of the encoded video data (404) (or the encoded video bitstream (404)), which can be stored on a streaming server (405) for future use. One or more streaming client subsystems, such as Figure 3 Client subsystems (406) and (408) can access a streaming server (405) to retrieve copies (407) and (409) of encoded video data (404). Client subsystem (406) may include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes the incoming copy (407) of the encoded video data and produces an output video picture stream (411) that can be displayed on a display (412) (e.g., a screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (404), video data (407), and video data (409) (e.g., video streams) may be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T H.265. In embodiments, the video coding standard under development is informally referred to as Versatile Video Coding (VVC), and this application can be used in the context of the VVC standard.

[0064] It should be noted that electronic devices (420) and (430) may include other components (not shown). For example, electronic device (420) may include a video decoder (not shown), and electronic device (430) may also include a video encoder (not shown).

[0065] Figure 5This is a block diagram of a video decoder (510) according to an embodiment disclosed in this application. The video decoder (510) may be disposed in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., receiving circuitry). The video decoder (510) may be used in place of... Figure 4 The video decoder (410) in the embodiment.

[0066] The receiver (531) may receive one or more encoded video sequences to be decoded by the video decoder (510); in the same embodiment or another embodiment, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the other encoded video sequences. The encoded video sequences may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (531) may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective user entities (not indicated). The receiver (531) may separate the encoded video sequences from other data. To prevent network jitter, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) is part of the video decoder (510). In other cases, the buffer memory (515) may be located external to the video decoder (510) (not indicated). In other cases, an external buffer (not shown) may be provided for the video decoder (510) to prevent network jitter, for example, and another buffer (515) may be configured internally for, for example, handling broadcast timing. When the receiver (531) receives data from a store / forward device with sufficient bandwidth and controllability or from an isochronous synchronization network, the buffer (515) may not be necessary, or it may be made smaller. Of course, for use on packet networks such as the Internet, a buffer (515) may be required; this buffer may be relatively large and adaptive in size, and may be at least partially implemented in the operating system or a similar component (not shown) external to the video decoder (510).

[0067] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the encoded video sequence. These symbols may include information for managing the operation of the video decoder (510) and potential information for controlling a display device (512) (e.g., a display screen), which is not part of the electronic device (530) but may be coupled to it, such as... Figure 5As shown in the diagram. The control information for the display device may be a parameter set fragment (not shown) of Supplemental Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (520) may parse / decode the received encoded video sequence. The encoding of the encoded video sequence may be based on video coding techniques or standards and may follow various principles, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) may extract a subgroup parameter set of at least one subgroup of pixels in the subgroup of pixels for use in the video decoder based on at least one parameter corresponding to a group. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), etc. The parser (520) can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0068] The parser (520) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to create symbols (521).

[0069] Depending on the type of encoded video frames or a subset of encoded video frames (e.g., inter-frame and intra-frame frames, inter-frame and intra-frame blocks) and other factors, the reconstruction of the symbol (521) may involve multiple different units. Which units are involved and how they are involved can be controlled by the subgroup control information parsed from the encoded video sequence by the parser (520). For brevity, the flow of such subgroup control information between the parser (520) and the various units described below is not described.

[0070] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually subdivided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and can be integrated with one another. However, for the purposes of describing the disclosed subject matter, it is appropriate to conceptually subdivide it into the functional units described below.

[0071] The first unit is the scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives quantization transform coefficients as symbols (521) and control information from the parser (520), including the transform method used, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (551) can output a block containing sample values, which can be input into the aggregator (555).

[0072] In some cases, the output samples of the scaler / inverse transform unit (551) may belong to intra-coded blocks; that is, blocks that do not use predictive information from previously reconstructed images but can use predictive information from previously reconstructed portions of the current image. Such predictive information may be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) uses reconstructed information extracted from the current picture buffer (558) to generate surrounding blocks of the same size and shape as the block being reconstructed. For example, the current picture buffer (558) buffers partially reconstructed and / or fully reconstructed current images. In some cases, the aggregator (555) adds the predictive information generated by the intra-picture prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) based on each sample.

[0073] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to inter-frame coding and latent motion compensation blocks. In this case, the motion compensation prediction unit (553) can access the reference image memory (557) to extract samples for prediction. After motion compensation is performed on the extracted samples according to the symbols (521), these samples can be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (referred to as residual samples or residual signals in this case) to generate output sample information. The motion compensation prediction unit (553) can obtain the predicted samples from the address in the reference image memory (557) under motion vector control, and the motion vector is available to the motion compensation prediction unit (553) in the form of the symbols (521), which, for example, include X, Y and reference image components. Motion compensation may also include interpolation of sample values ​​extracted from the reference image memory (557) when using subsample precise motion vectors, motion vector prediction mechanisms, etc.

[0074] The output samples of the aggregator (555) can be employed by various loop filtering techniques in the loop filter unit (556). Video compression techniques may include in-loop filtering techniques controlled by parameters included in the encoded video sequence (also referred to as the encoded video stream), and these parameters can be used as symbols (521) from the parser (520) in the loop filter unit (556). However, in other embodiments, the video compression techniques may also respond to metadata obtained during decoding of a previous (in decoding order) portion of the encoded image or encoded video sequence, and to previously reconstructed and loop-filtered sample values.

[0075] The output of the loop filter unit (556) can be a sample stream, which can be output to a display device (512) and stored in a reference image memory (557) for subsequent inter-frame image prediction.

[0076] Once fully reconstructed, some of the encoded images can be used as reference images for future predictions. For example, once the encoded images corresponding to the current image have been fully reconstructed and the encoded images (by, for example, the parser (520)) are identified as reference images, the current image buffer (558) can become part of the reference image memory (557), and a new current image buffer can be reallocated before the reconstruction of subsequent encoded images begins.

[0077] The video decoder (510) can perform decoding operations according to a predetermined video compression technique, such as that specified in the ITU-T H.265 standard. The encoded video sequence may conform to the syntax specified by the video compression technique or standard in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the configuration file recorded in the video compression technique or standard. Specifically, the configuration file may select certain tools from all available tools in the video compression technique or standard as the only tools available under said configuration file. For compliance, the complexity of the encoded video sequence is also required to be within the range defined by the hierarchy of the video compression technique or standard. In some cases, the hierarchy limits the maximum image size, maximum frame rate, maximum reconstruction sampling rate (measured in megasamples per second, for example), maximum reference image size, etc. In some cases, the limitations set by the hierarchy can be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata managed by the HRD buffer, which is represented by signals in the encoded video sequence.

[0078] In this embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be a portion of the encoded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, forward error correction codes, etc.

[0079] Figure 6 This is a block diagram of a video encoder (603) according to an embodiment disclosed in this application. The video encoder (603) is disposed in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used to replace... Figure 4 The video encoder (403) in the embodiment.

[0080] The video encoder (603) can obtain data from the video source (601) (not) Figure 6 In one embodiment, a portion of the electronic device (620) receives video samples, the video source being capable of capturing video images to be encoded by a video encoder (603). In another embodiment, the video source (601) is a portion of the electronic device (620).

[0081] A video source (601) can provide a sequence of source video samples encoded by a video encoder (603) in the form of a digital video sample stream, which can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601YCrCb, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (601) can be a storage device storing previously prepared video. In a video conferencing system, the video source (601) can be a camera that captures local image information as a video sequence. Video data can be provided as multiple individual pictures, which are given motion when viewed in sequence. The pictures themselves can be constructed as spatial pixel arrays, where each pixel can include one or more samples depending on the sampling structure, color space, etc., used. Those skilled in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.

[0082] According to an embodiment, the video encoder (603) can encode and compress images of a source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by the application. Implementing an appropriate encoding rate is a function of the controller (650). In some embodiments, the controller (650) controls and is functionally coupled to other functional units described below. For simplicity, coupling is not shown in the figures. Parameters set by the controller (650) may include rate control-related parameters (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) may be used with other suitable functions related to the video encoder (603) optimized for a particular system design.

[0083] In some embodiments, the video encoder (603) operates within an encoding loop. For simplicity, in an embodiment, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and a reference image) and a (local) decoder (633) embedded within the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data (because in the video compression techniques considered in this application, any compression between the symbols and the encoded video stream is lossless). The reconstructed sample stream (sample data) is input to a reference image memory (634). Since decoding of the symbol stream produces bit-precise results independent of the decoder's location (local or remote), the contents of the reference image memory (634) also correspond bit-precisely between the local and remote encoders. In other words, the reference image samples "seen" by the encoder's prediction portion are exactly the same sample values ​​that the decoder will "see" when using the prediction during decoding. This fundamental principle of reference image synchronization (and the drift that occurs when synchronization cannot be maintained, for example, due to channel errors) is also used in some related technologies.

[0084] The operation of the “local” decoder (633) can be combined with, for example, the above-mentioned Figure 5 The video decoder (510) is described in detail as the same as the "remote" decoder. However, a further brief reference is provided. Figure 5 When symbols are available and the entropy encoder (645) and parser (520) are able to encode / decode the symbols into an encoded video sequence without loss, the entropy decoding portion of the video decoder (510), including the buffer (515) and parser (520), may not be fully implemented in the local decoder (633).

[0085] It can be observed that any decoder technique other than parsing / entropy decoding present in the decoder must also exist in the corresponding encoder in essentially the same functional form. For this reason, this application focuses on decoder operation. The description of encoder techniques can be simplified because encoder techniques are inverses of the fully described decoder techniques. More detailed descriptions are only required in certain areas, and are provided below.

[0086] During operation, in some embodiments, the source encoder (630) may perform motion-compensated predictive coding. The motion-compensated predictive coding predictively encodes the input image, referencing one or more previously encoded images from the video sequence designated as "reference images." In this manner, the encoding engine (632) encodes the differences between pixel blocks of the input image and pixel blocks of the reference image, which may be selected as a predictive reference for the input image.

[0087] The local video decoder (633) can decode encoded video data of a picture that can be designated as a reference picture, based on symbols created by the source encoder (630). The operation of the encoding engine (632) can be a lossy process. When the encoded video data can be decoded by the video decoder (633), Figure 6 When the source video sequence (not shown) is decoded, the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process, which can be performed by the video decoder on the reference image, and allows the reconstructed reference image to be stored in a reference image cache (634). In this way, the video encoder (603) can locally store a copy of the reconstructed reference image that shares the same content (no transmission errors) as the reconstructed reference image to be obtained by the remote video decoder.

[0088] The predictor (635) can perform a prediction search against the encoding engine (632). That is, for a new image to be encoded, the predictor (635) can search in the reference image memory (634) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that can serve as appropriate prediction references for the new image. The predictor (635) can operate pixel-by-pixel based on the sample blocks to find suitable prediction references. In some cases, based on the search results obtained by the predictor (635), it can be determined that the input image may have prediction references obtained from multiple reference images stored in the reference image memory (634).

[0089] The controller (650) can manage the encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding video data.

[0090] The outputs of all the above-mentioned functional units can be entropy encoded in the entropy encoder (645). The entropy encoder (645) performs lossless compression on the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, and arithmetic coding, thereby converting the symbols into an encoded video sequence.

[0091] The transmitter (640) can buffer the encoded video sequence created by the entropy encoder (645) in preparation for transmission via a communication channel (660), which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter (640) can combine the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0092] The controller (650) manages the operation of the video encoder (603). During encoding, the controller (650) can assign a specific encoded image type to each encoded image, but this may affect the encoding techniques applicable to the corresponding images. For example, images can typically be assigned to any of the following image types:

[0093] An intra-frame picture (I-picture) is a picture that can be encoded and decoded without using any other pictures in the sequence as a prediction source. Some video codecs allow different types of intra-frame pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art will understand variations of I-pictures and their corresponding applications and characteristics.

[0094] A predictive image (P-image) can be an image that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses at most one motion vector and a reference index to predict sample values ​​for each block.

[0095] A bidirectional predictive image (B-image) can be an image that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses at most two motion vectors and a reference index to predict sample values ​​for each block. Similarly, multiple predictive images can use more than two reference images and associated metadata to reconstruct a single block.

[0096] Source images are typically spatially subdivided into multiple sample blocks (e.g., 4×4, 8×8, 4×8, or 16×16 sample blocks), and encoded block by block. These blocks can be predictively coded with reference to other (already coded) blocks, determined based on the coding assignments of the corresponding images applied to the blocks. For example, blocks of an I-image can be non-predictively coded, or the blocks can be predictively coded (spatial or intra-frame prediction) with reference to already coded blocks of the same image. Pixel blocks of a P-image can be predictively coded with reference to a previously coded reference image via spatial or temporal prediction. Blocks of a B-image can be predictively coded with reference to one or two previously coded reference images via spatial or temporal prediction.

[0097] The video encoder (603) can perform encoding operations according to a predetermined video coding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder (603) can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard used.

[0098] In this embodiment, the transmitter (640) may transmit additional data while transmitting encoded video. The source encoder (630) may include such data as part of the encoded video sequence. Additional data may include temporal / spatial / SNR enhancement layers, redundant images and slices, other forms of redundant data, SEI messages, VUI parameter set fragments, etc.

[0099] The acquired video can serve as multiple source images (video images) presented in a time series. Intra-frame image prediction (often simplified to intra-frame prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes (temporal or other) correlations between images. In an embodiment, a specific image being encoded / decoded is segmented into blocks, referred to as the current image. When a block in the current image resembles a reference block in a previously encoded and still buffered reference image in the video, the block in the current image can be encoded using a vector called a motion vector. This motion vector points to the reference block in the reference image, and when multiple reference images are used, the motion vector may have a third dimension that identifies the reference image.

[0100] In some embodiments, bidirectional prediction techniques can be used in inter-frame image prediction. According to bidirectional prediction, two reference images are used, such as a first reference image and a second reference image, both preceding the current image in the video in decoding order (but possibly past and future in display order). A block in the current image can be encoded using a first motion vector pointing to a first reference block in the first reference image and a second motion vector pointing to a second reference block in the second reference image. Specifically, the block can be predicted using a combination of the first and second reference blocks.

[0101] In addition, merging mode techniques can be used in inter-frame image prediction to improve coding efficiency.

[0102] According to some embodiments disclosed in this application, predictions such as inter-frame image prediction and intra-frame image prediction are performed on a block-by-block basis. For example, according to the HEVC standard, images in a video image sequence are segmented into coding tree units (CTUs) for compression. The CTUs in the images have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs): one luma CTB and two chroma CTBs. Furthermore, each CTU can be further subdivided into one or more coding units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be subdivided into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In embodiments, each CU is analyzed to determine the prediction type used for the CU, such as inter-frame prediction or intra-frame prediction. Furthermore, depending on temporal and / or spatial predictability, the CU is divided into one or more prediction units (PUs). Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In embodiments, prediction operations in encoding (encoding / decoding) are performed on a per-prediction-block basis. Taking a luma prediction block as an example, a prediction block includes a matrix of pixel values ​​(e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0103] Figure 7 This is a diagram of a video encoder (703) according to another embodiment disclosed in this application. The video encoder (703) is used to receive processing blocks (e.g., prediction blocks) of sample values ​​within a current video image in a video image sequence, and to encode the processing blocks into an encoded image that is part of an encoded video sequence. In this embodiment, the video encoder (703) is used instead of Figure 4The video encoder (403) in the embodiment.

[0104] In the HEVC embodiment, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as an 8×8 sample prediction block. The video encoder (703) uses, for example, rate-distortion (RD) optimization to determine whether to use an intra-frame mode, an inter-frame mode, or a bidirectional prediction mode to encode the processing block. When encoding the processing block in intra-frame mode, the video encoder (703) can use intra-frame prediction techniques to encode the processing block into an already encoded picture; and when encoding the processing block in inter-frame mode or bidirectional prediction mode, the video encoder (703) can use inter-frame prediction or bidirectional prediction techniques to encode the processing block into an already encoded picture, respectively. In some video coding techniques, the merging mode can be an inter-frame picture prediction sub-mode, in which motion vectors are derived from one or more motion vector prediction values ​​without relying on already encoded motion vector components outside the prediction values. In some other video coding techniques, motion vector components applicable to the subject block may exist. In the embodiment, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the processing block mode.

[0105] exist Figure 7 In one embodiment, the video encoder (703) includes, as shown below: Figure 7 The inter-frame encoder (730), intra-frame encoder (722), residual calculator (723), switch (726), residual encoder (724), general controller (721) and entropy encoder (725) are shown coupled together.

[0106] An inter-frame encoder (730) is configured to receive samples of the current block (e.g., the processing block), compare the block with one or more reference blocks in a reference image (e.g., blocks in previous and later images), generate inter-frame prediction information (e.g., redundancy information description, motion vectors, merging mode information based on inter-frame coding techniques), and calculate inter-frame prediction results (e.g., predicted blocks) based on the inter-frame prediction information using any suitable technique. In some embodiments, the reference image is a decoded reference image based on encoded video information.

[0107] The intra encoder (722) is used to receive samples of the current block (e.g., the processing block), in some cases compare the block with previously encoded blocks in the same image, generate quantization coefficients after transformation, and in some cases also (e.g., based on intra prediction direction information of one or more intra coding techniques) generate intra prediction information. In an embodiment, the intra encoder (722) also calculates intra prediction results (e.g., predicted blocks) based on the intra prediction information and reference blocks in the same image.

[0108] A general controller (721) determines general control data and controls other components of the video encoder (703) based on the general control data. In an embodiment, the general controller (721) determines the mode of a block and provides control signals to a switch (726) based on the mode. For example, when the mode is an intra-frame mode, the general controller (721) controls the switch (726) to select an intra-frame mode result for use by the residual calculator (723) and controls the entropy encoder (725) to select intra-frame prediction information and add the intra-frame prediction information to the bitstream; and when the mode is an inter-frame mode, the general controller (721) controls the switch (726) to select an inter-frame prediction result for use by the residual calculator (723) and controls the entropy encoder (725) to select inter-frame prediction information and add the inter-frame prediction information to the bitstream.

[0109] A residual calculator (723) is used to calculate the difference (residual data) between the received block and the prediction result selected from the intra encoder (722) or the inter encoder (730). A residual encoder (724) is used to operate on the residual data to encode the residual data to generate transform coefficients. In an embodiment, the residual encoder (724) is used to transform the residual data from the time domain to the frequency domain and generate transform coefficients. The transform coefficients are then quantized to obtain quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is used to perform an inverse transform and generate decoded residual data. The decoded residual data can be used appropriately by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and inter-frame prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and intra-frame prediction information. The decoded blocks are processed appropriately to generate a decoded image, and in some embodiments, the decoded image may be buffered in a memory circuit (not shown) and used as a reference image.

[0110] An entropy encoder (725) is used to format the bitstream to produce encoded blocks. The entropy encoder (725) generates various information according to a suitable standard such as the HEVC standard. In an embodiment, the entropy encoder (725) is used to obtain general control data, selected prediction information (e.g., intra-frame prediction information or inter-frame prediction information), residual information, and other suitable information in the bitstream. It should be noted that, according to the disclosed subject matter, residual information is not present when blocks are encoded in a merged sub-mode of inter-frame mode or bidirectional prediction mode.

[0111] Figure 8This is a diagram of a video decoder (810) according to another embodiment disclosed in this application. The video decoder (810) is used to receive an encoded image as part of an encoded video sequence and decode the encoded image to generate a reconstructed picture. In an embodiment, the video decoder (810) is used instead of Figure 4 The video decoder (410) in the embodiment.

[0112] exist Figure 8 In the embodiment, the video decoder (810) includes, as follows: Figure 8 The entropy decoder (871), inter-frame decoder (880), residual decoder (873), reconstruction module (874), and intra-frame decoder (872) are shown coupled together.

[0113] An entropy decoder (871) can be used to reconstruct certain symbols from an encoded image, these symbols representing the syntax elements constituting the encoded image. Such symbols may include, for example, a mode for encoding the block (e.g., intra-frame mode, inter-frame mode, bidirectional prediction mode, a combined sub-mode of the latter two, or another sub-mode), prediction information (e.g., intra-frame prediction information or inter-frame prediction information) that can respectively identify certain samples or metadata used by the intra-frame decoder (872) or the inter-frame decoder (880) for prediction, residual information in the form of, for example, quantized transform coefficients, and so on. In an embodiment, when the prediction mode is inter-frame or bidirectional prediction mode, inter-frame prediction information is provided to the inter-frame decoder (880); and when the prediction type is intra-frame prediction type, intra-frame prediction information is provided to the intra-frame decoder (872). Residual information may be provided to the residual decoder (873) via inverse quantization.

[0114] The inter-frame decoder (880) is used to receive inter-frame prediction information and generate inter-frame prediction results based on the inter-frame prediction information.

[0115] The intra-frame decoder (872) is used to receive intra-frame prediction information and generate prediction results based on the intra-frame prediction information.

[0116] The residual decoder (873) performs inverse quantization to extract the dequantized transform coefficients and processes the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require some control information (to obtain the quantizer parameters QP), and this information may be provided by the entropy decoder (871) (the data path is not indicated because this is only low-level control information).

[0117] The reconstruction module (874) is used to combine the residual output by the residual decoder (873) with the prediction result (which may be output by the inter-frame prediction module or the intra-frame prediction module) in the spatial domain to form a reconstructed block, which may be a part of a reconstructed image, which in turn may be a part of a reconstructed video. It should be noted that other suitable operations, such as deblocking, may be performed to improve visual quality.

[0118] It should be noted that any suitable technology can be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810). In one embodiment, one or more integrated circuits can be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810). In another embodiment, one or more processors executing software instructions can be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810).

[0119] This invention discloses video codec techniques related to the context design of entropy encoding and decoding of indices (such as quadratic transform indices). Indices can be used to identify one of a set of inseparable quadratic transforms used to encode and decode (e.g., encode and / or decode) blocks. The context design of entropy encoding and decoding of indices can be applied to any suitable video codec format or standard. Video codec formats can include open video codec formats designed for video transmission over the Internet, such as AOMedia Video 1 (AV1) or next-generation AOMedia video formats beyond AV1. Video codec standards can include the High Efficiency Video Codec (HEVC) standard or next-generation video codecs beyond HEVC (e.g., Universal Video Codec (VVC)).

[0120] Various intra-frame prediction modes can be used for intra-frame prediction, such as in AV1 and / or VVC. In embodiments, such as in AV1, directional intra-frame prediction is used. In directional intra-frame prediction, predicted samples of a block are generated by extrapolating from adjacent reconstructed samples along a direction. This direction corresponds to an angle. The mode used to predict predicted samples of a block in directional intra-frame prediction can be referred to as a directional mode (also known as a directional prediction mode, directional intra-frame mode, directional intra-frame prediction mode, angle mode). Each directional mode can correspond to a different angle or a different direction. In examples, such as in the open video codec format VP9, ​​eight directional modes correspond to eight angles from 45° to 207°, such as... Figure 9As shown in the diagram. The eight orientation modes can also be referred to as nominal modes (e.g., V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED). To utilize more types of spatial redundancy in oriented textures (e.g., in AV1), the orientation modes can be extended (e.g., beyond eight nominal modes) to an angle set with finer granularity and more angles (or directions), such as... Figure 9 As shown in the image.

[0121] Figure 9 An example of a nominal mode of a coded block (CB) (910) according to an embodiment of the present disclosure is shown. Certain angles (also referred to as nominal angles) may correspond to nominal modes. In the example, eight nominal angles (or nominal intra-frame angles) (901)-(908) correspond to eight nominal modes (e.g., V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED), respectively. The eight nominal angles (901)-(908) and the eight nominal modes may be referred to as V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED, respectively. A nominal mode index may indicate a nominal mode (e.g., one of the eight nominal modes). In the example, the nominal mode index is signaled.

[0122] Furthermore, each nominal angle can correspond to multiple finer angles (e.g., seven finer angles), and thus, for example, in AV1, 56 angles (or prediction angles) or 56 directional patterns (or angle patterns, directional intra-prediction patterns) can be used. Each prediction angle can be represented by a nominal angle and an angle offset (or angle increment). The angle offset can be obtained by multiplying an offset integer I (e.g., -3, -2, -1, 0, 1, 2, or 3) by a step size (e.g., 3°). In the example, the prediction angle is equal to the sum of the nominal angle and the angle offset. In the example, such as in AV1, the nominal patterns (e.g., eight nominal patterns (901)-(908)) can be signaled along with certain non-angle-smoothing patterns (e.g., five non-angle-smoothing patterns as described below, such as DC pattern, PAETH pattern, SMOOTH pattern, vertical SMOOTH pattern, and horizontal SMOOTH pattern). Subsequently, if the current prediction mode is a directional mode (or an angle mode), the index can be further signaled to indicate the angle offset corresponding to the nominal angle (e.g., an offset integer I). In the example, the directional mode (e.g., one of 56 directional modes) can be determined based on the nominal mode index and an index indicating the angle offset relative to the nominal mode. In the example, to implement the directional prediction mode in a general manner, such as the 56 directional modes used in AV1, a unified directional predictor is used, which projects each pixel to a reference subpixel location and interpolates the reference pixel through a 2-tap bilinear filter.

[0123] Non-directional smoothing intra-predictors (also known as non-directional smoothing intra-prediction modes, non-directional smoothing modes, or non-angular smoothing modes) can be used for intra-prediction of blocks (such as blocks). In some examples (e.g., in AV1), five non-directional smoothing intra-prediction modes include DC mode or DC predictor (e.g., DC), PAETH mode or PAETH predictor (e.g., PAETH), SMOOTH mode or SMOOTH predictor (e.g., SMOOTH), vertical SMOOTH mode (referred to as SMOOTH_V mode, SMOOTH_V predictor, or SMOOTH_V), and horizontal SMOOTH mode (referred to as SMOOTH_H mode, SMOOTH_H predictor, or SMOOTH_H).

[0124] Figure 10Examples of non-directional smooth intra-frame prediction modes (e.g., DC mode, PAETH mode, SMOOTH mode, SMOOTH_V mode, and SMOOTH_H mode) according to various aspects of this disclosure are shown. In order to predict a sample (1001) in CB (1000) based on the DC predictor, the average of the first value of the left neighboring sample (1012) and the second value of the upper neighboring sample (or top neighboring sample) (1011) can be used as the predictor.

[0125] To predict sample (1001) based on the PAETH predictor, we can obtain the first value of the left neighboring sample (1012), the second value of the top neighboring sample (1011), and the third value of the top-left neighboring sample (1013). Then, we use Equation 1 to obtain the reference value. Reference value = First value + Second value - Third value (Equation 1)

[0126] The value that is closest to the reference value among the first, second, and third values ​​can be set as the predictor (1001) for the sample.

[0127] The SMOOTH_V, SMOOTH_H, and SMOOTH modes can be used to predict CB(1000) using quadratic interpolation in the vertical direction, the horizontal direction, and the average direction of the vertical and horizontal directions, respectively. To predict sample (1001) based on the SMOOTH predictor, the average (e.g., a weighted combination) of the first value, the second value, the value of the right sample (1014), and the value of the bottom sample (1016) can be used. In various examples, the right sample (1014) and the bottom sample (1016) are not reconstructed, and therefore, the values ​​of the upper-right neighboring sample (1015) and the lower-left neighboring sample (1017) can replace the values ​​of the right sample (1014) and the bottom sample (1016), respectively. Accordingly, the average (e.g., a weighted combination) of the first value, the second value, the value of the upper-right neighboring sample (1015), and the value of the lower-left neighboring sample (1017) can be used as the SMOOTH predictor. To predict sample (1001) based on the SMOOTH_V predictor, the average of the second value of the top neighbor sample (1011) and the value of the bottom left neighbor sample (1017) can be used (e.g., a weighted combination). To predict sample (1001) based on the SMOOTH_H predictor, the average of the first value of the left neighbor sample (1012) and the value of the top right neighbor sample (1015) can be used (e.g., a weighted combination).

[0128] Figure 11An example of a recursive filtering-based intra-predictor (also referred to as a filter intra-mode or recursive filtering mode) according to an embodiment of the present disclosure is shown. To capture the attenuation spatial correlation with references on the edges, the filter intra-mode can be used for blocks such as CB (1100). In the example, CB (1100) is a luma block. The luma block (1100) can be divided into multiple patches (e.g., eight 4×2 patches B0-B7). Each of patches B0-B7 can have multiple neighboring samples. For example, patch B0 has seven neighboring samples (or seven neighbors) R00-R06, including four top neighboring samples R01-R04, two left neighboring samples R05-R06, and one top-left neighboring sample R00. Similarly, patch B7 has seven neighboring samples R70-R76, including four top neighboring samples R71-R74, two left neighboring samples R75-R76, and one top-left neighboring sample R70.

[0129] In some examples, such as for AV1, multiple (e.g., five) filter intra-frame patterns (or multiple recursive filter patterns) are pre-designed. Each filter intra-frame pattern can be represented by a set of eight 7-tap filters that reflect the correlation between a sample (or pixel) in the corresponding 4×2 patch (e.g., B0) and its seven neighbors (e.g., R00-R06) adjacent to the 4×2 patch B0. The weighting factors of the 7-tap filters can be position-dependent. For each of patches B0-B7, the seven neighbors (e.g., R00-R06 for B0, R70-R76 for B7) can be used to predict the sample in the corresponding patch. In the example, neighbors R00-R06 are used to predict the sample in patch B0. In the example, neighbors R70-R76 are used to predict the sample in patch B7. For some patches in CB(1100) (such as patch B0), all seven neighbors (e.g., R00-R06) have been reconstructed. For the other patches in CB(1100), at least one of the seven neighbors is not reconstructed, and therefore one or more predictions of one or more direct neighbors (or one or more prediction samples of one or more direct neighbors) can be used as a reference. For example, the seven neighbors R70-R76 of patch B7 are not reconstructed, so prediction samples of direct neighbors can be used.

[0130] Chroma samples can be predicted from luminance samples. In an embodiment, the chroma from a luminance mode (e.g., CfL mode, CfL predictor) is a chroma-only intra-frame predictor that can model chroma samples (or pixels) as a linear function of reconstructed luminance samples (or pixels). For example, CfL prediction can be represented using Equation 2 below. CfL(α)=αL A +D (Equation 2) Among them, L ALet α represent the AC contribution of the luminance component, α represent the scaling parameter of the linear model, and D represent the DC contribution of the chrominance component. In the example, the reconstructed luminance pixels are subsampled based on the chrominance resolution, and the average value is subtracted to form the AC contribution (e.g., L). A To approximate the chroma AC component based on the AC contribution, in some examples, such as AV1, without requiring the decoder to calculate the scaling parameter α, the CfL mode determines the scaling parameter α based on the original chroma pixels and signals the scaling parameter α in the bitstream, thus reducing decoder complexity and producing more accurate predictions. The DC contribution of the chroma component can be calculated using an intra-frame DC mode. Intra-frame DC modes are sufficient for most chroma content and have mature and fast implementations.

[0131] Multi-line intra-prediction can use more reference lines for intra-prediction. Reference lines can include multiple samples from the image. In the example, the reference lines include samples from one row and samples from one column. In the example, the encoder can determine and signal the reference lines used to generate the intra-predictor. The index of the reference lines (also called the reference line index) can be signaled before one or more intra-prediction modes. In the example, only MPM is allowed when a non-zero reference line index is signaled. Figure 12 An example of four reference lines for CB(1210) is shown. (Reference) Figure 12 Reference lines can include up to six segments (e.g., segments A through F) and a top-left reference sample. For example, reference line 0 includes segments B and E and a reference sample in the top-left corner. For example, reference line 3 includes segments A through F and a top-left reference sample. Segments A and F can be filled with the nearest samples from segments B and E, respectively. In some examples, such as in HEVC, only one reference line (e.g., reference line 0 adjacent to CB(1210)) is used for intra-frame prediction. In some examples, such as in VVC, multiple reference lines (e.g., reference lines 0, 1, and 3) are used for intra-frame prediction.

[0132] Typically, references such as those above can be used. Figures 9 to 12 The description refers to one or an appropriate combination of various intra-frame prediction modes to predict blocks.

[0133] Transform block partitioning (also known as transform partitioning or transform unit partitioning) can be implemented to partition a block into multiple TUs or multiple TBs. Figures 13 to 14Exemplary transform block partitioning according to embodiments of the present disclosure is illustrated. In some examples, such as in AV1, both intra-frame coded blocks and inter-frame coded blocks can be further partitioned into multiple transform units having a partitioning depth of up to several levels (e.g., two levels). The multiple transform units obtained from transform block partitioning can be referred to as TBs. Transformations (e.g., primary transforms and / or secondary transforms such as those described below) can be performed on each of the multiple TUs or multiple TBs. Accordingly, multiple transforms can be performed on the block partitioned into multiple transform units or multiple TBs.

[0134] For intra-coded blocks, transform partitioning can be performed so that transform blocks associated with the intra-coded block have the same size, and the transform blocks can be encoded in raster scan order. (Reference) Figure 13 Transform block partitioning can be performed on a block (e.g., an intra-coded block) (1300). The block (1300) can be partitioned into transform units, such as four transform units (e.g., TB) (1301)-(1304), and the partitioning depth is 1. The four transform units (e.g., TB) (1301)-(1304) can have the same size and can be encoded in raster scan order (1310) from transform unit (1301) to transform unit (1304). In examples, different transform kernels are used, for instance, to transform the four transform units (e.g., TB) (1301)-(1304) separately. In some examples, each of the four transform units (e.g., TB) (1301)-(1304) is further partitioned into four transform units. For example, transform unit (1301) is partitioned into transform units (1321), (1322), (1325), and (1326), transform unit (1302) is partitioned into transform units (1323), (1324), (1327), and (1328), transform unit (1303) is partitioned into transform units (1329), (1330), (1333), and (1334), and transform unit (1304) is partitioned into transform units (1331), (1332), (1335), and (1336). The partition depth is 2. Transform units (e.g., TB) (1321)-(1336) can have the same size and can be encoded in the raster scan order (1320) from transform unit (1321) to transform unit (1336).

[0135] For inter-frame coded blocks, transform partitioning can be performed recursively, with partitioning depths reaching multiple levels (e.g., two levels). Transform partitioning can support any suitable transform unit size and shape. Transform unit shapes can include square and non-square shapes (e.g., non-square rectangles) with any suitable aspect ratio. Transform unit sizes can range from 4×4 to 64×64. The aspect ratio of the transform unit (e.g., the ratio of the transform unit's width to its height) can be 1:1 (square), 1:2, 2:1, 1:4, or 4:1, etc. Transform partitioning can support transform unit sizes ranging from 4×4 to 64×64, including 1:1 (square), 1:2, 2:1, 1:4, and / or 4:1. Reference Figure 14 Transform block partitioning (1400) can be performed recursively on blocks (e.g., inter-coded blocks). For example, block (1400) is partitioned into transform units (1401)-(1407). Transform units (e.g., TBs) (1401)-(1407) can have different sizes and can be encoded in raster scan order (1410) from transform unit (1401) to transform unit (1407). In the example, the partition depth of transform units (1401), (1406), and (1407) is 1, and the partition depth of transform units (1402)-(1405) is 2.

[0136] In the example, if the coded block is less than or equal to 64×64, the transform partition can only be applied to the luma component. In the example, the coded block refers to the CTB.

[0137] If the block width W or block height H is greater than 64, the block can be implicitly divided into multiple TBs, where the block is a luma-coded block. The width of one of the TBs can be the minimum of W and 64, and the height of one of the TBs can be the minimum of H and 64.

[0138] If the block width W or block height H is greater than 64, the block can be implicitly divided into multiple TBs, where the block is a chroma-coded block. The width of one of the TBs can be the minimum of W and 32, and the height of one of the TBs can be the minimum of H and 32.

[0139] The following describes embodiments of master transforms such as those used in AOMedia Video 1 (AV1). A forward transform (e.g., in an encoder) can be performed on a transform block (TB) that includes residuals (e.g., residuals in the spatial domain) to obtain a TB that includes transform coefficients in the frequency domain (or spatial frequency domain). A TB that includes residuals in the spatial domain may be referred to as a residual TB, and a TB that includes transform coefficients in the frequency domain may be referred to as a coefficient TB. In the example, the forward transform includes a forward master transform that can transform the residual TB into the coefficient TB. In the example, the forward transform includes a forward master transform and a forward second-order transform, wherein the forward master transform can transform the residual TB into an intermediate coefficient TB, and the forward second-order transform can transform the intermediate coefficient TB into the coefficient TB.

[0140] An inverse transform can be performed on the coefficients TB in the frequency domain (e.g., in an encoder or decoder) to obtain the residual TB in the spatial domain. In the example, the inverse transform includes an inverse principal transform that can transform the coefficients TB into the residual TB. In the example, the inverse transform includes an inverse principal transform and an inverse quadratic transform, wherein the inverse quadratic transform can transform the coefficients TB into intermediate coefficients TB, and the inverse principal transform can transform the intermediate coefficients TB into the residual TB.

[0141] Typically, the principal transform can refer to a forward principal transform or an inverse principal transform, wherein the principal transform is performed between the residual TB and the coefficient TB. In some embodiments, the principal transform can be a separable transform, wherein the 2D principal transform can include a horizontal principal transform (also referred to as a horizontal transform) and a vertical principal transform (also referred to as a vertical transform). The quadratic transform can refer to a forward quadratic transform or an inverse quadratic transform, wherein the quadratic transform is performed between the intermediate coefficients TB and the coefficients TB.

[0142] To support extended coding block partitioning, as described in this disclosure, multiple transform sizes (e.g., ranging from 4 to 64 points for each dimension) and transform shapes (e.g., squares, rectangular shapes with aspect ratios of 2:1, 1:2, 4:1, or 1:4) can be used, such as in AV1.

[0143] The 2D transformation process can use a hybrid transform kernel, which can include different 1D transforms for each dimension of the encoded residual block. Primary 1D transforms can include (a) 4-point, 8-point, 16-point, 32-point, and 64-point DCT-2; (b) 4-point, 8-point, and 16-point asymmetric DST (ADST) (e.g., DST-4, DST-7) and corresponding flipped versions (e.g., flipped versions of ADST or FlipADST can be applied in reverse order); and / or (c) 4-point, 8-point, 16-point, and 32-point identity transform (IDTX). Figure 15 An example of the principal transformation basis function according to an embodiment of the present disclosure is shown. Figure 15The main transform basis functions in the example include the basis functions of DCT-2 with N-point input and the asymmetric DST (DST-4 and DST-7). Figure 15 The master transformation basis functions shown can be used in AV1.

[0144] The availability of hybrid transform kernels can depend on the transform block size and the prediction mode. Figure 16A Exemplary dependencies on the availability of various transform kernels (e.g., those shown in the first column and those described in the second column) based on transform block size (e.g., the size shown in the third column) and prediction mode (e.g., intra-frame prediction and inter-frame prediction shown in the third column) are illustrated. Exemplary hybrid transform kernels and availability based on prediction mode and transform block size can be used in AV1. References Figure 16A The symbols “→” and “↓” represent the horizontal dimension (also known as the horizontal direction) and the vertical dimension (also known as the vertical direction), respectively. The symbols “√” and “x” indicate the availability of the transform kernel for the corresponding block size and prediction mode. For example, the symbol “√” indicates that the transform kernel is available, and the symbol “x” indicates that the transform kernel is not available.

[0145] In the example, the transformation type (1610) is determined by... Figure 16A The first column shows the ADST_DCT representation. For example... Figure 16A As shown in the second column, the transformation type (1610) includes ADST in the vertical direction and DCT in the horizontal direction. According to Figure 16A In the third column, when the block size is less than or equal to 16×16 (e.g., 16×16 samples, 16×16 luminance samples), the transform type (1610) can be used for intra-frame prediction and inter-frame prediction.

[0146] In the example, the transformation type (1620) is determined by... Figure 16A The first column shows V_ADST as a representation. For example... Figure 16A As shown in the second column, the transformation type (1620) includes ADST in the vertical direction and IDTX (i.e., the identity matrix) in the horizontal direction. Therefore, the transformation type (1620) (e.g., V_ADST) is performed in the vertical direction but not in the horizontal direction. According to... Figure 16A In the third column, regardless of block size, transform type (1620) cannot be used for intra-frame prediction. When the block size is less than 16×16 (e.g., 16×16 samples, 16×16 luma samples), transform type (1620) can be used for inter-frame prediction.

[0147] In the example, Figure 16A It can be applied to the luma component. For the chroma component, transform type (or transform kernel) selection can be performed implicitly. In the example, for intra-prediction residuals, the transform type can be selected based on the intra-prediction mode, such as... Figure 16B As shown in the example. Figure 16B The transform type selection shown can be applied to the chroma component. For inter-frame prediction residuals, the transform type can be selected based on the transform type selection of luma blocks at the same location. Therefore, in the example, the transform type of the chroma component is not signaled in the bitstream.

[0148] For example, in AOMedia Video 2 (AV2), Line Graph Transform (LGT) can be used in transforms such as the master transform. 8-bit / 10-bit transform kernels can be used in AV2. In the examples, LGTs include various DCTs and Discrete Sine Transforms (DSTs), as described below. LGTs can include 32-point and 64-point one-dimensional (1D) DSTs.

[0149] A graph is a general mathematical structure consisting of a set of vertices and edges that can be used to model similarity relationships between objects of interest. A weighted graph, where a set of weights is assigned to edges and optionally to vertices, can provide a sparse representation for robust modeling of signals / data. LGTs can improve encoding / decoding efficiency by providing better adaptation to different block statistics. Separable LGTs can be designed and optimized by learning line graphs from the data to model the underlying row-by-row and column-by-column statistics of the residual signals of blocks, and the associated generalized graph Laplacian (GGL) matrix can be used to derive the LGT.

[0150] Figure 16C The embodiments of this disclosure are illustrated by self-circulating weights (e.g., v). c1 v c2 ) and edge weight w c An example of a general LGT representation. Given a weighted graph G(W,V), the GGL matrix can be defined as follows. L c =D - W + V (Equation 3) Where W can be the weight of the non-negative edge w c The adjacency matrix, D, can be a diagonal matrix, and V can be a matrix representing the self-circulating weights v. c1 and v c2 A diagonal matrix. Figure 16D The matrix L is shown c Examples.

[0151] LGT can be obtained through the GGL matrix L as follows: c The feature decomposition is derived from this. L c =UΦU T (Equation 4) In this context, the columns of the orthogonal matrix U can be the basis vectors of the LGT, and Φ can be the diagonal eigenvalue matrix.

[0152] In various examples, certain DCTs and DSTs (e.g., DCT-2, DCT-8, and DST-7) are subsets of the LGT set derived from some form of GGL. This can be achieved by... c1 Set to 0 (e.g., v) c1 =0) to derive DCT-2. This can be achieved by setting v c1 Set to w c (for example, v) c1 =w c To export DST-7, you can use v. c2 Set to w c (for example, v) c2 =w c To export DCT-8, you can use v. c1 Set to 2w c (for example, v) c1 =2w c To export DST-4, you can use v. c2 Set to 2w c (for example, v) c2 =2w c To export DCT-4.

[0153] In some examples, such as in AV2, LGT can be implemented as matrix multiplication. This can be achieved by using L... c Lieutenant General V c1 Set to 2w c To derive a 4-point (4p) LGT kernel, and therefore a 4p LGT kernel is DST-4. This can be achieved by... c Lieutenant General V c1 Set to 1.5w c To export an 8-point (8p) LGT kernel. In the example, LGT kernels such as 16-point (16p) LGT kernels, 32-point (32p) LGT kernels, or 64-point (64p) LGT kernels can be exported by v c1 Set to w c And v c2 Set to 0 for export, and the LGT kernel can be changed to DST-7.

[0154] Transformations such as principal and secondary transformations can be applied to blocks such as blocks of type CB. In the example, the transformation includes a combination of principal and secondary transformations. Transformations can be inseparable, separable, or a combination of inseparable and separable transformations.

[0155] Quadratic transforms can be performed, for example, in VVC. In some examples, such as in VVC, the Low-Frequency Inseparable Transform (LFNST) (also known as Reduced Quadratic Transform (RST)) can be applied between the forward master transform and quantization at the encoder side and between dequantization and inverse master transform at the decoder side, such as... Figures 17 to 18 As shown in the figure, this is to further correlate the principal transform coefficients.

[0156] A 4×4 input block (or input matrix) X can be used as an example (as shown in Equation 5) to illustrate the application of an inseparable transformation that can be used in LFNST, as described below. To apply a 4×4 inseparable transformation (e.g., LFNST), the 4×4 input block X can be derived from a vector... This is indicated as shown in Equation 5-6.

[0157] Inseparable transformations can be computed as in The vector indicates the transformation coefficients, and T is a 16×16 transformation matrix. The 16×1 coefficient vector can then be transformed using the scan order of the 4×4 input blocks (e.g., horizontal scan order, vertical scan order, zigzag scan order, or diagonal scan order). Reorganize into 4×4 output blocks (or output matrices, coefficient blocks). Transform coefficients with smaller indices can be placed in 4×4 coefficient blocks with smaller scan indices.

[0158] Inseparable quadratic transformations can be applied to blocks (e.g., CB). In some examples, such as in VVC, ... Figures 17 to 18 As shown, LFNST is applied between the forward master transform and quantization (e.g., on the encoder side) and between dequantization and inverse master transform (e.g., on the decoder side).

[0159] Figures 17 to 18 Examples of two transform encoding / decoding processes (1700) and (1800) are shown, respectively, using a 16×64 transform (or 64×16 transform, depending on whether the transform is a forward quadratic or inverse quadratic transform) and a 16×48 transform (or 48×16 transform, depending on whether the transform is a forward quadratic or inverse quadratic transform). Reference Figure 17In process (1700), on the encoder side, a forward master transform (1710) can first be performed on a block (e.g., a residual block) to obtain a coefficient block (1713). Subsequently, a forward quadratic transform (or forward LFNST) (1712) can be applied to the coefficient block (1713). In the forward quadratic transform (1712), the 64 coefficients of the 4×4 sub-block AD at the top left corner of the coefficient block (1713) can be represented by a 64-length vector, and the 64-length vector can be multiplied by a 64×16 (i.e., a width of 64 and a height of 16) transform matrix to produce a 16-length vector. The elements in the 16-length vector are filled back into the top left 4×4 sub-block A of the coefficient block (1713). The coefficients in the sub-block BD can be zero. Then, in the quantization step (1714), the coefficients obtained after the forward quadratic transform (1712) are quantized and entropy encoded to generate encoded bits in the bitstream (1716).

[0160] The encoded bits can be received at the decoder side and entropy decoded, followed by a dequantization step (1724) to generate a coefficient block (1723). An inverse quadratic transform (or inverse LFNST) (1722), such as an inverse RST8×8, can be performed to obtain 64 coefficients, for example, from the 16 coefficients at the top-left 4×4 subblock E. The 64 coefficients can be filled back into the 4×4 subblock EH. Further, the coefficients in the coefficient block (1723) after the inverse quadratic transform (1722) can be processed with an inverse master transform (1720) to obtain the recovered residual block.

[0161] Figure 18 The example procedure (1800) is similar to procedure (1700), except that fewer coefficients (i.e., 48) are processed during the forward quadratic transformation (1712). Specifically, the 48 coefficients in subblock AC are processed using a smaller transformation matrix of size 48×16. Using a smaller transformation matrix of 48×16 reduces the memory size required to store the transformation matrix and the amount of computation (e.g., multiplication, addition, and / or subtraction, etc.), and thus reduces computational complexity.

[0162] In the example, a 4×4 non-separable transformation (e.g., 4×4 LFNST) or an 8×8 non-separable transformation (e.g., 8×8 LFNST) is applied based on the block size (e.g., CB). The block size can include width, height, etc. For example, 4×4 LFNST is applied to blocks where the minimum width and height are less than a threshold such as 8 (e.g., min(width, height) < 8). For example, 8×8 LFNST is applied to blocks where the minimum width and height are greater than a threshold such as 4 (e.g., min(width, height) > 4)).

[0163] A non-separable transform (e.g., LFNST) can be based on a direct matrix multiplication method and can thus be implemented in a single pass without iteration. To reduce the dimension of the non-separable transform matrix and minimize the computational complexity and memory space for storing transform coefficients, a reduced non-separable transform method (or RST) can be used in the LFNST. Accordingly, in the reduced non-separable transform, an N-dimensional vector (e.g., for an 8×8 non-separable quadratic transform (NSST), N is 64) can be mapped to an R-dimensional vector in a different space, where N / R (R < N) is the reduction factor. Thus, instead of an N×N matrix, the RST matrix is an R×N matrix as described in Equation 7.

[0164] In Equation 7, the R rows of the R×N transform matrix are the R bases of the N-dimensional space. The inverse transform matrix can be the transpose of the transform matrix used in the forward transform (e.g., T RxN ). For an 8×8 LFNST, a reduction factor of 4 can be applied, and the 64×64 direct matrix used for the 8×8 non-separable transform can be reduced to a 16×64 direct matrix, as Figure 17 shown. Alternatively, a reduction factor greater than 4 can be applied, and the 64×64 direct matrix used in the 8×8 non-separable transform can be reduced to a 16×48 direct matrix, as Figure 18 shown. Thus, a 48×16 inverse RST matrix can be used on the decoder side to generate the kernel (primary) transform coefficients in the 8×8 upper left region.

[0165] Refer to Figure 18 , when a 16×48 matrix is applied instead of a 16×64 matrix with the same transform set configuration, the input of the 16×48 matrix includes 48 input data from three 4×4 blocks A, B, and C in the upper left 8×8 block except for the lower right 4×4 block D. With the reduction in size, the memory usage for storing the LFNST matrix can be reduced with a minimal performance degradation, e.g., from 10 KB to 8 KB.

[0166] To reduce complexity, the LFNST can be restricted to be applicable if the coefficients outside the first coefficient subgroup are unimportant. In an example, the LFNST can be restricted to be applicable only when all coefficients outside the first coefficient subgroup are non-effective. Refer to Figures 17 to 18 , the first coefficient subgroup corresponds to the upper left block E, and thus the coefficients outside block E are non-effective.

[0167] In the example, when LFNST is applied, only the principal transform coefficients are invalid (e.g., zero). In the example, when LFNST is applied, all principal transform coefficients are zero. Principal transform coefficients can refer to transform coefficients obtained from a principal transform without secondary transforms. Accordingly, the LFNST index signaling can be conditional on the end valid position, thereby avoiding extra coefficient scans in LFNST. In some examples, extra coefficient scans are used to check for valid transform coefficients at specific positions. In the example, the worst-case handling of LFNST (e.g., multiplication per pixel) restricts the inseparable transforms of 4×4 blocks and 8×8 blocks to 8×16 transforms and 8×48 transforms, respectively. In the above cases, the end valid scan position can be less than 8 when LFNST is applied. For other sizes, the end valid scan position can be less than 16 when LFNST is applied. For 4×N and N×4 CBs where N is greater than 8, this restriction can mean that LFNST is applied to the top-left 4×4 region in the CB. In the example, this restriction means that LFNST is applied only once to the top-left 4×4 region in the CB. In the example, when LFNST is applied, all primary coefficients are invalid (e.g., zero), reducing the number of operations required for the main transform. From the encoder's perspective, the quantization of the transform coefficients can be significantly simplified when testing the LFNST transform. Rate-distortion optimized quantization can be maximized for the first 16 coefficients; for example, the remaining coefficients can be set to zero according to the scan order.

[0168] The LFNST transform (e.g., transform kernel, transform core, or transform matrix) can be selected as described below. In embodiments, multiple transform sets can be used, and one or more inseparable transform matrices (or kernels) can be included in each of the multiple transform sets in the LFNST. According to aspects of this disclosure, transform sets can be selected from multiple transform sets, and inseparable transform matrices can be selected from one or more inseparable transform matrices in the transform sets.

[0169] Table 1 illustrates an exemplary mapping from intra-prediction modes to multiple transform sets according to embodiments of the present disclosure. The mapping indicates the relationship between the intra-prediction modes and the multiple transform sets. Relationships such as those indicated in Table 1 may be predefined and may be stored in the encoder and decoder. Table 1: Transform Set Selection Table

[0170] Referring to Table 1, multiple transform sets comprise four transform sets, for example, transform sets 0 through 3 represented by transform set indices from 0 to 3 (e.g., Tr.set index). An index (e.g., IntraPredMode) can indicate the intra-prediction mode, and the transform set index can be obtained based on the index and Table 1. Accordingly, the transform set can be determined based on the intra-prediction mode. In the example, if one of the three cross-component linear model (CCLM) modes (e.g., INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the CB (e.g., 81 <= IntraPredMode <= 83), then transform set 0 is selected for the CB.

[0171] As described above, each transform set may include one or more inseparable transform matrices. One of the one or more inseparable transform matrices may be selected by an LFNST index, which is, for example, explicitly signaled. For example, the LFNST index may be signaled once in the bitstream of each intraframe-encoded CU (e.g., CB) after the transform coefficients have been signaled. In embodiments, each transform set includes two inseparable transform matrices (kernels), and the selected inseparable quadratic transform candidate may be one of the two inseparable transform matrices. In some examples, LFNST is not applied to the CB (e.g., the CB is encoded with a transform skip mode or the number of non-zero coefficients of the CB is less than a threshold). In examples, when LFNST is not applied to the CB, the LFNST index of the CB is not signaled. The default value of the LFNST index may be zero and not signaled, indicating that LFNST is not applied to the CB.

[0172] In this embodiment, LFNST is restricted to application only when all coefficients outside the first coefficient subgroup are non-significant, and the encoding / decoding of the LFNST index may depend on the position of the last significant coefficient. The LFNST index may be context-coded. In this example, the context encoding / decoding of the LFNST index does not depend on the intra-prediction mode, and only the first binary number is context-coded. LFNST can be applied to intra-coded CUs in intra-slices or inter-slices, and for both the luma and chroma components. If dual-tree is enabled, the LFNST indexes for the luma and chroma components can be signaled separately. For inter-slices (e.g., dual-tree is disabled), a single LFNST index can be signaled and used for both the luma and chroma components.

[0173] Intra-Frame Sub-Partition (ISP) encoding / decoding mode can be used. In ISP encoding / decoding mode, depending on the block size, the luma intra-prediction block can be divided vertically or horizontally into 2 or 4 sub-partitions. In some examples, the performance improvement is not significant when RST is applied to each feasible sub-partition. Therefore, in some examples, when ISP mode is selected, LFNST is disabled and the LFNST index (or RST index) is not signaled. Disabling RST or LFNST for residues predicted by ISP can reduce encoding / decoding complexity. In some examples, when matrix-based intra-frame prediction mode (MIP) is selected, LFNST is disabled and the LFNST index is not signaled.

[0174] In some examples, due to the maximum transform size limit (e.g., 64×64), CUs larger than 64×64 are implicitly partitioned (TU tiling), and LFNST index search can quadruple the data buffer for a certain number of decoding pipeline stages. Therefore, the maximum size allowed by LFNST can be limited to 64×64. In this example, LFNST is enabled only via Discrete Cosine Transform (DCT) Type 2 (DCT-2) transform.

[0175] In some examples, separable transform schemes may not be effective for capturing oriented texture patterns (e.g., edges along 45° or 135° directions). For example, in the scenario above, non-separable transform schemes can improve encoding / decoding efficiency. To reduce computational complexity and memory usage, non-separable transform schemes can be used as secondary transforms applied to low-frequency transform coefficients obtained from the master transform. The secondary transform can be applied to blocks and can be signaled with information for the block based on prediction mode information, master transform type, neighboring reconstructed samples, transform block partitioning information, and / or block size and shape, such as an index indicating the secondary transform (e.g., a secondary transform index). One or a combination of prediction mode information, master transform type, and block size (or block size) can be used to derive context information. The context information can be used to determine one or more contexts for entropy encoding / decoding of the secondary transform index, which is used to identify the secondary transform within a set of secondary transforms (e.g., a set of non-separable secondary transforms). The secondary transform index can be used to decode the block.

[0176] In this disclosure, the term "block" may refer to PB, CB, coded block, coding unit (CU), transform block (TB), transform unit (TU), luma block (e.g., luma CB), chroma block (e.g., chroma CB), etc.

[0177] The size of a block can refer to its width, height, aspect ratio (e.g., the ratio of width to height, or height to width), area or size (e.g., width × height), minimum of width and height, or maximum of width and height.

[0178] Information indicating the quadratic transforms (e.g., quadratic transform kernels, quadratic transform cores, or quadratic transform matrices) used for encoding / decoding (e.g., encoding and / or decoding) of a block may include an index (e.g., a quadratic transform index). As discussed above, in embodiments, multiple transform sets may be used, and each of the multiple transform sets may include one or more quadratic transform matrices (or kernels). According to aspects of this disclosure, transform sets can be selected from multiple transform sets using any suitable method, including but not limited to those described with reference to Table 1, and the quadratic transforms (e.g., quadratic transform matrices) used for encoding / decoding (e.g., encoding and / or decoding) of a block can be determined (e.g., selected) from one or more quadratic transform matrices in the transform set by an index (e.g., a quadratic transform index). The index (e.g., a quadratic transform matrix) can be used to identify a quadratic transform in a transform set used for decoding a block. In an example, the transform set of quadratic transforms includes a set of inseparable quadratic transforms used for decoding a block. In an example, the index (e.g., a quadratic transform matrix) is denoted as stIdx. For example, the index (e.g., the quadratic transform index) can be explicitly signaled in the encoded video bitstream.

[0179] In the examples, the quadratic transform is LFNST, and the quadratic transform index refers to the LFNST index. In some examples, the quadratic transform is not applied to the block (e.g., skipping the CB encoded by the pattern with the transform or the number of non-zero coefficients of the CB is less than a threshold). In the examples, when the quadratic transform is not applied to the block, the quadratic transform index (e.g., the LFNST index) is not signaled for the block. The default value of the quadratic transform index can be zero and is not signaled, indicating that the quadratic transform was not applied to the block.

[0180] According to various aspects of this disclosure, coded information of a block (e.g., TB) can be decoded from an encoded video bitstream. The coded information may indicate one or more of the block's prediction mode information, the block size, and the main transform type used for the block. The main transform type may refer to any suitable main transform, such as a reference transform. Figure 15 and Figures 16A to 16D One or a combination of those transformations described.

[0181] In the example, the block is intra-coded or intra-predicted, and the prediction mode information is referred to as intra-prediction mode information. The block's prediction mode information (e.g., intra-prediction mode information) can indicate the intra-prediction mode used for the block. The intra-prediction mode can refer to the prediction mode used for intra-prediction of the block, such as... Figure 9 The directional pattern (or directional prediction pattern) described in the text Figure 10 The non-directional prediction patterns described in the document (e.g., DC pattern, PAETH pattern, SMOOTH pattern, SMOOTH_V pattern, or SMOOTH_H pattern) or Figure 11 The recursive filtering mode described herein. Intra-frame prediction mode can also refer to the prediction mode described herein, suitable variations of the prediction mode described herein, or suitable combinations of the prediction modes described herein. For example, intra-frame prediction mode can be combined with... Figure 12 The multi-line intra-frame prediction combination described in the text.

[0182] Encoding information (such as one or more of the block's prediction mode information, the block size, and the main transform type used for the block) can be used as context for entropy encoding and / or decoding of a quadratic transform index (e.g., stIdx). The context for entropy encoding and / or decoding of the quadratic transform index can be determined based on encoding information such as the block's prediction mode information, the block size, and the main transform type used for the block. The quadratic transform index (e.g., stIdx) can indicate a quadratic transform among a set of quadratic transforms to be performed on the block.

[0183] The context derivation process can be used to determine the context. The context derivation process used for entropy encoding and / or decoding (e.g., encoding and / or decoding) of the quadratic transform index can depend on the encoding information.

[0184] The context for entropy encoding / decoding (e.g., encoding and / or decoding) of the secondary transform index can be determined based on one or more of the block's prediction mode information (e.g., intra-frame prediction mode information), the block size, and the primary transform type used for the block. In the example, the context for entropy encoding / decoding (e.g., encoding and / or decoding) of the secondary transform index is determined based on one of the block's prediction mode information (e.g., intra-frame prediction mode information), the block size, and the primary transform type used for the block. Entropy encoding / decoding (e.g., encoding and / or decoding) of the secondary transform index can be performed based on the context. The secondary transform (e.g., inverse secondary transform) indicated by the secondary transform index can be performed on the block (e.g., TB).

[0185] In embodiments, the context derivation process for entropy encoding / decoding (e.g., encoding and / or decoding) of a quadratic transform index (e.g., stIdx) can depend on the block size. In an example, the maximum square block size associated with a block refers to the size of the largest square block within the block. The maximum square block size is less than or equal to the block size. In an example, when the block refers to TB, the block size is the TB size. The maximum square block size can be used in the context derivation process. In an example, the maximum square block size is used as context. For example, the block size can indicate the block width W and the block height H. The maximum square block size is L×L, where L is the minimum of W and H. The maximum square block size can also be represented by L. If W is 32 and H is 8, then L is 8. The maximum square block size can be indicated by 8 or 8×8.

[0186] In an embodiment, the block size may be indicated by one of the block's intra-frame prediction mode information, the block size, and the primary transform type used for the block. The context for entropy encoding / decoding (e.g., encoding and / or decoding) of the secondary transform index may be determined based on the block size. In the example, the block size indicates the block's width W and height H. The minimum of the block's width W and height H is denoted as L. The context may be determined based on the maximum square block size L or L×L. In the example, the maximum square block size L or L×L is used as the context.

[0187] In an embodiment, the context derivation process for entropy encoding and / or decoding (e.g., encoding and / or decoding) of the secondary transform index (e.g., stIdx) may depend on the block's prediction mode information (e.g., intra-frame prediction mode information). The context for entropy encoding and / or decoding (e.g., encoding and / or decoding) of the secondary transform index can be derived based on the block's prediction mode information (e.g., intra-frame prediction mode information).

[0188] In the example, one of several directional modes (also known as directional prediction modes) is used to perform intra-frame prediction of the block, such as see [link to example]. Figure 9 Described. Reference Figure 9 The nominal mode index can indicate a nominal mode (e.g., one of eight nominal modes). In the example, the orientation mode (e.g., one of 56 orientation modes) can be determined based on the nominal mode index and an index indicating the angular offset relative to the nominal mode. The block's prediction mode information (e.g., intra-frame prediction mode information) can indicate the nominal mode index. The nominal mode index can be used in the context derivation process. A context for entropy encoding / decoding (e.g., encoding and / or decoding) of the quadratic transform index can be derived based on the nominal mode index.

[0189] In an embodiment, one of the block's intra-prediction mode information, the block size, and the major transform type used for the block can indicate the block's intra-prediction mode information. The context for entropy encoding / decoding (e.g., encoding and / or decoding) of the secondary transform index can be determined based on the block's intra-prediction mode information. In the example, the block's intra-prediction mode information indicates the nominal mode index. As described above, the block can be predicted using a directional prediction mode determined based on the nominal mode index and an angle offset. The context for entropy encoding / decoding (e.g., encoding and / or decoding) of the secondary transform index can be determined based on the nominal mode index.

[0190] In this embodiment, intra-frame prediction of blocks is performed using the directional modes described above. The nominal mode index can be mapped to an index value different from the nominal mode index, and the index value can be used during context derivation. In this example, multiple intra-frame prediction modes (e.g., Figure 9 Multiple nominal patterns (or multiple nominal pattern indices corresponding to multiple nominal patterns) are mapped to the same index value. These multiple nominal patterns can be adjacent to each other. (See reference...) Figure 9 D67_PRED and D45_PRED can be mapped to the same index value.

[0191] In an embodiment, the intra-frame prediction mode information of a block indicates a nominal mode index, and the block can be predicted using a directional mode (or directional prediction mode) determined based on the nominal mode index and an angle offset. The context for entropy encoding / decoding (e.g., encoding and / or decoding) of the secondary transform index (e.g., mapping the nominal mode index to index values) can be determined based on the index value associated with the nominal mode index.

[0192] In an embodiment, intra-frame prediction of a block is performed using one of a plurality of non-directional prediction modes (also known as non-directional smooth intra-frame prediction modes), including, for example, DC mode, PAETH mode, SMOOTH mode, SMOOTH_V mode, and SMOOTH_H mode, as referenced. Figure 10 Described. A corresponding pattern index (also referred to as a non-directional prediction pattern index) indicating one of multiple non-directional prediction patterns can be used in the context derivation process. In the example, the same context can be applied to entropy encoding / decoding of the quadratic transformation index (e.g., stIdx) of one or more of the non-directional prediction patterns. The same context can correspond to multiple non-directional prediction patterns.

[0193] In an embodiment, the intra-frame prediction mode information of a block can indicate a non-directional prediction mode index. A block can be predicted using one of a plurality of non-directional prediction modes (e.g., non-directional prediction modes) indicated by the non-directional prediction mode index. The context for entropy encoding / decoding (e.g., encoding and / or decoding) of the quadratic transform index can be determined based on the non-directional prediction mode index.

[0194] In the example, one of the recursive filtering modes is used for intra-frame prediction of the block, such as reference. Figure 11 The description states that a mapping from one of the recursive filtering modes to a nominal mode index can be performed first, where the nominal mode index can indicate the nominal mode (e.g., ...). Figure 9 (One of the eight nominal modes in the example). The nominal mode index can then be used in the context derivation process. In the example, the same context can be applied to entropy encoding / decoding (e.g., encoding and / or decoding) of one or more quadratic transform indices (e.g., stIdx) in the recursive filtering modes. The same context can correspond to multiple recursive filtering modes.

[0195] In an embodiment, the intra-frame prediction mode information of a block can indicate a recursive filtering mode (one of a plurality of recursive filtering modes) used to predict the block. A nominal mode index indicating the nominal mode can be determined based on the recursive filtering mode. The context for entropy encoding / decoding (e.g., encoding and / or decoding) of the secondary transform index can be determined based on the nominal mode index.

[0196] In an embodiment, the context derivation process for entropy encoding / decoding the secondary transform index (e.g., stIdx) may depend on the master transform type information. The master transform type information may indicate the master transform type or the type of master transform used for the block. In the example, the context is determined based on the master transform type.

[0197] In an embodiment, one of the block's intra-prediction mode information, the block size, and the primary transform type used for the block indicates the primary transform type used for the block. The context for entropy encoding / decoding (e.g., encoding and / or decoding) of the secondary transform index can be determined based on the primary transform type used for the block.

[0198] In an embodiment, the primary transform of a block may include a horizontal primary transform (referred to as a horizontal transform) and a vertical primary transform (referred to as a vertical transform). The context for entropy encoding / decoding (e.g., encoding and / or decoding) of the secondary transform index can be determined based on the horizontal primary transform type and the vertical primary transform type. The horizontal primary transform type may indicate a horizontal transform, and the vertical primary transform type may indicate a vertical transform.

[0199] The context (also known as the context value) can depend on whether both the horizontal principal transform type and the vertical principal transform type are DCT or both are ADST.

[0200] In this disclosure, a combination of horizontal principal transform type and vertical principal transform type can be represented as {horizontal principal transform type, vertical principal transform type}. Therefore, {DCT, DCT} indicates that both the horizontal and vertical principal transform types are DCT. {ADST, ADST} indicates that both the horizontal and vertical principal transform types are ADST. {LGT, LGT} indicates that both the horizontal and vertical principal transform types are LGT. {DCT, LGT} indicates that the horizontal principal transform type is DCT and the vertical principal transform type is LGT. {LGT, DCT} indicates that the horizontal principal transform type is LGT and the vertical principal transform type is DCT.

[0201] In an embodiment, the primary transform indicated by the primary transform type includes a horizontal transform indicated by a horizontal primary transform type and a vertical transform indicated by a vertical primary transform type. The context for entropy encoding / decoding (e.g., encoding and / or decoding) of the secondary transform index can be determined based on whether both the horizontal and vertical primary transform types are DCTs or ADSTs.

[0202] In an embodiment, the context may depend on whether the combination of the horizontal principal transform type and the vertical principal transform type is {DCT,DCT} or {LGT,LGT}.

[0203] In an embodiment, the primary transform indicated by the primary transform type includes a horizontal transform indicated by a horizontal primary transform type and a vertical transform indicated by a vertical primary transform type. The context for entropy encoding / decoding (e.g., encoding and / or decoding) of the secondary transform index can be determined based on whether both the horizontal and vertical primary transform types are DCTs or LGTs.

[0204] In an embodiment, the context may depend on whether the combination of the horizontal principal transform type and the vertical principal transform type is {DCT,DCT}, {LGT,LGT}, {DCT,LGT}, or {LGT,DCT}.

[0205] In an embodiment, the primary transform indicated by the primary transform type includes a horizontal transform indicated by a horizontal primary transform type and a vertical transform indicated by a vertical primary transform type. The context for entropy encoding / decoding (e.g., encoding and / or decoding) of the secondary transform index can be determined based on whether the horizontal primary transform type and the vertical primary transform type are (i) both DCT, (ii) both LGT, (iii) DCT and LGT respectively, or (iv) LGT and DCT respectively.

[0206] In embodiments, the context may depend on whether the combination of the horizontal principal transform type and the vertical principal transform type is {DCT,DCT}, {LGT,LGT}, {DCT,IDTX}, or {IDTX,DCT}, where IDTX represents the identity transform. {DCT,IDTX} indicates that the horizontal principal transform type is DCT and the vertical principal transform type is IDTX. {IDTX,DCT} indicates that the horizontal principal transform type is IDTX and the vertical principal transform type is DCT.

[0207] In an embodiment, the primary transform indicated by the primary transform type includes a horizontal transform indicated by a horizontal primary transform type and a vertical transform indicated by a vertical primary transform type. The context for entropy encoding / decoding (e.g., encoding and / or decoding) of the quadratic transform index can be determined based on the horizontal primary transform type and the vertical primary transform type being (i) both DCT, (ii) both LGT, (iii) DCT and identity transform (IDTX) respectively, or (iv) IDTX and DCT respectively.

[0208] Figure 19 A flowchart outlining a process (1900) according to an embodiment of the present disclosure is shown. Process (1900) can be used for the reconstruction of blocks (such as CB, TB, luma CB, luma TB, chroma CB, chroma TB, etc.). In various embodiments, process (1900) is executed by processing circuitry, such as processing circuitry in terminal devices (310), (320), (330), and (340), processing circuitry performing the functions of a video encoder (403), processing circuitry performing the functions of a video decoder (410), processing circuitry performing the functions of a video decoder (510), processing circuitry performing the functions of a video encoder (603), and so on. In some embodiments, process (1900) is implemented with software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry executes process (1900). This begins at (S1901) and proceeds to (S1910).

[0209] At (S1910), the encoding information of the block (e.g., TB, Luminance TB, Intra-coded TB, CB) can be decoded from the encoded video bitstream. The encoding information may indicate one or more of the block's intra-prediction mode information, block size, and main transform type.

[0210] Intra-prediction mode information for a block can indicate the intra-prediction mode used for intra-prediction of the block, such as reference... Figures 9 to 11 The description specifies the directional, non-directional prediction, or recursive filtering mode. The main transform type indicates the main transform used for the block. Figure 15 and Figures 16A to 16DAn example of the main transformation type used for blocks is described in the document.

[0211] At (S1920), the context for entropy decoding of the secondary transform index can be determined based on one of the intra-prediction mode information for the block, the block size, and the main transform type for the block. The secondary transform index can indicate the secondary transform in a set of secondary transforms to be performed on the block.

[0212] In the example, the context for entropy decoding of the secondary transform index can be determined based on one or more of the block's intra-prediction mode information, the block size, and the main transform type used for the block.

[0213] In the example, the block's intra-prediction mode information, the block size, and the main transform type used for the block indicate the block size, and the context used for entropy decoding of the secondary transform index can be determined based on the block size.

[0214] In the example, one of the block's intra-prediction mode information, the block size, and the main transform type used for the block indicates the block's intra-prediction mode information, and the context used for entropy decoding of the secondary transform index can be determined based on the block's intra-prediction mode information.

[0215] In the example, one of the block's intra-prediction mode information, the block size, and the main transform type used for the block indicates the main transform type used for the block, and the context used for entropy decoding of the secondary transform index can be determined based on the main transform type used for the block.

[0216] At (S1930), the quadratic transform index can be entropy decoded based on the context determined at (S1920).

[0217] At (S1940), a second transformation indicated by the second transformation index can be performed on the block. For example, the second transformation is an inverse second transformation performed after the inverse principal transformation. In the example, the second transformation is LFNST. In the example, the second transformation is an inseparable second transformation. The process (1900) proceeds to (S1999) and ends.

[0218] Procedure (1900) can be modified as appropriate. One or more steps in procedure (1900) can be modified and / or omitted. One or more additional steps can be added. Any suitable implementation order can be used. For example, except for the reference Figures 9 to 11 In addition to the described directional, non-directional prediction, or recursive filtering modes, the process (1900) can also be applied to other intra-frame prediction modes. Besides the reference... Figure 15 and Figures 16A to 16DIn addition to the described primary transform type, the procedure (1900) can also be applied to other appropriate primary transform types or appropriate combinations of horizontal and vertical primary transform types.

[0219] The embodiments in this disclosure can be used individually or in any combination in any order. Furthermore, each of the method (or embodiment), encoder, and decoder can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-volatile computer-readable medium. The embodiments in this disclosure can be applied to luma blocks or chroma blocks.

[0220] The techniques described above can be implemented as computer software that uses computer-readable instructions and is physically stored on one or more computer-readable media. For example, Figure 20 A computer system (2000) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0221] Computer software can be encoded using any suitable machine code or computer language, which can be assembled, compiled, linked, or similarly processed to create code that includes instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), or through interpretation, microcode execution, or other means.

[0222] These instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0223] Figure 20 The components shown for the computer system (2000) are exemplary in nature and are not intended to impose any limitation on the scope or functionality of the computer software used to implement the embodiments of this disclosure. The configuration of the components should also not be construed as having any dependency or requirement on any component or combination thereof illustrated in the exemplary embodiments of the computer system (2000).

[0224] Computer systems (2000) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to one or more human users through, for example, tactile input (such as keystrokes, swipes, data glove movements), audio input (such as voice, tapping), visual input (such as gestures), and olfactory input (not depicted). Human-machine interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (such as voice, music, ambient sounds), images (such as scanned images, photographic images obtained from still image cameras), and video (such as two-dimensional video, three-dimensional video including stereoscopic video).

[0225] Input human-machine interface devices may include one or more of the following (only one of each is depicted): keyboard (2001), mouse (2002), touchpad (2003), touch screen (2010), data glove (not shown), joystick (2005), microphone (2006), scanner (2007), camera (2008).

[0226] The computer system (2000) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback via a touchscreen (2010), data gloves (not shown), or joystick (2005), but tactile feedback devices that are not used as input devices may also exist), audio output devices (such as speakers (2009), headphones (not depicted)), visual output devices (such as screens (2010) including CRT screens, LCD screens, plasma screens, OLED screens, each screen may or may not have touchscreen input capability, each screen may or may not have tactile feedback capability, some of which are capable of outputting two-dimensional visual output or more than three-dimensional output by means such as stereoscopic output; virtual reality glasses (not depicted), holographic displays and smoke canisters (not depicted), and printers (not depicted).

[0227] Computer systems (2000) may also include human-accessible storage devices and their associated media, such as optical media including media (2021) such as CD / DVD ROM / RW with CD / DVD (2020), thumb drives (2022), removable hard disk drives or solid-state drives (2023), conventional magnetic media such as magnetic tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.

[0228] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other volatile signals.

[0229] The computer system (2000) may also include interfaces (2054) to one or more communication networks (2055). Networks may be, for example, wireless, wired, or optical. Networks may further be local, wide-area, metropolitan, vehicular, industrial, real-time, latency-tolerant, etc. Examples of networks include local area networks (such as Ethernet), wireless LANs, cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), cable or wireless wide-area digital television networks (including cable television, satellite television, and terrestrial broadcast television), vehicle and industrial networks (including CANbus), etc. Some networks typically require external network interface adapters attached to certain general-purpose data ports or peripheral buses (2049) (such as the USB port of the computer system (2000); other networks are typically integrated into the core of the computer system (2000) by attaching to system buses as described below (e.g., an Ethernet interface integrated into a PC computer system or a cellular network interface integrated into a smartphone computer system). Using any of these networks, the computer system (2000) can communicate with other entities. This communication can be unidirectional, receive-only (e.g., broadcasting TV), transmit-only (e.g., CANbus to certain CANbus devices), or bidirectional (e.g., to other computer systems using a local area digital network or a wide area digital network). Certain protocols and protocol stacks can be used on each of the networks and network interfaces described above.

[0230] The aforementioned human-machine interface devices, human-accessible storage devices, and network interfaces can be attached to the kernel (2040) of a computer system (2000).

[0231] The kernel (2040) may include one or more central processing units (CPU) (2041), graphics processing units (GPUs) (2042), dedicated programmable processing units (FPGAs) in the form of field-programmable gate areas (FPGAs) (2043), hardware accelerators (2044) for certain tasks, graphics adapters (2050), etc. These devices, along with read-only memory (ROM) (2045), random access memory (2046), and internal mass storage devices such as internal non-user-accessible hard disk drives (SDs) and SSDs (2047), can be connected via a system bus (2048). In some computer systems, the system bus (2048) may be accessed as one or more physical connectors to allow for expansion by additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the kernel's system bus (2048) or attached to the system bus (2048) via a peripheral bus (2049). In the example, a screen (2010) may be connected to a graphics adapter (2050). Peripheral bus architectures include PCI, USB, etc.

[0232] The CPU (2041), GPU (2042), FPGA (2043), and accelerator (2044) can execute certain instructions, and combinations of these instructions can constitute the aforementioned computer code. This computer code can be stored in ROM (2045) or RAM (2046). Transient data can also be stored in RAM (2046), while permanent data can be stored, for example, in an internal mass storage device (2047). Fast storage and retrieval of any memory device can be enabled by using a cache memory, which can be closely associated with one or more CPUs (2041), GPUs (2042), mass storage devices (2047), ROMs (2045), RAMs (2046), etc.

[0233] Computer-readable media may contain computer code for performing operations of various computer implementations. The media and computer code may be those specifically designed and constructed for the purposes of this disclosure, or they may be of types known and available to those skilled in the art of computer software.

[0234] By way of example and not limitation, a computer system having an architecture (2000) and, in particular, a kernel (2040) can provide functionality as a result of executing software contained in one or more tangible computer-readable media (including CPUs, GPUs, FPGAs, accelerators, etc.). Such computer-readable media can be media associated with user-accessible mass storage devices as described above, as well as certain storage devices of the kernel (2040) having non-volatile properties (such as internal kernel mass storage (2047) or ROM (2045)). Software implementing various embodiments of this disclosure can be stored in such devices and executed by the internal kernel (2040). Depending on specific needs, the computer-readable medium may include one or more memory devices or chips. The software can cause the kernel (2040) and, in particular, the processors therein (including CPUs, GPUs, FPGAs, etc.) to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM (2046) and modifying such data structures according to software-defined processes. Alternatively or as an alternative, the computer system may provide functionality as a result of hard-wired logic or otherwise embodied in circuitry (e.g., an accelerator (2044)) that may replace or operate with software to perform a particular process or a particular portion of a particular process described herein. References to software may, where appropriate, include logic, and vice versa. References to computer-readable media may, where appropriate, include circuitry storing software for execution (such as integrated circuits (ICs)), circuitry containing logic for execution, or both. This disclosure includes any suitable combination of hardware and software. Appendix A: Acronyms JEM: Joint Exploration Model VVC: Universal Video Codec BMS: Benchmark Set MV: Motion Vector HEVC: High-efficiency video encoding and decoding SEI: Supplemental Enhancement Information VUI: Video Availability Information GOP: Image Group TU: Transformation Unit PU: Prediction Unit CTU: Coding Tree Unit CTB: Coded Tree Block PB: Prediction Block HRD: Assuming a reference decoder SNR: Signal-to-noise ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: Cathode Ray Tube LCD: Liquid Crystal Display OLED: Organic Light Emitting Diode CD: CD-ROM DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile Communications LTE: Long Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Device Interconnect FPGA: Field Programmable Gate Domain SSD: Solid State Drive IC: Integrated Circuit CU: Encoding Unit

[0235] While several exemplary embodiments have been described in this disclosure, there are changes, substitutions, and various alternative equivalents that fall within the scope of this disclosure. Therefore, it should be understood that those skilled in the art will be able to design numerous systems and methods that, while not expressly shown or described herein, embody the principles of this disclosure and are thus within its spirit and scope.

Claims

1. A video encoding method, characterized in that, include: Determine the main transform type to be used for the transform block TB to be encoded; The secondary transform index of the TB is entropy encoded based on the primary transform type used for the TB, the secondary transform index indicating the secondary transform in the set of secondary transforms to be performed on the TB; Generate encoding information indicating the main transform type for the TB; and Generate an encoded video stream that includes the encoded information and the entropy-encoded quadratic transform index.

2. The method according to claim 1, characterized in that, The encoding information also indicates one of the intra-prediction mode used for the TB and the size of the TB.

3. The method according to claim 1, characterized in that, The entropy coding also includes determining the context for entropy coding the secondary transform index based on the primary transform type used for the TB.

4. The method according to claim 3, characterized in that, The primary transform indicated by the primary transform type includes a horizontal transform indicated by a horizontal primary transform type and a vertical transform indicated by a vertical primary transform type, and The determination of the context includes determining the context for entropy encoding the quadratic transform index based on whether the horizontal principal transform type and the vertical principal transform type are both discrete cosine transforms (DCT) or both are asymmetric discrete sine transforms (ADST).

5. The method according to claim 3, characterized in that, The primary transform indicated by the primary transform type includes a horizontal transform indicated by a horizontal primary transform type and a vertical transform indicated by a vertical primary transform type, and The determination of the context includes determining the context for entropy encoding the quadratic transform index based on whether the horizontal principal transform type and the vertical principal transform type are both discrete cosine transforms (DCT) or line graph transforms (LGT).

6. The method according to claim 3, characterized in that, The primary transform indicated by the primary transform type includes a horizontal transform indicated by a horizontal primary transform type and a vertical transform indicated by a vertical primary transform type, and The determination of the context includes determining the context for entropy encoding the quadratic transform index based on whether the horizontal principal transform type and the vertical principal transform type are (i) both discrete cosine transform (DCT), (ii) both line graph transform (LGT), (iii) DCT and LGT respectively, or (iv) LGT and DCT respectively.

7. The method according to claim 3, characterized in that, The primary transform indicated by the primary transform type includes a horizontal transform indicated by a horizontal primary transform type and a vertical transform indicated by a vertical primary transform type, and The determination of the context includes determining the context for entropy encoding the quadratic transform index based on whether the horizontal principal transform type and the vertical principal transform type are (i) both discrete cosine transform (DCT), (ii) both line graph transform (LGT), (iii) DCT and identity transform (IDTX) respectively, or (iv) IDTX and DCT respectively.

8. The method according to claim 1, characterized in that, The entropy encoding of the secondary transform index of the TB includes: entropy encoding the value of the secondary transform index based on the primary transform type used for the TB.

9. A video decoding method, characterized in that, include: The encoding information of the transform block (TB) and the entropy-encoded secondary transform index of the TB are decoded from the encoded video bitstream, wherein the encoding information indicates the main transform type used for the TB; Entropy-encoded secondary transform indexes of the TB are entropy-encoded based on the primary transform type used for the TB, the secondary transform indexes indicating secondary transforms in the set of secondary transforms to be performed on the TB. as well as Perform the quadratic transformation indicated by the quadratic transformation index on the TB.

10. A video encoding apparatus, characterized in that, It includes a processor and a memory; the memory stores a computer program that, when executed by the processor, causes the processor to perform the method as described in any one of claims 1-8.

11. A video decoding device, characterized in that, Includes processing circuitry, the processing circuitry being configured to: The encoding information of the transform block (TB) and the entropy-encoded secondary transform index of the TB are decoded from the encoded video bitstream, wherein the encoding information indicates the main transform type used for the TB; Entropy-encoded secondary transform indexes of the TB are entropy-encoded based on the primary transform type used for the TB, the secondary transform indexes indicating secondary transforms in the set of secondary transforms to be performed on the TB. as well as Perform the quadratic transformation indicated by the quadratic transformation index on the TB.

12. A non-volatile computer-readable storage medium storing instructions, which, when executed by a processor, cause the processor to perform the method of any one of claims 1-9.

13. A method for storing video streams, characterized in that, Generate a video stream by performing the method according to any one of claims 1-8, and store the video stream.

14. A method for transmitting a video stream, characterized in that, The method of any one of claims 1-8 is used to generate a video stream and transmit the video stream.

15. A computer-readable storage medium storing a computer program / instructions and a video stream thereon, characterized in that, When executed by a processor, the computer program / instructions implement the steps of the method according to any one of claims 1-8 to generate the video stream.