Method and apparatus for video coding

JP2024088772A5Active Publication Date: 2025-05-22TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024063336
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-09-30
Filing Date
2024-04-10
Publication Date
2025-05-22
Estimated Expiration
2041-10-05

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently reducing redundancy in video data, particularly in intra-prediction modes, where less likely directions require more bits, leading to suboptimal compression ratios.

Method used

The method involves determining transform candidates for video blocks based on feature vectors extracted from neighboring reconstructed blocks, using statistical analysis and predefined thresholds to select optimal transform sets, thereby improving compression efficiency.

Benefits of technology

This approach enhances video coding efficiency by reducing the number of bits required to represent less likely intra-prediction directions, resulting in improved compression ratios and reduced data size without compromising video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide an apparatus including a processing circuit for video decoding / decoding which reduces redundancy in an input video signal through compression, and a method.SOLUTION: A method implemented by a processing circuit of a video encoder determines a transform candidate for a block in a current picture from a group of transform sets based on one of a feature vector or a feature scalar extracted from reconstructed samples in one or more neighboring blocks of the block. Each transform set includes one or more transform candidates for the block. The one or more neighboring blocks are in the current picture or a reconstructed picture different from the current picture. The method reconstructs samples of the block based on the determined transform candidate, selects a sub-group of transform sets from the group of transform sets based on a prediction mode for the block indicated in coded information for the block and determines the transform candidate from the sub-group of transform sets based on the reconstructed samples.SELECTED DRAWING: Figure 19
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] This application claims priority to U.S. Provisional Application No. 63 / 130,249, filed December 23, 2020, entitled "Feature based transform selection," which in turn claims the benefit of U.S. Provisional Application No. 17 / 490,967, filed September 30, 2021, entitled "METHOD AND APPARATUS FOR VIDEO CODING," the disclosure of which is incorporated herein by reference in its entirety.

[0002] [Technical field] The present disclosure generally relates to video coating embodiments. [Background technology]

[0003] The background discussion provided herein is intended to generally present the context of the present disclosure. The inventors' work, to the extent that it is described in this background section, and aspects of the discussion that may not be admitted as prior art at the time of filing, are not admitted expressly or impliedly as prior art to the present disclosure.

[0004] Video coding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video can include a sequence of pictures, each having spatial dimensions of, for example, 1920x1080 luma samples and associated chroma samples. The sequence of pictures can have a fixed or variable picture rate (also informally called "frame rate") of, for example, 60 pictures per second or 60Hz. Uncompressed video has certain bitrate requirements. For example, 1080p60 4:2:0 video (1920x1080 luma sample resolution at 60Hz frame rate) at 8 bits per sample requires a bandwidth approaching 1.5Gbit / s. One hour of such video requires more than 600GBytes of storage space.

[0005] One of the goals of video coding and decoding may be the reduction of redundancy in the input video signal through compression. Compression may help reduce the aforementioned bandwidth or storage space requirements, in some cases by more than one order of magnitude. Both lossless and lossy compression, as well as combinations thereof, may be used. Lossless compression refers to techniques where an exact replica of the original signal can be reconstructed from the compressed original signal. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signals is small enough to make the reconstructed signal useful for its intended application. For video, lossy compression has been widely adopted. The amount of acceptable distortion depends on the application, for example, users of a particular consumer streaming application may tolerate higher distortion than users of a television distribution application. The achievable compression ratio may reflect that a higher tolerable / acceptable distortion can result in a higher compression ratio.

[0006] Video encoders and decoders may utilize techniques from a number of broad categories, including, for example, motion compensation, transform, quantization, and entropy coding.

[0007] Video codec techniques may include a technique known as intra-coding. In intra-coding, sample values ​​are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, a picture is spatially subdivided into blocks of samples. When all blocks of samples are coded in intra mode, the picture may be an intra picture. Intra pictures and their derivatives, such as independent decoder refresh pictures, can be used to reset the decoder state and therefore can be used as the first picture of a coded video bitstream and video session or as still pictures. Samples of an intra block may be subjected to a transform and the transform coefficients may be quantized prior to entropy coding. Intra prediction may be a technique that minimizes sample values ​​in the pre-transformed domain. In some cases, the smaller the DC value and the smaller the AC coefficients after the transform, the fewer bits are required for a given quantization step size to represent the block after entropy coding.

[0008] Conventional intra-coding, for example as known from MPEG-2 generation coding techniques, does not use intra-prediction. However, some newer video compression techniques include techniques that attempt to do so from surrounding sample data and / or metadata obtained during encoding / decoding, for example, from spatially adjacent and earlier in the decoding order. Such techniques are hereafter referred to as "intra-prediction" techniques. It should be noted that, at least in some cases, intra-prediction uses only reference data from the current picture being reconstructed, and not from reference pictures.

[0009] Intra prediction can exist in various forms. If more than one such technique is available for a given video coding technique, the technique in use can be coded in an intra prediction mode. In some cases, the mode can have sub-modes and / or parameters that can be coded separately or included in the mode codeword. Which codeword is used for a given mode / sub-mode / parameter combination can affect the coding efficiency gains from intra prediction, as can the entropy coding technique used to convert the codeword into a bitstream.

[0010] Certain modes of intra prediction were presented in H.264, improved in H.265, and further improved in newer coding techniques such as the Joint Search Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). Predictor blocks can be formed using neighboring sample values ​​belonging to already available samples. The sample values ​​of the neighboring samples are replicated in the predictor block according to the direction. The reference to the direction in use can be coded in the bitstream or can itself be predicted.

[0011] Referring to FIG. 1A, shown at the bottom right is a subset of 9 predictor directions from the 33 possible predictor directions of H.265 (corresponding to the 33 angle modes of the 35 intra modes). The point where the arrows converge (101) represents the sample to be predicted. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples to the upper right at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples to the lower left at an angle of 22.5 degrees from the horizontal of sample (101).

[0012] Continuing with reference to FIG. 1A, at the top left, a square block (104) of 4×4 samples (indicated by a thick dashed line) is shown. The square block (104) contains 16 samples, each labeled with an “S”, its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (from the top) and the first sample in the X dimension (from the left). Similarly, sample S44 is the fourth sample in the block (104) in both the Y and X dimensions. Since the block is 4×4 samples in size, S44 is at the bottom right. Additionally, reference samples are shown that follow a similar numbering scheme. The reference samples are labeled with R, their Y position (e.g., row index) and X position (column index) relative to the block (104). In both H.264 and H.265, the predicted samples are adjacent to the block being reconstructed. Therefore, there is no need to use negative values.

[0013] Intra-picture prediction can be performed by copying reference sample values ​​from adjacent samples assigned by the signaled prediction direction. For example, assume that the coded video bitstream includes signaling indicating a prediction direction for this block that is consistent with the arrow (102) (i.e., predicted from one or more prediction samples to the upper right at an angle of 45 degrees from the horizontal). In this case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Sample S44 is then predicted from reference sample R08.

[0014] In some cases, values ​​of multiple reference samples may be combined, for example by interpolation, to calculate a reference sample, particularly when the orientation is not evenly divided by 45 degrees.

[0015] As video coding techniques develop, the number of predictable directions is also increasing. In H.264 (2003), 9 different directions can be represented. In H.265 (2013), this increases to 33, and JEM / VVC / BMS can support up to 65 directions at the time of disclosure. Experiments have been performed to identify the most likely directions, and certain techniques in entropy coding are used to represent more likely directions with fewer bits, accepting certain penalties for less likely directions. Furthermore, the direction itself may be predicted from neighboring directions used in neighboring, already decoded blocks.

[0016] FIG. 1B shows a schematic diagram (180) illustrating 65 intra prediction directions with JEM to illustrate the increasing number of prediction directions over time.

[0017] The mapping of intra-prediction direction bits in a coded video bitstream representing a direction may vary from one video coding technique to another and may range, for example, from a simple direct mapping of prediction directions to intra-prediction modes to complex adaptation schemes involving codewords, most likely modes, and similar techniques. In all cases, however, there may be certain directions that are statistically less likely to occur in the video content than other certain directions. Because the goal of video compression is to reduce redundancy, in a well-performing video coding technique, these less likely directions are represented by a greater number of bits than the more likely directions.

[0018] Motion compensation can be a lossy compression technique and can be related to a technique in which blocks of sample data from a previously reconstructed picture or part thereof (reference picture) are used for the prediction of a new reconstructed picture or part thereof after being spatially shifted in a direction indicated by a motion vector (hereinafter also referred to as MV). In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV can have two dimensions, X and Y, or three dimensions, the third dimension being an indication of the reference picture being used (the latter can indirectly be the temporal dimension).

[0019] In some video compression techniques, the MV applicable to an area of ​​sample data can be predicted from other MVs, for example from an MV associated with another region of sample data that is spatially adjacent to the region under reconstruction and precedes it in decoding order. In this way, the amount of data required for coding the MV can be significantly reduced, thereby removing redundancy and increasing the amount of compression. MV prediction can work efficiently because, for example, when coding an input video signal derived from a camera (called natural video), there is a statistical probability that regions larger than the region to which a single MV is applicable move along similar directions and therefore, in some cases, can be predicted using similar motion vectors derived from MVs of neighboring regions. As a result, the MV found for a given region will be similar or identical to the MV predicted from the surrounding MVs and, after being entropy coded, can be represented with a smaller number of bits than would be used when coding the MV directly. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., MV) derived from the original signal (i.e., sample stream). In other cases, the MV prediction itself can be lossy, for example, due to rounding errors in computing the predictor from some surrounding MVs.

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, High Efficiency Video Coding, December 2016). Among the many MV prediction mechanisms provided by H.265, a technique hereafter referred to as "spatial merging" will be described here.

[0021] Referring to Figure 2, the current block (201) contains samples that the encoder discovered during the motion search process to be predictable from a spatially shifted previous block of the same size. Instead of coding the MV directly, the MV can be derived from metadata associated with multiple reference pictures, e.g., from the most recent reference picture (in decoding order), using the MV associated with any of the five surrounding samples denoted A0, A1, and B0, B1, B2 (102 to 106, respectively). In H.265, the MV prediction can use a predictor from the same reference picture as the neighboring blocks use. Summary of the Invention [Means for solving the problem]

[0022] Aspects of the present disclosure provide methods and apparatus for video encoding and / or decoding. In some examples, an apparatus for video decoding includes a processing circuit that processes feature vectors extracted from reconstructed samples in one or more neighboring blocks of a block of a current picture.

number

[0023] In an embodiment, a subgroup of transform sets may be selected from the group of transform sets based on a prediction mode of the block indicated in the coded information of the block. A feature vector extracted from reconstructed samples in one or more neighboring blocks of the block.

number

[0024] In one example, a feature vector extracted from the reconstruction samples in one or more neighboring blocks of the block

number

[0025] In one example, a transform set may be selected from a subgroup of transform sets based on an index signaled in the coded information.

number

[0026] In one example, a feature vector extracted from the reconstruction samples in one or more neighboring blocks of the block

number

[0027] In one example, a feature vector,

number

number

number

number

[0028] In one example, the feature vector

number

[0029] In one example, a threshold set K S The processing circuit (i) calculates the moment of the variables and the threshold set K S the coding information index is used to determine a transformation set from the subgroup of transformation sets based on a threshold from the coding information index; (ii) determining candidate transformations from the subgroup of transformation sets based on moments of the variables and a threshold; or (iii) selecting a transformation set from the subgroup of transformation sets based on an index of the coded information and determining candidate transformations from the selected transformation set based on moments of the variables and a threshold.

[0030] In one example, the moment of the variable is one of a first moment of the variable, a second moment of the variable, or a third moment of the variable. The prediction mode of the block is one of a plurality of prediction modes, each of which is determined by a threshold set K. S A unique threshold subset K in S ', a unique threshold subset K S ' is a set of multiple prediction modes and a threshold set K S 4 shows an injective mapping between multiple threshold subsets in

[0031] In one example, the processing circuitry may select a threshold set K based on one of: (i) a block size of the block; (ii) a quantization parameter; or (iii) a prediction mode of the block. SSelect a threshold value from

[0032] In one example, the feature vector

number

number

number

[0033] In one example, the classification vector set

number

number

number

number

[0034] Aspects of the present disclosure further provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to perform a method for video decoding and / or encoding.

[0035] Further features, the nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. [Brief description of the drawings]

[0036] [Figure 1A] FIG. 2 is a schematic diagram of an example subset of intra-prediction modes. [Figure 1B] FIG. 2 is a diagram of an example intra-prediction direction. [Diagram 2] FIG. 2 is a schematic diagram of a current block and its surrounding spatial merge candidates in one example. [Diagram 3] FIG. 3 is a schematic diagram of a simplified block diagram of a communication system (300) according to one embodiment. [Figure 4] FIG. 4 is a schematic diagram of a simplified block diagram of a communication system (400) according to one embodiment. [Diagram 5] FIG. 2 is a schematic diagram of a simplified block diagram of a decoder according to an embodiment; [Figure 6] FIG. 2 is a schematic diagram of a simplified block diagram of an encoder according to an embodiment; [Figure 7] 4 shows a block diagram of an encoder according to another embodiment; [Figure 8]4 shows a block diagram of a decoder according to another embodiment; [Figure 9] 4 illustrates an example of a nominal mode for a coding block according to an embodiment of the present disclosure. [Figure 10] 1 illustrates an example of a non-directional smooth intra-prediction mode according to aspects of the present disclosure. [Figure 11] 1 illustrates an example of an intra predictor based on recursive filtering according to an embodiment of the present disclosure. [Figure 12] 1 illustrates an example of multiple reference lines for a coding block according to an embodiment of the present disclosure. [Figure 13] 4 illustrates an example of a linear transformation basis function according to an embodiment of the present disclosure. [Figure 14A] 1 illustrates an example dependency of availability of various transform kernels based on transform block size and prediction mode according to an embodiment of the present disclosure. [Figure 14B] 1 illustrates an example transform type selection based on intra-prediction mode according to one embodiment of this disclosure. [Figure 15] 17 shows two example transform coding processes (1700) and (1800) using a 16x64 transform and a 16x48 transform, respectively, according to an embodiment of the present disclosure. [Figure 16] 17 shows two example transform coding processes (1700) and (1800) using a 16x64 transform and a 16x48 transform, respectively, according to an embodiment of the present disclosure. [Figure 17] 17(A)-17(D) show example residual patterns (grayscale) observed for intra prediction modes according to embodiments of this disclosure. [Figure 18] 4 illustrates example spatially adjacent samples of a block according to an embodiment of the present disclosure. [Figure 19] 19 shows a flowchart outlining a process (1900) according to an embodiment of the present disclosure. [Figure 20] FIG. 1 is a schematic diagram of a computer system according to an embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0037] FIG. 3 shows a schematic diagram of a simplified block diagram of a communication system (300) according to an embodiment of the present disclosure. The communication system (300) includes a plurality of terminal devices capable of communicating with each other, for example, via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). In the example of FIG. 3, the first pair of terminal devices (310) and (320) perform unidirectional transmission of data. For example, the terminal device (310) may code video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to the other terminal device (320) via the network (350). The coded video data may be transmitted in the form of one or more coded video bitstreams. The terminal device (320) may receive the coded video data from the network (350), decode the coded video data to reconstruct the video pictures, and display the video pictures according to the reconstructed video data. Unidirectional data transmission is common, such as in media distribution applications.

[0038] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) performing bidirectional transmission of coded video data, such as may occur during a video conference. For bidirectional transmission of data, in one example, each of the terminal devices (330) and (340) may code video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) over the network (350). Each of the terminal devices (330) and (340) may receive the coded video data transmitted by the other of the terminal devices (330) and (340), decode the coded video data to reconstruct the video pictures, and cause the video pictures to be displayed on an accessible display device according to the reconstructed video data.

[0039] In the example of FIG. 3, the terminal devices (310), (320), (330), and (340) may be depicted as a server, a personal computer, and a smartphone, but the principles of the present disclosure are not limited thereto. The embodiments of the present disclosure may be applied to laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (350) represents any number of networks that transmit coded video data between the terminal devices (310), (320), (330), and (340), including, for example, wired and / or wireless communication networks. The communication network (350) may exchange data over circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this discussion, the architecture and topology of the network (350) may not be important to the operation of the present disclosure, unless described below.

[0040] 4 shows an arrangement of a video encoder and a video decoder in a streaming environment as an example of an application of the disclosed subject matter. The disclosed subject matter is similarly applicable to other video function applications including, for example, video conferencing, digital TV, and storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0041] The streaming system may include a capture subsystem (413), which may include a video source (401), such as a digital camera, that creates an uncompressed video picture stream (402). In one example, the video picture stream (402) includes samples captured by a digital camera. The video picture stream (402), shown in bold to emphasize its high amount of data compared to the coded video data (404) (or coded video bitstream), may be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The coded video data (404) (or coded video bitstream (404)), shown in thin to emphasize its low amount of data compared to the stream of video pictures (402), may be stored in a streaming server (405) for future use. One or more streaming client subsystems, such as client subsystems (406) and (408) in FIG. 4, can access the streaming server (405) to retrieve copies (407) and (409) of the coded video data (404). The client subsystem (406) can include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes an input copy (407) of the coded video data and creates an output stream of video pictures (411) that can be displayed on a display (412) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the coded video data (404), (407), and (409) (e.g., a video bitstream) can be coded according to a particular video coding / compression standard. Examples of these standards include ITU-T Recommendation H.265.In one example, the developing video coding standard is informally known as Versatile Video Coding (VVC), and the disclosed subject matter may be used in the context of VVC.

[0042] It should be noted that the electronics (420) and (430) may include other components (not shown). For example, the electronics (420) may include a video decoder (not shown), and the electronics (430) may include a video encoder (not shown).

[0043] 5 shows a block diagram of a video decoder (510) according to an embodiment of the present disclosure. The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., a receiving circuit). The video decoder (510) can be used in place of the video decoder (410) in the example of FIG. 4.

[0044] The receiver (531) can receive one or more coded video sequences decoded by the video decoder (510), in the same or another embodiment, one coded video sequence at a time, with the decoding of each coded video sequence being independent of the other coded video sequences. The coded video sequences can be received from a channel (501), which can be a hardware / software link to a storage device that stores the coded video data. The receiver (531) can receive the coded video data with other data, e.g., coded audio data and / or auxiliary data streams, which can be forwarded to a respective using entity (not shown). The receiver (531) can separate the coded video sequences from the other data. To prevent network jitter, a buffer memory (515) can be coupled between the receiver (531) and the entropy decoder / parser (520), hereafter referred to as the "parser (520)". In certain applications, the buffer memory (515) is part of the video decoder (510). In other embodiments, the buffer memory may be external to the video decoder (510) (not shown). In yet another embodiment, there may be a buffer memory (not shown) external to the video decoder (510) to, for example, prevent network jitter, and another buffer memory (515) internal to the video decoder (510) to, for example, handle playback timing. When the receiver (531) receives data from a store / forward device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory (515) may not be needed or may be small. For use with best-effort packet networks such as the Internet, the buffer memory (515) may be needed, and may be relatively large and advantageously adaptively sized, and may be implemented at least in part in an operating system or similar element (not shown) external to the video decoder (510).

[0045] The video decoder (510) may include a parser (520) that reconstructs symbols (521) from the coded video sequence. These categories of symbols include information used to manage the operation of the video decoder (510) and potential information to control a rendering device, such as a rendering device (512) (e.g., a display screen) that is not part of the electronics (530) but may be coupled to the electronics (530) as shown in FIG. 5. The control information for the rendering device may be in the form of a supplemental enhancement information (SEI message) or a video usability information (VUI) parameter set fragment (not shown). The parser (520) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context dependency, etc. The parser (520) may extract a set of subgroup parameters for at least one subgroup of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroups may include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser (520) may also extract information such as transform coefficients, quantization parameter values, motion vectors, etc. from the coded video sequence.

[0046] The parser (520) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to produce symbols (521).

[0047] The reconstruction of the symbols (521) can involve several different units, depending on the type of coded video picture or part thereof (e.g., inter and intra pictures, inter and intra blocks), and other factors. Which units are involved and how may be controlled by subgroup control information parsed from the coded video sequence by the parser (520). For clarity, the flow of such subgroup control information between the parser (520) and the following units is not shown.

[0048] In addition to the functional blocks already mentioned, the video decoder (510) may be conceptually subdivided into several functional units as described below. In practical implementations subject to commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate:

[0049] The first unit is a scalar / inverse transform unit (551), which receives control information from the parser (520) including the transform to be used, block size, quantization factor, quantization scaling matrix, etc., as well as the quantized transform coefficients as symbols (521). The scalar / inverse transform unit (551) can output blocks containing sample values ​​that can be input to an aggregator (555).

[0050] In some cases, the output samples of the scaler / inverse transform (551) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture, but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates blocks of the same size and shape of the block being reconstructed using surrounding already reconstructed information retrieved from a current picture buffer (558). The current picture buffer (558) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. The aggregator (555) adds, possibly on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).

[0051] In other cases, the output samples of the scalar / inverse transform unit (551) may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion compensated prediction unit (553) may access the reference picture memory (557) to retrieve samples for prediction. After motion compensating the retrieved samples according to the symbols (521) related to the block, these samples may be added to the output of the scalar / inverse transform unit (551) by the aggregator (555) to generate output sample information (in this case referred to as residual samples or residual signals). The addresses in the reference picture memory (557) from which the motion compensated prediction unit (553) retrieves prediction samples may be controlled by motion vectors available to the motion compensated prediction unit (553), for example in the form of symbols (521) that may have X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​retrieved from the reference picture memory (557) when sub-sample accurate motion vectors are in use, motion vector prediction mechanisms, etc.

[0052] The output samples of the aggregator (555) may be subjected to various loop filtering techniques in a loop filter unit (556). The video compression techniques are controlled by parameters made available to the loop filter unit (556) as symbols (521) from the parser (520) contained in the coded video sequence (also called a coded video bitstream), and may include loop filter techniques that are responsive to meta-information obtained during the progress of decoding of a coded picture or previous parts of the coded video sequence (in decoding order), as well as to previously reconstructed loop filtered sample values.

[0053] The output of the loop filter unit (556) may be a sample stream that can be output to a rendering device (512) as well as stored in a reference picture memory (557) for use in future inter-picture prediction.

[0054] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future predictions. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a new current picture buffer can be reallocated before beginning reconstruction of the next coded picture.

[0055] The video decoder (510) may perform decoding operations according to a given video compression technique in a standard such as ITU-T Rec. H.265. The coded video sequence may comply with the syntax specified by the video compression technique or standard being used, in the sense that the coded video sequence complies with both the syntax of the video compression technique or standard and the profile documented in the video compression technique or standard. Specifically, the profile may select certain tools from all tools available in the video compression technique or standard as unique tools available in that profile. Compliance also requires that the complexity of the coded video sequence is within a range defined by the level of the video compression technique or standard. In some cases, the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. are limited by the level. The limits set by the level may be further limited in some cases by a hypothetical reference decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.

[0056] In an embodiment, the receiver (531) can receive additional (redundant) data along with the coded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0057] 6 shows a block diagram of a video encoder (603) according to an embodiment of the present disclosure. The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) may be used in place of the video encoder (403) in the example of FIG. 4.

[0058] The video encoder (603) can receive video samples from a video source (601) (not part of the electronics (620) in the example of FIG. 6) that can capture video images to be coded by the video encoder (603). In another example, the video source (601) is part of the electronics (620).

[0059] The video source (601) may provide a source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream that may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...) and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media distribution system, the video source (601) may be a storage device that stores pre-prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as multiple individual pictures that are given motion when viewed in succession. The pictures themselves may be organized as a spatial array of pixels, each of which may contain one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following discussion focuses on samples.

[0060] According to one embodiment, the video encoder (603) can code and compress pictures of a source video sequence into a coded video sequence (643) in real time or under any other time constraint required by the application. Enforcing the appropriate coding rate is one of the functions of the controller (650). In some embodiments, the controller (650) controls and is functionally coupled to other functional units described below. For clarity, the couplings are not depicted. Parameters set by the controller (650) can include rate control related parameters (picture skip, quantization, lambda values ​​for rate distortion optimization techniques, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured with other appropriate functions for the video encoder (603) optimized for a particular system design.

[0061] In some embodiments, the video encoder (603) is configured to operate in a coding loop. As an oversimplified explanation, in one example, the coding loop may include a source coder (630) (e.g., responsible for creating symbols, such as a symbol stream based on the input picture to be coded and reference pictures) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a similar manner as the (remote) decoder would create them (since any compression between the symbols and the coded video bitstream is lossless in the video compression techniques considered in the disclosed subject matter). The reconstructed sample stream (sample data) is input to a reference picture memory (634). Since the decoding of the symbol stream leads to a bit-exact result regardless of the location of the decoder (local or remote), the contents of the reference picture memory (634) are also bit-exact between the local and remote encoders. In other words, the prediction part of the encoder "sees" as reference picture samples exactly the same sample values ​​that the decoder "sees" when using the prediction during decoding. Such basic principles of reference picture synchrony (and the drift that occurs when synchrony cannot be maintained due to, for example, channel errors) are also used in several related fields.

[0062] The operation of the "local" decoder (633) may be similar to the operation of a "remote" decoder, such as the video decoder (510), already described in detail above with reference to Figure 5. However, with brief reference to Figure 5, because symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder (645) and parser (520) may be lossless, the entropy decoding portion of the video decoder (510), including the buffer memory (515) and parser (520), may not be fully implemented in the local decoder (633).

[0063] As can be seen from this point, any decoder technique other than analysis / entropy decoding present in a decoder must necessarily be present in a corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter focuses on the decoder operation. A description of the encoder techniques can be omitted since they are the opposite of the decoder techniques described generically. Only in certain areas is more detailed description required and is provided below.

[0064] During operation, in some examples, the source coder (630) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as “reference pictures.” In this manner, the coding engine (632) codes differences between pixel blocks of the input picture and pixel blocks of the reference pictures that may be selected as prediction references for the input picture.

[0065] The local video decoder (633) may decode the coded video data of pictures that may be designated as reference pictures based on symbols created by the source coder (630). The operation of the coding engine (632) may advantageously be a lossy process. When the coded video data can be decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence may be a replica of the source video sequence, usually with some errors. The local video decoder (633) may replicate the decoding process that may be performed by the video decoder for the reference pictures and store the reconstructed reference pictures in a reference picture cache (634). In this way, the video encoder (603) may locally store copies of reconstructed reference pictures that have common content (no transmission errors) with the reconstructed reference pictures obtained by the far-end video decoder.

[0066] The predictor (635) can perform a prediction search for the coding engine (632). That is, for a new picture to be coded, the predictor (635) can search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or specific metadata that can serve as suitable prediction criteria for the new picture, such as the reference picture's motion vectors, block shapes, etc. The predictor (635) can operate on a sample block / pixel block basis to find a suitable prediction criteria. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have prediction criteria drawn from multiple reference pictures stored in the reference picture memory (634).

[0067] The controller (650) can manage the coding operations of the source coder (630), including, for example, setting the parameters and subgroup parameters used to code the video data.

[0068] The output of all of the aforementioned functional units may undergo entropy coding in an entropy coder (645), which converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0069] The transmitter (640) may buffer the coded video sequence created by the entropy coder (645) in preparation for transmission over a communication channel (660), which may be a hardware / software link to a storage device that stores the coded video data. The transmitter (640) may merge the coded video data from the video coder (603) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0070] The controller (650) may manage the operation of the video encoder (603). During coding, the controller (650) may assign a particular coding picture type to each coded picture, which may affect the coding technique that may be applied to the respective picture. For example, pictures may generally be assigned to one of the following picture types:

[0071] An Intra picture (I-picture) may be one that can be coded and decoded without using any other picture in a sequence as a source of prediction. Some video codecs allow different types of Intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures. Those skilled in the art are aware of these variations of I-pictures and their respective uses and characteristics.

[0072] A predictive picture (P picture) may be one that can be coded and decoded by intra- or inter-prediction using at most one motion vector and reference index to predict sample values ​​of each block.

[0073] Bidirectionally predicted pictures (B-pictures) may be those that can be coded and decoded by intra- or inter-prediction using up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predicted pictures may use more than two reference pictures and associated metadata to reconstruct a single block.

[0074] A source picture may generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples, respectively) and coded block by block. A block may be predictively coded with reference to other (already coded) blocks determined by a coding assignment applied to the block's respective picture. For example, a block of an I picture may be non-predictively coded or may be predictively coded with reference to already coded blocks of the same picture (spatial or intra prediction). A pixel block of a P picture may be predictively coded via spatial prediction or via temporal prediction with reference to one pre-coded reference picture. A block of a B picture may be predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.

[0075] The video encoder (603) may perform coding operations according to a pre-defined video coding technique or standard, such as ITU-T Rec. H.265. During operation, the video encoder (603) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard used.

[0076] In an embodiment, the transmitter (640) can transmit additional data along with the coded video. The source coder (630) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, redundant data in other forms such as redundant pictures or slices, SEI messages, VUI parameter set fragments, etc.

[0077] A video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as "intra prediction") exploits spatial correlation in a given picture, while inter-picture prediction exploits correlation (temporal or other) between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. If a block in the current picture is similar to a reference block of a previously coded and still buffered reference picture in the video, the block in the current picture may be coded by a vector called a motion vector. The motion vector points to a reference block of the reference picture, and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0078] In some embodiments, bidirectional prediction may be used in interpicture prediction. Bidirectional prediction uses two reference pictures, such as a first reference picture and a second reference picture, each of which is earlier in decoding order than the current picture in a video (but may be in the past and future in display order, respectively). A block in the current picture may be coded by a first motion vector that points to a first reference block in the first reference picture and a second motion vector that points to a second reference block in the second reference picture. A block may be predicted by a combination of the first reference block and the second reference block.

[0079] Furthermore, merge mode techniques can be applied to inter-picture prediction to improve coding efficiency.

[0080] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed on a block-by-block basis. For example, according to the HEVC standard, a picture in a sequence of video pictures is divided into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. In general, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU may be recursively quadtree partitioned into one or more coding units (CUs). For example, a CTU of 64×64 pixels may be partitioned into one CU of 64×64 pixels, four CUs of 32×32 pixels, or sixteen CUs of 16×16 pixels. In one example, each CU is analyzed to determine the prediction type of the CU, such as an inter-prediction type or an intra-prediction type. The CU is partitioned into one or more prediction units (PUs) depending on the temporal and / or spatial predictability. Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In an embodiment, prediction operations in coding (encoding / decoding) are performed in units of prediction blocks. Taking the luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values ​​(e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0081] 7 shows a diagram of a video encoder (703) according to another embodiment of the present disclosure. The video encoder (703) is configured to receive a processed block (e.g., a predictive block) of sample values ​​in a current video picture in a video picture sequence and to code the processed block into a coded picture that is part of the coded video sequence. In one example, the video encoder (703) is used instead of the video encoder (403) in the example of FIG. 4.

[0082] In an HEVC example, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as a predictive block, such as 8×8 samples. The video encoder (703) determines whether the processing block is best coded in an intra mode, an inter mode, or a bi-predictive mode, for example using rate-distortion optimization. If the processing block is to be coded in an intra mode, the video encoder (703) may code the processing block into a coded picture using an intra prediction method. Also, if the processing block is to be coded in an inter mode or a bi-predictive mode, the video encoder (703) may code the processing block into a coded picture using an inter prediction or bi-predictive method, respectively. In certain video coding techniques, the merge mode may be an inter-picture prediction sub-mode that derives motion vectors from one or more motion vector predictors without going through a coded motion vector component outside the predictor. In certain other video coding techniques, there may be motion vector components applicable to the current block. In one example, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the mode of the processing block.

[0083] In the example of FIG. 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725), all coupled together as shown in FIG.

[0084] The inter-encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block to one or more reference blocks in a reference picture (e.g., blocks in previous preceding and following pictures), generate inter-prediction information (e.g., a description of redundant information due to an inter-coding scheme, motion vectors, merge mode information), and calculate an inter-prediction result (e.g., a predicted block) using any suitable technique based on the inter-prediction information. In some examples, the reference picture is a decoded reference picture based on the coded video information.

[0085] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), possibly compare the block with previously coded blocks in the same picture, generate quantized coefficients after transformation, and possibly also generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques). In one example, the intra encoder (722) also calculates intra prediction results (e.g., predicted blocks) based on reference blocks and intra prediction information in the same picture.

[0086] The generic controller (721) is configured to determine generic control data and control other components of the video encoder (703) based on the generic control data. In one example, the generic controller (721) determines a mode of the block and provides a control signal to the switch (726) based on the mode. For example, if the mode is an intra mode, the generic controller (721) controls the switch (726) to select an intra mode result for use by the residual calculator (723) and controls the entropy encoder (725) to select intra prediction information and include the intra prediction information in the bitstream. If the mode is an inter mode, the generic controller (721) controls the switch (726) to select an inter prediction result for use by the residual calculator (723) and controls the entropy encoder (725) to select inter prediction information and include the inter prediction information in the bitstream.

[0087] The residual calculation unit (723) is configured to calculate a difference (residual data) between the received block and a prediction result selected from the intra-encoder (722) or the inter-encoder (730). The residual encoder (724) is configured to operate on the residual data and code the residual data to generate transform coefficients. In one example, the residual encoder (724) is configured to transform the residual data from a spatial domain to a frequency domain to generate transform coefficients. Then, a quantization process is performed on the transform coefficients to obtain quantized transform coefficients. In various embodiments, the video encoder (703) further includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data may be used by the intra-encoder (722) and the inter-encoder (730) as appropriate. For example, the inter-encoder (730) can generate a decoding block based on the decoded residual data and the inter-prediction information, and the intra-encoder (722) can generate a decoding block based on the decoded residual data and the intra-prediction information. In some examples, the decoding block is suitably processed to generate a decoding picture, which may be buffered in a memory circuit (not shown) and used as a reference picture.

[0088] The entropy encoder (725) is configured to format the bitstream to include the coded block. The entropy encoder (725) is configured to include various information according to an appropriate standard, such as the HEVC standard. In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other appropriate information in the bitstream. It should be noted that, according to the disclosed subject matter, there is no residual information when coding a block in an inter mode or a merged sub-mode of a bi-predictive mode.

[0089] 8 shows a diagram of a video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive coded pictures that are part of a coded video sequence and decode the coded pictures to generate reconstructed pictures. In one example, the video decoder (810) is used instead of the video decoder (410) in the example of FIG. 4.

[0090] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter-decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-decoder (872), coupled to each other as shown in FIG.

[0091] The entropy decoder (871) may be configured to reconstruct from the coded picture certain symbols that represent syntax elements that make up the coded picture. Such symbols may include, for example, prediction information (e.g., intra- or inter-prediction information) that may identify the mode in which the block is coded (e.g., intra-mode, inter-mode, bi-prediction mode, later merged submode of both, or other submode), certain samples or metadata used for prediction by the intra-decoder (872) or inter-decoder (880), respectively, residual information, for example, in the form of quantized transform coefficients, etc. In one example, if the prediction mode is an inter- or bi-prediction mode, the inter-prediction information is provided to the inter-decoder (880). Also, if the prediction type is an intra-prediction type, the intra-prediction information is provided to the intra-decoder (872). The residual information may be dequantized and provided to the residual decoder (873).

[0092] The inter decoder (880) is configured to receive the inter prediction information and to generate inter prediction results based on the inter prediction information.

[0093] The intra decoder (872) is configured to receive the intra prediction information and to generate a prediction result based on the intra prediction information.

[0094] The residual decoder (873) is configured to perform inverse quantization to extract inverse quantized transform coefficients, and process the inverse quantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) may require certain control information (such as to include quantizer parameters (QP)), which may be provided by the entropy decoder (871) (a data path not shown, as there may only be a small amount of control information).

[0095] The reconstruction module (874) is configured to combine, in the spatial domain, the residual output by the residual decoder (873) and a prediction result (possibly output by an inter- or intra-prediction module) to form a reconstructed block that may be part of a reconstructed picture that may be part of the reconstructed video. Note that other suitable operations, such as a deblocking operation, may be performed to improve visual quality.

[0096] It should be noted that the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using any suitable technology. In an embodiment, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using one or more processors executing software instructions.

[0097] Disclosed are video coding techniques related to a transform set or transform kernel selection scheme based on reconstructed samples (e.g., feature indicators (e.g., feature vectors or feature scalars) of one or more neighboring blocks of a block. The video coding format may include an open video coding format designed for video transmission over the Internet, such as AOMedia Video 1 (AV1) or a next-generation AOMedia Video format beyond AV1. Also, the video coding standard may include the High Efficiency Video Coding (HEVC) standard, a next-generation video coding beyond HEVC (e.g., Versatile Video Coding (VVC)), etc.

[0098] In intra prediction, e.g., AV1, VVC, etc., various intra prediction modes may be used. In an embodiment, e.g., in AV1, directional intra prediction is used. In directional intra prediction, the predicted samples of a block may be generated by extrapolating from neighboring reconstructed samples along one direction. The direction corresponds to an angle. The modes used in directional intra prediction to predict the predicted samples of a block may be called directional modes (also called directional prediction modes, directional intra modes, directional intra prediction modes, angle modes). Each directional mode may correspond to a different angle or a different direction. In one example, e.g., in the open video coding format VP9, ​​eight directional modes are used, corresponding to eight angles from 45° to 207°. The eight directional modes may also be called nominal modes (e.g., V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED). To take advantage of more diverse spatial redundancy in directional textures (e.g., AV1), the directional modes can be extended, for example, beyond the eight nominal modes to angle sets with finer granularity and more angles (or directions), as shown in FIG. 9.

[0099] FIG. 9 illustrates an example of nominal modes of a coding block (CB) (910) according to an embodiment of the present disclosure. A particular angle (referred to as a nominal angle) may correspond to a nominal mode. In one example, eight nominal angles (or nominal intra-angles) (901)-(908) correspond to eight nominal modes (e.g., V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED). Also, the eight nominal angles (901)-(908) and the eight nominal modes may be referred to as V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED, respectively. The nominal mode index may indicate a nominal mode (e.g., one of eight nominal modes). In one example, the nominal mode index is signaled.

[0100] Further, each nominal angle can correspond to multiple finer angles (e.g., seven finer angles), and thus, for example, in AV1, 56 angles (or predicted angles) or 56 directional modes (or angle modes, directional intra prediction modes) can be used. Each predicted angle can be represented by a nominal angle and an angle offset (or angle delta). The angle offset is obtained by multiplying an offset integer I (e.g., -3, -2, -1, 0, 1, 2, or 3) by a step size (e.g., 3°). In one example, the predicted angle is equal to the sum of the nominal angle and the angle offset. In one example, for example, in AV1, the nominal modes (e.g., eight nominal modes (901)-(908)) can be signaled along with a particular non-angular smooth mode (e.g., DC mode, PAETH mode, SMOOTH mode, vertical SMOOTH mode, and horizontal SMOOTH mode, which will be described later). Subsequently, if the current prediction mode is a directional mode (or an angular mode), an index indicating an angular offset (e.g., offset integer I) corresponding to the nominal angle may be further signaled. In one example, the directional mode (e.g., one of 56 directional modes) may be determined based on the nominal mode index and an index indicating an angular offset from the nominal mode. In one example, to realize the directional prediction mode via a generic method, the 56 directional modes as used in AV1 are realized with a unified directional predictor that can project each pixel to a reference sub-pixel position and interpolate the reference pixel by a 2-tap bilinear filter.

[0101] A non-directional smooth intra predictor (also referred to as a non-directional smooth intra prediction mode, a non-directional smooth mode, or a non-angular smooth mode) may be used for intra prediction of CB. In some examples (e.g., in AV1), the five non-directional smooth intra prediction modes include a DC mode or DC predictor (e.g., DC), a PAETH mode or PAETH predictor (e.g., PAETH), a SMOOTH mode or SMOOTH predictor (e.g., SMOOTH), a vertical SMOOTH mode (referred to as a SMOOTH_V mode, a SMOOTH_V predictor, or SMOOTH_V), and a horizontal SMOOTH mode (referred to as a SMOOTH_H mode, a SMOOTH_H predictor, or SMOOTH_H).

[0102] 10 illustrates examples of non-directional smooth intra-prediction modes (e.g., DC mode, PAETH mode, SMOOTH mode, SMOOTH_V mode, and SMOOTH_H mode) according to an embodiment of the present disclosure. To predict a sample (1001) in CB (1000) based on a DC predictor, an average value of a first value of a left neighboring sample (1012) and a second value of an above neighboring sample (or top neighboring sample) (1011) can be used as a predictor.

[0103] To predict a sample (1001) based on the PAETH predictor, a first value of the left adjacent sample (1012), a second value of the top adjacent sample (1011), and a third value of the top-left adjacent sample (1013) can be obtained. Then, a reference value is obtained using Equation 1. Reference value = first value + second value - third value (Equation 1)

[0104] One of the first value, the second value, and the third value that is closest to the reference value can be set as a predictor for the sample (1001).

[0105] The SMOOTH_V mode, the SMOOTH_H mode, and the SMOOTH mode can predict the CB(1000) using quadratic interpolation of the average values ​​in the vertical direction, the horizontal direction, and the vertical and horizontal directions, respectively. To predict the sample (1001) based on the SMOOTH predictor, an average value (e.g., a weighted combination) of the first value, the second value, the value of the right sample (1014), and the value of the bottom sample (1016) can be used. In various examples, since the right sample (1014) and the bottom sample (1016) are not reconstructed, the value of the upper right neighboring sample (1015) and the value of the lower left neighboring sample (1017) can replace the values ​​of the right sample (1014) and the bottom sample (1016), respectively. Thus, an average value (e.g., a weighted combination) of the first value, the second value, the value of the upper right neighboring sample (1015), and the value of the lower left neighboring sample (1017) can be used as the SMOOTH predictor. To predict the sample (1001) based on the SMOOTH_V predictor, an average (e.g., weighted combination) of the second value of the top neighboring sample (1011) and the value of the bottom-left neighboring sample (1017) can be used. To predict the sample (1001) based on the SMOOTH_H predictor, an average (e.g., weighted combination) of the first value of the left neighboring sample (1012) and the value of the top-right neighboring sample (1015) can be used.

[0106] FIG. 11 illustrates an example of an intra predictor based on recursive filtering (also referred to as filter intra mode, or recursive filtering mode) according to an embodiment of the present disclosure. Filter intra mode can be used for CB (1100) to capture the decaying spatial correlation with the reference on the edge. In one example, CB (1100) is a luma block. The luma block (1100) can be divided into multiple patches (e.g., eight 4×2 patches B0-B7). Each of the patches B0-B7 can have multiple neighboring samples. For example, patch B0 has seven neighboring samples (or seven neighboring elements) R00-R06, including four top neighboring samples R01-R04, two left neighboring samples R05-R06, and a top-left neighboring sample R00. Similarly, patch B7 has four top neighboring samples R71-R74, two left neighboring samples R75-R76, and seven neighboring samples R70-R76, including the top-left neighbor sample R70.

[0107] In some examples, multiple (e.g., five) filter intra modes (or multiple recursive filtering modes) are pre-designed, e.g., for AV1. Each filter intra mode may be represented by a set of eight 7-tap filters that reflect the correlation between a sample (or pixel) in a corresponding 4×2 patch (e.g., B0) and seven neighbors (e.g., R00-R06) adjacent to the 4×2 patch B0. The weighting coefficients of the 7-tap filters may be position-dependent. For each of the patches B0-B7, the seven neighbors (e.g., R00-R06 for B0 and R70-R76 for B7) may be used to predict a sample in the corresponding patch. In one example, the neighboring elements R00-R06 are used to predict a sample in patch B0. In one example, the neighboring elements R70-R76 are used to predict a sample in patch B7. For a particular patch in CB (1100), such as patch B0, all of the seven neighboring elements (e.g., R00-R06) have already been reconstructed. For other patches in CB(1100), at least one of the seven neighbors has not been reconstructed, so the predicted value of the nearest neighbor (or the predicted sample of the nearest neighbor) can be used as a reference. For example, the seven neighbors R70-R76 of patch B7 have not been reconstructed, so the predicted sample of the nearest neighbor can be used.

[0108] Chroma samples may be predicted from luma samples. In one embodiment, a chroma from luma mode (e.g., CfL mode, CfL predictor) is a chroma-only intra predictor that can model a chroma sample (or pixel) as a linear function of the corresponding reconstructed luma sample (or pixel). For example, CfL prediction can be expressed using Equation 2 as follows: CfL(α)=αL A +D (Formula 2) Here, L Arepresents the AC contribution of the luma component, α represents a scaling parameter of the linear model, and D represents the DC contribution of the chroma component. In one example, the reconstructed luma pixels are subsampled based on the chroma resolution and the mean value is subtracted to reduce the AC contribution (e.g., L A ) is obtained. Instead of the decoder calculating the scaling parameter α to approximate the chroma AC components from the AC contributions, in some instances, e.g., in AV1, the CfL mode determines the scaling parameter α based on the original chroma pixels and signals the scaling parameter α in the bitstream, resulting in a more accurate prediction while reducing decoder complexity. The DC contributions of the chroma components are obtained using an intra DC mode, which is sufficient for most chroma content and has a mature, high-speed implementation.

[0109] Multi-line intra prediction may use more reference lines for intra prediction. A reference line may include multiple samples in a picture. In one example, a reference line includes row samples and column samples. In one example, an encoder may determine and signal a reference line used to generate an intra predictor. An index indicating the reference line (also called a reference line index) may be signaled before an intra prediction mode. In one example, if a non-zero reference line index is signaled, only MPM is allowed. FIG. 12 shows an example of four reference lines for CB (1210). With reference to FIG. 12, a reference line may include up to six segments, e.g., segments A-F and a top-left reference sample. For example, reference line 0 includes segments B and E and a top-left reference sample. For example, reference line 3 includes segments A-F and a top-left reference sample. Segments A and F may be padded with nearest neighbor samples from segments B and E, respectively. In some examples, for example, in HEVC, only one reference line (e.g., reference line 0 adjacent to CB (1210)) is used for intra prediction. In some examples, for example, in VVC, multiple reference lines (e.g., reference lines 0, 1, 3) are used for intra prediction.

[0110] Typically, a block may be predicted using one or a suitable combination of various intra-prediction modes, such as those described above with reference to Figures 9-12.

[0111] In the following, an embodiment of a linear transform as used in AOMedia Video 1 (AV1) is described. A forward transform (e.g., in an encoder) may be performed on a transform block (TB) including a residual (e.g., a residual in a spatial domain) so that a TB including a transform coefficient in a frequency domain (or a spatial frequency domain) is obtained. The TB including the residual in the spatial domain is called a residual TB, and the TB including the transform coefficient in the frequency domain is called a coefficient TB. In one example, the forward transform includes a forward linear transform that can transform the residual TB into a coefficient TB. In one example, the forward transform includes a forward linear transform and a forward secondary transform, in which the forward linear transform can transform the residual TB into an intermediate coefficient TB, and the forward secondary transform can transform the intermediate coefficient TB into a coefficient TB.

[0112] An inverse transform (e.g., in an encoder or decoder) may be performed on the coefficients TB in the frequency domain to obtain residuals TB in the spatial domain. In one example, the inverse transform includes an inverse primary transform that can transform the coefficients TB to residuals TB. In one example, the inverse transform includes an inverse primary transform and an inverse secondary transform, where the inverse secondary transform can transform the coefficients TB to intermediate coefficients TB, and the inverse primary transform can transform the intermediate coefficients TB to residuals TB.

[0113] Generally, the primary transform can refer to a forward primary transform or an inverse primary transform, where the primary transform is performed between the residual TB and the coefficient TB. In some embodiments, the primary transform can be a separable transform. Here, the 2D primary transform may include a horizontal primary transform (also called horizontal transform) and a vertical primary transform (also called vertical transform). The secondary transform can refer to a forward secondary transform or an inverse secondary transform, where the secondary transform is performed between the intermediate coefficient TB and the coefficient TB.

[0114] To support extended coding block partitions as described in this disclosure, multiple transform sizes (e.g., ranging from 4 points to 64 points for each dimension) and transform shapes (e.g., square, rectangle with width to height ratios of 2:1, 1:2, 4:1 or 1:4) may be used, such as in AV1.

[0115] In the 2D transform process, a hybrid transform kernel may be used that may include different 1D transforms for each dimension of the coded residual block. The linear 1D transform may include (a) 4-point, 8-point, 16-point, 32-point, 64-point DCT-2; (b) 4-point, 8-point, 16-point Asymmetric DST (ADST) (e.g., DST-4, DST-7, etc.) and corresponding flipped versions (e.g., a flipped version of ADST or FlipADST can apply ADST in a reverse order) and / or (c) 4-point, 8-point, 16-point, 32-point Identity Transform (IDTX). Figure 13 shows an example of linear transform basis functions according to an embodiment of the present disclosure. The linear transform basis functions in the example of Figure 13 include basis functions of DCT-2 and Asymmetric DST (DST- and DST-7) with N-point input. The linear transform basis functions shown in Figure 13 may be used for AV1.

[0116] The availability of hybrid transform kernels may depend on transform block size and prediction mode. Figure 14A illustrates an example dependency of the availability of various transform kernels (e.g., transform types shown in the first column and described in the second column) based on transform block size (e.g., sizes shown in the third column) and prediction mode (e.g., intra prediction and inter prediction shown in the third column). Exemplary hybrid transform kernels and their availability based on prediction mode and transform block size can be used in AV1. Referring to Figure 14A, the symbol

number

number

number

[0117] In one example, the transform type (1410) is represented by ADST_DCT as shown in the first column of Figure 14A. The transform type (1410) includes ADST in the vertical direction and DCT in the horizontal direction as shown in the second column of Figure 14A. According to the third column of Figure 14A, the transform type (1410) is available for intra prediction and inter prediction when the block size is 16x16 (e.g., 16x16 samples, 16x16 luma samples) or less.

[0118] In one example, the transform type (1420) is indicated by V_ADST, as shown in the first column of FIG. 14A. The transform type (1420) includes ADST in the vertical direction and IDTX (i.e., identity matrix) in the horizontal direction, as shown in the second column of FIG. 14A. Thus, the transform type (1420) (e.g., V_ADST) is performed vertically and not horizontally. According to the third column of FIG. 14A, the transform type (1420) is not available for intra prediction, regardless of the block size. The transform type (1420) is available for inter prediction when the block size is less than 16x16 (e.g., 16x16 samples, 16x16 luma samples).

[0119] In one example, Figure 14A is applicable to the luma component. For the chroma components, the selection of the transform type (or transform kernel) may be performed implicitly. In one example, for intra-prediction residuals, the transform type may be selected according to the intra-prediction mode, as shown in Figure 14B. In one example, the selection of the transform type shown in Figure 14B is applicable to the chroma components. For inter-prediction residuals, the transform type may be selected according to the transform type of the co-located luma block. Thus, in one example, the transform type for the chroma components is not signaled in the bitstream.

[0120] A transform, such as a linear transform, a secondary transform, etc., may be applied to a block, such as CB. In one example, the transform includes a combination of a linear transform, a secondary transform, etc. Also, the transform may be a non-separable transform, a separable transform, or a combination of a non-separable transform and a separable transform.

[0121] The secondary transform can be performed in VVC, etc. In some examples, as shown in Figures 15-16, for example in VVC, a low-frequency non-separable transform (LFNST) (also called a reduced secondary transform (RST)) can be applied between the forward primary transform and quantization at the encoder side and the inverse quantization and inverse primary transform at the decoder side to further decorrelate the primary transform coefficients.

[0122] The application of a non-separable transform available in LFNST can be described as follows, taking a 4×4 input block (or input matrix) X as an example (shown in Equation 3). To apply a 4×4 non-separable transform, the 4×4 input block X is divided into vectors

number

number

[0123] A non-separable transformation is

number

number

number

[0124] A non-separable secondary transform may be applied to a block such as CB. In some examples, for example, in VVC, LFNST is applied between the forward primary transform and the quantization (e.g., at the encoder side) and between the inverse quantization and the inverse primary transform (e.g., at the decoder side), as shown in Figures 15-16.

[0125] 15-16 show examples of two transform coding processes (1700) and (1800) using a 16×64 transform (or a 64×16 transform depending on whether the transform is a forward secondary transform or an inverse secondary transform) and a 16×48 transform (or a 48×16 transform depending on whether the transform is a forward secondary transform or an inverse secondary transform), respectively. Referring to FIG. 15, in the process (1700), the encoder side may first perform a forward primary transform (1710) on a block (e.g., a residual block) to obtain a coefficient block (1713). Then, a forward secondary transform (or a forward LFNST) (1712) may be applied to the coefficient block (1713). In the forward secondary transform (1712), the 64 coefficients of the 4×4 sub-blocks A through D in the top left corner of the coefficient block (1713) can be represented by a 64-length vector, which can be multiplied by a 64×16 (i.e., 64 widths and 16 heights) transform matrix to obtain a 16-length vector. The elements in the 16-length vector are backfilled into the top left 4×4 sub-block A of the coefficient block (1713). The coefficients in sub-blocks B through D can be zero. In a quantization step (1714), the resulting coefficients after the forward secondary transform (1712) are quantized and entropy coded to generate coded bits in a bitstream (1716).

[0126] The coded bits may be received at the decoder side and entropy decoded, followed by an inverse quantization step (1724) to generate a coefficient block (1723). An inverse secondary transform (or inverse LFNST) (1722), such as an inverse RST8×8, may be performed to obtain, for example, 64 coefficients from the 16 coefficients in the top-left 4×4 sub-block E. The 64 coefficients may also be backfilled into 4×4 sub-blocks E-H. The coefficients in the coefficient block (1723) after the inverse secondary transform (1722) may then be processed with an inverse primary transform (1720) to obtain a reconstructed residual block.

[0127] The process (1800) of the example in FIG. 16 is similar to the process (1700), except that fewer (i.e., 48) coefficients are processed during the forward secondary transform (1712). Specifically, the 48 coefficients within sub-blocks A - C are processed with a smaller transform matrix of size 48×16. By using the smaller 48×16 transform matrix, the memory size for storing the transform matrix and the number of calculations (e.g., multiplications, additions, subtractions, etc.) can be reduced, thus reducing the complexity of the calculations.

[0128] In one example, depending on the block size such as block CB, a 4×4 non-separable transform (e.g., 4×4 LFNST) or an 8×8 non-separable transform (e.g., 8×8 LFNST) is applied. The block size of CB can include the width, height, etc. For example, the 4×4 LFNST is applied to a CB where the minimum of the width and height is less than a threshold such that the minimum (width, height) < 8. For example, the 8×8 LFNST is applied to a CB where the minimum of the width and height is greater than a threshold such that the minimum (width, height) > 4.

[0129] The non-separable transform (e.g., LFNST) can be performed based on a direct matrix multiplication approach and thus can be implemented in a single pass without iteration. To reduce the dimension of the non-separable transform matrix and minimize the computational complexity and the memory space for storing the transform coefficients, in LFNST, a reduced non-separable transform method (or RST) can be used. Thus, in the reduced non-separable transform, a vector of dimension N (e.g., N is 64 for an 8×8 non-separable secondary transform (NSST)) can be mapped to a vector of dimension R in a different space, where N / R (R < N) is the reduction factor. Thus, the RST matrix becomes an R×N matrix as explained in Equation 5 instead of an N×N matrix.

Number

[0130] In Equation 5, the R rows of the R×N transformation matrix are the R basis for the N-dimensional space. The inverse transformation matrix is ​​the transformation matrix used in the forward transformation (e.g., T R×N ) may be the transpose of 8×8 LFNST. In the case of 8×8 LFNST, a reduction factor of 4 may be applied, and the 64×64 direct matrix used in the 8×8 non-separable transform may be reduced to a 16×64 direct matrix, as shown in FIG. 15. Alternatively, a reduction factor greater than 4 may be applied, and the 64×64 direct matrix used in the 8×8 non-separable transform may be reduced to a 16×48 direct matrix, as shown in FIG. 16. Thus, a 48×16 inverse RST matrix may be used at the decoder side to generate the core (primary) transform coefficients of the 8×8 top-left region.

[0131] 16, when a 16×48 matrix is ​​applied instead of a 16×64 matrix with the same transform set configuration, the input to the 16×48 matrix contains 48 input data from three 4×4 blocks A, B, and C in the top left 8×8 block excluding the bottom right 4×4 block D. The dimensionality reduction can reduce the memory usage for storing the LFNST matrix, for example, from 10 KB to 8 KB with minimal performance degradation.

[0132] To reduce complexity, LFNST may be restricted so that it is applicable only when the coefficients outside the first coefficient subgroup are non-significant. In one example, LFNST may be restricted so that it is applicable only when all coefficients outside the first coefficient subgroup are non-significant. With reference to Figures 15-16, the first coefficient subgroup corresponds to the top-left block E, and therefore the coefficients outside block E are non-significant.

[0133] In one example, when LFNST is applied, only the primary transform coefficients are non-significant (e.g., zero). In one example, when LFNST is applied, only all the primary transform coefficients are zero. Only the primary transform coefficients can refer to transform coefficients resulting from a primary transform without a secondary transform. Thus, the LFNST index signaling can be conditional on the last significant position, so that an extra coefficient scan in LFNST can be avoided. In some examples, the extra coefficient scan is used to check for significant transform coefficients at specific positions. In one example, the worst case scenario of LFNST, for example, in terms of per-pixel multiplications, limits non-separable transforms of 4×4 and 8×8 blocks to 8×16 and 8×48 transforms, respectively. In the above cases, when LFNST is applied, the last significant scan position can be smaller than 8. For other sizes, when LFNST is applied, the last significant scan position can be smaller than 16. For CBs of 4×N and N×4 (where N is greater than 8), this restriction can mean that LFNST is applied to the top-left 4×4 region of the CB. In one example, this restriction means that LFNST is applied only once, to the top left 4×4 region of the CB. In one example, when LFNST is applied, the number of operations for the primary transform is reduced, since only all primary coefficients are non-significant (e.g., zero). From the encoder's perspective, quantization of the transform coefficients can be greatly simplified if the LFNST transform is tested. Rate-distortion optimized quantization can be done at a maximum, for example, for the first 16 coefficients in scan order, and the remaining coefficients can be zeroed.

[0134] An LFNST transform (e.g., a transform kernel, transform core, or transform matrix) may be selected as described below. In one embodiment, multiple transform sets may be used, and one or more non-separable transform matrices (or kernels) may be included in each of the multiple transform sets in the LFNST. According to aspects of the present disclosure, a transform set may be selected from the multiple transform sets, and a non-separable transform matrix may be selected from one or more non-separable transform matrices in the transform set.

[0135] Table 1 shows an example mapping from intra-prediction modes to multiple transform sets according to one embodiment of the present disclosure. This mapping shows the relationship between intra-prediction modes and multiple transform sets. The relationship as shown in Table 1 can be predefined and stored in the encoder and the decoder. [Table 1]

[0136] Referring to Table 1, the multiple transform sets include four transform sets, for example, transform sets 0 to 3, represented by transform set indexes (e.g., Tr.set index) from 0 to 3, respectively. The index (e.g., IntraPredMode) can indicate an intra prediction mode, and the transform set index can be obtained based on the index and Table 1. Thus, the transform set can be determined based on the intra prediction mode. In one example, when one of three cross-component linear model (CCLM) modes (e.g., INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for CB (e.g., 81≦IntraPredMode≦83), transform set 0 is selected for CB.

[0137] As mentioned above, each transform set may include one or more non-separable transform matrices. One of the one or more non-separable transform matrices may be selected by an explicitly signaled LFNST index. The LFNST index may be signaled in the bitstream, for example, once per intra-coded CU (e.g., CB) after signaling the transform coefficients. In one embodiment, each transform set includes two non-separable transform matrices (kernels), and the selected non-separable secondary transform candidate may be one of the two non-separable transform matrices. In some examples, LFNST is not applied to the CB (e.g., a CB coded in transform skip mode, or the number of non-zero coefficients of the CB is less than a threshold). In one example, if LFNST is not applied to the CB, an LFNST index is not signaled for the CB. The default value of the LFNST index is zero and is not signaled, indicating that LFNST is not applied to the CB.

[0138] In one embodiment, LFNST is restricted to be applicable only when all coefficients outside the first coefficient subgroup are non-significant, and the coding of the LFNST index may be determined to the position of the last significant coefficient. The LFNST index may be context coded. In one example, the context coding of the LFNST index does not depend on the intra prediction mode, and only the first bin is context coded. LFNST may be applied to intra-coded CUs in intra slices or inter slices, and may be applied to both luma and chroma components. If dual tree is enabled, the LFNST indexes for the luma and chroma components may be signaled separately. In the case of inter slices (e.g., if dual tree is disabled), a single LFNST index may be signaled and used for both luma and chroma components.

[0139] An intra sub-partition (ISP) coding mode may be used. In the ISP coding mode, a luma intra prediction block may be divided into two or four sub-partitions vertically or horizontally, depending on the block size. In some examples, the performance improvement reaches a limit when RST is applied to all feasible sub-partitions. Therefore, in some examples, when an ISP mode is selected, LFNST is disabled and an LFNST index (or an RST index) is not signaled. Disabling RST or LFNST for ISP predicted residuals can reduce coding complexity. In some examples, when a matrix-based intra prediction mode (MIP) is selected, LFNST is disabled and an LFNST index is not signaled.

[0140] In some examples, due to a maximum transform size limitation (e.g., 64x64), CUs larger than 64x64 are implicitly split (TU tiling) and the LFNST index lookup can increase data buffering by 4x for a certain number of decoding pipeline stages. Thus, the maximum size allowed for LFNST can be limited to 64x64. In one example, LFNST is enabled only for discrete cosine transform (DCT) type 2 (DCT-2) transforms.

[0141] In some examples, a separable transform scheme may not be efficient for capturing directional texture patterns (e.g., edges along 45° or 135° directions). A non-separable transform scheme can improve coding efficiency, for example, in the above cases. To reduce computational complexity and memory usage, a non-separable transform scheme can be used as a secondary transform applied to the low-frequency transform coefficients obtained from the primary transform.

[0142] In some implementations, for example, selecting the transform kernel to be used from the grouped transform kernels is based on prediction mode information. The prediction mode information can indicate a prediction mode.

[0143] In some examples, prediction mode information alone provides only a rough representation of the entire space of residual patterns observed in the prediction modes. FIGS. 17A-17D show that, according to an embodiment, neighboring reconstructed samples can provide additional information to more efficiently represent the residual patterns. Thus, methods of transform set selection and / or transform kernel selection based on neighboring reconstructed samples in addition to prediction mode information are disclosed. For example, in addition to prediction mode information, feature indicators of the neighboring reconstructed samples (e.g., feature vectors

number

[0144] In this disclosure, the term block may refer to PB, CB, coded block, coding unit (CU), transform block (TB), transform unit (TU), luma block (e.g., luma CB), chroma block (e.g., chroma CB), etc.

[0145] In this disclosure, the size of a block may refer to a block width, a block height, a block aspect ratio (e.g., the ratio of block width to block height, the ratio of block height to block width), a block area size or block area (e.g., block width x block height), a minimum of the block width and block height, a maximum of the block width and block height, etc.

[0146] In this disclosure, the transform kernel of a block may be used for a linear transform, a quadratic transform, a cubic transform, or a transform scheme beyond a cubic transform. The transform kernel may be used for a separated or non-separable transform. The transform kernel may be used for a luma block, a chroma block, inter prediction, intra prediction, etc. The transform kernel may also be referred to as a transform core, a transform candidate, a transform kernel option, etc. In one example, the transform kernel is a transform matrix. The transform of a block may be performed based on at least the transform kernel. Thus, the methods of this disclosure may be applied to a linear transform, a quadratic transform, a cubic transform, any transform scheme beyond a cubic transform, a separated transform, a non-separable transform, a luma block, a chroma block, inter prediction, intra prediction, etc.

[0147] In this disclosure, a transform set may refer to a group of transform kernels or transform kernel options. A transform set may include one or more transform kernels or transform kernel options. In one example, a transform kernel for a block may be selected from a transform set.

[0148] According to an aspect of the present disclosure, a transform kernel can be determined from a group of transform sets using neighboring reconstructed samples of a block under reconstruction of a current picture. A transform kernel for the block can be determined from a group of transform sets using neighboring reconstructed samples of the block. The group of transform sets may be pre-defined. In one example, the group of transform sets is pre-stored in the encoder and / or decoder.

[0149] Neighboring reconstructed samples of a block (e.g., a set of neighboring reconstructed samples) may refer to reconstructed samples from previously decoded neighboring blocks in the current picture (e.g., a group of reconstructed samples) or reconstructed samples in a previously decoded picture. The previously decoded neighboring blocks may include one or more neighboring blocks of the block, and the neighboring reconstructed samples are also referred to as reconstructed samples in one or more neighboring blocks of the block. The one or more neighboring blocks of the block may be included in the current picture or a reconstructed picture different from the current picture (e.g., a reference picture). The neighboring reconstructed samples of the block may include spatially neighboring samples of the block of the current picture and / or temporally neighboring samples of a block of another picture different from the current picture (e.g., a previously decoded picture). Figure 18 illustrates exemplary spatially neighboring samples of a block according to an embodiment of the present disclosure. The block is a current block (1851) of a current picture. Samples 1-4 and A-X are spatially neighboring samples of the current block (1851) and have already been reconstructed. In one example, samples 1-4 and A-X are in one or more reconstructed neighboring blocks of the current block (1851). According to aspects of the present disclosure, one or more of samples 1-4 and A-X are used to determine a transform kernel for the current block (1851). In one example, the neighboring reconstructed samples of the current block (1851) include an upper left neighboring sample 4, a left neighboring column (1852) including samples A-F, and an upper neighboring row (1854) including samples G-L, and are used to determine a transform kernel for the current block (1851). In one example, the neighboring reconstructed samples of the current block (1851) include an upper left neighboring sample 1-4, a left neighboring column (1852)-(1853) including samples A-F and M-R, and an upper neighboring row (1854)-(1855) including samples G-L, and are used to determine a transform kernel for the current block (1851).

[0150] The group of transform sets may include one or more transform sets. Each of the one or more transform sets may include any suitable one or more transform kernels. Thus, the group of transform sets may include transform kernels used in a linear transform, a secondary transform, a cubic transform, or a transform scheme beyond a cubic transform. The group of transform sets may include transform kernels that are separable transforms and / or non-separable transforms. The group of transform sets may include transform kernels used for luma blocks, chroma blocks, inter prediction, intra prediction, etc.

[0151] According to an aspect of the present disclosure, a transformation kernel (or transformation candidate) for a block of a current picture may be determined from a group of transformation sets based on reconstructed samples in one or more neighboring blocks of the block of the current picture. Feature indicators (e.g., feature vectors) extracted from the reconstructed samples in one or more neighboring blocks of the block of the current picture may be used to determine a transformation kernel (or transformation candidate) for the block of the current picture.

number

[0152] In one embodiment, feature indicators of the adjacent reconstructed samples (e.g., feature vectors

number

number

[0153] In one embodiment, a transform kernel of a block (e.g., the current block (1851)) may be determined based on neighboring reconstructed samples (e.g., feature indicators) and prediction mode information of the block. The prediction mode information of the block may indicate information used to predict the block, such as inter-prediction or intra-prediction of the block. According to aspects of the present disclosure, the prediction mode information of the block may indicate a prediction mode of the block (e.g., intra-prediction mode, inter-prediction mode). In one example, the transform kernel of the block is determined based on neighboring reconstructed samples (e.g., feature indicators) and the prediction mode of the block.

[0154] In one example, the block is intra-coded or intra-predicted, and the prediction mode information is referred to as intra-prediction mode information. The prediction mode information (e.g., intra-prediction mode information) of the block may indicate the intra-prediction mode used for the block. The intra-prediction mode may refer to a prediction mode used for intra-prediction of the block, such as a directional mode (or a directional prediction mode) described in FIG. 9, a non-directional prediction mode (e.g., a DC mode, a PAETH mode, a SMOOTH mode, a SMOOTH_V mode, or a SMOOTH_H mode) described in FIG. 10, or a recursive filtering mode described in FIG. 11. The intra-prediction mode may also refer to a prediction mode described in this disclosure, an appropriate variation of a prediction mode described in this disclosure, or an appropriate combination of a prediction mode described in this disclosure. For example, the intra-prediction mode may be combined with a multi-line intra-prediction described in FIG. 12.

[0155] More specifically, the subgroup of transform sets may be selected from the group of transform sets based on, for example, coded information of the block from the coded video bitstream. The coded information of the block may include prediction mode information indicating a prediction mode of the block (e.g., an intra-prediction mode or an inter-prediction mode). In one embodiment, the subgroup of transform sets may be selected from the group of transform sets based on the prediction mode information (e.g., a prediction mode).

[0156] According to aspects of the present disclosure, a subgroup of transform sets may be selected from the group of transform sets based on a prediction mode of the block indicated in the coded information of the block. Furthermore, the transform candidates from the subgroup of transform sets may be determined based on reconstructed samples in one or more neighboring blocks of the block. For example, the transform candidates from the subgroup of transform sets may be determined based on feature indicators (e.g., feature vectors) extracted from the reconstructed samples in one or more neighboring blocks of the block.

number

[0157] After selecting a subgroup of transform sets from the group of transform sets, a transform kernel for the block may be further determined based on at least adjacent reconstructed samples (e.g., feature indicators) using any suitable method described below.

[0158] In one embodiment, neighboring reconstructed samples of the block (or reconstructed samples in one or more neighboring blocks) are used to identify or select a transform set of the selected sub-group of transform sets from the sub-group of selected transform sets. In one example, feature indicators (e.g., feature vectors) extracted from the reconstructed samples in one or more neighboring blocks are

number

[0159] In one embodiment, a transform set of the subgroup of selected transform sets is selected from the subgroup of selected transform sets based on a second index. The coded information may indicate (e.g., include) the second index. The second index may be signaled in the coded video bitstream. Furthermore, a transform kernel (or transform candidate) of a block from the selected transform set may be determined (e.g., selected) based on neighboring reconstructed samples of the block (or reconstructed samples in one or more neighboring blocks of the block) (e.g., feature indicators). In one example, the transform kernel (or transform candidate) in the selected transform set may be determined (e.g., selected) based on feature indicators of the block (e.g., feature vector

number

[0160] In one embodiment, a transformation kernel (or transformation candidate) for a block is implicitly determined (or identified) from a subgroup of the selected transformation set based on neighboring reconstructed samples of the block. In one example, the transformation kernel (or transformation candidate) is determined based on feature indicators (e.g., feature vectors) of the block from the subgroup of the selected transformation set.

number

[0161] Block feature indicators (e.g., feature vectors)

number

number

number

number

number

[0162] In one embodiment, a single variable X may be used to indicate the neighboring reconstructed samples, and the feature indicator of the block is the feature scalar S of the block. The variable X may indicate the sample values ​​of the neighboring reconstructed samples. In one example, the variable X is an array containing the sample values ​​of the neighboring reconstructed samples, reflecting the distribution of the sample values ​​in the neighboring reconstructed samples. The feature scalar S may indicate statistical information of the sample values. The feature scalar S may include, but is not limited to, scalar quantitative measurements of the variable X obtained from the neighboring reconstructed samples, such as the mean (or first moment) of the variable X, the variance (or second moment) of the variable X, and the skewness (or third moment) of the variable X. In one example, the variable X is referred to as a random variable X.

[0163] In one example, the feature indicator is a feature scalar S. The feature scalar S is determined as a moment (e.g., first moment, second moment, third moment, etc.) of a variable X that is indicative of sample values ​​of reconstructed samples in one or more neighboring blocks of the block.

[0164] An example of variable X is shown in Figure 18. In one example, variable X includes reconstructed neighboring samples adjacent to the current block (1851), including sample 4, the left neighboring column (1852), and the above neighboring row (1854). In one example, variable X includes reconstructed neighboring samples 1-4 and A-X.

[0165] In one embodiment, multiple variables (e.g., two variables Y and Z) can be used to separately indicate different sets of adjacent reconstruction samples (if available) (e.g., a first set corresponding to Y and a second set corresponding to Z), and the feature indicator of a block is the feature vector of this block.

number

number

[0166] Feature Vector

number

[0167] 18 shows an example of multiple variables Y and Z. In one example, variable Y includes the left adjacent column (1852) and variable Z includes the above adjacent row (1854). In one example, one of variables Y and Z includes sample 4 at the top left.

[0168] In one example, the feature indicators are

number

number

[0169] In one example, variable Y includes the left adjacent columns 1852 through 1853, and variable Z includes the top adjacent rows 1854 through 1855. In one example, one of variables Y and Z includes the top left samples 1 through 4.

[0170] According to aspects of the present disclosure, when the feature indicator is used in determining the transform kernel of the block or the transform set including the transform kernel of the block, the transform kernel or the transform set may be determined based on the feature indicator and a threshold value. The threshold value may be selected from a predefined threshold set. The transform kernel or the transform set may be determined based on the feature indicator and the threshold value using any suitable method. Some example methods are described below. As described above, according to each aspect of the present disclosure, a subgroup of transform sets is selected from the group of transform sets, for example, based on the prediction mode information. Then, in one example, the transform set is identified from the selected subgroup of transform sets using the feature indicator of the block and a threshold value. In another example, the transform set of the selected subgroup of transform sets is selected using an index (e.g., a second index) signaled in the coded video bitstream, and the transform kernel of the block from the selected transform set may be determined using the feature indicator of the block and a threshold value. In another example, the transform kernel is implicitly identified using the feature indicator of the selected subgroup of transform sets and a threshold value.

[0171] The threshold set (K s , which are predefined, for example for classification purposes. In one example, (i) the moments of a variable X and a predefined set of thresholds K s the coding information index may include: (i) determining a transformation set from the subgroup of transformation sets based on a threshold selected from; (ii) determining candidate transformations from the subgroup of transformation sets based on moments of variable X and a threshold; or (iii) selecting a transformation set from the subgroup of transformation sets based on an index of the coded information and determining candidate transformations from the selected transformation set based on moments of variable X and a threshold.

[0172] In one embodiment, the feature scalar S is used to identify a transformation kernel for the block, or to identify a transformation set that includes the transformation kernel for the block, for example from a subgroup of transformation sets selected as described above. s may include one or more first thresholds (or first threshold values).

[0173] In one example, the feature scalar S is a quantitative measurement of a variable X, such as, for example, the mean, variance, skewness, etc. of the variable X. For example, the feature scalar S is a moment of the variable X. The moment of a variable can be either the first moment (or mean) of the variable X, the second moment (or variance) of the variable X, or the third moment (such as skewness) of the variable X.

[0174] Each prediction mode (e.g., intra prediction mode or inter prediction mode) may, for example, include multiple prediction modes and multiple threshold sets K s A threshold set K that shows an injective mapping between s A unique threshold subset K of s The prediction mode of the block may be one of a plurality of prediction modes. Each prediction mode and a corresponding threshold subset K s ' is an injective mapping. For example, for a threshold set K s a threshold subset K corresponding to the first prediction mode s1 ' and the threshold subset K corresponding to the second prediction mode. s2 ', and also the threshold subset K s1 ' contains the threshold subset K s2 ' has no elements or thresholds that are identical to the elements or thresholds in

[0175] In one example, the feature scalar S is a quantitative measure of a variable X, such as its mean, variance, or skewness. Each prediction mode is determined by a threshold set K s Any threshold subset in K s '. For each prediction mode and threshold set Ks The corresponding threshold subset K in s ' is a non-injective mapping. In one example, the multiple prediction modes (e.g., the first prediction mode and the second prediction mode) are mapped to a threshold set K s The same threshold subset K in s In one example, the threshold set K s is the threshold subset K corresponding to the first prediction mode. s3 ' and the threshold subset K corresponding to the second prediction mode s4 ' and the threshold subset K s3 The elements (or thresholds) in ' are the threshold subset K s4 ' is the same as an element (or threshold) in the threshold set K s is a single element set.

[0176] In an embodiment, the threshold set K s Threshold subset K in s The elements of ' are thresholds that depend on the quantization index (or the corresponding quantization step size associated with the quantization index). Typically, each quantization index can correspond to a unique quantization step size. The above discussion applies to each prediction mode and to the threshold set K s The corresponding threshold subset in K s ' is non-injective or injective.

[0177] In one embodiment, a threshold set K corresponding to a particular quantization index (or corresponding quantization step size) is s Threshold subset K in s ' is defined. s The remaining elements in ', if any, may be derived using a mapping function. The mapping function may be linear or non-linear. The mapping function may be one that is predefined and used by the encoder and / or decoder.s The remaining elements in ' can correspond to different quantization indexes or (or corresponding quantization step sizes). The above description is based on the assumption that for each prediction mode and threshold set K s The corresponding threshold subset K in s ' is non-injective or injective.

[0178] In one embodiment, a lookup table is used to determine the threshold subset K s For example, the lookup table may determine the threshold subset K based on the prediction mode and / or the quantization index. s In one example, the lookup table includes a relationship between a prediction mode, a quantization index (or a corresponding quantization step size), and a threshold value. The lookup table is traversed to select an element for the threshold subset K using the prediction mode (e.g., intra prediction mode or inter prediction mode) and / or the quantization index (or the corresponding quantization step size). s ' (e.g., threshold subset K s The above description is based on the assumption that for each prediction mode and threshold set K s The corresponding threshold subset in K s ' is non-injective or injective.

[0179] In an embodiment, the mapping from the quantization index (or the corresponding quantization step size) to the threshold value is a linear mapping. The parameters used for the linear mapping, such as the slope and intercept, may be predefined or derived using coded information of the block. The coded information may include, but is not limited to, the block size, the quantization index (or the corresponding quantization step size), and the prediction mode (e.g., intra prediction mode or inter prediction mode). The linear mapping (e.g., the parameters used for the linear mapping) may be predefined or derived based on one or a combination of the block size, the quantization index (or the corresponding quantization step size), and the prediction mode. If the mapping from the quantization index (or the corresponding quantization step size) to the threshold value is a non-linear mapping, this description can be adjusted appropriately. The above description is based on the following: s The corresponding threshold subset K in s ' is non-injective or injective.

[0180] As described above, the feature scalar S can be used to identify a transformation set that includes the transformation kernel for the block, or to identify the transformation kernel for the block, and a threshold set K s may include one or more first thresholds. In an embodiment, for example, the threshold set K s The selection of the threshold from depends on the block size of the block. s The selection of the threshold from depends on the quantization parameter used in the quantization, such as the quantization index (or the corresponding quantization step size). s The selection of the threshold from depends on the prediction mode (eg, intra prediction mode and / or inter prediction mode).

[0181] In one embodiment, the feature vector

number

number

number

[0182] In one example,

number

[0183] In one example, each prediction mode (e.g., an intra prediction mode or an inter prediction mode) may be

number

[0184] In an embodiment, the threshold set K v Threshold subset K in v The elements of ' are thresholds that depend on the quantization index (or the corresponding quantization step size associated with the quantization index).

number

[0185] In an embodiment, corresponding to a particular quantization index (or corresponding quantization step size),

number

[0186] In one embodiment, the lookup table is a threshold subset K v ' Select elements for threshold subset K v ' and determine the classification vector subset

number

number

number

number

number

[0187] In one embodiment,

number

[0188] In one example, (i) the distance and a threshold set K v A threshold subset K contained in v ' the selected threshold (e.g., {K v (ii) determining a set of transformations from a subgroup of transformation sets based on a comparison with a distance and a threshold (e.g., {K v (iii) determining candidate transformations from the subgroup of transformation sets based on a comparison with the coded information index and a distance and a threshold (e.g., {K v '}) based on a comparison with the selected set of transformations.

number

[0189] In one example, this distance is calculated by dividing the two vectors

number

[0190] In one embodiment, the comparison is: (i) whether the distance is less than or equal to a threshold;

number

[0046] In the above, the {} represents an element of the corresponding set, but is not limited to this.

[0191] In some embodiments, neighboring reconstructed samples, such as feature indicators of neighboring reconstructed samples, can be used to constrain the selection of transform candidates from a subgroup of transform sets. For some subgroups of transform sets selected using coded information such as the prediction mode of a block (e.g., intra-prediction mode, inter-prediction mode), the identification process of transform candidates or transform kernels can include feature indicators of neighboring reconstructed samples (e.g., feature vectors

number

number

number

[0192] FIG. 19 shows a flow chart outlining a process (1900) according to one embodiment of the disclosure. The process (1900) can be used to reconstruct blocks such as CB, CU, PB, TB, TU, luma blocks (e.g., luma CB or luma TB), chroma blocks (e.g., chroma CB or chroma TB), etc. In various embodiments, the process (1900) is performed by processing circuits in the terminal devices (310), (320), (330), and (340), processing circuits performing the functions of a video encoder (403), processing circuits performing the functions of a video decoder (410), processing circuits performing the functions of a video decoder (510), processing circuits performing the functions of a video encoder (603), etc. In some embodiments, the process (1900) is implemented in software instructions, such that the processing circuits perform the process (1900) when the processing circuits execute the software instructions. The process starts at (S1901) and proceeds to (S1910).

[0193] In (S1910), feature indicators (e.g., feature vectors) extracted from reconstructed samples in one or more neighboring blocks of the block of the current picture, e.g., reconstructed samples in one or more neighboring blocks of the block of the current picture, are

number

[0194] In an embodiment, a subgroup of transform sets may be selected from the group of transform sets based on a prediction mode (e.g., intra prediction mode, inter prediction mode) of the block signaled in the coded information of the block. Feature indicators (e.g., feature vectors) extracted from reconstructed samples in one or more neighboring blocks of the block may be used.

number

[0195] In one example, feature indicators (e.g., feature vectors) extracted from the reconstruction samples in one or more neighboring blocks of the block are

number

[0196] In one example, a transform set from the subgroup of transform sets may be selected based on a second index signaled in the coded information.

number

[0197] In one example, feature indicators (e.g., feature vectors) extracted from the reconstruction samples in one or more neighboring blocks of the block are

number

[0198] In one example, based on a statistical analysis of the reconstruction samples in one or more neighboring blocks of the block,

number

[0199] In one example, the feature indicators are feature scalars S and are determined as moments of variables indicative of sample values ​​of the reconstructed samples in one or more neighboring blocks of the block. S is predefined. Therefore, (i) the moment of the variables and the threshold set K S the second index of the coded information and determining candidate transforms from the subgroup of transformation sets based on the moments of the variables and the threshold value; or (ii) determining candidate transforms from the subgroup of transformation sets based on the moments of the variables and the threshold value; or (iii) selecting a transformation set from the subgroup of transformation sets based on a second index of the coded information and determining candidate transforms from the selected transformation set based on the moments of the variables and the threshold value.

[0200] In one example, the moment of the variable is one of a first moment of the variable, a second moment of the variable, or a third moment of the variable. The prediction mode of the block is one of a plurality of prediction modes, each of which is determined by a threshold set K. S A unique threshold subset K in S' corresponds to this unique threshold subset K S ' is a set of multiple prediction modes and a threshold set K S 4 shows an injective mapping between multiple threshold subsets in

[0201] A threshold set K is determined based on one of: (i) the block size of the block; (ii) the quantization parameter; or (iii) the prediction mode of the block. S Select a threshold value from .

[0202] In one example, the feature indicators are

number

[0203] In (S1920), the samples of the block can be reconstructed based on the determined candidate transforms.

[0204] The process (1900) may be adjusted as appropriate. One or more steps in the process (1900) may be modified and / or omitted. One or more additional steps may be added. Any suitable order of performance may be used.

[0205] The embodiments of the present disclosure may be used alone or in combination in any order. Furthermore, each of the methods (or embodiments), the encoder, and the decoder may be realized by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium. The embodiments of the present disclosure may be applied to a luma block or a chroma block.

[0206] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 20 illustrates a computer system (2000) suitable for implementing certain embodiments of the disclosed subject matter.

[0207] The computer software can be coded using any suitable machine code or computer language that can be assembled, compiled, linked, or such mechanisms to create code containing instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or by interpretation, microcode execution, etc.

[0208] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, and the like.

[0209] 20 for computer system (2000) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement regarding any one or combination of components shown in the exemplary embodiment of computer system (1700).

[0210] The computer system (2000) may include certain human interface input devices. Such human interface input devices may respond to input by one or more users, for example, by tactile input (keystrokes, swipes, data glove movements, etc.), audio input (voice, clapping, etc.), visual input (gestures, etc.), and olfactory input (not shown). Human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (voice, music, ambient sounds, etc.), images (scanned images, photographic images obtained from still image cameras, etc.), and video (two-dimensional video, three-dimensional video including stereoscopic video, etc.).

[0211] The input human interface devices may include one or more of a keyboard (2001), a mouse (2002), a trackpad (2003), a touch screen (2010), a data glove (not shown), a joystick (2005), a microphone (2006), a scanner (2007), a camera (2008) (only one of each type is shown).

[0212] The computer system (2000) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the user's senses, for example, through haptic output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., a touch screen (2010), data gloves (not shown), or a haptic feedback device that has haptic feedback via a joystick (2005) but does not function as an input device), audio output devices (speakers (2009), headphones (not shown), etc.), visual output devices (screens (2010) including CRT screens, LCD screens, plasma screens, OLED screens (each with or without touch screen input capability, each with or without haptic feedback capability, some of which may output two-dimensional visual output or three or more dimensional output, such as through means such as stereographic output), virtual reality glasses (not shown), holographic displays and smoke tanks (not shown), etc.), and printers (not shown).

[0213] The computer system (2000) may also include human accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (2020) with media (2021) such as CDs / DVDs, thumb drives (2022), removable hard drives or solid state drives (2023), traditional magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD based devices (not shown) such as security dongles, etc.

[0214] It should be understood by those skilled in the art that the term "computer-readable medium" as used herein with respect to the disclosed subject matter does not include transmission media, carrier waves, or other transitory signals.

[0215] The computer system (2000) may further include an interface (2054) to one or more communication networks (2055). The networks may be, for example, wireless, wired, optical. The networks may further be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, WLAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular and industrial including CANBus, etc. Certain networks generally require an external network interface adapter connected to a particular general purpose data port or peripheral bus (2049) (e.g., a USB port of the computer system (2000)). Others are generally integrated into the core of the computer system (2000) by connecting to a system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (2000) can communicate with other entities. Such communication may be one-way receive only (e.g., broadcast TV), one-way transmit only (e.g., CANbus to a specific CANbus device), or two-way, such as transmit to other computer systems using local or wide area digital networks. Specific protocols and protocol stacks may be used with each of these networks and network interfaces described above.

[0216] The aforementioned human interface devices, human accessible storage devices, and network interfaces may be connected to the core (2040) of the computer system (2000).

[0217] The cores (2040) may include one or more central processing units (CPUs) (2041), graphics processing units (GPUs) (2042), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (2043), task-specific hardware accelerators (2044), graphics adapters (2050), and the like. These devices may be connected via a system bus (2048), along with read-only memory (ROM) (2045), random access memory (2046), and internal mass storage devices (2047), such as non-user-accessible internal hard drives, SSDs, and the like. In some computer systems, the system bus (2048) may be accessible in the form of one or more physical plugs, allowing expansion with additional CPUs, GPUs, and the like. Peripherals may be connected directly to the core's system bus (2048) or via a peripheral bus (2049). In one example, a display (2010) may be connected to the graphics adapter (2050). Peripheral bus architectures include PCI, USB, and the like.

[0218] The CPU (2041), GPU (2042), FPGA (2043), and accelerator (2044) may combine to execute certain instructions that may constitute the aforementioned computer code. That computer code may be stored in ROM (2045) or RAM (2046). Persistent data may be stored, for example, in internal mass storage (2047), while transitory data may also be stored in RAM (1746). The use of cache memory, which may be closely associated with one or more CPUs (2041), GPUs (2042), mass storage (2047), ROM (2045), RAM (2046), etc., allows for fast storage and retrieval in any memory device.

[0219] The computer-readable medium can comprise computer code for performing various computer-implemented operations. The media and computer code can be those specially designed and constructed for the purposes of the present disclosure, or they can be of the kind well known and available to those skilled in the computer software arts.

[0220] By way of example and not limitation, a computer system (2000) having the architecture, and in particular the core (2040), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be mass storage accessible to a user as described above, as well as media associated with a particular storage of the core (2040) that is of a non-transitory nature, such as the core internal mass storage (2047) or ROM (2045). Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (2040). The computer-readable media can include one or more memory devices or chips, as appropriate. The software can cause the core (2040), and in particular the processor therein (including a CPU, GPU, FPGA, etc.) to execute a particular process or a particular portion of a particular process as described herein, including defining data structures stored in RAM (2046) and modifying such data structures according to a software-defined process. Additionally or alternatively, the computer system may provide functionality as a result of logic hardwired or embedded in circuitry (e.g., accelerator (2044)) that may operate in place of or together with software to perform certain processes or certain portions of certain processes described herein. References to software may include logic, where appropriate, and vice versa. References to computer-readable media may include circuitry (such as integrated circuits (ICs)) that store software for execution, circuitry that embodies logic for execution, or both, where appropriate. The present disclosure includes any suitable combination of hardware and software.

[0221] Appendix A: Acronyms JEM: Joint Exploration Model VVC: Versatile Video Coding BMS:Benchmark Set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplementary Enhancement Information VUI: Video Usability Information GOP: Groups of Pictures TU: Transform Units PU: Prediction Units CTU: Coding Tree Units CTB: Coding Tree Blocks PB: Prediction Blocks HRD: Hypothetical Reference Decoder SNR: Signal Noise Ratio CPU: Central Processing Units GPU: Graphics Processing Units CRT:Cathode Ray Tube LCD: Liquid-Crystal Display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile Communications LTE:Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Array SSD: Solid-State Drive IC: Integrated Circuit CU: Coding Unit

[0222] While this disclosure has described several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of this disclosure. It should thus be understood that those skilled in the art may devise various systems and methods that, although not expressly or specifically described herein, embody the principles of the present disclosure and are included within its spirit and scope.

Claims

1. A method for video encoding by an encoder, comprising: Sampling blocks of the current picture into a video bitstream based on an intra-prediction mode or an inter-prediction mode, A feature vector is calculated from samples in one or more neighboring blocks of the block of the current picture. [0010] or extracting a feature scalar S; Based on a statistical analysis of the samples in one or more neighboring blocks of the block, the feature vector [0025] or determining one of said feature scalars S; selecting a sub-group of transform sets from a group of transform sets based on information indicating a prediction mode of the block, where each transform set of the group of transform sets includes one or more transform candidates for the block, and one or more neighboring blocks are in the current picture or a picture different from the current picture; (i) the feature vector [0030] or determining a transformation set from a subgroup of said transformation sets based on one of said feature scalars S; (ii) the feature vector [0045] or determining the candidate transformations from a subgroup of the set of transformations based on one of the feature scalars S; or (iii) selecting the transformation set from the subgroup of transformation sets based on an index of information indicating a prediction mode of the block; and [0050] or determining the candidate transformations from a selected set of transformations based on one of the feature scalars S; Sampling the block based on the determined candidate transforms; and coding information indicating a prediction mode of the block into the video bitstream. method.

2. The feature vector [006] or one of the feature scalars S is the feature scalar S; The feature vector [0070] or determining one of the feature scalars S further comprises determining the feature scalar S as a moment of a variable indicative of sample values ​​of the samples in one or more neighboring blocks of the block. The method of claim 1.

3. A threshold set K S is predefined, The implementation includes: (i) determining the set of transformations from a subgroup of the set of transformations based on the moments of the variables and a threshold value from the threshold set K S ; (ii) determining the candidate transformations from a subgroup of the set of transformations based on the moments of the variables and the threshold set; or (iii) selecting the transform set from the subgroup of transform sets based on an index of information indicating a prediction mode of the block; and determining the transform candidate from the selected transform set based on moments of the variables and the threshold value. The method of claim 2.

4. The moment of the variable is one of a first moment of the variable, a second moment of the variable, or a third moment of the variable; a prediction mode for the block is one of a plurality of prediction modes; Each of the plurality of prediction modes corresponds to a unique threshold subset K S ′ in the threshold set K S , the unique threshold subset K S ′ indicating a single-vehicle mapping between the plurality of prediction modes and a plurality of threshold subsets in the threshold set K S . The method according to claim 3.

5. The method of claim 4, further comprising the step of selecting the threshold from the threshold set K S based on one of: (i) a block size of the block; (ii) a quantization parameter; or (iii) a prediction mode of the block. The method according to claim 3.

6. The feature vector [0080] Or one of the feature scalars S is the feature vector [0090] and The feature vector [0010] Or one of the feature scalars S is the feature vector ##EQU00011## including in the block a covariance or second moment of a variable indicative of the sample values ​​of samples in adjacent columns to the left of the block and indicative of the sample values ​​of samples in adjacent rows above the block, The method of claim 1.

7. Classification vector set ##EQU00012## , and the classification vector set ##EQU00013## A threshold set K v associated with the co-variation of said variables and said classification vector set ##EQU14## A subset of classification vectors in ##EQU00015## Calculating the distance between the classification vector selected from The implementation includes: (i) determining the set of transformations from a subgroup of the sets of transformations based on a comparison of the distance to a threshold selected from a threshold subset K v ′ contained in the set of thresholds K v ; (ii) determining the candidate transformations from a subgroup of the set of transformations based on a comparison of the distance to the threshold; or (iii) selecting the transform set from the subgroup of transform sets based on an index of information indicating a prediction mode of the block; and determining the transform candidate from the selected transform set based on a comparison between the distance and the threshold. The method according to claim 6.

8. A video encoding device comprising: Sampling a block of a current picture into a video bitstream based on an intra-prediction mode or an inter-prediction mode, A feature vector is calculated from samples in one or more neighboring blocks of the block of the current picture. ##EQU00016## or extracting a feature scalar S; Based on a statistical analysis of the samples in one or more neighboring blocks of the block, the feature vector ##EQU00017## or determining one of said feature scalars S; selecting a sub-group of transform sets from a group of transform sets based on information indicating a prediction mode of the block, where each transform set of the group of transform sets includes one or more transform candidates for the block, and one or more neighboring blocks are in the current picture or a picture different from the current picture; (i) the feature vector [0018] or determining a transformation set from a subgroup of said transformation sets based on one of said feature scalars S; (ii) the feature vector [0019] or determining the candidate transformations from a subgroup of the set of transformations based on one of the feature scalars S; or (iii) selecting the transformation set from the subgroup of transformation sets based on an index of information indicating a prediction mode of the block; and [0020] or determining the candidate transformations from a selected set of transformations based on one of the feature scalars S; Sampling the block based on the determined candidate transforms; and coding information indicating a prediction mode of the block into the video bitstream. a processing circuit configured to Device.

9. The feature vector ##EQU00021## or one of the feature scalars S is the feature scalar S; the processing circuitry is configured to determine the feature scalar S as a moment of a variable indicative of sample values ​​of the reconstructed samples in one or more neighboring blocks of the block; 9. The apparatus of claim 8. A threshold set K S is predefined; The processing circuitry includes: (i) determining the set of transformations from a subgroup of the set of transformations based on the moments of the variables and a threshold value from the threshold set K S ; (ii) determining the candidate transformations from a subgroup of the set of transformations based on the moments of the variables and the threshold; or (iii) selecting the transform set from the subgroup of transform sets based on an index of information indicating a prediction mode of the block; and determining the transform candidate from the selected transform set based on moments of the variables and the threshold value.

10. The apparatus of claim 9.

11. The feature vector [0022] Or one of the feature scalars S is the feature vector [0023] and The processing circuitry includes: The feature vector ##EQU00024## includes covariances or second moments of variables indicative of sample values ​​of samples in adjacent columns to the left of said block and sample values ​​of samples in adjacent rows to the top of said block, respectively.

9. The apparatus of claim 8.

12. Classification vector set [0025] , and the classification vector set [0026] A threshold set K v associated with The processing circuitry includes: the co-variation of said variables and said classification vector set [0027] A subset of classification vectors in [0028] Calculate the distance between the classification vector selected from (i) determining the set of transformations from a subgroup of the sets of transformations based on a comparison of the distance to a threshold selected from a threshold subset K v ′ contained in the set of thresholds K v ; (ii) determining the candidate transformations from a subgroup of the set of transformations based on a comparison of the distance to the threshold; or (iii) selecting the transform set from the subgroup of transform sets based on an index of information indicating a prediction mode of the block; and determining the transform candidate from the selected transform set based on a comparison between the distance and the threshold.

12. The apparatus of claim 11.

13. The moment of the variable is one of a first moment of the variable, a second moment of the variable, or a third moment of the variable; a prediction mode for the block is one of a plurality of prediction modes; each of the plurality of prediction modes corresponds to a unique threshold subset K S ′ in the threshold set K S , the unique threshold subset K S ′ indicating an injective mapping between the plurality of prediction modes and a plurality of threshold subsets in the threshold set K S ; 13. The apparatus of claim 12.