Method and apparatus for video coating

By employing a processing circuit to generate feature vectors and transform sets based on intra-prediction and motion compensation, the solution addresses inefficiencies in existing video coding technologies, achieving enhanced compression efficiency and adaptability to diverse video content.

JP7789829B2Active Publication Date: 2025-12-22TENCENT AMERICA LLC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024063336
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-09-30
Filing Date
2024-04-10
Publication Date
2025-12-22
Estimated Expiration
2041-10-05

AI Technical Summary

Technical Problem

Existing video coding technologies face inefficiencies in reducing redundancy and achieving optimal compression ratios due to the varying likelihood of intra-prediction direction and motion compensation techniques, particularly in the use of intra-prediction and motion compensation, without sufficient use of intra-prediction, and motion compensation, without sufficient use of innovative methods to adapt to diverse video content.

Method used

The proposed solution involves implementing a processing circuit that generates feature vectors from reconstructed samples in neighboring blocks of a block of a current picture or a reconstructed picture, using a processing circuit that determines specific measures to generate feature vectors from a group of transform sets based on feature scalars and transforms, including the use of specific combinations of these techniques, such as intra-prediction and motion compensation, to enhance compression efficiency.

Benefits of technology

This approach allows for improved video coding efficiency by adapting to the statistical likelihood of intra-prediction directions, reducing redundancy, and enhancing compression ratios through the use of feature vectors and transform sets, thereby optimizing video encoding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007789829000100
    Figure 0007789829000100
  • Figure 0007789829000101
    Figure 0007789829000101
  • Figure 0007789829000102
    Figure 0007789829000102
Patent Text Reader

Abstract

To provide an apparatus including a processing circuit for video decoding / decoding which reduces redundancy in an input video signal through compression, and a method.SOLUTION: A method implemented by a processing circuit of a video encoder determines a transform candidate for a block in a current picture from a group of transform sets based on one of a feature vector or a feature scalar extracted from reconstructed samples in one or more neighboring blocks of the block. Each transform set includes one or more transform candidates for the block. The one or more neighboring blocks are in the current picture or a reconstructed picture different from the current picture. The method reconstructs samples of the block based on the determined transform candidate, selects a sub-group of transform sets from the group of transform sets based on a prediction mode for the block indicated in coded information for the block and determines the transform candidate from the sub-group of transform sets based on the reconstructed samples.SELECTED DRAWING: Figure 19
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority to U.S. Provisional Application No. 63 / 130,249, filed December 23, 2020, entitled "Feature based transform selection," which in turn claims priority to U.S. Provisional Application No. 17 / 490,967, filed September 30, 2021, entitled "METHOD AND APPARATUS FOR VIDEO CODING," the disclosure of which is incorporated herein by reference in its entirety.

[0002] [Technical field] The present disclosure generally relates to video coating embodiments. [Background technology]

[0003] The background discussion provided herein is intended to generally present the context for the present disclosure. The inventors' work, to the extent that it is described in this background section, and aspects of the discussion that may not be admitted as prior art at the time of filing, are not admitted expressly or implicitly as prior art to the present disclosure.

[0004] Video coding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video can include a series of pictures, each having spatial dimensions of, for example, 1920x1080 luma samples and associated chroma samples. The series of pictures can have a fixed or variable picture rate (also informally called a "frame rate"), for example, 60 pictures per second or 60 Hz. Uncompressed video has specific bitrate requirements. For example, 1080p60 4:2:0 video (1920x1080 luma sample resolution at a 60 Hz frame rate) at 8 bits per sample requires a bandwidth approaching 1.5 Gbit / s. One hour of such video requires more than 600 GBytes of storage space.

[0005] One goal of video coding and decoding is to reduce redundancy in the input video signal through compression. Compression can help reduce the aforementioned bandwidth or storage space requirements, sometimes by more than two orders of magnitude. Both lossless and lossy compression, as well as combinations of them, can be used. Lossless compression refers to techniques that allow an exact replica of the original signal to be reconstructed from a compressed version of the original signal. When lossy compression is used, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signal is small enough to make the reconstructed signal useful for its intended application. For video, lossy compression is widely adopted. The amount of acceptable distortion depends on the application; for example, users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio can reflect that higher tolerable / acceptable distortion can result in a higher compression ratio.

[0006] Video encoders and decoders may utilize techniques from several broad categories, including, for example, motion compensation, transform, quantization, and entropy coding.

[0007] Video codec technology may include a technique known as intra-coding. In intra-coding, sample values ​​are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, pictures are spatially subdivided into blocks of samples. When all blocks of samples are coded in intra mode, the picture may become an intra-picture. Intra-pictures and their derivatives, such as independent decoder refresh pictures, can be used to reset the decoder state and therefore can be used as the first picture in a coded video bitstream and video session or as still pictures. Samples in intra-blocks may be subjected to a transform, and the transform coefficients may be quantized before entropy coding. Intra-prediction may be a technique that minimizes sample values ​​in the pre-transform domain. In some cases, smaller DC values ​​and smaller AC coefficients after the transform require fewer bits for a given quantization step size to represent the block after entropy coding.

[0008] Conventional intra-coding, such as that known from MPEG-2 generation coding techniques, does not use intra-prediction. However, some newer video compression techniques include techniques that attempt to predict intra-prediction from surrounding sample data and / or metadata obtained during encoding / decoding, for example, from spatially adjacent and earlier-in-decoding order samples. Such techniques are hereinafter referred to as "intra-prediction" techniques. Note that, at least in some cases, intra-prediction uses only reference data from the current picture being reconstructed, and does not use reference data from reference pictures.

[0009] Intra-prediction can exist in various forms. If multiple such techniques are available for a given video coding technique, the technique in use may be coded as an intra-prediction mode. In some cases, modes may have sub-modes and / or parameters, which may be coded separately or included in a mode codeword. Which codeword is used for a given mode / sub-mode / parameter combination may affect the coding efficiency achieved by intra-prediction, as may the entropy coding technique used to convert the codeword into a bitstream.

[0010] Specific modes of intra prediction were introduced in H.264, refined in H.265, and further refined in newer coding techniques such as the Joint Search Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). Predictor blocks can be formed using neighboring sample values ​​belonging to already available samples. The sample values ​​of the neighboring samples are replicated in the predictor block according to the direction. A reference to the direction in use can be coded in the bitstream or can itself be predicted.

[0011] Referring to FIG. 1A, shown at the bottom right is a subset of nine predictor directions from the 33 possible predictor directions in H.265 (corresponding to the 33 angle modes of the 35 intra modes). The point where the arrows converge (101) represents the sample being predicted. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples to the upper right at an angle of 45 degrees from horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples to the lower left at an angle of 22.5 degrees from horizontal of sample (101).

[0012] Continuing with FIG. 1A , a square block (104) of 4×4 samples (indicated by a thick dashed line) is shown in the upper left. The square block (104) contains 16 samples, each labeled with “S,” its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (from the top) and the first sample in the X dimension (from the left). Similarly, sample S44 is the fourth sample in the block (104) in both the Y and X dimensions. Because the block is 4×4 samples in size, S44 is located in the lower right. Also shown are reference samples that follow a similar numbering scheme. The reference samples are labeled R, their Y position (e.g., row index) and X position (column index) relative to the block (104). In both H.264 and H.265, predicted samples are adjacent to the block being reconstructed. Therefore, negative values ​​do not need to be used.

[0013] Intra-picture prediction can be performed by copying reference sample values ​​from neighboring samples assigned by the signaled prediction direction. For example, suppose the coded video bitstream includes signaling indicating a prediction direction for this block that matches the arrow (102) (i.e., predicted from one or more prediction samples to the upper right at a 45-degree angle from horizontal). In this case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.

[0014] In some cases, the values ​​of multiple reference samples can be combined, for example by interpolation, to calculate a reference sample, especially when the directions are not evenly divided by 45 degrees.

[0015] As video coding technology evolves, the number of predictable directions is also increasing. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013), and JEM / VVC / BMS can support up to 65 directions at launch. Experiments have been conducted to identify the most likely directions, and specific techniques in entropy coding are used to represent more likely directions with fewer bits, accepting a specific penalty for less likely directions. Furthermore, the direction itself may be predicted from neighboring directions used in neighboring, already decoded blocks.

[0016] FIG. 1B shows a schematic diagram (180) illustrating 65 intra-prediction directions with JEM to illustrate the increasing number of prediction directions over time.

[0017] The mapping of intra-prediction direction bits in a coded video bitstream representing a direction can vary from one video coding technique to another and can range, for example, from a simple direct mapping of prediction directions to intra-prediction modes to complex adaptation schemes involving codewords, most likely modes, and similar techniques. However, in all cases, there may be certain directions that are statistically less likely to occur in the video content than certain other directions. Because the goal of video compression is to reduce redundancy, a well-performing video coding technique will represent these less likely directions with more bits than more likely directions.

[0018] Motion compensation can be a lossy compression technique and can be related to a technique in which blocks of sample data from a previously reconstructed picture or part thereof (reference picture) are spatially shifted in a direction indicated by a motion vector (hereinafter also referred to as MV) and then used for the prediction of a new reconstructed picture or part thereof. In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV can have two dimensions, X and Y, or three dimensions, where the third dimension is an indication of the reference picture in use (the latter may indirectly be the temporal dimension).

[0019] In some video compression techniques, the MV applicable to an area of ​​sample data can be predicted from other MVs, e.g., from an MV associated with another region of sample data that is spatially adjacent to the region being reconstructed and precedes that MV in decoding order. In this way, the amount of data required for coding the MV can be significantly reduced, thereby eliminating redundancy and increasing the amount of compression. MV prediction can work efficiently because, for example, when coding an input video signal derived from a camera (called natural video), regions larger than the region to which a single MV is applicable have a statistical likelihood of moving along similar directions and therefore, in some cases, can be predicted using similar motion vectors derived from MVs of neighboring regions. As a result, the MV found for a given region will be similar or identical to the MV predicted from surrounding MVs and, after entropy coding, can be represented with fewer bits than would be used to code the MV directly. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., an MV) derived from the original signal (i.e., a sample stream). In other cases, the MV prediction itself can be lossy, for example, due to rounding errors in computing a predictor from several surrounding MVs.

[0020] H.265 / HEVC (ITU-T Rec. H.265, High Efficiency Video Coding, December 2016) describes various MV prediction mechanisms. Among the many MV prediction mechanisms provided by H.265, this section describes a technique called "spatial merging" hereafter.

[0021] Referring to Figure 2, the current block (201) contains samples that the encoder discovered during the motion search process to be predictable from a spatially shifted previous block of the same size. Instead of coding the MV directly, the MV can be derived from metadata associated with multiple reference pictures, e.g., from the most recent reference picture (in decoding order) using the MV associated with one of five surrounding samples denoted A0, A1, and B0, B1, B2 (102 to 106, respectively). In H.265, MV prediction can use a predictor from the same reference picture as the neighboring blocks. Summary of the Invention [Means for solving the problem]

[0022] Aspects of the present disclosure provide methods and apparatus for video encoding and / or decoding. In some examples, the apparatus for video decoding includes a processing circuit that generates feature vectors extracted from reconstructed samples in one or more neighboring blocks of a block of a current picture.

number

[0023] In an embodiment, a subgroup of transform sets may be selected from the group of transform sets based on the prediction mode of the block indicated in the coded information of the block. Feature vectors extracted from reconstructed samples in one or more neighboring blocks of the block.

number

[0024] In one example, a feature vector extracted from reconstruction samples in one or more neighboring blocks of a block

number

[0025] In one example, a transform set can be selected from a subgroup of transform sets based on an index signaled in the coded information. Feature vectors extracted from reconstructed samples in one or more neighboring blocks of the block.

number

[0026] In one example, a feature vector extracted from reconstruction samples in one or more neighboring blocks of a block

number

[0027] In one example, a feature vector is generated based on a statistical analysis of the reconstructed samples in one or more neighboring blocks of the block.

number

number

number

number

[0028] In one example, the feature vector

number

[0029] In one example, a threshold set K S The processing circuit (i) calculates the moments of the variables and the threshold set K S (ii) determining a transformation set from the subgroup of transformation sets based on a threshold value from the coded information; (ii) determining candidate transformations from the subgroup of transformation sets based on moments of the variables and a threshold value; or (iii) selecting a transformation set from the subgroup of transformation sets based on an index of the coded information and determining candidate transformations from the selected transformation set based on moments of the variables and a threshold value.

[0030] In one example, the moment of the variable is one of a first moment of the variable, a second moment of the variable, or a third moment of the variable. The prediction mode of the block is one of a plurality of prediction modes, each of which is based on a threshold set K. S Unique threshold subset K in S ' corresponds to a unique threshold subset K S ' is a set of multiple prediction modes and a threshold set K S 1 shows an injective mapping between multiple threshold subsets in .

[0031] In one example, the processing circuitry determines a threshold set K based on one of: (i) a block size of the block; (ii) a quantization parameter; or (iii) a prediction mode of the block. SSelect a threshold from the

[0032] In one example, the feature vector

number

number

number

[0033] In one example, the classification vector set

number

number

number

number

[0034] Aspects of the present disclosure further provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to perform a method for video decoding and / or encoding.

[0035] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Brief explanation of the drawings]

[0036] [Figure 1A] FIG. 2 is a schematic diagram of an example subset of intra-prediction modes. [Figure 1B] FIG. 1 is a diagram of an exemplary intra-prediction direction. [Figure 2] FIG. 1 is a schematic diagram of a current block and its surrounding spatial merge candidates in one example. [Figure 3] FIG. 3 is a simplified block diagram schematic of a communication system (300) according to one embodiment. [Figure 4] FIG. 4 is a schematic diagram of a simplified block diagram of a communication system (400) according to one embodiment. [Figure 5] FIG. 2 is a schematic diagram of a simplified block diagram of a decoder according to an embodiment. [Figure 6] FIG. 2 is a schematic diagram of a simplified block diagram of an encoder according to an embodiment. [Figure 7] 4 shows a block diagram of an encoder according to another embodiment; [Figure 8]4 shows a block diagram of a decoder according to another embodiment; [Figure 9] 1 illustrates an example of a nominal mode of a coding block according to an embodiment of the present disclosure. [Figure 10] 1 illustrates an example of a non-directional smooth intra prediction mode according to aspects of the present disclosure. [Figure 11] 1 illustrates an example of an intra predictor based on recursive filtering according to an embodiment of the present disclosure. [Figure 12] 1 illustrates an example of multiple reference lines for a coding block according to an embodiment of the present disclosure. [Figure 13] 1 illustrates an example of a linear transformation basis function according to an embodiment of the present disclosure. [Figure 14A] 10 illustrates an example dependency of availability of various transform kernels based on transform block size and prediction mode according to an embodiment of the present disclosure. [Figure 14B] 1 illustrates an exemplary transform type selection based on intra-prediction mode according to one embodiment of this disclosure. [Figure 15] 17 shows two example transform coding processes (1700) and (1800) using a 16x64 transform and a 16x48 transform, respectively, according to an embodiment of the present disclosure. [Figure 16] 17 shows two example transform coding processes (1700) and (1800) using a 16x64 transform and a 16x48 transform, respectively, according to an embodiment of the present disclosure. [Figure 17] 17(A)-17(D) show exemplary residual patterns (grayscale) observed for intra prediction modes according to embodiments of the present disclosure. [Figure 18] 1 illustrates exemplary spatially adjacent samples of a block according to an embodiment of the present disclosure. [Figure 19] 19 shows a flowchart outlining a process (1900) according to an embodiment of the present disclosure. [Figure 20] FIG. 1 is a schematic diagram of a computer system according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0037] Figure 3 shows a simplified block diagram of a communication system (300) according to an embodiment of the present disclosure. The communication system (300) includes multiple terminal devices capable of communicating with each other, e.g., via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). In the example of Figure 3, the first pair of terminal devices (310) and (320) perform unidirectional transmission of data. For example, the terminal device (310) may code video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to the other terminal device (320) via the network (350). The coded video data may be transmitted in the form of one or more coded video bitstreams. The terminal device (320) may receive the coded video data from the network (350), decode the coded video data to reconstruct the video pictures, and display the video pictures according to the reconstructed video data. Unidirectional data transmission is common in media distribution applications, for example.

[0038] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of coded video data, such as may occur during a video conference. In the case of bidirectional transmission of data, in one example, each of the terminal devices (330) and (340) may code video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) over the network (350). Each of the terminal devices (330) and (340) may receive the coded video data transmitted by the other of the terminal devices (330) and (340), decode the coded video data to reconstruct the video pictures, and display the video pictures on an accessible display device according to the reconstructed video data.

[0039] In the example of FIG. 3 , terminal devices 310, 320, 330, and 340 may be depicted as a server, a personal computer, and a smartphone, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure may be applied to laptop computers, tablet computers, media players, and / or dedicated videoconferencing equipment. Network 350 represents any number of networks that transmit coded video data between terminal devices 310, 320, 330, and 340, including, for example, wired and / or wireless communication networks. Communication network 350 may exchange data over circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this discussion, the architecture and topology of network 350 may not be important to the operation of the present disclosure, unless described below.

[0040] 4 illustrates the arrangement of a video encoder and a video decoder in a streaming environment as an example of an application of the disclosed subject matter. The disclosed subject matter is equally applicable to other video function applications, including, for example, video conferencing, digital TV, and storage of compressed video on digital media, including CDs, DVDs, memory sticks, etc.

[0041] The streaming system may also include a capture subsystem (413), which may include a video source (401), such as a digital camera, that creates an uncompressed video picture stream (402). In one example, the video picture stream (402) includes samples captured by the digital camera. The video picture stream (402), shown with a thick line to emphasize its high data volume compared to the coded video data (404) (or coded video bitstream), may be processed by electronics (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The coded video data (404) (or coded video bitstream (404)), shown with a thin line to emphasize its low data volume compared to the stream of video pictures (402), may be stored on a streaming server (405) for future use. One or more streaming client subsystems, such as the client subsystems (406) and (408) in FIG. 4, can access the streaming server (405) to retrieve copies (407) and (409) of the coded video data (404). The client subsystem (406) can include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes an input copy (407) of the coded video data and creates an output stream of video pictures (411) that can be displayed on a display (412) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the coded video data (404), (407), and (409) (e.g., a video bitstream) can be coded according to a particular video coding / compression standard. Examples of these standards include ITU-T Recommendation H.265.In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC), and the disclosed subject matter may be used in the context of VVC.

[0042] It should be noted that the electronic devices 420 and 430 may include other components (not shown). For example, the electronic device 420 may include a video decoder (not shown), and the electronic device 430 may include a video encoder (not shown).

[0043] 5 shows a block diagram of a video decoder (510) according to an embodiment of the present disclosure. The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., receiving circuitry). The video decoder (510) can be used in place of the video decoder (410) in the example of FIG. 4.

[0044] The receiver (531) can receive one or more coded video sequences decoded by the video decoder (510), in the same or another embodiment, one coded video sequence at a time, with the decoding of each coded video sequence being independent of the other coded video sequences. The coded video sequences can be received from a channel (501), which can be a hardware / software link to a storage device that stores the coded video data. The receiver (531) can receive the coded video data with other data, such as coded audio data and / or auxiliary data streams, that can be forwarded to a respective using entity (not shown). The receiver (531) can separate the coded video sequences from other data. To prevent network jitter, a buffer memory (515) can be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In certain applications, the buffer memory (515) is part of the video decoder (510). In other embodiments, the buffer memory may be external to the video decoder (510) (not shown). In yet another embodiment, there may be a buffer memory (not shown) external to the video decoder (510), for example, to prevent network jitter, and another buffer memory (515) internal to the video decoder (510), for example, to handle playback timing. When the receiver (531) receives data from a store-and-forward device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory (515) may not be required or may be smaller. For use with best-effort packet networks such as the Internet, the buffer memory (515) may be required, and may be relatively large and advantageously adaptively sized, and may be implemented at least in part in an operating system or similar element (not shown) external to the video decoder (510).

[0045] The video decoder (510) may include a parser (520) that reconstructs symbols (521) from the coded video sequence. These symbol categories include information used to manage the operation of the video decoder (510) and potential information to control a rendering device, such as a rendering device (512) (e.g., a display screen) that is not a component of the electronics (530) but may be coupled to the electronics (530) as shown in FIG. 5. The control information for the rendering device may be in the form of a supplemental enhancement information (SEI) message or a video usability information (VUI) parameter set fragment (not shown). The parser (520) can parse / entropy decode the received coded video sequence. The coding of the coded video sequence can follow a video coding technique or standard and can follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context-sensitivity, etc. The parser (520) can extract a set of subgroup parameters for at least one subgroup of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroups may include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (520) may also extract information such as transform coefficients, quantization parameter values, motion vectors, etc. from the coded video sequence.

[0046] The parser (520) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to create symbols (521).

[0047] The reconstruction of the symbols (521) can involve several different units, depending on the type of coded video picture or portion thereof (e.g., inter and intra picture, inter and intra block), and other factors. Which units are involved and how may be controlled by subgroup control information parsed from the coded video sequence by the parser (520). For clarity, the flow of such subgroup control information between the parser (520) and the following units is not shown.

[0048] In addition to the functional blocks already mentioned, the video decoder (510) may be conceptually subdivided into several functional units, as described below. In actual implementations subject to commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate:

[0049] The first unit is a scalar / inverse transform unit (551), which receives control information from the parser (520) including the transform to be used, block size, quantization factor, quantization scaling matrix, etc., as well as quantized transform coefficients as symbols (521). The scalar / inverse transform unit (551) can output blocks containing sample values ​​that can be input to an aggregator (555).

[0050] In some cases, the output samples of the scaler / inverse transform (551) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates blocks of the same size and shape as the block being reconstructed using surrounding, already reconstructed information retrieved from a current picture buffer (558). The current picture buffer (558), for example, buffers the partially reconstructed and / or fully reconstructed current picture. The aggregator (555) optionally adds, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).

[0051] In other cases, the output samples of the scalar / inverse transform unit (551) may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion-compensated prediction unit (553) may access a reference picture memory (557) to retrieve samples for prediction. After motion-compensating the retrieved samples according to the symbols (521) related to the block, these samples may be added to the output of the scalar / inverse transform unit (551) by an aggregator (555) to generate output sample information (in this case, referred to as residual samples or residual signals). The addresses in the reference picture memory (557) from which the motion-compensated prediction unit (553) retrieves prediction samples may be controlled by motion vectors available to the motion-compensated prediction unit (553), for example, in the form of symbols (521) that may have X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​retrieved from the reference picture memory (557) when sub-sample accurate motion vectors are in use, motion vector prediction mechanisms, etc.

[0052] The output samples of the aggregator (555) may be subjected to various loop filtering techniques in a loop filter unit (556). The video compression techniques are controlled by parameters made available to the loop filter unit (556) as symbols (521) from the parser (520) contained in the coded video sequence (also called a coded video bitstream), and may include loop filter techniques that are responsive to meta-information obtained during the decoding of a coded picture or previous portions of the coded video sequence (in decoding order), as well as to previously reconstructed, loop-filtered sample values.

[0053] The output of the loop filter unit (556) may be a sample stream that can be output to the rendering device (512) and stored in a reference picture memory (557) for use in future inter-picture prediction.

[0054] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a new current picture buffer can be reallocated before beginning reconstruction of the next coded picture.

[0055] The video decoder (510) can perform decoding operations according to a predetermined video compression technology in a standard such as ITU-T Rec. H.265. A coded video sequence may comply with the syntax specified by the video compression technology or standard being used, in the sense that the coded video sequence conforms to both the syntax of the video compression technology or standard and the profile documented in the video compression technology or standard. Specifically, a profile may select specific tools from all tools available in the video compression technology or standard as unique tools available for use in that profile. Compliance also requires that the complexity of the coded video sequence be within a range defined by the level of the video compression technology or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level may be further limited by a hypothetical reference decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.

[0056] In an embodiment, the receiver (531) can receive additional (redundant) data along with the coded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0057] 6 shows a block diagram of a video encoder (603) according to an embodiment of the present disclosure. The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) may be used in place of the video encoder (403) in the example of FIG. 4.

[0058] The video encoder (603) can receive video samples from a video source (601) (not part of the electronics (620) in the example of FIG. 6) that can capture video images to be coded by the video encoder (603). In another example, the video source (601) is part of the electronics (620).

[0059] The video source (601) may provide a source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream, which may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 Y CrCB, RGB, etc.), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media distribution system, the video source (601) may be a storage device that stores prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as multiple individual pictures that, when viewed sequentially, impart motion. The pictures themselves may be organized as a spatial array of pixels, each of which may contain one or more samples depending on the sampling structure, color space, etc., in use. Those skilled in the art will readily understand the relationship between pixels and samples. The following discussion will focus on samples.

[0060] According to one embodiment, the video encoder (603) can code and compress pictures of a source video sequence into a coded video sequence (643) in real time or under any other time constraints required by the application. Enforcing an appropriate coding rate is one of the functions of the controller (650). In some embodiments, the controller (650) controls and is operatively coupled to other functional units described below. For clarity, couplings are not depicted. Parameters set by the controller (650) can include rate control-related parameters (picture skip, quantization, lambda values ​​for rate-distortion optimization techniques, etc.), picture size, group-of-picture (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured with other appropriate functions for the video encoder (603) optimized for a particular system design.

[0061] In some embodiments, the video encoder (603) is configured to operate in a coding loop. As an overly simplified explanation, in one example, the coding loop may include a source coder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be coded and reference pictures) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to that created by the (remote) decoder (since any compression between the symbols and the coded video bitstream is lossless in the video compression techniques considered in the disclosed subject matter). The reconstructed sample stream (sample data) is input to a reference picture memory (634). Because decoding of the symbol stream leads to bit-exact results regardless of the location (local or remote) of the decoder, the contents of the reference picture memory (634) are also bit-exact between the local and remote encoders. In other words, the prediction part of the encoder "sees" as reference picture samples exactly the same sample values ​​that the decoder "sees" when using the prediction during decoding. This basic principle of reference picture synchrony (and the drift that occurs when synchrony cannot be maintained, e.g., due to channel errors) is also used in several related fields.

[0062] The operation of the "local" decoder (633) may be similar to the operation of a "remote" decoder, such as the video decoder (510), already described in detail above with reference to Figure 5. However, with brief reference to Figure 5, because symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder (645) and parser (520) may be lossless, the entropy decoding portion of the video decoder (510), including the buffer memory (515) and parser (520), may not be fully implemented in the local decoder (633).

[0063] As can be seen from this point, any decoder techniques other than analysis / entropy decoding present in a decoder must necessarily be present in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter focuses on decoder operation. A description of the encoder techniques can be omitted, as they are the opposite of the decoder techniques described generically. Only in certain areas will more detailed descriptions be required, and these will be provided below.

[0064] During operation, in some examples, the source coder (630) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as “reference pictures.” In this manner, the coding engine (632) codes differences between pixel blocks of the input picture and pixel blocks of reference pictures that may be selected as prediction references for the input picture.

[0065] The local video decoder (633) can decode coded video data of pictures that may be designated as reference pictures based on symbols created by the source coder (630). The operation of the coding engine (632) can advantageously be a lossy process. When the coded video data can be decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence may be a replica of the source video sequence, typically with some errors. The local video decoder (633) can replicate the decoding process that may be performed on reference pictures by the video decoder and store the reconstructed reference pictures in a reference picture cache (634). In this way, the video encoder (603) can locally store copies of reconstructed reference pictures that have content in common (no transmission errors) with reconstructed reference pictures obtained by the far-end video decoder.

[0066] The predictor (635) can perform a prediction search for the coding engine (632). That is, for a new picture to be coded, the predictor (635) can search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or specific metadata that can serve as an appropriate prediction basis for the new picture, such as the reference picture's motion vectors, block shapes, etc. The predictor (635) can operate on a sample block / pixel block basis to find an appropriate prediction basis. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have prediction basis drawn from multiple reference pictures stored in the reference picture memory (634).

[0067] The controller (650) can manage the coding operations of the source coder (630), including, for example, setting the parameters and subgroup parameters used to code the video data.

[0068] The output of all of the aforementioned functional units may undergo entropy coding in an entropy coder (645), which converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, or arithmetic coding.

[0069] The transmitter (640) can buffer the coded video sequence created by the entropy coder (645) in preparation for transmission over a communication channel (660), which can be a hardware / software link to a storage device that stores the coded video data. The transmitter (640) can merge the coded video data from the video coder (603) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0070] The controller (650) may manage the operation of the video encoder (603). During coding, the controller (650) may assign a particular coding picture type to each coded picture, which may affect the coding technique that may be applied to the respective picture. For example, pictures may generally be assigned one of the following picture types:

[0071] An intra picture (I-picture) may be one that can be coded and decoded without using any other picture in a sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures. Those skilled in the art are aware of these variations of I-pictures and their respective uses and characteristics.

[0072] A predictive picture (P picture) may be one that can be coded and decoded by intra- or inter-prediction using at most one motion vector and reference index to predict the sample values ​​of each block.

[0073] Bidirectionally predicted pictures (B-pictures) may be coded and decoded by intra- or inter-prediction using up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predicted pictures may use more than two reference pictures and associated metadata to reconstruct a single block.

[0074] A source picture is generally spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks determined by the coding assignment applied to the block's respective picture. For example, blocks of an I-picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P-picture may be predictively coded via spatial prediction or via temporal prediction with reference to one pre-coded reference picture. Blocks of a B-picture may be predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.

[0075] The video encoder (603) may perform coding operations according to a predetermined video coding technique or standard, such as ITU-T Rec. H.265. During operation, the video encoder (603) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard used.

[0076] In an embodiment, the transmitter (640) can transmit additional data along with the coded video. The source coder (630) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, redundant data in other forms such as redundant pictures or slices, SEI messages, VUI parameter set fragments, etc.

[0077] Video may be captured as a time sequence of multiple source pictures (video pictures). Intra-picture prediction (often abbreviated as "intra-prediction") exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or otherwise) between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. If a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, the block in the current picture may be coded by a vector called a motion vector. A motion vector points to a reference block in the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0078] In some embodiments, bidirectional prediction may be used in inter-picture prediction. Bidirectional prediction uses two reference pictures, such as a first reference picture and a second reference picture, each of which is earlier in decoding order than a current picture in a video (but may be earlier and later in display order, respectively). A block in the current picture may be coded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. A block may be predicted by a combination of the first reference block and the second reference block.

[0079] Furthermore, merge mode techniques can be applied to inter-picture prediction to improve coding efficiency.

[0080] According to some embodiments of the present disclosure, prediction, such as inter-picture prediction and intra-picture prediction, is performed block-by-block. For example, according to the HEVC standard, pictures in a sequence of video pictures are divided into coding tree units (CTUs) for compression, and the CTUs within a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU may be recursively quadtree-decomposed into one or more coding units (CUs). For example, a 64x64 pixel CTU may be divided into one CU of 64x64 pixels, four CUs of 32x32 pixels, or 16 CUs of 16x16 pixels. In one example, each CU is analyzed to determine the prediction type of the CU, such as inter-prediction or intra-prediction. The CU is then divided into one or more prediction units (PUs) according to temporal and / or spatial predictability. Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In an embodiment, prediction operations in coding (encoding / decoding) are performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values ​​(e.g., luma values), such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.

[0081] 7 shows a diagram of a video encoder (703) according to another embodiment of this disclosure. The video encoder (703) is configured to receive a processed block (e.g., a predictive block) of sample values ​​in a current video picture in a video picture sequence and to code the processed block into a coded picture that is part of the coded video sequence. In one example, the video encoder (703) is used in place of the video encoder (403) in the example of FIG. 4.

[0082] In an HEVC example, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as a prediction block, such as an 8x8 sample. The video encoder (703) determines, for example, using rate-distortion optimization, whether to best code the processing block in intra mode, inter mode, or bi-prediction mode. If the processing block is to be coded in intra mode, the video encoder (703) can code the processing block into a coded picture using intra prediction. If the processing block is to be coded in inter mode or bi-prediction mode, the video encoder (703) can code the processing block into a coded picture using inter prediction or bi-prediction, respectively. In certain video coding techniques, merge mode may be an inter-picture prediction sub-mode that derives motion vectors from one or more motion vector predictors without relying on coded motion vector components outside the predictors. In certain other video coding techniques, there may be motion vector components applicable to the current block. In one example, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the mode of the processing block.

[0083] In the example of Figure 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculation unit (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725), which are coupled together as shown in Figure 7.

[0084] The inter-encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous preceding picture and a subsequent picture), generate inter-prediction information (e.g., a description of redundant information due to an inter-coding method, motion vectors, merge mode information), and calculate an inter-prediction result (e.g., a predicted block) using any suitable technique based on the inter-prediction information. In some examples, the reference picture is a decoded reference picture based on coded video information.

[0085] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), possibly compare the block with previously coded blocks in the same picture, and, after transformation, generate quantized coefficients and possibly also generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques). In one example, the intra encoder (722) also calculates intra prediction results (e.g., predicted blocks) based on reference blocks and intra prediction information in the same picture.

[0086] The general-purpose controller (721) is configured to determine general-purpose control data and control other components of the video encoder (703) based on the general-purpose control data. In one example, the general-purpose controller (721) determines the mode of the block and provides a control signal to the switch (726) based on the mode. For example, if the mode is intra mode, the general-purpose controller (721) controls the switch (726) to select an intra-mode result for use by the residual calculation unit (723) and controls the entropy encoder (725) to select intra-prediction information and include the intra-prediction information in the bitstream. If the mode is inter mode, the general-purpose controller (721) controls the switch (726) to select an inter-prediction result for use by the residual calculation unit (723) and controls the entropy encoder (725) to select inter-prediction information and include the inter-prediction information in the bitstream.

[0087] The residual calculation unit (723) is configured to calculate the difference (residual data) between the received block and a prediction result selected from the intra-encoder (722) or the inter-encoder (730). The residual encoder (724) operates based on the residual data and is configured to code the residual data to generate transform coefficients. In one example, the residual encoder (724) is configured to transform the residual data from the spatial domain to the frequency domain to generate transform coefficients. Then, the transform coefficients are quantized to obtain quantized transform coefficients. In various embodiments, the video encoder (703) further includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform to generate decoded residual data. The decoded residual data may be used appropriately by the intra-encoder (722) and the inter-encoder (730). For example, the inter-encoder (730) can generate a decoding block based on the decoded residual data and the inter-prediction information, and the intra-encoder (722) can generate a decoding block based on the decoded residual data and the intra-prediction information. In some examples, the decoding block is appropriately processed to generate a decoding picture, which can be buffered in a memory circuit (not shown) and used as a reference picture.

[0088] The entropy encoder (725) is configured to format the bitstream to include the coded block. The entropy encoder (725) is configured to include various information in accordance with an appropriate standard, such as the HEVC standard. In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other appropriate information in the bitstream. Note that, according to the disclosed subject matter, there is no residual information when coding a block in an inter mode or a merged sub-mode of a bi-prediction mode.

[0089] 8 shows a diagram of a video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive coded pictures that are part of a coded video sequence and decode the coded pictures to generate reconstructed pictures. In one example, the video decoder (810) is used in place of the video decoder (410) in the example of FIG. 4.

[0090] In the example of Figure 8, the video decoder (810) includes an entropy decoder (871), an inter-decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-decoder (872), which are coupled together as shown in Figure 8.

[0091] The entropy decoder (871) can be configured to reconstruct, from a coded picture, certain symbols that represent the syntax elements that make up the coded picture. Such symbols can include, for example, prediction information (e.g., intra-prediction information or inter-prediction information) that can identify the mode in which the block is coded (e.g., intra-mode, inter-mode, bi-prediction mode, the later merged submode of both, or other submodes), certain samples or metadata used for prediction by the intra-decoder (872) or inter-decoder (880), respectively, residual information in the form of quantized transform coefficients, etc. In one example, if the prediction mode is an inter- or bi-prediction mode, the inter-prediction information is provided to the inter-decoder (880). Also, if the prediction type is an intra-prediction type, the intra-prediction information is provided to the intra-decoder (872). The residual information can be dequantized and provided to the residual decoder (873).

[0092] The inter decoder (880) is configured to receive the inter prediction information and generate an inter prediction result based on the inter prediction information.

[0093] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.

[0094] The residual decoder (873) is configured to perform inverse quantization to extract dequantized transform coefficients and process the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) may require certain control information (such as quantizer parameters (QP)), which may be provided by the entropy decoder (871) (a data path not shown, as there may only be a small amount of control information).

[0095] The reconstruction module (874) is configured to combine, in the spatial domain, the residual output by the residual decoder (873) and the prediction results (possibly output by the inter- or intra-prediction module) to form reconstructed blocks that may be part of a reconstructed picture that may be part of the reconstructed video. Note that other appropriate operations, such as deblocking operations, may be performed to improve visual quality.

[0096] It should be noted that the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using any suitable technology. In some embodiments, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using one or more integrated circuits. In other embodiments, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using one or more processors executing software instructions.

[0097] Video coding techniques related to a transform set or transform kernel selection scheme based on reconstructed samples of one or more neighboring blocks of a block (e.g., feature indicators (e.g., feature vectors or feature scalars of reconstructed samples of one or more neighboring blocks of the block) are disclosed. The video coding format may include an open video coding format designed for video transmission over the Internet, such as AOMedia Video 1 (AV1) or a next-generation AOMedia Video format beyond AV1. The video coding standard may also include the High Efficiency Video Coding (HEVC) standard, a next-generation video coding beyond HEVC (e.g., Versatile Video Coding (VVC)), etc.

[0098] Intra prediction, e.g., AV1, VVC, etc., may use various intra prediction modes. In an embodiment, e.g., AV1, directional intra prediction is used. In directional intra prediction, predicted samples of a block may be generated by extrapolating from neighboring reconstructed samples along one direction. The direction corresponds to an angle. The modes used in directional intra prediction to predict predicted samples of a block may be referred to as directional modes (also referred to as directional prediction modes, directional intra modes, directional intra prediction modes, or angle modes). Each directional mode may correspond to a different angle or a different direction. In one example, e.g., the open video coding format VP9 uses eight directional modes corresponding to eight angles from 45° to 207°. The eight directional modes may also be referred to as nominal modes (e.g., V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED). To take advantage of more diverse spatial redundancy in directional textures (e.g., AV1), the directional modes can be extended beyond the eight nominal modes to angle sets with finer granularity and more angles (or directions), as shown in Figure 9, for example.

[0099] 9 illustrates an example of nominal modes for a coding block (CB) (910) according to an embodiment of the present disclosure. A particular angle (referred to as a nominal angle) can correspond to a nominal mode. In one example, eight nominal angles (or nominal intra-angles) (901) through (908) correspond to eight nominal modes (e.g., V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED). The eight nominal angles (901) through (908) and the eight nominal modes can be referred to as V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED, respectively. The nominal mode index may indicate the nominal mode (e.g., one of eight nominal modes). In one example, the nominal mode index is signaled.

[0100] Furthermore, each nominal angle can correspond to multiple finer angles (e.g., seven finer angles), and thus, for example, in AV1, 56 angles (or prediction angles) or 56 directional modes (or angle modes, directional intra-prediction modes) can be used. Each prediction angle can be represented by a nominal angle and an angle offset (or angle delta). The angle offset is determined by multiplying an offset integer I (e.g., −3, −2, −1, 0, 1, 2, or 3) by a step size (e.g., 3°). In one example, the prediction angle is equal to the sum of the nominal angle and the angle offset. In one example, for example, in AV1, the nominal modes (e.g., eight nominal modes (901) to (908)) can be signaled along with specific non-angular smoothing modes (e.g., DC mode, PAETH mode, SMOOTH mode, vertical SMOOTH mode, and horizontal SMOOTH mode, which will be described later). Subsequently, if the current prediction mode is a directional mode (or angular mode), an index indicating an angular offset (e.g., offset integer I) corresponding to the nominal angle can be further signaled. In one example, the directional mode (e.g., one of 56 directional modes) can be determined based on the nominal mode index and an index indicating the angular offset from the nominal mode. In one example, to realize the directional prediction mode through a generic method, the 56 directional modes used in AV1 are realized with a unified directional predictor that can project each pixel to a reference sub-pixel position and interpolate the reference pixel by a 2-tap bilinear filter.

[0101] An omnidirectional smooth intra predictor (also referred to as an omnidirectional smooth intra prediction mode, an omnidirectional smooth mode, or a non-angular smooth mode) may be used for intra prediction of CB. In some examples (e.g., in AV1), five omnidirectional smooth intra prediction modes include a DC mode or DC predictor (e.g., DC), a PAETH mode or PAETH predictor (e.g., PAETH), a SMOOTH mode or SMOOTH predictor (e.g., SMOOTH), a vertical SMOOTH mode (referred to as a SMOOTH_V mode, a SMOOTH_V predictor, or SMOOTH_V), and a horizontal SMOOTH mode (referred to as a SMOOTH_H mode, a SMOOTH_H predictor, or SMOOTH_H).

[0102] 10 illustrates examples of non-directional smooth intra-prediction modes (e.g., DC mode, PAETH mode, SMOOTH mode, SMOOTH_V mode, and SMOOTH_H mode) according to aspects of the present disclosure. To predict a sample (1001) in CB (1000) based on a DC predictor, an average value of a first value of a left-neighboring sample (1012) and a second value of an above-neighboring sample (or top-neighboring sample) (1011) can be used as a predictor.

[0103] To predict sample (1001) based on the PAETH predictor, the first value of the left adjacent sample (1012), the second value of the top adjacent sample (1011), and the third value of the upper-left adjacent sample (1013) can be obtained. Then, the reference value is calculated using Equation 1. Reference value = First value + Second value - Third value (Equation 1)

[0104] One of the first value, the second value, and the third value that is closest to the reference value can be set as the predictor for sample (1001).

[0105] The SMOOTH_V mode, the SMOOTH_H mode, and the SMOOTH mode can predict CB(1000) using quadratic interpolation of the average values ​​in the vertical direction, the horizontal direction, and the vertical and horizontal directions, respectively. To predict the sample (1001) based on the SMOOTH predictor, an average value (e.g., a weighted combination) of the first value, the second value, the value of the right sample (1014), and the value of the bottom sample (1016) can be used. In various examples, because the right sample (1014) and the bottom sample (1016) are not reconstructed, the value of the upper right adjacent sample (1015) and the value of the lower left adjacent sample (1017) can replace the values ​​of the right sample (1014) and the bottom sample (1016), respectively. Therefore, an average value (e.g., a weighted combination) of the first value, the second value, the value of the upper right adjacent sample (1015), and the value of the lower left adjacent sample (1017) can be used as the SMOOTH predictor. To predict sample (1001) based on the SMOOTH_V predictor, an average value (e.g., a weighted combination) of the second value of the top neighboring sample (1011) and the value of the lower-left neighboring sample (1017) can be used. To predict sample (1001) based on the SMOOTH_H predictor, an average value (e.g., a weighted combination) of the first value of the left neighboring sample (1012) and the value of the upper-right neighboring sample (1015) can be used.

[0106] FIG. 11 illustrates an example of an intra predictor based on recursive filtering (also referred to as a filter intra mode or recursive filtering mode) according to an embodiment of the present disclosure. The filter intra mode can be used for the CB (1100) to capture the decaying spatial correlation with the reference on the edge. In one example, the CB (1100) is a luma block. The luma block (1100) can be divided into multiple patches (e.g., eight 4×2 patches B0-B7). Each of the patches B0-B7 can have multiple neighboring samples. For example, patch B0 has seven neighboring samples (or seven neighboring elements) R00-R06, including four top neighboring samples R01-R04, two left neighboring samples R05-R06, and an upper-left neighboring sample R00. Similarly, patch B7 has four top-neighboring samples R71-R74, two left-neighboring samples R75-R76, and seven neighboring samples R70-R76, including the top-left neighbor sample R70.

[0107] In some examples, multiple (e.g., five) filter intra modes (or multiple recursive filtering modes) are pre-designed, e.g., for AV1. Each filter intra mode may be represented by a set of eight 7-tap filters that reflect the correlation between a sample (or pixel) in a corresponding 4×2 patch (e.g., B0) and seven neighbors (e.g., R00-R06) adjacent to the 4×2 patch B0. The weighting coefficients of the 7-tap filters may be position-dependent. For each of the patches B0-B7, the seven neighbors (e.g., R00-R06 for B0 and R70-R76 for B7) may be used to predict the sample in the corresponding patch. In one example, neighboring elements R00-R06 are used to predict the sample in patch B0. In one example, neighboring elements R70-R76 are used to predict the sample in patch B7. For a particular patch in CB (1100), such as patch B0, all of the seven neighboring elements (e.g., R00-R06) have already been reconstructed. For other patches in CB(1100), at least one of the seven neighboring elements has not been reconstructed, so the predicted value (or predicted sample of the nearest neighboring element) can be used as a reference. For example, the seven neighboring elements R70-R76 of patch B7 have not been reconstructed, so the predicted sample of the nearest neighboring element can be used.

[0108] Chroma samples may be predicted from luma samples. In one embodiment, a chroma from luma mode (e.g., CfL mode, CfL predictor) is a chroma-only intra predictor that can model chroma samples (or pixels) as linear functions of the corresponding reconstructed luma samples (or pixels). For example, CfL prediction can be expressed using Equation 2 as follows: CfL(α)=αL A +D (Formula 2) where L Arepresents the AC contribution of the luma component, α represents a scaling parameter of the linear model, and D represents the DC contribution of the chroma component. In one example, the reconstructed luma pixels are subsampled based on the chroma resolution and the mean value is subtracted to remove the AC contribution (e.g., L A ) is obtained. Instead of the decoder calculating the scaling parameter α to approximate the chroma AC components from the AC contributions, in some cases, for example in AV1, the CfL mode determines the scaling parameter α based on the original chroma pixels and signals the scaling parameter α in the bitstream, resulting in a more accurate prediction while reducing decoder complexity. The DC contributions of the chroma components are obtained using intra DC mode, which is sufficient for most chroma content and has mature, high-speed implementations.

[0109] Multi-line intra prediction can use more reference lines for intra prediction. A reference line may include multiple samples within a picture. In one example, a reference line includes row samples and column samples. In one example, an encoder can determine and signal the reference line used to generate an intra predictor. An index indicating the reference line (also referred to as a reference line index) may be signaled before the intra prediction mode. In one example, if a non-zero reference line index is signaled, only MPM is allowed. Figure 12 shows an example of four reference lines for CB (1210). Referring to Figure 12, the reference line may include up to six segments, such as segments A through F, and an upper-left reference sample. For example, reference line 0 includes segments B and E and an upper-left reference sample. For example, reference line 3 includes segments A through F and an upper-left reference sample. Segments A and F may be padded with nearest samples from segments B and E, respectively. In some examples, for example, in HEVC, only one reference line (e.g., reference line 0 adjacent to CB(1210)) is used for intra prediction. In some examples, for example, in VVC, multiple reference lines (e.g., reference lines 0, 1, and 3) are used for intra prediction.

[0110] Generally, a block may be predicted using one or a suitable combination of various intra-prediction modes, such as those described above with reference to FIGS.

[0111] The following describes an embodiment of a linear transform such as that used in AO Media Video 1 (AV1). A forward transform (e.g., in an encoder) may be performed on a transform block (TB) containing a residual (e.g., a residual in the spatial domain) to obtain a TB containing transform coefficients in the frequency domain (or spatial frequency domain). The TB containing the residual in the spatial domain is referred to as a residual TB, and the TB containing transform coefficients in the frequency domain is referred to as a coefficient TB. In one example, the forward transform includes a forward linear transform that can convert the residual TB into coefficients TB. In one example, the forward transform includes a forward linear transform and a forward secondary transform, in which the forward linear transform can convert the residual TB into intermediate coefficients TB, and the forward secondary transform can convert the intermediate coefficients TB into coefficients TB.

[0112] An inverse transform (e.g., in an encoder or decoder) may be performed on the coefficients TB in the frequency domain to obtain residuals TB in the spatial domain. In one example, the inverse transform includes an inverse linear transform that can transform the coefficients TB into residuals TB. In one example, the inverse transform includes an inverse linear transform and an inverse secondary transform, where the inverse secondary transform can transform the coefficients TB into intermediate coefficients TB, and the inverse primary transform can transform the intermediate coefficients TB into residuals TB.

[0113] Typically, the linear transform may refer to a forward linear transform or an inverse linear transform, where the linear transform is performed between the residual TB and the coefficients TB. In some embodiments, the linear transform may be a separable transform. Here, the 2D linear transform may include a horizontal linear transform (also referred to as a horizontal transform) and a vertical linear transform (also referred to as a vertical transform). The secondary transform may refer to a forward secondary transform or an inverse secondary transform, where the secondary transform is performed between the intermediate coefficients TB and the coefficients TB.

[0114] To support extended coding block partitions as described in this disclosure, multiple transform sizes (e.g., ranging from 4 points to 64 points for each dimension) and transform shapes (e.g., square, rectangle with width to height ratio of 2:1, 1:2, 4:1 or 1:4) may be used, such as in AV1.

[0115] The 2D transform process may use hybrid transform kernels that can include different 1D transforms for each dimension of the coded residual block. Linear 1D transforms may include (a) 4-point, 8-point, 16-point, 32-point, and 64-point DCT-2; (b) 4-point, 8-point, and 16-point asymmetric DST (ADST) (e.g., DST-4, DST-7, etc.) and corresponding inverted versions (e.g., an inverted version of ADST, or FlipADST, can apply ADST in reverse order); and / or (c) 4-point, 8-point, 16-point, and 32-point identity transform (IDTX). Figure 13 shows example linear transform basis functions according to an embodiment of the present disclosure. The linear transform basis functions in the example of Figure 13 include basis functions of DCT-2 and asymmetric DST (DST- and DST-7) with N-point input. The linear transform basis functions shown in Figure 13 may be used for AV1.

[0116] The availability of hybrid transform kernels may depend on the transform block size and prediction mode. Figure 14A illustrates an exemplary dependency of the availability of various transform kernels (e.g., the transform types shown in the first column and described in the second column) based on the transform block size (e.g., the sizes shown in the third column) and the prediction mode (e.g., intra-prediction and inter-prediction shown in the third column). Exemplary hybrid transform kernels and their availability based on the prediction mode and transform block size can be used in AV1. Referring to Figure 14A, the symbol

number

number

number

[0117] In one example, the transform type (1410) is represented by ADST_DCT as shown in the first column of Figure 14A. The transform type (1410) includes a vertical ADST and a horizontal DCT as shown in the second column of Figure 14A. According to the third column of Figure 14A, the transform type (1410) is available for intra prediction and inter prediction when the block size is 16x16 (e.g., 16x16 samples, 16x16 luma samples) or less.

[0118] In one example, the transform type (1420) is indicated by V_ADST, as shown in the first column of FIG. 14A. The transform type (1420) includes ADST in the vertical direction and IDTX (i.e., identity matrix) in the horizontal direction, as shown in the second column of FIG. 14A. Thus, the transform type (1420) (e.g., V_ADST) is performed vertically but not horizontally. According to the third column of FIG. 14A, the transform type (1420) is unavailable for intra prediction, regardless of the block size. The transform type (1420) is available for inter prediction when the block size is less than 16x16 (e.g., 16x16 samples, 16x16 luma samples).

[0119] In one example, Figure 14A is applicable to the luma component. For the chroma components, the selection of the transform type (or transform kernel) may be performed implicitly. In one example, for intra-prediction residuals, the transform type may be selected according to the intra-prediction mode, as shown in Figure 14B. In one example, the selection of the transform type shown in Figure 14B is applicable to the chroma components. For inter-prediction residuals, the transform type may be selected according to the transform type of the co-located luma block. Thus, in one example, the transform type for the chroma components is not signaled in the bitstream.

[0120] A transform, such as a linear transform, a secondary transform, etc., may be applied to a block, such as CB. In one example, the transform includes a combination of a linear transform, a secondary transform, etc. The transform may also be a non-separable transform, a separable transform, or a combination of a non-separable transform and a separable transform.

[0121] The secondary transform can be performed in VVC, etc. In some examples, as shown in Figures 15-16, for example in VVC, a low frequency non-separable transform (LFNST) (also called a reduced secondary transform (RST)) can be applied between the forward primary transform and quantization at the encoder side and the inverse quantization and inverse primary transform at the decoder side to further decorrelate the primary transform coefficients.

[0122] The application of non-separable transforms available in LFNST can be explained as follows, using a 4x4 input block (or input matrix) X as an example (shown in Equation 3). To apply a 4x4 non-separable transform, the 4x4 input block X is divided into vectors X, X = 1, X = 2, X = 3, X = 4, X = 5, X = 6, X = 7, X = 8, X = 9, X = 10, X = 11, X = 12, X = 13, X = 14, X = 15, X = 16, X = 17, X = 18, X = 19, X = 20, X = 21, X = 22, X = 23, X = 24, X = 25, X = 26, X = 27, X = 28, X = 29, X = 30, X = 31,

number

number

[0123] The non-separable transformation is

number

number

number

[0124] A non-separable secondary transform may be applied to a block such as CB. In some examples, for example, in VVC, LFNST is applied between the forward primary transform and quantization (e.g., at the encoder side) and between the inverse quantization and the inverse primary transform (e.g., at the decoder side), as shown in Figures 15-16.

[0125] 15-16 show examples of two transform coding processes (1700) and (1800), respectively, using a 16x64 transform (or a 64x16 transform, depending on whether the transform is a forward or inverse secondary transform) and a 16x48 transform (or a 48x16 transform, depending on whether the transform is a forward or inverse secondary transform). Referring to FIG. 15, in process (1700), the encoder side may first perform a forward linear transform (1710) on a block (e.g., a residual block) to obtain a coefficient block (1713). Subsequently, a forward secondary transform (or forward LFNST) (1712) may be applied to the coefficient block (1713). In the forward secondary transform (1712), the 64 coefficients of the 4x4 sub-blocks A through D in the upper left corner of the coefficient block (1713) can be represented by a 64-length vector, and a 16-length vector can be obtained by multiplying the 64-length vector by a 64x16 (i.e., 64 widths and 16 heights) transform matrix. The elements in the 16-length vector are backfilled into the 4x4 sub-block A in the upper left corner of the coefficient block (1713). The coefficients in sub-blocks B through D can be zero. In the quantization step (1714), the coefficients obtained after the forward secondary transform (1712) are quantized and entropy coded to generate coded bits in the bitstream (1716).

[0126] The coded bits are received at the decoder side and can be entropy decoded, followed by an inverse quantization step (1724) to generate a coefficient block (1723). An inverse secondary transform (or inverse LFNST) (1722), such as an inverse RST8x8, can be performed to obtain, for example, 64 coefficients from the 16 coefficients in the upper-left 4x4 sub-block E. The 64 coefficients can then be backfilled into 4x4 sub-blocks E-H. The coefficients in the coefficient block (1723) after the inverse secondary transform (1722) can then be processed with an inverse primary transform (1720) to obtain a reconstructed residual block.

[0127] The process (1800) of the example in FIG. 16 is similar to the process (1700), except that fewer (i.e., 48) coefficients are processed during the forward secondary transform (1712). Specifically, the 48 coefficients in sub-blocks A to C are processed with a smaller transform matrix of size 48×16. By using the smaller 48×16 transform matrix, the memory size for storing the transform matrix and the number of calculations (e.g., multiplications, additions, subtractions, etc.) can be reduced, thus reducing the complexity of the calculation.

[0128] In one example, a 4×4 non-separable transform (e.g., 4×4 LFNST) or an 8×8 non-separable transform (e.g., 8×8 LFNST) is applied according to the block size such as block CB. The block size of CB can include the width, height, etc. For example, the 4×4 LFNST is applied to a CB where the minimum value of the width and height is smaller than a threshold such as 8 (e.g., min(width, height) < 8). For example, the 8×8 LFNST is applied to a CB where the minimum value of the width and height is greater than a threshold such as 4 (e.g., min(width, height) > 4).

[0129] The non-separable transform (e.g., LFNST) can be performed based on a direct matrix multiplication approach and thus can be implemented in a single pass without iteration. To reduce the dimension of the non-separable transform matrix and minimize the computational complexity and the memory space for storing the transform coefficients, in LFNST, a reduced non-separable transform method (or RST) can be used. Thus, in the reduced non-separable transform, a vector of dimension N (e.g., N is 64 in the 8×8 non-separable secondary transform (NSST)) can be mapped to a vector of dimension R in a different space, where N / R (R < N) is the reduction factor. Thus, the RST matrix becomes an R×N matrix as described in Equation 5 instead of an N×N matrix.

Number

[0130] In Equation 5, the R rows of the R×N transformation matrix are the R basis of the N-dimensional space. The inverse transformation matrix is ​​the transformation matrix used in the forward transformation (e.g., T R×N ) can be the transpose of the 8x8 LFNST. In the case of an 8x8 LFNST, a reduction factor of 4 can be applied, reducing the 64x64 direct matrix used in the 8x8 non-separable transform to a 16x64 direct matrix, as shown in Figure 15. Alternatively, a reduction factor greater than 4 can be applied, reducing the 64x64 direct matrix used in the 8x8 non-separable transform to a 16x48 direct matrix, as shown in Figure 16. Therefore, a 48x16 inverse RST matrix can be used at the decoder side to generate the core (primary) transform coefficients of the 8x8 top-left region.

[0131] 16, when a 16x48 matrix is ​​applied instead of a 16x64 matrix with the same transform set configuration, the input to the 16x48 matrix contains 48 input data from three 4x4 blocks A, B, and C in the upper left 8x8 block, excluding the lower right 4x4 block D. Dimensionality reduction can reduce the memory usage for storing the LFNST matrix, for example, from 10 KB to 8 KB with minimal performance degradation.

[0132] To reduce complexity, LFNST may be restricted to be applicable only when coefficients outside the first coefficient subgroup are insignificant. In one example, LFNST may be restricted to be applicable only when all coefficients outside the first coefficient subgroup are insignificant. Referring to Figures 15-16, the first coefficient subgroup corresponds to the upper left block E, and therefore, coefficients outside block E are insignificant.

[0133] In one example, when LFNST is applied, only primary transform coefficients are non-significant (e.g., zero). In another example, when LFNST is applied, all primary transform coefficients are zero. Primary transform coefficients only may refer to transform coefficients resulting from a primary transform without a secondary transform. Thus, LFNST index signaling may be conditional on the last significant position, thereby avoiding an extra coefficient scan in LFNST. In some examples, extra coefficient scans are used to check for significant transform coefficients at specific positions. In one example, the worst-case scenario for LFNST, for example, limits non-separable transforms of 4x4 and 8x8 blocks to 8x16 and 8x48 transforms, respectively, in terms of per-pixel multiplications. In the above cases, when LFNST is applied, the last significant scan position can be less than 8. For other sizes, when LFNST is applied, the last significant scan position can be less than 16. For 4xN and Nx4 (N greater than 8) CBs, this restriction may mean that LFNST is applied to the top-left 4x4 region of the CB. In one example, this restriction means that LFNST is applied only once to the top-left 4x4 region of the CB. In one example, when LFNST is applied, the number of operations for the linear transform is reduced because all linear coefficients are non-significant (e.g., zero). From the encoder's perspective, quantization of the transform coefficients can be significantly simplified when the LFNST transform is tested. Rate-distortion optimized quantization can be performed at a maximum on, for example, the first 16 coefficients in scan order, and the remaining coefficients can be zeroed.

[0134] An LFNST transform (e.g., a transform kernel, transform core, or transform matrix) may be selected as described below. In one embodiment, multiple transform sets may be used, and one or more non-separable transform matrices (or kernels) may be included in each of the multiple transform sets in the LFNST. According to aspects of the present disclosure, a transform set may be selected from multiple transform sets, and a non-separable transform matrix may be selected from one or more non-separable transform matrices in the transform set.

[0135] Table 1 shows an example mapping from intra-prediction modes to multiple transform sets according to one embodiment of the present disclosure. This mapping indicates the relationship between intra-prediction modes and multiple transform sets. The relationship as shown in Table 1 can be predefined and stored in the encoder and decoder. [Table 1]

[0136] Referring to Table 1, the multiple transform sets include four transform sets, for example, transform sets 0 to 3, represented by transform set indices (e.g., Tr.set index) from 0 to 3, respectively. The index (e.g., IntraPredMode) may indicate an intra-prediction mode, and the transform set index may be obtained based on the index and Table 1. Thus, the transform set may be determined based on the intra-prediction mode. In one example, when one of three cross-component linear model (CCLM) modes (e.g., INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the CB (e.g., 81≦IntraPredMode≦83), transform set 0 is selected for the CB.

[0137] As described above, each transform set may include one or more non-separable transform matrices. One of the one or more non-separable transform matrices may be selected by an explicitly signaled LFNST index. The LFNST index may be signaled in the bitstream once for each intra-coded CU (e.g., CB), for example, after signaling the transform coefficients. In one embodiment, each transform set includes two non-separable transform matrices (kernels), and the selected non-separable secondary transform candidate may be one of the two non-separable transform matrices. In some examples, LFNST is not applied to the CB (e.g., a CB coded in transform skip mode or the number of non-zero coefficients of the CB is less than a threshold). In one example, if LFNST is not applied to the CB, an LFNST index is not signaled for the CB. The default value of the LFNST index is zero and is not signaled, indicating that LFNST is not applied to the CB.

[0138] In one embodiment, LFNST is restricted to be applicable only when all coefficients outside the first coefficient subgroup are non-significant, and the coding of the LFNST index may be determined at the position of the last significant coefficient. The LFNST index may be context coded. In one example, the context coding of the LFNST index is independent of the intra prediction mode, and only the first bin is context coded. LFNST may be applied to intra-coded CUs in intra slices or inter slices, and may be applied to both the luma and chroma components. When dual trees are enabled, LFNST indexes for the luma and chroma components may be signaled separately. In the case of inter slices (e.g., when dual trees are disabled), a single LFNST index may be signaled and used for both the luma and chroma components.

[0139] An intra sub-partition (ISP) coding mode can be used. In the ISP coding mode, a luma intra prediction block can be divided into two or four sub-partitions vertically or horizontally, depending on the block size. In some examples, performance improvement reaches a limit when RST is applied to all feasible sub-partitions. Therefore, in some examples, when the ISP mode is selected, LFNST is disabled and the LFNST index (or RST index) is not signaled. Disabling RST or LFNST for the ISP predicted residual can reduce coding complexity. In some examples, when a matrix-based intra prediction mode (MIP) is selected, LFNST is disabled and the LFNST index is not signaled.

[0140] In some examples, due to a maximum transform size limitation (e.g., 64x64), CUs larger than 64x64 are implicitly split (TU tiling), and the LFNST index search can increase data buffering by a factor of four for a certain number of decoding pipeline stages. Therefore, the maximum size allowed for LFNST can be limited to 64x64. In one example, LFNST is enabled only for discrete cosine transform (DCT) type 2 (DCT-2) transforms.

[0141] In some examples, a separable transform scheme may not be efficient at capturing directional texture patterns (e.g., edges along 45° or 135° directions). A non-separable transform scheme can improve coding efficiency, for example, in the above cases. To reduce computational complexity and memory usage, a non-separable transform scheme can be used as a secondary transform applied to the low-frequency transform coefficients obtained from the primary transform.

[0142] In some implementations, for example, selecting a transform kernel to be used from the grouped transform kernels is based on prediction mode information, which can indicate a prediction mode.

[0143] In some examples, prediction mode information alone provides only a rough representation of the entire space of residual patterns observed in the prediction mode. Figures 17A-17D show that, according to embodiments, neighboring reconstructed samples can provide additional information for more efficiently representing the residual patterns. Thus, methods of transform set selection and / or transform kernel selection based on neighboring reconstructed samples in addition to prediction mode information are disclosed. For example, in addition to prediction mode information, feature indicators (e.g., feature vectors) of neighboring reconstructed samples may be used.

number

[0144] In this disclosure, the term block may refer to PB, CB, coded block, coding unit (CU), transform block (TB), transform unit (TU), luma block (e.g., luma CB), chroma block (e.g., chroma CB), etc.

[0145] In this disclosure, block size may refer to block width, block height, block aspect ratio (e.g., ratio of block width to block height, ratio of block height to block width), block area size or block area (e.g., block width x block height), minimum values ​​of block width and block height, maximum values ​​of block width and block height, etc.

[0146] In this disclosure, a transform kernel for a block can be used for a linear transform, a quadratic transform, a cubic transform, or a transform scheme beyond a cubic transform. The transform kernel can be used for a separable transform or a non-separable transform. The transform kernel can be used for luma blocks, chroma blocks, inter prediction, intra prediction, etc. The transform kernel may also be referred to as a transform core, a transform candidate, a transform kernel option, etc. In one example, the transform kernel is a transform matrix. The transformation of a block can be performed based on at least the transform kernel. Therefore, the methods of this disclosure can be applied to a linear transform, a quadratic transform, a cubic transform, any transform scheme beyond a cubic transform, a separable transform, a non-separable transform, a luma block, a chroma block, inter prediction, intra prediction, etc.

[0147] In this disclosure, a transform set may refer to a group of transform kernels or transform kernel options. A transform set may include one or more transform kernels or transform kernel options. In one example, a transform kernel for a block may be selected from a transform set.

[0148] According to aspects of the present disclosure, a transform kernel can be determined from a group of transform sets using neighboring reconstructed samples. Neighboring reconstructed samples of a block being reconstructed in a current picture can be used to determine a transform kernel for the block from the group of transform sets. The group of transform sets may be predetermined. In one example, the group of transform sets is pre-stored in the encoder and / or decoder.

[0149] Neighboring reconstructed samples of a block (e.g., a set of neighboring reconstructed samples) may refer to reconstructed samples from previously decoded neighboring blocks in the current picture (e.g., a group of reconstructed samples) or reconstructed samples in a previously decoded picture. The previously decoded neighboring blocks may include one or more neighboring blocks of this block, and neighboring reconstructed samples are also referred to as reconstructed samples in one or more neighboring blocks of this block. The one or more neighboring blocks of this block may be included in the current picture or a reconstructed picture different from the current picture (e.g., a reference picture). Neighboring reconstructed samples of this block may include spatially neighboring samples of the block in the current picture and / or temporally neighboring samples of a block in another picture different from the current picture (e.g., a previously decoded picture). Figure 18 shows exemplary spatially neighboring samples of a block according to an embodiment of the present disclosure. This block is the current block (1851) of the current picture. Samples 1-4 and A-X are spatially neighboring samples of the current block (1851) that have already been reconstructed. In one example, samples 1-4 and A-X are in one or more reconstructed neighboring blocks of the current block (1851). According to aspects of the present disclosure, one or more of samples 1-4 and A-X are used to determine a transform kernel for the current block (1851). In one example, the neighboring reconstructed samples of the current block (1851) include an upper-left neighboring sample 4, a left neighboring column (1852) including samples A-F, and an upper neighboring row (1854) including samples G-L, and are used to determine a transform kernel for the current block (1851). In one example, the neighboring reconstructed samples of the current block (1851) include an upper-left neighboring sample 1-4, a left neighboring column (1852) including samples A-F and M-R, and an upper neighboring row (1854) including samples G-L, and are used to determine a transform kernel for the current block (1851).

[0150] A group of transform sets may include one or more transform sets. Each of the one or more transform sets may include any suitable one or more transform kernels. Thus, a group of transform sets may include transform kernels used in a linear transform, a quadratic transform, a cubic transform, or a transform scheme beyond cubic. A group of transform sets may include transform kernels that are separable transforms and / or non-separable transforms. A group of transform sets may include transform kernels used for luma blocks, chroma blocks, inter-prediction, intra-prediction, etc.

[0151] According to aspects of the present disclosure, a transform kernel (or transform candidate) for a block of a current picture can be determined from a group of transform sets based on reconstructed samples in one or more neighboring blocks of the block of the current picture. Feature indicators (e.g., feature vectors) extracted from the reconstructed samples in one or more neighboring blocks of the block of the current picture can be used to determine the transform kernel (or transform candidate).

number

[0152] In one embodiment, feature indicators (e.g., feature vectors) of the adjacent reconstructed samples are extracted from the adjacent reconstructed samples.

number

number

[0153] In one embodiment, a transform kernel for a block (e.g., the current block (1851)) may be determined based on neighboring reconstructed samples (e.g., feature indicators) and prediction mode information of the block. The prediction mode information of the block may indicate information used to predict the block, such as inter-prediction or intra-prediction of the block. According to aspects of the present disclosure, the prediction mode information of the block may indicate a prediction mode of the block (e.g., intra-prediction mode, inter-prediction mode). In one example, the transform kernel for the block is determined based on neighboring reconstructed samples (e.g., feature indicators) and the prediction mode of the block.

[0154] In one example, the block is intra-coded or intra-predicted, and the prediction mode information is referred to as intra-prediction mode information. The prediction mode information (e.g., intra-prediction mode information) of a block may indicate the intra-prediction mode used for the block. The intra-prediction mode may refer to a prediction mode used for intra-prediction of the block, such as the directional mode (or directional prediction mode) described in FIG. 9, the omni-directional prediction mode (e.g., DC mode, PAETH mode, SMOOTH mode, SMOOTH_V mode, or SMOOTH_H mode) described in FIG. 10, or the recursive filtering mode described in FIG. 11. The intra-prediction mode may also refer to a prediction mode described in this disclosure, an appropriate variation of a prediction mode described in this disclosure, or an appropriate combination of a prediction mode described in this disclosure. For example, the intra-prediction mode may be combined with multi-line intra-prediction described in FIG. 12.

[0155] More specifically, for example, a subgroup of transform sets may be selected from a group of transform sets based on coded information of a block from a coded video bitstream. The coded information of the block may include prediction mode information indicating a prediction mode (e.g., an intra-prediction mode or an inter-prediction mode) of the block. In one embodiment, a subgroup of transform sets may be selected from a group of transform sets based on the prediction mode information (e.g., a prediction mode).

[0156] According to aspects of the present disclosure, a subgroup of transform sets may be selected from a group of transform sets based on a prediction mode of the block indicated in the coded information of the block. Furthermore, the transform candidates from the subgroup of transform sets may be determined based on reconstructed samples in one or more neighboring blocks of the block. For example, the transform candidates from the subgroup of transform sets may be determined based on feature indicators (e.g., feature vectors) extracted from the reconstructed samples in one or more neighboring blocks of the block.

number

[0157] After selecting a subgroup of transform sets from the group of transform sets, the transform kernel for the block may be further determined based on at least adjacent reconstructed samples (e.g., feature indicators) using any suitable method described below.

[0158] In one embodiment, neighboring reconstructed samples of the block (or reconstructed samples in one or more neighboring blocks) are used to identify or select a transform set of the sub-group of selected transform sets from the sub-group of selected transform sets. In one example, feature indicators (e.g., feature vectors) extracted from the reconstructed samples in one or more neighboring blocks are used to identify or select a transform set of the sub-group of selected transform sets.

number

[0159] In one embodiment, a transform set of the subgroup of selected transform sets is selected from the subgroup of selected transform sets based on a second index. The coded information may indicate (e.g., include) the second index. The second index may be signaled in the coded video bitstream. Furthermore, a transform kernel (or transform candidate) for a block from the selected transform set may be determined (e.g., selected) based on neighboring reconstructed samples of the block (or reconstructed samples in one or more neighboring blocks of the block) (e.g., feature indicators). In one example, a transform kernel (or transform candidate) in the selected transform set is determined (e.g., selected) based on feature indicators of the block (e.g., feature vectors).

number

[0160] In one embodiment, a transformation kernel (or candidate transformations) for a block is implicitly determined (or identified) from a subgroup of the selected transformation set based on neighboring reconstructed samples of the block. In one example, the transformation kernel (or candidate transformations) is determined based on feature indicators (e.g., feature vectors) of the block from the subgroup of the selected transformation set.

number

[0161] Block feature indicators (e.g., feature vectors)

number

number

number

number

number

[0162] In one embodiment, a single variable X may be used to indicate adjacent reconstructed samples, and the feature indicator of the block is a feature scalar S for the block. The variable X may indicate sample values ​​of adjacent reconstructed samples. In one example, the variable X is an array containing sample values ​​of adjacent reconstructed samples, reflecting the distribution of sample values ​​in the adjacent reconstructed samples. The feature scalar S may indicate statistical information of the sample values. The feature scalar S may include, but is not limited to, a scalar quantitative measure of the variable X obtained from adjacent reconstructed samples, such as the mean (or first moment) of the variable X, the variance (or second moment) of the variable X, or the skewness (or third moment) of the variable X. In one example, the variable X is referred to as a random variable X.

[0163] In one example, the feature indicator is a feature scalar S. The feature scalar S is determined as a moment (e.g., first moment, second moment, third moment, etc.) of the variable X that indicates the sample values ​​of the reconstructed samples in one or more neighboring blocks of the block.

[0164] An example of variable X is shown in Figure 18. In one example, variable X contains reconstructed neighboring samples adjacent to the current block (1851), including sample 4, the left neighboring column (1852), and the above neighboring row (1854). In one example, variable X contains reconstructed neighboring samples 1-4 and A-X.

[0165] In one embodiment, multiple variables (e.g., two variables Y and Z) can be used to separately indicate different sets of adjacent reconstruction samples (if available) (e.g., a first set corresponding to Y and a second set corresponding to Z), and the feature indicators of a block are represented by the feature vectors of this block.

number

number

[0166] Feature Vector

number

[0167] 18 shows an example of multiple variables Y and Z. In one example, variable Y includes the left adjacent column (1852) and variable Z includes the above adjacent row (1854). In one example, one of variables Y and Z includes sample 4, which is at the top left.

[0168] In one example, the feature indicators are the feature vectors

number

number

[0169] In one example, variable Y contains the left adjacent columns (1852)-(1853) and variable Z contains the above adjacent rows (1854)-(1855). In one example, one of variables Y and Z contains samples 1-4 in the upper left.

[0170] According to aspects of the present disclosure, when a feature indicator is used in determining a transform kernel for a block or a transform set including a transform kernel for a block, the transform kernel or transform set may be determined based on the feature indicator and a threshold. The threshold may be selected from a predefined threshold set. The transform kernel or transform set may be determined based on the feature indicator and a threshold using any appropriate method. Some example methods are described below. As described above, according to aspects of the present disclosure, a subgroup of transform sets is selected from a group of transform sets based on, for example, prediction mode information. Then, in one example, a transform set is identified from the selected subgroup of transform sets using the feature indicator for the block and a threshold. In another example, a transform set of the selected subgroup of transform sets is selected using an index (e.g., a second index) signaled in the coded video bitstream, and a transform kernel for the block from the selected transform set may be determined using the feature indicator for the block and a threshold. In another example, the transform kernel is implicitly identified using the feature indicator from the selected subgroup of transform sets and a threshold.

[0171] The threshold set (K s ) are predefined, e.g., for classification purposes. In one example, (i) the moments of a variable X and a predefined threshold set K s the coding information index may be used to determine a transformation set from the subgroup of transformation sets based on a threshold value selected from the coded information index, (ii) determine candidate transformations from the subgroup of transformation sets based on moments of the variable X and a threshold value, or (iii) select a transformation set from the subgroup of transformation sets based on an index of the coded information and determine candidate transformations from the selected transformation set based on moments of the variable X and a threshold value.

[0172] In one embodiment, the feature scalar S is used to identify a transformation kernel for a block, or to identify a transformation set that includes a transformation kernel for the block, for example, from a subgroup of transformation sets selected as described above. s may include one or more first thresholds (or first threshold values).

[0173] In one example, the feature scalar S is a quantitative measure of variable X, such as the mean, variance, or skewness of variable X. For example, the feature scalar S is a moment of variable X. The moment of a variable can be either the first moment (or mean) of variable X, the second moment (or variance) of variable X, or the third moment (such as skewness) of variable X.

[0174] Each prediction mode (e.g., intra prediction mode or inter prediction mode) may be, for example, a set of multiple prediction modes and multiple threshold sets K s A threshold set K, which shows an injective mapping between s A unique threshold subset K of s The prediction mode of the block can be one of a plurality of prediction modes. Each prediction mode corresponds to a threshold subset K s ' is an injective mapping. For example, the mapping between threshold sets K s a threshold subset K corresponding to the first prediction mode s1 ' and the threshold subset K corresponding to the second prediction mode. s2 ', and also includes the threshold subset K s1 ' contains the threshold subset K s2 There are no elements or thresholds that are identical to the elements or thresholds in '.

[0175] In one example, the feature scalar S is a quantitative measure of a variable X, such as its mean, variance, or skewness. Each prediction mode is based on a threshold set K s Any threshold subset K in s ' can be used to calculate the prediction mode and threshold set Ks the corresponding threshold subset K in s ' is a non-injective mapping. In one example, the multiple prediction modes (e.g., the first prediction mode and the second prediction mode) are mapped to a threshold set K s the same threshold subset K in s '. In one example, the threshold set K s is the threshold subset K corresponding to the first prediction mode. s3 ' and the threshold subset K corresponding to the second prediction mode s4 ' and the threshold subset K s3 The elements (or thresholds) in ' are the threshold subset K s4 ' is the same as an element (or threshold) in the threshold set K s is a single element set.

[0176] In an embodiment, the threshold set K s Threshold subset K in s The elements of ' are thresholds that depend on the quantization index (or the corresponding quantization step size associated with the quantization index). Typically, each quantization index can correspond to a unique quantization step size. The above description is based on the threshold set K s the corresponding threshold subset K in s ' applies when the mapping between ' and ' is non-injective or injective.

[0177] In one embodiment, a threshold set K corresponding to a particular quantization index (or corresponding quantization step size) is s Threshold subset K in s Only one element of ' is defined. s The remaining elements in ' (if any) may be derived using a mapping function. The mapping function may be linear or non-linear. The mapping function may be one that is predefined and used by the encoder and / or decoder.s The remaining elements in ' can correspond to different quantization indices or (or corresponding quantization step sizes). The above description is based on the s the corresponding threshold subset K in s ' applies when the mapping between ' and ' is non-injective or injective.

[0178] In one embodiment, a lookup table is utilized to determine the threshold subset K. s For example, the lookup table may determine the threshold subset K based on the prediction mode and / or quantization index. s ' is used to select an element for the threshold subset K. In one example, the lookup table includes a relationship between a prediction mode, a quantization index (or a corresponding quantization step size), and a threshold. The lookup table is traversed to select a threshold subset K using the prediction mode (e.g., intra prediction mode or inter prediction mode) and / or the quantization index (or a corresponding quantization step size). s ' (e.g., threshold subset K s The above description is based on the threshold set K s the corresponding threshold subset K in s ' applies when the mapping between ' and ' is non-injective or injective.

[0179] In an embodiment, the mapping from quantization indexes (or corresponding quantization step sizes) to thresholds is a linear mapping. Parameters used for the linear mapping, such as a slope and an intercept, may be predefined or derived using coded information of the block. The coded information may include, but is not limited to, a block size, a quantization index (or corresponding quantization step size), and a prediction mode (e.g., an intra-prediction mode or an inter-prediction mode). The linear mapping (e.g., parameters used for the linear mapping) may be predefined or derived based on one or a combination of the block size, the quantization index (or corresponding quantization step size), and the prediction mode. If the mapping from quantization indexes (or corresponding quantization step sizes) to thresholds is a non-linear mapping, this description can be adjusted appropriately. The above description applies to each prediction mode and a threshold set K s the corresponding threshold subset K in s ' applies when the mapping between ' and ' is non-injective or injective.

[0180] As mentioned above, the feature scalar S can be used to identify a transformation set including the transformation kernel for the block, or to identify the transformation kernel for the block, and the threshold set K s may include one or more first thresholds. In an embodiment, for example, the threshold set K s The selection of the threshold from K depends on the block size of the block. s The selection of thresholds from the threshold set K depends on the quantization parameters used in the quantization, such as the quantization index (or corresponding quantization step size). s The selection of the threshold from depends on the prediction mode (eg, intra prediction mode and / or inter prediction mode).

[0181] In one embodiment, the feature vector

number

number

number

[0182] In one example,

number

[0183] In one example, each prediction mode (e.g., an intra prediction mode or an inter prediction mode) can be

number

[0184] In an embodiment, the threshold set K v Threshold subset K in v The elements of ' are thresholds that depend on the quantization index (or the corresponding quantization step size associated with the quantization index).

number

[0185] In an embodiment, corresponding to a particular quantization index (or corresponding quantization step size),

number

[0186] In one embodiment, the lookup table is a threshold subset K v 'Select elements for threshold subset K v ' and classify vector subsets

number

number

number

number

number

[0187] In one embodiment,

number

[0188] In one example, (i) distance and a threshold set K v Threshold subset K contained in v ', a threshold value selected from {K v (ii) determining a set of transformations from a subgroup of transformation sets based on a comparison with a distance and a threshold (e.g., {K v (iii) determining candidate transformations from the subgroup of transformation sets based on a comparison with the distance and a threshold (e.g., {K v '}) based on a comparison of the selected set of transformations.

number

[0189] In one example, this distance is calculated by dividing the two vectors

number

[0190] In one embodiment, the comparison is made by determining whether (i) the distance is less than or equal to a threshold value;

number

[0191] In some embodiments, neighboring reconstructed samples, such as feature indicators of neighboring reconstructed samples, can be used to constrain the selection of transform candidates from a subgroup of transform sets. For some subgroups of transform sets selected using coded information such as the prediction mode of the block (e.g., intra-prediction mode, inter-prediction mode), the process of identifying transform candidates or transform kernels can include feature indicators of neighboring reconstructed samples (e.g., feature vectors

number

number

number

[0192] FIG. 19 shows a flowchart outlining a process (1900) according to one embodiment of the present disclosure. The process (1900) can be used to reconstruct blocks such as CB, CU, PB, TB, TU, luma blocks (e.g., luma CB or luma TB), and chroma blocks (e.g., chroma CB or chroma TB). In various embodiments, the process (1900) is performed by a processing circuit, such as a processing circuit within the terminal devices (310), (320), (330), and (340), a processing circuit performing the functions of the video encoder (403), a processing circuit performing the functions of the video decoder (410), a processing circuit performing the functions of the video decoder (510), or a processing circuit performing the functions of the video encoder (603). In some embodiments, the process (1900) is implemented by software instructions, such that the processing circuit performs the process (1900) when the processing circuit executes the software instructions. The process starts at (S1901) and proceeds to (S1910).

[0193] In (S1910), reconstructed samples in one or more neighboring blocks of the block of the current picture, for example, feature indicators (e.g., feature vectors) extracted from the reconstructed samples in one or more neighboring blocks of the block, are

number

[0194] In an embodiment, a subgroup of transform sets may be selected from the group of transform sets based on the prediction mode (e.g., intra prediction mode, inter prediction mode) of the block signaled in the coded information of the block. Feature indicators (e.g., feature vectors) extracted from reconstructed samples in one or more neighboring blocks of the block may be used.

number

[0195] In one example, feature indicators (e.g., feature vectors) extracted from reconstruction samples in one or more neighboring blocks of a block are used.

number

[0196] In one example, one transform set from the subgroup of transform sets may be selected based on a second index signaled in the coded information. The feature indicators (e.g., feature vectors) extracted from the reconstructed samples in one or more neighboring blocks of the block may be used.

number

[0197] In one example, feature indicators (e.g., feature vectors) extracted from reconstruction samples in one or more neighboring blocks of a block are used.

number

[0198] In one example, based on a statistical analysis of the reconstructed samples in one or more neighboring blocks of the block,

number

[0199] In one example, the feature indicators are feature scalars S and are determined as moments of variables indicative of the sample values ​​of the reconstructed samples in one or more neighboring blocks of the block. S is predefined. Therefore, (i) the moments of the variables and the threshold set K S the coded information, (ii) determining a transformation set from the subgroup of transformation sets based on a threshold selected from the coded information, (iii) selecting a transformation set from the subgroup of transformation sets based on a second index of the coded information and determining a transformation candidate from the selected transformation set based on the moments of the variables and a threshold.

[0200] In one example, the moment of the variable is one of a first moment of the variable, a second moment of the variable, or a third moment of the variable. The prediction mode of the block is one of a plurality of prediction modes, each of which is determined by a threshold set K. S Unique threshold subset K in S' corresponds to this unique threshold subset K S ' is a set of multiple prediction modes and a threshold set K S 1 shows an injective mapping between multiple threshold subsets in .

[0201] A threshold set K is determined based on one of (i) the block size of the block, (ii) the quantization parameter, or (iii) the prediction mode of the block. S Select a threshold from the

[0202] In one example, the feature indicators are

number

[0203] In (S1920), the samples of the block can be reconstructed based on the determined transformation candidates.

[0204] The process 1900 may be adjusted as appropriate. One or more steps in the process 1900 may be modified and / or omitted. One or more additional steps may be added. Any suitable order of implementation may be used.

[0205] The embodiments of the present disclosure can be used alone or in combination in any order. Furthermore, each of the method (or embodiment), encoder, and decoder can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium. The embodiments of the present disclosure can be applied to luma blocks or chroma blocks.

[0206] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 20 illustrates a computer system (2000) suitable for implementing certain embodiments of the disclosed subject matter.

[0207] Computer software can be coded using any suitable machine code or computer language that can be assembled, compiled, linked, or similar mechanisms to create code containing instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or by interpretation, microcode execution, etc.

[0208] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.

[0209] 20 for computer system (2000) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing embodiments of the present disclosure, nor should the arrangement of components be construed as having any dependency or requirement regarding any one or combination of components shown in the exemplary embodiment of computer system (1700).

[0210] The computer system (2000) may include certain human interface input devices that can respond to input by one or more users, for example, through tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input (not shown). The human interface devices may also be used to capture certain media that do not necessarily involve direct conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video, including stereoscopic video).

[0211] The input human interface devices may include one or more of a keyboard (2001), a mouse (2002), a trackpad (2003), a touchscreen (2010), a data glove (not shown), a joystick (2005), a microphone (2006), a scanner (2007), and a camera (2008) (only one of each type is shown).

[0212] The computer system (2000) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., a touchscreen (2010), data gloves (not shown), or a haptic feedback device that has haptic feedback via a joystick (2005) but does not function as an input device), audio output devices (speakers (2009), headphones (not shown), etc.), visual output devices (screens (2010) including CRT screens, LCD screens, plasma screens, and OLED screens (each with or without touchscreen input capability and each with or without haptic feedback capability, some of which may output two-dimensional visual output or three-dimensional or higher-dimensional output via means such as stereographic output), virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown), etc.), and printers (not shown).

[0213] The computer system (2000) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (2020) with media (2021) such as CD / DVD, thumb drives (2022), removable hard drives or solid state drives (2023), conventional magnetic media such as tape or floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices (not shown) such as security dongles, etc.

[0214] Those skilled in the art will understand that the term "computer-readable medium" as used herein in conjunction with the disclosed subject matter does not include transmission media, carrier waves, or other transitory signals.

[0215] The computer system (2000) may further include an interface (2054) to one or more communication networks (2055). Networks may be, for example, wireless, wired, or optical. Networks may also be local, wide-area, metropolitan, vehicular, industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet and WLAN; cellular networks including GSM, 3G, 4G, 5G, LTE, etc.; TV wired or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; and vehicular and industrial networks including CANBus. Certain networks generally require an external network interface adapter connected to a particular general-purpose data port or peripheral bus (2049) (e.g., a USB port on the computer system (2000)). Others are generally integrated into the core of the computer system (2000) by connecting to a system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (2000) can communicate with other entities. Such communication can be one-way receive only (e.g., broadcast TV), one-way transmit only (e.g., CANbus to a specific CANbus device), or two-way, e.g., transmit to other computer systems using local or wide-area digital networks. Specific protocols and protocol stacks can be used with each of these networks and network interfaces described above.

[0216] The aforementioned human interface devices, human-accessible storage devices, and network interfaces may be connected to the core (2040) of the computer system (2000).

[0217] The cores (2040) may include one or more central processing units (CPUs) (2041), graphics processing units (GPUs) (2042), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (2043), task-specific hardware accelerators (2044), graphics adapters (2050), etc. These devices, along with read-only memory (ROM) (2045), random access memory (2046), and internal mass storage devices (2047), such as non-user-accessible internal hard drives or SSDs, may be connected via a system bus (2048). In some computer systems, the system bus (2048) is accessible in the form of one or more physical plugs, allowing expansion with additional CPUs, GPUs, etc. Peripheral devices may be connected directly to the core's system bus (2048) or via a peripheral bus (2049). In one example, a display (2010) may be connected to the graphics adapter (2050). Peripheral bus architectures include PCI, USB, and the like.

[0218] The CPU (2041), GPU (2042), FPGA (2043), and accelerator (2044) can combine to execute specific instructions that may constitute the aforementioned computer code. That computer code may be stored in ROM (2045) or RAM (2046). Transient data may also be stored in RAM (1746), while persistent data may be stored in, for example, internal mass storage (2047). The use of cache memory, which may be closely associated with one or more CPUs (2041), GPUs (2042), mass storage (2047), ROM (2045), RAM (2046), etc., allows for fast storage and retrieval from any memory device.

[0219] The computer-readable medium may comprise computer code for performing various computer-implemented operations. The media and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.

[0220] By way of example and not limitation, a computer system (2000) having the architecture, and in particular the core (2040), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be user-accessible mass storage as described above, as well as media associated with the core's (2040) specific storage of a non-transitory nature, such as the core's internal mass storage (2047) or ROM (2045). Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (2040). The computer-readable media can include one or more memory devices or chips, as appropriate. The software can cause the core (2040), and in particular the processor (including a CPU, GPU, FPGA, etc.) therein, to perform specific processes or portions of specific processes described herein, including defining data structures stored in RAM (2046) and modifying such data structures in accordance with the software-defined processes. Additionally, or alternatively, a computer system may provide functionality as a result of logic hardwired or embedded in circuitry (e.g., accelerator (2044)) that can operate in place of or together with software to perform particular processes or portions of particular processes described herein. References to software may, where appropriate, include logic, and vice versa. References to computer-readable media may, where appropriate, include circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.

[0221] Appendix A: Acronyms JEM: Joint Exploration Model VVC: Versatile Video Coding BMS:Benchmark Set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplementary Enhancement Information VUI: Video Usability Information GOP: Groups of Pictures TU: Transform Units PU: Prediction Units CTU: Coding Tree Units CTB: Coding Tree Blocks PB: Prediction Blocks HRD: Hypothetical Reference Decoder SNR: Signal-to-Noise Ratio CPU: Central Processing Units GPU: Graphics Processing Units CRT:Cathode Ray Tube LCD: Liquid-Crystal Display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile Communications LTE:Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Array SSD: Solid-State Drive IC: Integrated Circuit CU: Coding Unit

[0222] While this disclosure has described several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of this disclosure. It should thus be understood that those skilled in the art will be able to devise various systems and methods that, although not explicitly or specifically described herein, embody the principles of the present disclosure and are included within its spirit and scope.

Claims

1. 1. A method for video encoding by an encoder, comprising: sampling blocks of the current picture into a video bitstream based on an intra-prediction mode or an inter-prediction mode, Feature vectors from samples in one or more neighboring blocks of the block of the current picture [Equation 1] or extracting a feature scalar S; Based on a statistical analysis of the samples in one or more neighboring blocks of the block, the feature vector [Equation 2] or determining one of the feature scalars S; selecting a sub-group of transform sets from a group of transform sets based on information indicating a prediction mode of the block, wherein each transform set in the group of transform sets includes one or more candidate transforms for the block, and one or more neighboring blocks are in the current picture or a picture different from the current picture; (i) the feature vector [Equation 3] or determining a transformation set from a subgroup of said transformation sets based on one of said feature scalars S; (ii) the feature vector [Equation 4] or determining the candidate transformations from a subgroup of the set of transformations based on one of the feature scalars S; or (iii) selecting the transform set from the subgroup of transform sets based on an index of information indicating a prediction mode of the block; and [Equation 5] or determining the candidate transformations from a selected set of transformations based on one of the feature scalars S; sampling the block based on the determined candidate transforms; and coding information indicating a prediction mode of the block into the video bitstream; method.

2. The feature vector [Equation 6] or one of the feature scalars S is the feature scalar S; The feature vector [Equation 7] or determining one of the feature scalars S further comprises determining the feature scalar S as a moment of a variable indicative of sample values ​​of the sample in one or more neighboring blocks of the block. The method of claim 1.

3. Threshold set K S is predefined, The execution (i) the moments of the variables and the threshold set K S determining the transformation set from a subgroup of the transformation sets based on a threshold from (ii) determining the candidate transformations from a subgroup of the set of transformations based on the moments of the variables and the threshold set; or (iii) selecting the transform set from the subgroup of transform sets based on an index of information indicating a prediction mode of the block; and determining the transform candidate from the selected transform set based on moments of the variables and the threshold value. The method of claim 2.

4. the moment of the variable is one of a first moment of the variable, a second moment of the variable, or a third moment of the variable; the prediction mode of the block is one of a plurality of prediction modes; Each of the plurality of prediction modes is S A unique threshold subset K in S ', and the unique threshold subset K S ' is a combination of the plurality of prediction modes and the threshold set K S Denotes an injective mapping between multiple threshold subsets in The method of claim 3.

5. The threshold set K is determined based on one of (i) the block size of the block, (ii) the quantization parameter, or (iii) the prediction mode of the block. S further comprising selecting the threshold value from The method of claim 3.

6. The feature vector [Equation 8] or one of the feature scalars S is the feature vector [Equation 9] and The feature vector [Equation 10] Or one of the feature scalars S is the feature vector [0011] further comprising the step of including in the block a covariance or second moment of a variable indicative of the sample values ​​of samples in adjacent columns to the left of the block and sample values ​​of samples in adjacent rows above the block, respectively. The method of claim 1.

7. Classification Vector Set [0012] , and the classification vector set [0013] A threshold set K associated with v is predefined, the co-variation of the variables and the classification vector set [0014] The classification vector subset contained in [Equation 15] and calculating a distance between the classification vector selected from The execution (i) the distance and the threshold set K v The threshold subset K included in v determining the set of transformations from a subgroup of the set of transformations based on a comparison with a threshold selected from (ii) determining the candidate transformations from a subgroup of the set of transformations based on a comparison of the distance to the threshold; or (iii) selecting the transform set from the subgroup of transform sets based on an index of information indicating a prediction mode of the block; and determining the transform candidate from the selected transform set based on a comparison between the distance and the threshold value. The method of claim 6.

8. 1. A video encoding device, comprising: Sampling blocks of a current picture into a video bitstream based on an intra-prediction mode or an inter-prediction mode, Feature vectors from samples in one or more neighboring blocks of the block of the current picture [0016] or extracting a feature scalar S; Based on a statistical analysis of the samples in one or more neighboring blocks of the block, the feature vector [Equation 17] or determining one of the feature scalars S; selecting a sub-group of transform sets from a group of transform sets based on information indicating a prediction mode of the block, wherein each transform set of the group of transform sets includes one or more candidate transforms for the block, and one or more neighboring blocks are in the current picture or a picture different from the current picture; (i) the feature vector [Equation 18] or determining a transformation set from a subgroup of said transformation sets based on one of said feature scalars S; (ii) the feature vector [Equation 19] or determining the candidate transformations from a subgroup of the set of transformations based on one of the feature scalars S; or (iii) selecting the transform set from the subgroup of transform sets based on an index of information indicating a prediction mode of the block; and [Equation 20] or determining the candidate transformations from a selected set of transformations based on one of the feature scalars S; sampling the block based on the determined candidate transform; and coding information indicating a prediction mode of the block into the video bitstream. a processing circuit configured to: Device.

9. The feature vector [Equation 21] or one of the feature scalars S is the feature scalar S; the processing circuitry is configured to determine the feature scalar S as a moment of a variable indicative of sample values ​​of the samples in one or more neighboring blocks of the block; 9. The apparatus of claim 8.

10. Threshold set K S is predefined, The processing circuitry (i) the moments of the variables and the threshold set K S determining the transformation set from a subgroup of the transformation sets based on a threshold from (ii) determining the candidate transformations from a subgroup of the set of transformations based on the moments of the variables and the threshold; or (iii) selecting the transform set from the subgroup of transform sets based on an index of information indicating a prediction mode of the block; and determining the candidate transform from the selected transform set based on moments of the variables and the threshold value.

10. The apparatus of claim 9.

11. The feature vector [Equation 22] Or one of the feature scalars S is the feature vector [Equation 23] and The processing circuitry The feature vector [0000] includes the covariance or second moment of a variable indicating the sample values ​​of samples in an adjacent column to the left of the block and the sample values ​​of samples in an adjacent row above the block, respectively.

9. The apparatus of claim 8.

12. Classification Vector Set [Equation 25] , and the classification vector set [Equation 26] A threshold set K associated with v is predefined, The processing circuitry the co-variation of the variables and the classification vector set [0000] The classification vector subset contained in [0000] Calculate the distance between the classification vector selected from (i) the distance and the threshold set K v The threshold subset K included in v determining the set of transformations from a subgroup of the set of transformations based on a comparison with a threshold selected from (ii) determining the candidate transformations from a subgroup of the set of transformations based on a comparison of the distance to the threshold; or (iii) selecting the transform set from the subgroup of transform sets based on an index of information indicating a prediction mode of the block; and determining the transform candidate from the selected transform set based on a comparison between the distance and the threshold.

12. The apparatus of claim 11.

13. the moment of the variable is one of a first moment of the variable, a second moment of the variable, or a third moment of the variable; the prediction mode of the block is one of a plurality of prediction modes; Each of the plurality of prediction modes is S A unique threshold subset K in S ', and the unique threshold subset K S ' is a combination of the plurality of prediction modes and the threshold set K S Denotes an injective mapping between multiple threshold subsets in 13. The apparatus of claim 12.

Citation Information

Patent Citations

  • Transform selection for video coding

    JP2019534624A

  • Video coding method and apparatus based on selective transformation

    JP2021507634A

  • Determining a set of candidate transforms for video coding.

    JP2021519546A

  • Transform selection for video coding

    US20180098081A1

  • Method for coding image on basis of selective transform and device therefor

    WO2019125035A1