Adaptive Multiple Transformation Set Selection
Adaptive multiple transform selection in video coding optimizes transform selection based on statistical analysis of coefficient blocks, addressing inefficiencies in existing technologies and enhancing compression efficiency.
Patent Information
- Application Number
- JP2023548344
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-09-30
- Filing Date
- 2022-10-04
- Publication Date
- 2025-08-04
- Estimated Expiration
- 2042-10-04
AI Technical Summary
Existing video coding technologies face inefficiencies in reducing redundancy and optimizing transform selection, particularly in intra and inter-picture prediction, leading to suboptimal compression ratios and increased bit usage for less likely prediction directions and motion vectors.
Implementing methods and apparatuses for adaptive multiple transform selection (MTS) in video encoding and decoding, utilizing threshold information and statistical analysis of transform coefficient blocks to determine optimal MTS candidate subsets, reducing redundancy and improving compression efficiency.
Enhances video compression efficiency by optimizing transform selection, reducing bit usage for less likely prediction directions and motion vectors, thereby improving compression ratios and reducing storage and bandwidth requirements.
Smart Images

Figure 0007717823000004 
Figure 0007717823000005 
Figure 0007717823000006
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of priority to U.S. Patent Application No. 17 / 957,959, filed on September 30, 2022, "ADAPTIVE MULTIPLE TRANSFORM SET SELECTION", which claims the benefit of priority to U.S. Provisional Application No. 63 / 255,365, filed on October 13, 2021, "ADAPTIVE MULTIPLE TRANSFORM SET SELECTION", and U.S. Provisional Application No. 63 / 289,110, filed on December 13, 2021, "METHOD AND APPARATUS FOR ADAPTIVE MULTIPLE TRANSFORM SET SELECTION", the entire disclosures of which are hereby incorporated by reference.
[0002] This disclosure generally describes embodiments related to video coding.
Background Art
[0003] The description of the background art provided herein is for the purpose of generally presenting the context of the present disclosure. The work of the inventors, to the extent that it is described in this background art section, and aspects of the description that may not be considered prior art at the time of filing and may not be considered prior art to the present disclosure, either expressly or implicitly, are not admitted as prior art to the present disclosure.
[0004] Uncompressed digital images and / or videos can include a series of pictures, each picture having, for example, spatial dimensions of 1920×1080 luminance samples and associated chrominance samples. The series of pictures can have a fixed or variable (informally also known as frame rate) picture rate of, for example, 60 pictures per second or 60 Hz. Uncompressed images and / or videos have specific bitrate requirements. For example, 1080p60 4:2:0 video at 8 bits per sample (1920×1080 luminance sample resolution at a frame rate of 60 Hz) requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires storage space exceeding 600 GB.
[0005] One purpose of the coding and decoding of images and / or videos can be the reduction of redundancy within the input image and / or video signal by compression. Compression can help reduce the aforementioned bandwidth and / or storage space requirements, in some cases by more than two orders of magnitude. The description herein uses video encoding / decoding as an example for illustration, but the same techniques can be applied to image encoding / decoding in a similar manner without departing from the spirit of the present disclosure. Both reversible compression and irreversible compression, as well as combinations thereof, can be employed. Reversible compression refers to techniques that can restore an exact copy of the original signal from the compressed original signal. When using irreversible compression, the restored signal may not be identical to the original signal, but the distortion between the original signal and the restored signal is small enough to make the restored signal useful for the intended application. In the case of video, irreversible compression is widely adopted. The amount of allowable distortion depends on the application. For example, users of certain consumer streaming applications can tolerate higher distortion than users of television distribution applications. The achievable compression ratio can reflect that the higher the allowable / tolerable distortion, the higher the compression ratio can be.
[0006] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transform processing, quantization, and entropy coding.
[0007] Video codec technology can include techniques known as intra coding. In intra coding, sample values are represented without reference to samples from previously reconstructed reference pictures or other data. In some video codecs, a picture is spatially subdivided into blocks of samples. When all blocks of samples are coded in an intra mode, that picture can be an intra picture. Intra pictures, and their derivatives such as independent decoder refresh pictures, can be used to reset the decoder state and thus can be used as the first picture in a coded video bitstream and video session or as a still image. Samples of an intra block can be subject to a transform, and the transform coefficients can be quantized prior to entropy coding. Intra prediction can be a technique to minimize sample values in a pre-transform region. In some cases, the smaller the post-transform DC value and the smaller the AC coefficients, the fewer bits are required to represent the block at a given quantization step size after entropy coding.
[0008] For example, conventional intra coding used in MPEG-2 production coding technology does not use intra prediction. However, some newer video compression technologies include techniques that attempt to perform prediction based on, for example, surrounding sample data and / or metadata obtained during encoding / decoding of a block of data. Such techniques are hereinafter referred to as "intra prediction" techniques. It should be noted that in at least some cases, intra prediction uses only reference data from the current picture being reconstructed and not from reference pictures.
[0009] There can be many different forms of intra prediction. When two or more of such techniques can be used in a given video coding technology, the specific technique in use can be coded as a specific intra prediction mode that uses the specific technique. In certain cases, the intra prediction mode can have sub-modes and / or parameters, and the sub-modes and / or parameters can be coded individually or can be included in a mode codeword that defines the prediction mode being used. Which codeword to use for a given combination of mode, sub-mode, and / or parameter can affect the coding efficiency improvement by intra prediction, and thus can affect the entropy coding technology used to convert the codeword into the bitstream.
[0010] A specific mode of intra prediction was introduced from H.264, improved in H.265, and further improved in more recent coding technologies such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). The predictor block can be formed using adjacent sample values of already available samples. The sample values of the adjacent samples are copied into the predictor block according to the direction. The reference to the direction in use can be coded within the bitstream or can itself be predicted.
[0011] Referring to FIG. 1A, depicted at the lower right is a subset of nine predictor directions known from 33 possible predictor directions (corresponding to 33 of the 35 intra modes defined in H.265). The point (101) where the arrows converge represents the sample being predicted. The arrows represent the direction from which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples that are at a 45-degree angle from the horizontal and above the right of sample (101). Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples that are at a 22.5-degree angle from the horizontal and below the left of sample (101).
[0012] Referring further to FIG. 1A, at the upper left, a square block (104) of 4×4 samples (indicated by the thick dashed line) is depicted. The square block (104) contains 16 samples, each labeled with an "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample from the top in the Y dimension and the first sample from the left in the X dimension. Similarly, sample S44 is the fourth sample in both the Y and X dimensions within the block (104). Since the block is of size 4×4 samples, S44 is at the lower right. Further reference samples are shown following a similar numbering scheme. The reference samples are labeled with an "R" relative to the block (104), its Y position (e.g., row index), and its X position (column index). In both H.264 and H.265, since the predicted samples are adjacent to the block being reconstructed, negative values need not be used.
[0013] Intra-picture prediction can function by copying the reference sample values from adjacent samples indicated by the signaled prediction direction. For example, assume that the coded video bitstream contains signaling indicating a prediction direction that matches the arrow (102) for this block, i.e., the samples are predicted from samples at a 45-degree angle from the horizontal and towards the upper right. In that case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.
[0014] In certain cases, in order to calculate the reference sample, especially when the direction is not evenly divisible by 45 degrees, the values of multiple reference samples may be combined, for example, by interpolation.
[0015] The number of possible directions has been increasing as video coding technology develops. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013). Currently, JEM / VVC / BMS can support up to 65 directions. Experiments have been conducted to identify the most likely directions, and specific techniques of entropy coding are used to represent those likely directions with a small number of bits, accepting a certain penalty for less likely directions. Further, the directions themselves can sometimes be predicted from the neighboring directions used in adjacent already-decoded blocks.
[0016] FIG. 1B shows a schematic diagram (110) depicting 65 intra prediction directions by JEM to show an increasing number of prediction directions over time.
[0017] The mapping of intra prediction direction bits representing directions within the coded video bitstream can vary for each video coding technology. Such mappings can range from a simple direct mapping to complex adaptive schemes including codewords, the most likely modes, and similar techniques. However, in most cases, there may be specific directions that are statistically less likely to occur within the video content than certain other directions. Since the purpose of video compression is to reduce redundancy, those less likely directions are represented by a larger number of bits than the likely directions in well-functioning video coding technologies.
[0018] The coding and decoding of images and / or videos can be performed using inter-picture prediction with motion compensation. Motion compensation can be an irreversible compression technique, and a block of sample data from a previously reconstructed picture or a part thereof (reference picture) is spatially shifted in the direction indicated by a motion vector (hereinafter, MV) and then used for the prediction of a newly reconstructed picture or a part of the picture. In some cases, the reference picture can be the same as the picture currently being reconstructed. The MV can have two dimensions X and Y, or three dimensions, where the third dimension is an indication of the reference picture in use (the latter can be indirectly the time dimension).
[0019] In some video compression techniques, the MV applicable to a particular region of sample data can be predicted from other MVs, for example, from an MV associated with another region of sample data that is spatially adjacent to the region being reconstructed and that precedes that MV in the decoding order. By doing so, the amount of data required for coding the MV can be significantly reduced, thereby eliminating redundancy and increasing the compression ratio. For example, when coding an input video signal derived from a camera (known as natural video), there is a statistical likelihood that regions larger than the region to which a single MV is applicable move in a similar direction, so MV prediction can function effectively and thus, in some cases, can use a similar motion vector derived from the MVs of adjacent regions for prediction. As a result, the MV detected for a given region is similar or identical to the MV predicted from surrounding MVs and can be represented with fewer bits than the number of bits used when directly coding the MV after entropy coding. In some cases, MV prediction can be an example of reversible compression of a signal (i.e., the MV) derived from the original signal (i.e., the sample stream). In other cases, MV prediction itself can be irreversible, for example, due to rounding errors when calculating the predictor from several surrounding MVs.
[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). Among the many MV prediction mechanisms provided by H.265, the one described with reference to FIG. 2 is a technique hereinafter referred to as "spatial merge".
[0021] Referring to FIG. 2, the current block (201) includes samples that have been found to be predictable by the encoder from a previous block of the same size that has been spatially shifted during the motion search process. Instead of directly coding the MV, the MV can be derived from metadata associated with one or more reference pictures using the MV associated with any one of five surrounding samples denoted as A0, A1, and B0, B1, B2 (202 to 206 respectively), for example, from the latest reference picture (in decoding order). In H.265, MV prediction can use predictors from the same reference picture that adjacent blocks are using. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0022] Aspects of the present disclosure provide methods and apparatuses for video encoding and decoding. In some examples, an apparatus for video decoding includes a processing circuit. The processing circuit is configured to determine multiple transform selection (MTS) selection information for a plurality of transform coefficient blocks in a coded video bitstream. The MTS selection information indicates (i) threshold information or (ii) at least one of a plurality of MTS candidate subsets for a transform coefficient block among the plurality of transform coefficient blocks. The processing circuit determines, based on the MTS selection information, which MTS candidate subset among the plurality of MTS candidate subsets is selected for the transform coefficient block, and can inverse-transform the transform coefficient block based on the MTS candidates included in the determined MTS candidate subset.
[0023] In one embodiment, the MTS selection information is signaled within the coded video bitstream.
[0024] In one example, the MTS selection information indicates threshold information including at least one threshold. The processing circuit can determine an MTS candidate subset based on at least one threshold and one of (i) the number of non-zero coefficients within a transform coefficient block or (ii) the position of the last significant coefficient in the scan order within the transform coefficient block.
[0025] In one example, the number of multiple MTS candidate subsets is the sum of the number of at least one threshold and 1. The processing circuit can determine multiple MTS candidate subsets based on the number of multiple MTS candidate subsets. The processing circuit can further determine an MTS candidate subset based on multiple MTS candidate subsets.
[0026] In one example, the MTS selection information includes one or more numbers. Each of the one or more numbers can be the number of one or more MTS candidates in each of one or more of the multiple MTS candidate subsets. The processing circuit can determine multiple MTS candidate subsets based on the one or more numbers and the MTS candidate set, and can determine an MTS candidate subset based on multiple MTS candidate subsets.
[0027] In one example, the multiple MTS candidate subsets include a last MTS candidate subset not included in one or more MTS candidate subsets, and the last MTS candidate subset is the MTS candidate set.
[0028] In one example, the multiple MTS candidate subsets include a first MTS candidate subset not included in one or more MTS candidate subsets, and the first MTS candidate subset is composed of default MTS candidates within the MTS candidate set.
[0029] In one example, a high-level syntax header is a slice header, a picture header, a picture parameter set (PPS), a video parameter set (VPS), an adaptive parameter set (APS), or a sequence parameter set (SPS).
[0030] In one example, the MTS selection information is determined based on a plurality of previously decoded transform coefficient blocks within a previously decoded region.
[0031] In one example, the MTS selection information indicates threshold information including at least one threshold of a plurality of transform coefficient blocks. The processing circuit can determine at least one threshold based on the coefficient information of a plurality of previously decoded transform coefficient blocks. The processing circuit can determine an MTS candidate subset based on at least one threshold and one of (i) the number of non-zero coefficients within the transform coefficient block or (ii) the position of the last significant coefficient in the scan order within the transform coefficient block.
[0032] In one example, the coefficient information of a plurality of previously decoded transform coefficient blocks indicates (i) the average number of non-zero coefficients of a plurality of previously decoded transform coefficient blocks, or (ii) the average position of the last significant coefficient in the scan order of a plurality of previously decoded transform coefficient blocks.
[0033] In one example, the MTS selection information indicates threshold information including at least one threshold of a plurality of transform coefficient blocks. The plurality of coefficient information is associated with a plurality of previously decoded transform coefficient blocks. Each of the plurality of coefficient information corresponds to each type of a plurality of types of block sizes within the previously decoded region. The processing circuit can determine at least one threshold based on the coefficient information corresponding to each type of block size to which the transform coefficient block belongs. The processing circuit can determine an MTS candidate subset based on at least one threshold and one of (i) the number of non-zero coefficients within the transform coefficient block or (ii) the position of the last significant coefficient in the scan order within the transform coefficient block.
[0034] In one example, the MTS selection information indicates a plurality of MTS candidate subsets. The processing circuit can determine, for a plurality of transform coefficient blocks, the MTS candidates and the order of the MTS candidates to be used to form a plurality of MTS candidate subsets from the MTS candidate set based on the statistical information of the transform types of the plurality of previously decoded transform coefficient blocks. The processing circuit can determine a plurality of MTS candidate subsets from the MTS candidate set based on the MTS candidates and the order of the MTS candidates, and can determine the MTS candidate subset to be one of the plurality of MTS candidate subsets.
[0035] In one example, the MTS selection information indicates a plurality of MTS candidate subsets. The plurality of statistical information of the transform types is associated with a plurality of previously decoded transform coefficient blocks. Each of the plurality of statistical information of the transform types corresponds to each type of a plurality of types of block sizes within the previously decoded region. The processing circuit can determine, for a transform coefficient block, the MTS candidates and the order of the MTS candidates to be used to form a plurality of MTS candidate subsets from the MTS candidate set based on one statistical information of the transform type corresponding to the type of block size to which the transform coefficient block belongs. The processing circuit can determine a plurality of MTS candidate subsets from the MTS candidate set based on the MTS candidates and the order of the MTS candidates, and can determine the MTS candidate subset to be one of the plurality of MTS candidate subsets.
[0036] Aspects of the present disclosure also provide a non-transitory computer-readable storage medium storing a program executable by at least one processor to perform a method for video decoding.
[0037] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0038]
Figure 1A
Figure 1B
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
DETAILED DESCRIPTION OF THE INVENTION
[0039] FIG. 3 shows an exemplary block diagram of a communication system (300). The communication system (300) includes, for example, a plurality of terminal devices that can communicate with each other via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via a network (350). In the example of FIG. 3, the first pair of terminal devices (310) and (320) perform unidirectional transmission of data. For example, the terminal device (310) can encode video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The encoded video data can be transmitted in the form of one or more encoded video bitstreams. The terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to restore video pictures, and display the video pictures according to the restored video data. Unidirectional data transmission can be common in media serving applications and the like.
[0040] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data, for example, during a video conference. In the case of bidirectional data transmission, in one example, each of the terminal devices (330) and (340) can encode video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) via the network (350). Each of the terminal devices (330) and (340) can also receive the encoded video data transmitted by the other of the terminal devices (330) and (340), can decode the encoded video data to restore the video picture, and can display the video picture on a display device accessible according to the restored video data.
[0041] In the example of FIG. 3, the terminal devices (310), (320), (330), and (340) are shown as a server, a personal computer, and a smartphone, but the principles of the present disclosure may not be so limited. Embodiments of the present disclosure find applications using laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (350) represents any number of networks that transmit encoded video data between the terminal devices (310), (320), (330), and (340), including, for example, wire (wired) and / or wireless communication networks. The communication network (350) can exchange data over circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this description, the architecture and topology of the network (350) may not be important for the operation of the present disclosure, unless otherwise described in the following specification.
[0042] FIG. 4 shows a video encoder and a video decoder in a streaming environment as an example of an application for the disclosed subject matter. The disclosed subject matter may be equally applicable to other video-enabled applications including, for example, storage of compressed video on digital media including video conferencing, digital television, CDs, DVDs, memory sticks, and the like.
[0043] A streaming system may include, for example, a video source (401) that creates a stream (402) of uncompressed video pictures, and a capture subsystem (413) that may include, for example, a digital camera. In one example, the stream (402) of video pictures includes samples taken by a digital camera. The stream (402) of video pictures depicted as a thick line to emphasize a larger data volume compared to the encoded video data (404) (or encoded video bitstream) can be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (404) (or encoded video bitstream) depicted as a thin line to emphasize a smaller data volume compared to the stream (402) of video pictures can be stored in a streaming server (405) for future use. One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include, for example, a video decoder (410) within an electronic device (430). The video decoder (410) decodes an input copy (407) of the encoded video data and creates an output stream (411) of video pictures that can be rendered on a display (412) (e.g., a display screen) or another rendering device (not depicted). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., video bitstreams) can be encoded according to specific video coding / compression standards. Examples of those standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC).The disclosed subject matter may be used in the context of VVC.
[0044] It should be noted that the electronic devices (420) and (430) can include other components (not shown). For example, the electronic device (420) can include a video decoder (not shown), and the electronic device (430) can also include a video encoder (not shown).
[0045] FIG. 5 shows an exemplary block diagram of a video decoder (510). The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., a receiving circuit). The video decoder (510) can be used instead of the video decoder (410) in the example of FIG. 4.
[0046] The receiver (531) can receive one or more encoded video sequences decoded by the video decoder (510). In one embodiment, one encoded video sequence is received at a time, and the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. The encoded video sequence can be received from a channel (501), which can be a hardware / software link to a storage device storing the encoded video data. The receiver (531) can receive the encoded video data along with other data, such as encoded audio data and / or auxiliary data streams, that can be transferred to their respective usage entities (not depicted). The receiver (531) can separate the encoded video sequence from the other data. To counter network jitter, a buffer memory (515) can be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter, “parser (520)”). In certain applications, the buffer memory (515) is part of the video decoder (510). In other applications, it can be outside the video decoder (510) (not depicted). In still other applications, for example, to counter network jitter, there is a buffer memory (not depicted) outside the video decoder (510), and in addition, for example, to process playout timing, there can be another buffer memory (515) inside the video decoder (510). When the receiver (531) is receiving data from a store-and-forward device with sufficient bandwidth and controllability, or from an isochronous network, the buffer memory (515) may not be needed or may have a low probability of being needed. For use in a best-effort packet network such as the Internet, the buffer memory (515) may be needed, may have a relatively high probability of being needed, may advantageously be of an adaptable size, and may be at least partially implemented in an operating system or similar elements outside the video decoder (510) (not depicted).
[0047] The video decoder (510) may include a parser (520) to recover symbols (521) from the coded video sequence. The categories of these symbols include information used to manage the operation of the video decoder (510), and potentially information for controlling rendering devices such as a rendering device (512) (e.g., a display screen) that can be coupled to the electronic device (530) but is not an essential part of the electronic device (530) as shown in FIG. 5. The control information for the rendering device may be in the form of supplementary enhancement information (SEI message) or a parameter set fragment (not depicted) of video user utility information (VUI). The parser (520) can perform syntax analysis / entropy decoding on the received coded video sequence. The coding of the coded video sequence can follow a video coding technology or standard and can follow various principles including variable length coding, Huffman coding, arithmetic coding, etc., with or without context sensitivity. The parser (520) can extract a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroups can include picture groups (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (520) can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the coded video sequence.
[0048] The parser (520) can perform an entropy decoding / syntax analysis operation on the video sequence received from the buffer memory (515) to create symbols (521).
[0049] The restoration of symbol (521) can involve multiple different units depending on the type of the coded video picture or a part thereof (such as inter-picture and intra-picture, inter-block and intra-block), and other factors. How each unit is involved can be controlled by subgroup control information parsed from the coded video sequence by the parser (520). Such a flow of subgroup control information between the parser (520) and the following multiple units is not depicted for clarity.
[0050] In addition to the function blocks already described, the video decoder (510) can be conceptually subdivided into several functional units as described below. In an actual implementation operating under commercial constraints, many of these units interact closely with each other and can at least partially be integrated with each other. However, for the purpose of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.
[0051]
[0052] The first unit is the scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives, as symbol (521), quantization transform coefficients and control information including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc. from the parser (520). The scaler / inverse transform unit (551) can output a block including sample values that can be input to the aggregator (555).In some cases, the output samples of the scaler / inverse transform (551) may be related to the intra-coded block. An intra-coded block is a block that does not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information can be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) uses the surrounding already reconstructed information fetched from the current picture buffer (558) to generate a block of the same size and shape as the block being reconstructed. The current picture buffer (558) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. The aggregator (555) may, in some cases, add, for each sample, the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).
[0053] In other cases, the output samples of the scaler / inverse transform unit (551) may be related to inter-coded and potentially motion-compensated blocks. In such cases, the motion compensation prediction unit (553) can access the reference picture memory (557) to fetch the samples used for prediction. After motion-compensating the samples fetched according to the symbol (521) related to the block, these samples can be added to the output of the scaler / inverse transform unit (551) by the aggregator (555) to generate output sample information (in this case, called residual samples or residual signals). The address in the reference picture memory (557) from which the motion compensation prediction unit (553) fetches the prediction samples can be controlled, for example, by the motion vectors available to the motion compensation prediction unit (553) in the form of a symbol (521) that can have X, Y, and reference picture components. Motion compensation can also include interpolation of the sample values fetched from the reference picture memory (557) when exact sub-sample motion vectors are used, a motion vector prediction mechanism, and the like.
[0054] The output samples of the aggregator (555) can undergo various loop filtering techniques in the loop filter unit (556). The video compression technology is controlled by the parameters included in the coded video sequence (also called the coded video bitstream) and can include in-loop filter techniques made available to the loop filter unit (556) as symbols (521) from the parser (520). Video compression can also respond not only to the meta information obtained during the decoding of the previous part (in decoding order) of the coded picture or coded video sequence, but also to the previously restored and loop-filtered sample values.
[0055] The output of the loop filter unit (556) can be a sample stream that is not only output to the rendering device (512) but also stored in the reference picture memory (557) for use in future inter-picture prediction.
[0056] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, when the coded picture corresponding to the current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and the unused current picture buffer can be reallocated before starting the reconstruction of the next coded picture.
[0057] The video decoder (510) can perform decoding operations according to a predetermined video compression technology or standard such as ITU-T Rec. H.265. In the sense that the coded video sequence conforms to both the syntax of the video compression technology or standard and the profile documented in the video compression technology, the coded video sequence can conform to the syntax specified by the used video compression technology or standard. Specifically, the profile can select some tools from all the tools available in the video compression technology or standard as the only tools available for use under that profile. Also, what is required for compliance can be that the complexity of the coded video sequence is within the range defined by the level of the video compression technology or standard. In some cases, the level restricts, for example, the maximum picture size, the maximum frame rate, the maximum reconstruction sample rate (measured in megasamples per second), the maximum reference picture size, etc. The restrictions set by the level can be further restricted in some cases by the specifications of the hypothetical reference decoder (HRD) and the metadata for HRD buffer management signaled within the coded video sequence.
[0058] In one embodiment, the receiver (531) can receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, and the like.
[0059] FIG. 6 shows an exemplary block diagram of a video encoder (603). The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used in place of the video encoder (403) of the example of FIG. 4.
[0060] The video encoder (603) can receive video samples from a video source (601) that can capture the video images to be encoded by the video encoder (603) (not part of the electronic device (620) in the example of FIG. 6). In another example, the video source (601) is part of the electronic device (620).
[0061] The video source (601) can provide a source video sequence to be encoded by a video encoder (603) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, …), any color space (e.g., BT.601 Y CrCB, RGB, …), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (601) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that convey motion when viewed in sequence. Each picture itself may be organized as a spatial array of pixels, and each pixel can contain one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.
[0062] According to one embodiment, the video encoder (603) can encode and compress the pictures of the source video sequence into an encoded video sequence (643) in real time or under any other required time constraints. Implementing an appropriate coding rate is one function of the controller (650). In some embodiments, the controller (650) controls and is functionally coupled to other functional units described below. For clarity, the couplings are not depicted. The parameters set by the controller (650) can include rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization techniques, …), picture size, layout of picture groups (GOP), maximum motion vector search range, etc. The controller (650) can be configured to have other appropriate functions related to the video encoder (603) optimized for a particular system design.
[0063] In some embodiments, the video encoder (603) is configured to operate in a coding loop. As a grossly simplified explanation, in one example, the coding loop can include a source coder (630) (involved in creating symbols such as a symbol stream, based on, for example, an input picture to be coded and reference pictures), and a (local) decoder (633) incorporated in the video encoder (603). The decoder (633) restores the symbols to create sample data in the same manner as also done by a (remote) decoder. The restored sample stream (sample data) is input into the reference picture memory (634). Since the decoding of the symbol stream leads to bit-exact results regardless of the location of the decoder (local or remote), the content in the reference picture memory (634) is also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the same sample values as the reference picture samples that the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained, for example, due to channel errors) is also used in some related arts.
[0064] The operation of the "local" decoder (633) can be the same as that of a "remote" decoder such as the video decoder (510), which has already been described in detail above with reference to FIG. 5. However, again referring briefly to FIG. 5, since the symbols are available and the encoding / decoding of the symbols into the coded video sequence by the entropy coder (645) and the parser (520) can be reversible, the entropy decoding part of the video decoder (510) including the buffer memory (515), and the parser (520) may not be fully implemented in the local decoder (633).
[0065] In one embodiment, decoder techniques other than syntax analysis / entropy decoding that exist within the decoder exist in the same or substantially the same functional form within the corresponding encoder. Accordingly, the disclosed subject matter focuses on the operation of the decoder. The description of encoder techniques can be omitted since it is the reverse of the decoder techniques described comprehensively. In certain areas, more detailed descriptions are provided below.
[0066] During operation, in some examples, the source coder (630) can perform motion-compensated predictive coding that predictively codes an input picture by referring to one or more previously coded pictures from a video sequence designated as a "reference picture". In this way, the coding engine (632) codes the difference between a pixel block of the input picture and a pixel block of a reference picture that can be selected as a predictive reference to the input picture.
[0067] The local video decoder (633) can decode the coded video data of a picture that can be designated as a reference picture based on the symbols created by the source coder (630). The operation of the coding engine (632) may advantageously be a non-invertible process. When the coded video data can be decoded by a video decoder (not shown in FIG. 6), the restored video sequence can typically be a replica of the source video sequence with some errors. The local video decoder (633) can replicate the decoding process that can be performed by the video decoder for the reference picture and cause the restored reference picture to be stored in the reference picture memory (634). In this way, the video encoder (603) can locally store a copy of the restored reference picture that has common content as the restored reference picture obtained by a remote video decoder (in the absence of transmission errors).
[0068] The predictor (635) can perform a predictive search for the coding engine (632). That is, in the case of a new picture to be coded, the predictor (635) can search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors and block shapes that can serve as appropriate predictive references for the new picture. The predictor (635) can operate on the sample blocks for each pixel block to find an appropriate predictive reference. Optionally, the input picture can have a predictive reference drawn from a plurality of reference pictures stored in the reference picture memory (634) as determined by the search results obtained by the predictor (635).
[0069] The controller (650) can manage the coding operations of the source coder (630), including, for example, setting the parameters and subgroup parameters used to encode video data.
[0070] The outputs of all the foregoing functional units can undergo entropy coding within the entropy coder (645). The entropy coder (645) converts the symbols generated by the various functional units into a coded video sequence by applying reversible compression to the symbols according to techniques such as Huffman coding, variable-length coding, arithmetic coding, etc.
[0071] The transmitter (640) can buffer the coded video sequence created by the entropy coder (645) to prepare for transmission via the communication channel (660), which may be a hardware / software link to a storage device storing the encoded video data. The transmitter (640) can merge the encoded video data from the video coder (603) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).
[0072] The controller (650) can manage the operation of the video encoder (603). During coding, the controller (650) can assign a specific coded picture type to each coded picture, which may affect the coding technique applicable to each picture. For example, a picture may often be assigned as one of the following picture types.
[0073] An intra picture (I picture) can be a picture that can be coded and decoded without using any other picture in the sequence as a source of prediction. Some video codecs enable various types of intra pictures, including, for example, independent decoder refresh (「IDR」) pictures. Those skilled in the art are aware of those variants of I pictures, as well as their respective uses and characteristics.
[0074] A predicted picture (P picture) can be a picture that can be coded and decoded using intra prediction or inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.
[0075] A bi - directionally predicted picture (B picture) can be a picture that can be coded and decoded using intra prediction or inter prediction that uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predicted pictures can use three or more reference pictures and associated metadata for the restoration of a single block.
[0076] The source picture is typically spatially subdivided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and may be coded block by block. The blocks may be coded predictively by referring to other (already coded) blocks, as determined by the coding assignment applied to each picture of the block. For example, blocks of an I picture may be coded non-predictively or they may be coded predictively by referring to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture may be coded predictively via spatial prediction or via temporal prediction by referring to one previously coded reference picture. Blocks of a B picture may be coded predictively via spatial prediction or via temporal prediction by referring to one or two previously coded reference pictures.
[0077] The video encoder (603) can perform coding operations according to a predetermined video coding technology or standard such as ITU-T Rec. H.265. In its operation, the video encoder (603) can perform various compression operations including predictive coding operations that utilize the temporal and spatial redundancies in the input video sequence. Thus, the coded video data can conform to the syntax specified by the video coding technology or standard being used.
[0078] In one embodiment, the transmitter (640) can transmit additional data along with the encoded video. The source coder (630) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, and the like.
[0079] Video may be captured as a plurality of source pictures (video pictures) in a time series. Intra-picture prediction (often abbreviated as intra prediction) utilizes the spatial correlation within a given picture, while inter-picture prediction utilizes the (temporal or other) correlation between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. When a block within the current picture is similar to a reference block within a reference picture that has been previously encoded and is still buffered within the video, the block within the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block within the reference picture and can have a third dimension to identify the reference picture when multiple reference pictures are being used.
[0080] In some embodiments, a dual prediction technique can be used in inter-picture prediction. According to the dual prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture, both of which are before the current picture in the decoding order (but can be in the past and future in the display order respectively) within the video. A block within the current picture can be encoded by a first motion vector pointing to a first reference block within the first reference picture and a second motion vector pointing to a second reference block within the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.
[0081] Furthermore, to improve coding efficiency, a merge mode technique can be used in inter-picture prediction.
[0082] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard, pictures in a sequence of video pictures are divided into coding tree units (CTUs) for compression, and CTUs within a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU contains three coding tree blocks (CTBs), which are one luminance CTB and two chrominance CTBs. Each CTU can be recursively quadtree-divided into one or more coding units (CUs). For example, a 64×64 pixel CTU can be divided into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine a prediction type of the CU, such as an inter prediction type or an intra prediction type. The CU is divided into one or more prediction units (PUs) according to the temporal and / or spatial predictability. Generally, each PU contains one luminance prediction block (PB) and two chrominance PBs. In one embodiment, the prediction operation in coding (encoding / decoding) is performed in units of prediction blocks. Using a luminance prediction block as an example of a prediction block, the prediction block contains a matrix of pixel values (e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels.
[0083] FIG. 7 shows an exemplary diagram of a video encoder (703). The video encoder (703) is configured to receive a processing block (e.g., a prediction block) of sample values within a current video picture in a sequence of video pictures and encode the processing block into a coded picture that is part of a coded video sequence. In one example, the video encoder (703) is used instead of the video encoder (403) of the example in FIG. 4.
[0084] In an example of HEVC, a video encoder (703) receives a matrix of sample values for a processing block, such as an 8×8 sample prediction block. The video encoder (703) determines whether the processing block is best coded using an intra mode, an inter mode, or a bi-prediction mode, for example, using rate-distortion optimization. When the processing block is coded in the intra mode, the video encoder (703) can encode the processing block into the coded picture using an intra prediction technique, and when the processing block is coded in the inter mode or the bi-prediction mode, the video encoder (703) can encode the processing block into the coded picture using an inter prediction technique or a bi-prediction technique, respectively. In certain video coding techniques, the merge mode can be an inter-picture prediction sub-mode in which a motion vector is derived from one or more motion vector predictors without the advantage of coded motion vector components outside the predictor. In certain other video coding techniques, there may be motion vector components applicable to the target block. In one example, the video encoder (703) includes other components, such as a mode decision module (not shown), to determine the mode of the processing block.
[0085] In the example of FIG. 7, the video encoder (703) includes an inter encoder (730), an intra encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general-purpose controller (721), and an entropy encoder (725) that are coupled together as shown in FIG. 7.
[0086] The inter-encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks (e.g., blocks in a previous picture and a subsequent picture) in a reference picture, generate inter-prediction information (e.g., description of redundant information, motion vectors, merge mode information by an inter-encoding technique), and calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture decoded based on the encoded video information.
[0087] The intra-encoder (722) is configured to receive samples of a current block (e.g., a processing block), compare the block with already encoded blocks in the same picture in some cases, generate quantized coefficients after transformation, and also generate intra-prediction information (e.g., intra-prediction direction information by one or more intra-encoding techniques) in some cases. In one example, the intra-encoder (722) also calculates an intra-prediction result (e.g., a predicted block) based on the intra-prediction information and reference blocks in the same picture.
[0088] The general-purpose controller (721) is configured to determine general-purpose control data and control other components of the video encoder (703) based on the general-purpose control data. In one example, the general-purpose controller (721) determines a mode of a block and provides a control signal to a switch (726) based on the mode. For example, when the mode is an intra mode, the general-purpose controller (721) controls the switch (726) to select an intra-mode result for use by the residual calculator (723), controls the entropy encoder (725) to select intra-prediction information, includes the intra-prediction information in the bitstream, and when the mode is an inter mode, the general-purpose controller (721) controls the switch (726) to select an inter-prediction result for use by the residual calculator (723), controls the entropy encoder (725) to select inter-prediction information, and includes the inter-prediction information in the bitstream.
[0089] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) is configured to operate based on the residual data to encode the residual data and generate transform coefficients. In one example, the residual encoder (724) is configured to convert the residual data from the spatial domain to the frequency domain and generate transform coefficients. The transform coefficients then undergo quantization processing to obtain quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform inverse transformation and generate decoded residual data. The decoded residual data can be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and intra prediction information. The decoded block is appropriately processed to generate a decoded picture, and the decoded picture is buffered in a memory circuit (not shown) and can be used as a reference picture in some examples.
[0090] The entropy encoder (725) is configured to format the bitstream to include the encoded block. The entropy encoder (725) is configured to include various information in the bitstream according to an appropriate standard such as the HEVC standard. In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other appropriate information in the bitstream. Note that according to the disclosed subject matter, there is no residual information when encoding a block in either the merge submode of the inter mode or the bi-prediction mode.
[0091] FIG. 8 shows an exemplary diagram of a video decoder (810). The video decoder (810) is configured to receive a coded picture that is part of a coded video sequence and decode the coded picture to generate a restored picture. In one example, the video decoder (810) is used in place of the video decoder (410) of the example of FIG. 4.
[0092] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a restoration module (874), and an intra decoder (872) that are coupled together as shown in FIG. 8.
[0093] The entropy decoder (871) can be configured to restore from the coded picture specific symbols that represent the syntax elements that the coded picture is composed of. Such symbols can include, for example, the mode in which a block (such as the latter two of intra mode, inter mode, bi-prediction mode, merge sub-mode, or another sub-mode) is coded, and prediction information (such as intra prediction information or inter prediction information) that can identify specific samples or metadata used for prediction by the intra decoder (872) or the inter decoder (880), respectively. The symbols can also include, for example, residual information in the form of quantized transform coefficients. In one example, when the prediction mode is inter mode or bi-prediction mode, the inter prediction information is provided to the inter decoder (880), and when the prediction type is intra prediction type, the intra prediction information is provided to the intra decoder (872). The residual information can undergo inverse quantization and is provided to the residual decoder (873).
[0094] The inter decoder (880) is configured to receive inter prediction information and generate an inter prediction result based on the inter prediction information.
[0095] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.
[0096] The residual decoder (873) is configured to perform inverse quantization to extract inverse quantization transform coefficients, process the inverse quantization transform coefficients, and convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (including the quantization parameter (QP)), and that information may be provided by the entropy decoder (871) (since this may be only a small amount of control information, the data path is not depicted).
[0097] The restoration module (874) is configured to combine, in the spatial domain, the residual information output by the residual decoder (873) and the prediction result (optionally output by the inter prediction module or the intra prediction module) to form a restored block that can be part of the restored picture, and the restored picture can be part of the restored video. Note that other appropriate operations, such as a deblocking operation, can be performed to improve the appearance.
[0098] Note that the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) can be implemented using any suitable technique. In one embodiment, the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (603), and the video decoders (410), (510), and (810) can be implemented using one or more processors that execute software instructions.
[0099] In intra prediction, for example, in AOMedia Video 1 (AV1), Versatile Video Coding (VVC), etc., various intra prediction modes can be used. In embodiments such as AV1, directional intra prediction is used. In examples such as the open video coding format VP9, eight directional modes correspond to eight angles from 45° to 207°. In order to utilize more spatial redundancy in the directional texture, for example, in AV1, the directional mode (also called the directional intra mode, directional intra prediction mode, or angle mode) can be extended to an angle set with finer granularity as shown in FIG. 9.
[0100] FIG. 9 shows an example of a nominal mode for a coding block (CB) (910) according to an embodiment of the present disclosure. A specific angle (referred to as a nominal angle) can correspond to the nominal mode. In one example, eight nominal angles (or nominal intra-angles) (901)-(908) each correspond to eight nominal modes (e.g., V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED). The eight nominal angles (901)-(908) and the eight nominal modes can be referred to as V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED, respectively. Further, each nominal angle can correspond to a plurality of finer angles (e.g., seven finer angles), and thus, for example, in AV1, 56 angles (or prediction angles) or 56 directional modes (or angle modes, directional intra-prediction modes) can be used. Each prediction angle can be presented by a nominal angle and an angle offset (or angle delta). The angle offset can be obtained by multiplying an offset integer I (e.g., -3, -2, -1, 0, 1, 2, or 3) by a step size (e.g., 3°). In one example, the prediction angle is equal to the sum of the nominal angle and the angle offset. In examples such as AV1, the nominal modes (e.g., the eight nominal modes (901)-(908)) can be signaled together with a specific non-angle smoothing mode (e.g., five non-angle smoothing modes such as the DC mode, PAETH mode, SMOOTH mode, vertical SMOOTH mode, and horizontal SMOOTH mode described below). Thereafter, if the current prediction mode is a directional mode (or angle mode), an index can be further signaled to indicate the angle offset (e.g., offset integer I) corresponding to the nominal angle.In one example, to implement the directional prediction mode by a general method, the 56 directional modes as used in AV1 are implemented using a unified directional predictor that projects each pixel to a reference sub-pixel position and can interpolate the reference pixels by a 2-tap bilinear filter.
[0101] For intra prediction of blocks such as CB, a non-directional smooth intra predictor (also called non-directional smooth intra prediction mode, non-directional smooth mode, or non-angular smooth mode) can be used. In some examples (e.g., in AV1), the five non-directional smooth intra prediction modes include the DC mode or DC predictor (e.g., DC), the PAETH mode or PAETH predictor (e.g., PAETH), the SMOOTH mode or SMOOTH predictor (e.g., SMOOTH), the vertical SMOOTH mode (referred to as SMOOTH_V mode, SMOOTH_V predictor, or SMOOTH_V), and the horizontal SMOOTH mode (referred to as SMOOTH_H mode, SMOOTH_H predictor, or SMOOTH_H).
[0102] FIG. 10 shows an example of non-directional smooth intra prediction modes (e.g., DC mode, PAETH mode, SMOOTH mode, SMOOTH_V mode, and SMOOTH_H mode) according to an aspect of the present disclosure. To predict a sample (1001) within a CB (1000) based on a DC predictor, the average value of the first value of the left adjacent sample (1012) and the second value of the upper adjacent sample (or top adjacent sample) (1011) can be used as the predictor.
[0103] To predict a sample (1001) based on a PAETH predictor, the first value of the left adjacent sample (1012), the second value of the upper adjacent sample (1011), and the third value of the upper left adjacent sample (1013) can be obtained. Then, a reference value is obtained using Equation 1. Reference value = First value + Second value - Third value (Equation 1)
[0104] One of the first value, the second value, and the third value that is closest to the reference value can be set as the predictor for the sample (1001).
[0105] The SMOOTH_V mode, the SMOOTH_H mode, and the SMOOTH mode can predict the CB (1000) using quadratic interpolation of the average value in the vertical direction, the horizontal direction, and both the vertical and horizontal directions, respectively. To predict the sample (1001) based on the SMOOTH predictor, the average value (e.g., weighted combination) of the first value, the second value, the value of the right sample (1014), and the value of the lower sample (1016) can be used. In various examples, the right sample (1014) and the lower sample (1016) are not restored, and thus, the value of the upper-right adjacent sample (1015) and the value of the lower-left adjacent sample (1017) can replace the value of the right sample (1014) and the value of the lower sample (1016), respectively. Therefore, the average value (e.g., weighted combination) of the first value, the second value, the value of the upper-right adjacent sample (1015), and the value of the lower-left adjacent sample (1017) can be used as the SMOOTH predictor. To predict the sample (1001) based on the SMOOTH_V predictor, the average value (e.g., weighted combination) of the second value of the upper adjacent sample (1011) and the value of the lower-left adjacent sample (1017) can be used. To predict the sample (1001) based on the SMOOTH_H predictor, the average value (e.g., weighted combination) of the first value of the left adjacent sample (1012) and the value of the upper-right adjacent sample (1015) can be used.
[0106] Embodiments of primary transformation such as those used in AV1 are described below. In AV1 and the like, a plurality of transform sizes (e.g., ranging from 4 points to 64 points for each dimension) and various transform shapes (e.g., square, rectangular shapes with a width-to-height ratio of 2:1, 1:2, 4:1, or 1:4) can be used.
[0107] The 2D transformation process can use a hybrid transformation kernel that can include different 1D transformations for each dimension of the coded residual block. The primary 1D transformation can include (a) a 4-point, 8-point, 16-point, 32-point, 64-point DCT-2 (or DCT2), (b) a 4-point, 8-point, 16-point asymmetric DST (ADST) (e.g., DST-4 or DST4, DST-7 or DST7) and the corresponding inverse version (e.g., the inverse version of ADST or FlipADST, which can apply ADST in reverse order), and / or (c) a 4-point, 8-point, 16-point, 32-point identity transform (IDTX or IDT). FIG. 11 shows examples of primary transformation basis functions. The primary transformation basis functions in the example of FIG. 11 include the basis functions for DCT-2 as well as asymmetric DSTs (e.g., DST-4 and DST-7) with N-point inputs. The primary transformation basis functions shown in FIG. 11 can be used in AV1.
[0108] The availability of the hybrid transformation kernel can depend on the transformation block size and the prediction mode. FIG. 12 shows exemplary dependencies of the availability of various transformation kernels (e.g., the transformation types shown in the first column and explained in the second column) based on the transformation block size (e.g., the sizes shown in the third column) and the prediction mode (e.g., the intra prediction and inter prediction shown in the third column). The exemplary hybrid transformation kernels, as well as the availability based on the prediction mode and transformation block size, can be used, for example, in AV1. Referring to FIG. 12, the symbols "→" and "↓" mean the horizontal dimension (also called the horizontal direction) and the vertical dimension (also called the vertical direction), respectively. The symbol
Number
Number
[0109] In one example, the conversion type indicated by ADST_DCT in the first column of FIG. 12 includes vertical ADST and horizontal DCT as shown in the second column of FIG. 12. According to the third column of FIG. 12, when the block size is 16×16 or less (for example, 16×16 samples, 16×16 luminance samples), the conversion type ADST_DCT is available for intra prediction and inter prediction.
[0110] In one example, the conversion type indicated by V_ADST as shown in the first column of FIG. 12 includes vertical ADST and horizontal IDTX (i.e., identity matrix) as shown in the second column of FIG. 12. Therefore, the conversion type V_ADST is executed vertically and not horizontally. According to the third column of FIG. 12, the conversion type V_ADST is not available for intra prediction regardless of the block size. The conversion type V_ADST is available for inter prediction when the block size is less than 16×16 (for example, 16×16 samples, 16×16 luminance samples).
[0111] In one example, FIG. 12 is applicable to the luminance component. For the chrominance component, the selection of the conversion type (or conversion kernel) can be implicitly performed. In one example, for the intra prediction residual, as shown in FIG. 13, the conversion type can be selected according to the intra prediction mode. FIG. 13 shows an exemplary selection of the conversion type based on the intra prediction mode. In one example, the selection of the conversion type shown in FIG. 13 is applicable to the chrominance component. In the case of the inter prediction residual, the conversion type can be selected according to the selection of the conversion type of the luminance block in the same location. Therefore, in one example, the conversion type for the chrominance component is not signaled in the bitstream.
[0112] The present disclosure includes embodiments related to adaptive multi - transform set selection.
[0113] To demonstrate a reference implementation form of an encoding technique and a decoding process for Joint Video Experts Team (JVET) extended compression beyond the multi - purpose video coding (VVC) capability exploration work, an Extended Compression Model (ECM) - 2.0 reference software can be provided.
[0114] In one embodiment, such as ECM - 2.0, in addition to DCT2, multiple (e.g., four) different multi - transform selection (MTS) candidates are used. The transform pairs associated with each MTS candidate (including, for example, a vertical transform along the vertical direction and a horizontal transform along the horizontal direction) may depend on prediction modes such as TU size (e.g., TB size or transform block size) and / or intra - mode (e.g., intra - prediction mode) as shown in FIGS. 12 - 13. In one embodiment, the transform pairs can be constructed using non - DCT2 transform kernels such as DST7, DCT8, DCT5, DST4, DST1, and identity transform (IDT).
[0115] (For example, the MTS index, denoted as mts_idx) can indicate which MTS candidate is selected from a plurality of MTS candidates. The MTS index (e.g., mts_idx) can be signaled or inferred. The MTS index (e.g., mts_idx) can be signaled when a block (e.g., TB) contains at least one non-DC coefficient, for example, when the position of the last significant coefficient in the scan order is greater than 0. A TB containing coefficients can be called a transform coefficient block. In one example, the MTS index (e.g., mts_idx) can be signaled when a block (e.g., TB) contains at least one non-DC coefficient, for example, when the position of the last significant coefficient in the scan order is greater than 0. Otherwise, the MTS index (e.g., mts_idx) is not signaled, and the transform can be inferred to be DCT2 (e.g., when mts_idx is 0), for example, when a block (e.g., TB) contains at least one non-DC coefficient. To signal the MTS index (e.g., mts_idx), the first bin of the MTS index can indicate whether the MTS index is greater than 0. If the MTS index is greater than 0, an additional M bits (e.g., 2 bits) using a fixed-length code can be signaled to indicate the signaled MTS candidate among a plurality (e.g., 4) of MTS candidates. M can be a positive integer.
[0116] In some embodiments, such as related MTS methods, a fixed number (e.g., 4) of MTS candidates are used. Using a fixed number of MTS candidates may not be optimal as the residual characteristics of a block (e.g., a TB) are not considered. For example, a block (e.g., a TB) having less residual energy or a smaller number of coefficients (e.g., a smaller number of non-DC coefficients) can benefit from a reduced number of MTS candidates as reducing the number of MTS candidates can reduce overhead signaling. A block (e.g., a TB) having higher residual energy or a larger number of coefficients (e.g., a larger number of non-DC coefficients) can benefit from a larger number of MTS candidates as a larger number of MTS candidates can provide more diversity in MTS candidate selection.
[0117] In some embodiments, a variable number (e.g., 1, 4, or 6) of MTS candidates can be selected according to the position of the last significant coefficient (e.g., lastScanPos) in the scan order. A smaller MTS candidate set having a smaller number of MTS candidates can be a subset of a larger MTS candidate set (e.g., the next larger MTS candidate set) having a larger number of MTS candidates, as shown in FIG. 14.
[0118] FIG. 14 shows an example of MTS candidates and the selection of subsets of MTS candidates according to the position of the last significant coefficient (e.g., lastScanPos) in the scanning order. The MTS candidate set (e.g., the full MTS candidate set) includes a plurality of MTS candidates. The plurality of MTS candidates can include any transformation or any type of transformation. The number of the plurality of MTS candidates can be any positive number. In one example, the number of the plurality of MTS candidates in the MTS candidate set is greater than 1. In the example of FIG. 14, the MTS candidate set (e.g., the full MTS candidate set) includes a plurality of MTS candidates T0 to T5. For example, the plurality of MTS candidates T0 to T5 includes six candidates: DCT2, DST7, DCT8, DCT5, DST4, DST1, and IDT. In one example, the plurality of MTS candidates T0 to T5 respectively correspond to DCT2, DST7, DCT8, DCT5, DST4, and DST1. In one example, the plurality of MTS candidates T0 to T5 respectively correspond to DST7, DCT8, DCT5, DST4, DST1, and IDT.
[0119] A plurality of subsets of MTS candidates can be determined based on the MTS candidate set. For example, one or more thresholds (TH), such as a first threshold TH[0] (e.g., 6) and a second threshold TH[1] (e.g., 32), are used to determine the plurality of subsets of MTS candidates. In the example of FIG. 14, one or more thresholds, such as TH[0] being 6 and TH[1] being 32, are fixed thresholds. Referring to FIG. 14, the plurality of subsets of MTS candidates includes a first subset of MTS candidates, a second subset of MTS candidates, and a third subset of MTS candidates. The first subset of MTS candidates includes one MTS candidate such as T0. The second subset of MTS candidates includes four MTS candidates such as T0 to T3. The third subset of MTS candidates includes six MTS candidates such as T0 to T5. In FIG. 14, the third subset of MTS candidates is the full MTS candidate set.
[0120] Referring to FIG. 14, when the position of the last significant coefficient (e.g., lastScanPos) in the scanning order of a block (e.g., TB) is less than or equal to the first threshold TH[0], the MTS candidate subset of the block is the first MTS candidate subset. When the position of the last significant coefficient (e.g., lastScanPos) in the scanning order of a block (e.g., TB) is (i) less than or equal to the second threshold TH[1] and (ii) greater than the first threshold TH[0], the MTS candidate subset of the block is the second MTS candidate subset. When the position of the last significant coefficient (e.g., lastScanPos) in the scanning order of a block (e.g., TB) is greater than or equal to the second threshold TH[1], the MTS candidate subset of the block is the third MTS candidate subset.
[0121] In the example of FIG. 14, no additional non-DCT2 transform kernels are used compared to ECM-2.0. The TU shape and intra-mode dependency remain unchanged from ECM-2.0.
[0122] In various embodiments, determining the MTS candidate subset based on fixed thresholds (e.g., TH[0] in FIG. 14 is 6 and TH[1] is 32) may not be optimally adapted to (i) the content of the image and / or video and / or (ii) the coding conditions. The present disclosure includes embodiments of adaptive MTS subset selection that can use an adaptive method to determine the MTS subset selection for a block (e.g., TB).
[0123] The embodiments described in the present disclosure can be applied individually or in any form of combination. The transforms or transform types included in the MTS candidate set can include any transform and are not limited to T0 to T5 described with reference to FIG. 14. The term "block" may be interpreted as a prediction block, a transform block, a coding block, a coding unit (CU), etc.
[0124] According to one embodiment, a method for determining (e.g., selecting) an MTS candidate subset for a TB (e.g., a transform coefficient block) can be signaled at a high level, such as in a high-level syntax (HLS) header. A plurality of TBs (e.g., transform coefficient blocks) can refer to the HLS header. The MTS selection information for the plurality of TBs can be signaled within the HLS header. The MTS selection information can indicate which MTS candidates in the MTS candidate set are included in the MTS candidate subset for the TB.
[0125] The high level can be higher than the block level. The HLS header can include, but is not limited to, a slice header, a picture header, a picture parameter set (PPS), a video parameter set (VPS), an adaptation parameter set (APS), a sequence parameter set (SPS), etc.
[0126] The MTS candidate set (e.g., a full MTS candidate set) can be predefined. In one example, the MTS candidate set is agreed upon by or pre-determined for the encoder and decoder. In the example of FIG. 14, the full MTS candidate set is {T0, T1, T2, T3, T4, T5}. The possible transforms used to form the MTS candidate subsets (e.g., the first MTS candidate subset, the second MTS candidate subset, and the third MTS candidate subset in FIG. 14) can be derived from the full MTS candidate set. In one example, the MTS selection information is used by a plurality of TBs to form the MTS candidate subset.
[0127] At the block level (e.g., the TB level), the MTS candidate subset for the TB can be determined based on actively used information such as (i) threshold information indicating one or more thresholds for counting coefficients, and (ii) an MTS candidate subset that can be formed from the full MTS candidate set. The threshold and the MTS candidate subset can be determined (e.g., derived) from the HLS header that the block (e.g., the TB) refers to.
[0128] In one embodiment, threshold information indicating one or more thresholds (e.g., TH[0], TH[1]) is signaled. For example, the MTS selection information in the HLS header indicates one or more thresholds. As described in FIG. 14, the MTS candidate subset for the TB can be determined based on one or more thresholds of the TB and coefficient information. The coefficient information of the TB can indicate the number of non-zero coefficients in the TB or the position of the last significant coefficient in the scan order of the TB (e.g., lastScanPos). In one example, the MTS candidate subset for the TB is determined based on a comparison of one or more thresholds with (i) the number of non-zero coefficients in the TB, or (ii) the position of the last significant coefficient in the scan order of the TB (e.g., lastScanPos).
[0129] In one example, the number of one or more thresholds (e.g., 2 in the case of TH[0] and TH[1]) is signaled. In one example, the number of one or more thresholds (e.g., 2 in the case of TH[0] and TH[1]) is not signaled. The number of different MTS candidate subsets can be derived from the number of one or more thresholds. For example, the number of different MTS candidate subsets is equal to the number of one or more thresholds plus one. Referring to FIG. 14, two thresholds TH[0] and TH[1] are signaled and three MTS candidate subsets of MTS can be used.
[0130] In one embodiment, MTS selection information indicating a conversion selection within one or more MTS candidate subsets is signaled. The MTS selection information may be signaled within the HLS header. The MTS selection information can indicate the number of conversions (e.g., MTS candidates) in each of the one or more MTS candidate subsets. In one example, the number of conversions in each of the one or more MTS candidate subsets is signaled. In the example of FIG. 14, the numbers 1, 4, and 6 indicating that three MTS candidate subsets are formed are signaled. The three MTS candidate subsets each include one, four, and six conversions (e.g., MTS candidates), respectively. The three MTS candidate subsets can be determined based on (i) the numbers 1, 4, and 6, and (ii) the MTS candidate set (e.g., T0 to T5). In one example, T0 to T5 are ranked in descending order of selection. Thus, the first MTS candidate subset having one conversion is {T0}, the second MTS candidate subset having four conversions is {T0, T1, T2, T3}, and the third MTS candidate subset having six conversions is {T0, T1, T2, T3, T4, T5}.
[0131] If it is pre-determined that the MTS candidate subset (e.g., the last MTS candidate subset) is the full MTS candidate set, the corresponding number of MTS candidates (the last number such as the maximum number 6) need not be signaled and can be inferred to be the number of candidates within the full MTS candidate set. If it is pre-determined that the MTS candidate subset (e.g., the first MTS candidate subset) includes only the default conversion (e.g., T0 is DCT2), the corresponding number of MTS candidates need not be signaled and can be inferred to be 1.
[0132] According to one embodiment of the present disclosure, the MTS selection information can be determined based on a previously encoded region.
[0133] In one embodiment, MTS selection information indicating threshold information including one or more thresholds indicating which MTS candidate subset is used for the TBs within a current region (e.g., a current CTU, a current slice, a current picture, or a current picture group (GOP)) can be determined (e.g., derived). The one or more thresholds may be determined based on previously encoded residual information within a previously encoded region (e.g., a previously encoded CTU, a previously encoded slice, a previously encoded picture, or a previously encoded GOP). The previously encoded region may include a plurality of previously encoded blocks (e.g., TBs). The previously encoded region may correspond to the current region. In one example, the previously encoded region and the current region are slices. In one example, the previously encoded region and the current region are pictures. The previously encoded residual information may indicate coefficient information of a plurality of previously encoded blocks.
[0134] In one embodiment, the previously encoded residual information indicates the encoded residuals (e.g., coefficients) within the previously encoded region. The previously encoded residual information (e.g., the previously encoded coefficient information) can indicate the average number of non-zero coefficients per block (e.g., per TB) within the previously encoded region, or the average number of the last significant coefficients in the scan order per block (e.g., per TB) within the previously encoded region. In one example, the average number of non-zero coefficients per TB within the previously encoded region, or the average number of the last significant coefficients in the scan order per TB within the previously encoded region is calculated and / or stored.
[0135] The previously encoded residual information can be used to determine one or more thresholds for the current region. For example, one or more thresholds (e.g., TH[0] and TH[1]) for a plurality of TBs within the current region are determined based on the previously encoded residual information and are not signaled.
[0136] In one example, previously coded residual information from a previously coded region is determined and / or stored based on different types of block sizes (e.g., TB size) within the previously coded region. The block size can be measured by, for example, the number of samples in the block, the block width, the block height, etc. The selection of the MTS candidate subset for a TB within the current region can be determined based on the previously coded residual information of the type of block size to which the TB belongs, as described below.
[0137] The previously coded residual information can include a plurality of coefficient information. Each of the plurality of coefficient information can correspond to each type of a plurality of types of block sizes within the previously coded region. In one example, the plurality of types of block sizes within the previously coded region includes a first type (e.g., the number of samples within the TB is N1 or less) and a second type (e.g., the number of samples within the TB is greater than N1). N1 is a positive integer. The plurality of coefficient information includes first coefficient information and second coefficient information corresponding to the first type and the second type, respectively. The first coefficient information can be determined based on a first subset of a plurality of previously coded TBs belonging to the first type. The second coefficient information can be determined based on a second subset of a plurality of previously coded TBs belonging to the second type.
[0138] The current region includes a first TB and a second TB. The first number of samples within the first TB is less than N1, and the second number of samples within the second TB is greater than N1. The first TB belongs to the first type, and the second TB belongs to the second type. The first threshold for the first TB can be determined based on the first coefficient information, and the second threshold for the second TB can be determined based on the second coefficient information. Then, based on the first threshold and the second threshold, the MTS candidate subsets of the first TB and the second TB can be determined.
[0139] According to one embodiment of the present disclosure, the MTS selection information determined based on previously coded regions indicates a subset of MTS candidates. The subset of MTS candidates selected for the TBs within the current region (e.g., the three subsets of MTS candidates in FIG. 14) can be determined (e.g., derived) based on the previously coded MTS information within the previously coded regions.
[0140] In one embodiment, the previously coded MTS information includes statistical information (e.g., statistical data) of the conversion types used within the previously coded regions. In one example, the frequency of each conversion type (e.g., DCT2, DST7, DCT8, DCT5, DST4, DST1, and IDT) used within the previously coded regions can be determined. For example, the most frequently used conversion type per block from the previously coded regions can be calculated and / or stored. The previously coded MTS information (e.g., the frequency of each conversion type) can be used to determine the conversions (e.g., conversion types) within the subset of MTS candidates. The previously coded MTS information can be used to determine the order of the MTS candidates within the subset of MTS candidates in the current region.
[0141] Referring to FIG. 14, the frequencies of the respective conversion types used within the previously coded regions are T0, T1, T2, T3, T4, and T5 in descending order. Thus, when subsets of MTS candidates having one, four, and six MTS candidates are to be determined for the TBs within the current region, the subsets of MTS candidates are the three subsets of MTS candidates shown in FIG. 14. For example, the second subset of MTS candidates includes T0, T1, T2, and T3, and the most frequently used conversion T0 within the previously coded region is the first within the second subset of MTS candidates.
[0142] In one example, previously encoded MTS information from a previously encoded region is determined and / or stored based on different types of block sizes (e.g., TB size) within the previously encoded region as described above. The selection of the MTS candidate subset for TBs within the current region and / or the order of MTS candidates or conversions within the MTS candidate subset can be determined based on the previously encoded residual information of the type of block size to which the TB belongs, as described below.
[0143] The previously encoded MTS information can include multiple MTS information. Each of the multiple MTS information can correspond to each type of multiple types of block sizes within the previously encoded region. The multiple MTS information can include first MTS information and second MTS information corresponding to a first type and a second type, respectively. The first MTS information, such as the first frequency of each conversion type within the first subset of multiple previously encoded TBs belonging to the first type, can be determined based on the first subset. The second MTS information, such as the second frequency of each conversion type within the second subset of multiple previously encoded TBs belonging to the second type, can be determined based on the second subset.
[0144] The current region includes a first TB and a second TB as described above. The first conversion within the first MTS candidate subset for the first TB can be determined based on the first MTS information, and the second conversion within the second MTS candidate subset for the second TB can be determined based on the second MTS information.
[0145] Embodiments of the present disclosure may be used separately or combined in any order. For example, threshold information indicating thresholds for multiple TBs within a current region can be signaled within a video bitstream, derived based on previously encoded regions, or pre-determined. Multiple MTS candidate subsets can be determined based on signaled information and / or derived information such as the aforementioned thresholds, the number of MTS candidates within each MTS candidate subset, previously encoded residual information, and / or previously encoded MTS information. The MTS candidate subset for a TB can be determined based on signaled information, derived information, and / or pre-determined information such as (e.g., signaled or derived) thresholds, (e.g., derived or pre-determined) multiple MTS candidate subsets, and coefficient information of the TB. Multiple derived MTS candidate subsets can be determined from a full MTS candidate set pre-determined for the encoder and decoder.
[0146] FIG. 15 shows a flowchart outlining an encoding process (1500) according to an embodiment of the present disclosure. In various embodiments, process (1500) is executed by a processing circuit such as a processing circuit within terminal devices (310), (320), (330), and (340), or a processing circuit that performs the functions of a video encoder (e.g., (403), (603), (703)). In some embodiments, process (1500) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes process (1500). The process starts at (S1501) and proceeds to (S1510).
[0147] In (S1510), the multiple transformation selection (MTS) selection information for a plurality of transformation blocks (TBs) to be transformed can be determined based on a previously encoded region, such as a plurality of previously encoded TBs in the previously encoded region. The MTS selection information can indicate (i) threshold information for a TB among the plurality of transformation blocks (TBs) or (ii) at least one of a plurality of MTS candidate subsets.
[0148] In (S1520), which MTS candidate subset among the plurality of MTS candidate subsets is selected for a TB can be determined based on the MTS selection information.
[0149] In (S1530), the TB can be transformed based on the MTS candidates included in the determined MTS candidate subset.
[0150] In one example, the MTS selection information is encoded. The encoded MTS selection information can be included in the video bitstream.
[0151] The process (1500) then proceeds to (S1599) and ends.
[0152] The process (1500) can be appropriately adapted to various scenarios, and the steps within the process (1500) can be adjusted accordingly. One or more of the steps within the process (1500) can be adapted, omitted, repeated, and / or combined. Any suitable order can be used to implement the process (1500). Additional steps can be added.
[0153] FIG. 16 shows a flowchart outlining a decoding process (1600) according to an embodiment of the present disclosure. In various embodiments, process (1600) is executed by processing circuitry within terminal devices (310), (320), (330), and (340), processing circuitry that executes the functions of video encoder (403), processing circuitry that executes the functions of video decoder (410), processing circuitry that executes the functions of video decoder (510), processing circuitry that executes the functions of video encoder (603), and the like. In some embodiments, process (1600) is implemented by software instructions, and thus, when the processing circuitry executes the software instructions, the processing circuitry executes process (1600). The process begins at (S1601) and proceeds to (S1610).
[0154] (S1610), multiple transform selection (MTS) selection information for a plurality of transform coefficient blocks within the coded video bitstream may be determined. The MTS selection information can indicate threshold information for the transform coefficient blocks among the plurality of transform coefficient blocks and / or a plurality of MTS candidate subsets. In one example, the MTS selection information is applicable to a plurality of transform coefficient blocks.
[0155] In one embodiment, the MTS selection information is signaled within the coded video bitstream.
[0156] (S1620), which of the plurality of MTS candidate subsets is selected for a transform coefficient block may be determined based on the MTS selection information.
[0157] In one example, the MTS selection information signaled within the coded video bitstream indicates threshold information including at least one threshold. The MTS candidate subset can be determined based on at least one threshold and one of (i) the number of non-zero coefficients within the transform coefficient block or (ii) the position of the last significant coefficient in the scan order within the transform coefficient block.
[0158] In one example, the number of a plurality of MTS candidate subsets is the sum of the number of at least one threshold and 1. The plurality of MTS candidate subsets can be determined based on the number of the plurality of MTS candidate sets, and the MTS candidate subset can be determined based on the plurality of MTS candidate subsets.
[0159] In one embodiment, the MTS selection information signaled in the coded video bitstream includes one or more numbers. Each of the one or more numbers is the number of one or more MTS candidates in each one of one or more MTS candidate subsets among the plurality of MTS candidate subsets. The plurality of MTS candidate subsets can be determined based on the one or more numbers and the MTS candidate sets. The MTS candidate subset can be determined based on the plurality of MTS candidate sets. In one example, the plurality of MTS candidate subsets includes the last MTS candidate subset not included in the one or more MTS candidate subsets, and the last MTS candidate subset is an MTS candidate set. In one example, the plurality of MTS candidate subsets includes the first MTS candidate subset not included in the one or more MTS candidate subsets, and the first MTS candidate subset is composed of default MTS candidates in the MTS candidate set.
[0160] In one embodiment, the MTS selection information is determined based on a plurality of previously decoded transform coefficient blocks in a previously decoded region. The MTS selection information indicates threshold information including at least one threshold of the plurality of transform coefficient blocks. The at least one threshold can be determined based on the coefficient information of the plurality of previously decoded transform coefficient blocks. The MTS candidate subset can be determined based on the at least one threshold and one of (i) the number of non-zero coefficients in the transform coefficient block or (ii) the position of the last significant coefficient in the scan order in the transform coefficient block.
[0161] In one example, the coefficient information of a plurality of previously decoded transform coefficient blocks indicates (i) the average value of the non-zero coefficients of the plurality of previously decoded transform coefficient blocks, or (ii) the average position of the last significant coefficient in the scan order of the plurality of previously decoded transform coefficient blocks.
[0162] In one example, the MTS selection information indicates threshold information including at least one threshold for a plurality of transform coefficient blocks. The plurality of coefficient information is associated with a plurality of previously decoded transform coefficient blocks. Each of the plurality of coefficient information corresponds to each type of a plurality of types of block sizes within the previously decoded region. The at least one threshold can be determined based on the coefficient information corresponding to each type of block size to which the transform coefficient block belongs. The MTS candidate subset can be determined based on the at least one threshold and one of (i) the number of non-zero coefficients within the transform coefficient block or (ii) the position of the last significant coefficient in the scan order within the transform coefficient block.
[0163] In one example, the MTS selection information indicates a plurality of MTS candidate subsets. For a plurality of transform coefficient blocks, the MTS candidates and the order of the MTS candidates used to form the plurality of MTS candidate subsets from the MTS candidate set can be determined based on the statistical information of the transform types of the plurality of previously decoded transform coefficient blocks. The plurality of MTS candidate subsets can be determined from the MTS candidate set based on the MTS candidates and the order of the MTS candidates. The MTS candidate subset can be determined to be one of the plurality of MTS candidate sets.
[0164] In one example, the MTS selection information indicates a plurality of MTS candidate subsets. A plurality of statistical information of the conversion type is associated with a plurality of previously decoded conversion coefficient blocks. Each of the plurality of statistical information of the conversion type corresponds to each type of a plurality of types of block sizes within the previously decoded region. For a conversion coefficient block, the MTS candidates and the order of the MTS candidates used to form a plurality of MTS candidate subsets from the MTS candidate set can be determined based on one statistical information of the conversion type corresponding to the block size of the type to which the conversion coefficient block belongs. The plurality of MTS candidate subsets from the MTS candidate set can be determined based on the MTS candidates and the order of the MTS candidates. The MTS candidate subset can be determined to be one of the plurality of MTS candidate sets.
[0165] (In S1630), the conversion coefficient block can be inversely transformed based on the MTS candidates included in the determined MTS candidate subset.
[0166] The process (1600) proceeds to (S1699) and ends.
[0167] The process (1600) can be suitably adapted to various scenarios, and the steps within the process (1600) can be adjusted accordingly. One or more of the steps within the process (1600) can be adapted, omitted, repeated, and / or combined. Any suitable order can be used to implement the process (1600). Additional steps can be added.
[0168] FIG. 17 shows a flowchart outlining a decoding process (1700) according to an embodiment of the present disclosure. In various embodiments, process (1700) is executed by processing circuitry such as in terminal devices (310), (320), (330), and (340), processing circuitry that executes the functions of video encoder (403), processing circuitry that executes the functions of video decoder (410), processing circuitry that executes the functions of video decoder (510), processing circuitry that executes the functions of video encoder (603), and the like. In some embodiments, process (1700) is implemented by software instructions, and thus, when the processing circuitry executes the software instructions, the processing circuitry executes process (1700). The process begins at (S1701) and proceeds to (S1710).
[0169] (S1710), a coded video bitstream including a high-level syntax header is received. The high-level syntax header includes multiple transform selection (MTS) selection information for a plurality of transform coefficient blocks.
[0170] (S1720), at least one of (i) threshold information or (ii) MTS selection information indicating a plurality of MTS candidate subsets for a transform coefficient block among the plurality of transform coefficient blocks is determined.
[0171] (S1730), which MTS candidate subset among the plurality of MTS candidate subsets is selected for the transform coefficient block is determined based on the MTS selection information.
[0172] (S1740), the transform coefficient block can be inverse-transformed based on the MTS candidates included in the determined MTS candidate subset.
[0173] Process (1700) proceeds to (S1799) and ends.
[0174] Process (1700) can be appropriately adapted to various scenarios, and the steps within process (1700) can be adjusted accordingly. One or more of the steps within process (1700) can be adapted, omitted, repeated, and / or combined. Any suitable order can be used to implement process (1700). Additional steps can be added.
[0175] FIG. 18 shows a flowchart outlining a decoding process (1800) according to an embodiment of the present disclosure. In various embodiments, process (1800) is executed by processing circuits such as those within terminal devices (310), (320), (330), and (340), a processing circuit that executes the functions of video encoder (403), a processing circuit that executes the functions of video decoder (410), a processing circuit that executes the functions of video decoder (510), and a processing circuit that executes the functions of video encoder (603). In some embodiments, process (1800) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes process (1800). The process starts at (S1801) and proceeds to (S1810).
[0176] (S1810), the multiple transform selection (MTS) selection information for a plurality of transform coefficient blocks in the coded video bitstream can be determined based on a plurality of previously decoded transform coefficient blocks in a previously decoded region. The MTS selection information indicates (i) threshold information or (ii) at least one of a plurality of MTS candidate subsets for a transform coefficient block among the plurality of transform coefficient blocks.
[0177] (S1820), which MTS candidate subset among the plurality of MTS candidate subsets is selected for a transform coefficient block is determined based on the MTS selection information.
[0178] (S1830), the transform coefficient block can be inverse-transformed based on the MTS candidates included in the determined MTS candidate subset.
[0179] The process (1800) proceeds to (S1899) and ends.
[0180] The process (1800) can be appropriately adapted to various scenarios, and the steps within the process (1800) can be adjusted accordingly. One or more of the steps within the process (1800) can be adapted, omitted, repeated, and / or combined. Any suitable order can be used to implement the process (1800). Additional steps can be added.
[0181] Embodiments of the present disclosure may be used separately or combined in any order. Further, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.
[0182] The techniques described above can be implemented as computer software physically stored on one or more computer-readable media using computer-readable instructions. For example, FIG. 19 shows a computer system (1900) suitable for implementing a particular embodiment of the disclosed subject matter.
[0183] The computer software can be coded using any suitable machine language or computer language that can undergo an assembly, compilation, linking, or similar mechanism to create code that includes instructions executable directly, or via interpretation, microcode execution, etc., by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc.
[0184] The command can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, game devices, Internet of Things devices, and the like.
[0185] The components shown in FIG. 19 with respect to the computer system (1900) are exemplary in nature and do not imply any limitation regarding the use or functionality of the computer software implementing the embodiments of the present disclosure. The configuration of the components should not be construed as having any dependency or requirement regarding any one or combination of the components shown in the exemplary embodiments of the computer system (1900).
[0186] The computer system (1900) may include a specific human interface input device. Such a human interface input device can respond to input by one or more human users via, for example, tactile input (such as keystrokes, swipes, movements of a data glove), audio input (such as voice, clapping), visual input (such as gestures), and olfactory input (not depicted). The human interface device can also be used to capture specific media that is not necessarily directly related to conscious human input, such as audio (such as voice, music, ambient sound), images (such as scanned images, photographic images obtained from a still camera), and video (such as 2D video, 3D video including stereoscopic video).
[0187] The input human interface device may include one or more of a keyboard (1901), a mouse (1902), a trackpad (1903), a touch screen (1910), a data glove (not shown), a joystick (1905), a microphone (1906), a scanner (1907), and a camera (1908) (only one of each is depicted).
[0188] The computer system (1900) may also include certain human interface output devices. Such human interface output devices may, for example, stimulate the senses of one or more human users via tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., tactile feedback by a touch screen (1910), a data glove (not shown), or a joystick (1905), although there may also be tactile feedback devices that do not function as input devices), audio output devices (such as a speaker (1909), headphones (not depicted)), visual output devices (such as screens (1910) including CRT screens, LCD screens, plasma screens, OLED screens, some of which may output two-dimensional visual output or three-dimensional or higher output via means such as stereographic output, virtual reality glasses (not depicted), holographic displays, and smoke tanks (not depicted), regardless of whether they have a touch screen input function and regardless of whether they have a tactile feedback function), and a printer (not depicted) may also be included.
[0189] The computer system (1900) can also include storage devices accessible by humans and media related thereto, such as optical media including a CD / DVD ROM / RW (1920) having a CD / DVD or similar medium (1921), a thumb drive (1922), a removable hard drive or solid state drive (1923), legacy magnetic media such as tapes and floppy disks (not depicted), and special ROM / ASIC / PLD-based devices such as security dongles (not depicted).
[0190] One of ordinary skill in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transient signals.
[0191] The computer system (1900) can also include an interface (1954) to one or more communication networks (1955). The network can be, for example, wireless, wired, optical. The network can further be local, wide area, metropolitan, vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., wired or wireless wide area digital networks for TV including cable TV, satellite TV, and terrestrial broadcast TV, vehicle and industrial including CANBus, etc. A particular network typically requires an external network interface adapter attached to a particular general-purpose data port or peripheral bus (1949) (such as a USB port of the computer system (1900)), and other networks are typically integrated into the core of the computer system (1900) by attaching to the system bus described below (such as an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (1900) can communicate with other entities. Such communication can be only unidirectional reception (such as broadcast TV), only unidirectional transmission (such as CANbus to a particular CANbus device), or bidirectional with other computer systems using, for example, local or wide area digital networks. Particular protocols and protocol stacks can be used with each of those networks and network interfaces described above.
[0192] The aforementioned human interface device, human-accessible storage device, and network interface can be attached to the core (1940) of the computer system (1900).
[0193] The core (1940) can include special programmable processing devices in the form of one or more central processing units (CPUs) (1941), graphics processing units (GPUs) (1942), field programmable gate arrays (FPGAs) (1943), hardware accelerators for specific tasks (1944), graphics adapters (1950), etc. These devices may be connected via a system bus (1948) together with read-only memory (ROM) (1945), random access memory (1946), internal mass storage such as an internal hard drive or SSD that is not accessible to the user (1947). In some computer systems, the system bus (1948) may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the core's system bus (1948) or via a peripheral bus (1949). In one example, a screen (1910) can be connected to a graphics adapter (1950). Architectures for peripheral buses include PCI, USB, etc.
[0194] The CPU (1941), GPU (1942), FPGA (1943), and accelerator (1944) can, in combination, execute specific instructions that can make up the aforementioned computer code. That computer code can be stored in ROM (1945) or RAM (1946). Migration data can also be stored in RAM (1946), while persistent data can be stored, for example, in internal mass storage (1947). Fast storage and retrieval for any of the memory devices can be enabled using cache memory that can be closely associated with one or more CPUs (1941), GPUs (1942), mass storage (1947), ROM (1945), RAM (1946), etc.
[0195] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code can be specially designed and constructed for the purposes of this disclosure, or they can be of the kind well-known and available to persons having skill in the computer software arts.
[0196] As an example, rather than as a limitation, a computer system (1900) having an architecture, specifically a core (1940), can provide functionality as a result of software embodied within one or more tangible computer-readable media being executed by a processor (including a CPU, GPU, FPGA, accelerator, etc.). Such a computer-readable media can be the user-accessible mass storage introduced above, as well as media associated with specific storage of the core (1940) of a non-transitory nature, such as core internal mass storage (1947) or ROM (1945). Software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core (1940). The computer-readable media can include one or more memory devices or chips, depending on specific needs. Software can cause the core (1940), and specifically the processor (including a CPU, GPU, FPGA, etc.) therein, to define data structures stored in RAM (1946) and modify such data structures according to processes defined by the software, thereby executing specific processes or specific parts of specific processes described herein. Additionally, or alternatively, the computer system can provide functionality as a result of logic wired or otherwise embodied within a circuit (e.g., an accelerator (1944)) that can operate instead of or in conjunction with software to execute specific processes or specific parts of specific processes described herein. Optionally, references to software can include logic and vice versa. Optionally, references to computer-readable media can include circuits (such as integrated circuits (ICs)) that store software for execution, circuits that embody logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.
[0197] Appendix A: Acronyms JEM: joint exploration model, collaborative exploration model VVC: Versatile Video Coding, Multi-purpose Video Coding BMS: Benchmark Set, Benchmark Set MV: Motion Vector, Motion Vector HEVC: High Efficiency Video Coding, High Efficiency Video Coding SEI: Supplementary Enhancement Information, Supplementary Enhancement Information VUI: Video Usability Information, Video Usability Information GOP: Group of Pictures, Picture Group TU: Transform Unit, Transform Unit PU: Prediction Unit, Prediction Unit CTU: Coding Tree Unit, Coding Tree Unit CTB: Coding Tree Block, Coding Tree Block PB: Prediction Block, Prediction Block HRD: Hypothetical Reference Decoder, Hypothetical Reference Decoder SNR: Signal Noise Ratio, Signal-to-Noise Ratio CPU: Central Processing Unit, Central Processing Unit GPU: Graphics Processing Unit, Graphics Processing Unit CRT: Cathode Ray Tube, Cathode Ray Tube LCD: Liquid-Crystal Display, Liquid-Crystal Display OLED: Organic Light-Emitting Diode, Organic Light-Emitting Diode CD: Compact Disc, Compact Disc DVD: Digital Video Disc, Digital Video Disc ROM: Read-Only Memory, Read-Only Memory RAM: Random Access Memory, Random Access Memory ASIC: Application-Specific Integrated Circuit, Application-Specific Integrated Circuit PLD: Programmable Logic Device, Programmable Logic Device LAN: Local Area Network, Local Area Network GSM: Global System for Mobile communications, Global System for Mobile Communications LTE: Long-Term Evolution, Long-Term Evolution CANBus: Controller Area Network Bus, Controller Area Network Bus USB: Universal Serial Bus, Universal Serial Bus PCI: Peripheral Component Interconnect, Peripheral Component Interconnect FPGA: Field Programmable Gate Array, Field Programmable Gate Array SSD: solid-state drive, solid-state drive IC: Integrated Circuit, Integrated Circuit CU: Coding Unit, Coding Unit R-D: Rate-Distortion, Rate-Distortion
[0198] This disclosure describes several exemplary embodiments, but there are changes, substitutions, and various alternative equivalents that fall within the scope of this disclosure. Therefore, it will be understood that those skilled in the art can devise numerous systems and methods that embody the principles of this disclosure and are thus within its spirit and scope, although not explicitly illustrated or described herein.
Explanation of Signs
[0199] 300 Communication System 310, 320, 330, 340 Terminal Devices 350 Communication Network 400 Communication System 401 Video Source 402 Video Picture Stream 403 Video Encoder 404 Video Data 405 Streaming Server 406 Client Subsystem 407 Copy of Video Data, Input Copy 408 Client Subsystem 409 Copy of Video Data 410 Video Decoder 411 Output Stream of Video Picture 412 Display 413 Capture Subsystem 420, 430 Electronic Devices 501 Channel 510 Video Decoder 512 Rendering Device 515 Buffer Memory 520 Parser 521 Symbol 530 Electronic Device 531 Receiver 551 Scaler / Inverse Conversion Unit 552 Intra-Picture Prediction Unit 553 Motion Compensation Prediction Unit 555 Aggregator 556 Loop Filter Unit 557 Reference Picture Memory 558 Current Picture Buffer 601 Video Source 603 Video Encoder 620 Electronic Device 630 Source Coder 632 Coding Engine 633 Local Video Decoder 634 Reference Picture Memory 635 Predictor 640 Transmitter 643 Encoded Video Sequence 645 Entropy Encoder 650 Controller 660 Communication Channel 703 Video Encoder 721 General-Purpose Controller 722 Intra Encoder 723 Residual Calculator 724 Residual Encoder 725 Entropy Encoder 726 Switch 728 Residual Decoder 730 Inter Encoder 810 Video Decoder 871 Entropy Decoder 872 Intra Decoder 873 Residual Decoder 874 Restoration Module 880 Inter Decoder 1900 Computer System 1901 Keyboard 1902 Mouse 1903 Track Pad 1905 Joystick 1906 Microphone 1907 Scanner 1908 Camera 1909 Speaker 1910 Touch Screen 1920 CD / DVD ROM / RW 1921 CD / DVD or Similar Medium 1922 Thumb Drive 1923 Removable Hard Drive or Solid State Drive 1940 Core 1941 Central Processing Unit (CPU) 1942 Graphics Processing Unit (GPU) 1943 Field Programmable Gate Array (FPGA) 1944 Hardware Accelerator 1945 Read Only Memory (ROM) 1946 Random Access Memory (RAM) 1947 On - Chip Large - Capacity Storage 1948 System Bus 1949 Peripheral Bus 1950 Graphics Adapter 1954 Network Interface 1955 Communication Network
Claims
1. A method for video decoding in a video decoder, comprising: receiving a coded video bitstream including a high-level syntax header, the high-level syntax header including multiple transform coefficient block multiple transform selection (MTS) selection information; determining the MTS selection information indicating (i) threshold information or (ii) at least one of a plurality of MTS candidate subsets for a transform coefficient block among the plurality of transform coefficient blocks; determining, based on the MTS selection information, which MTS candidate subset among the plurality of MTS candidate subsets is selected for the transform coefficient block, further comprising determining the MTS candidate subset based on the at least one threshold and one of (i) the number of non-zero coefficients in the transform coefficient block or (ii) the position of the last significant coefficient in the scan order within the transform coefficient block; inverse-transforming the transform coefficient block based on the MTS candidates included in the determined MTS candidate subset; and the MTS selection information indicates the threshold information including at least one threshold. A method.
2. The number of the plurality of MTS candidate subsets is the sum of the number of the at least one threshold and 1, the step of determining which MTS candidate subset, comprises determining the plurality of MTS candidate subsets based on the number of the plurality of MTS candidate subsets, and further comprises determining the MTS candidate subset based on the plurality of MTS candidate subsets. The method according to claim 1.
3. The MTS selection information includes one or more numbers, each of the one or more numbers being the number of one or more MTS candidates in one of one or more MTS candidate subsets among the plurality of MTS candidate subsets, the step of determining which MTS candidate subset, comprises determining the plurality of MTS candidate subsets based on the one or more numbers and the MTS candidate set, and further comprises determining the MTS candidate subset based on the plurality of MTS candidate subsets. The method according to claim 1.
4. The plurality of MTS candidate subsets includes a last MTS candidate subset not included in the one or more MTS candidate subsets, The method according to claim 3, wherein the last MTS candidate subset is the MTS candidate set.
5. The plurality of MTS candidate subsets includes a first MTS candidate subset not included in the one or more MTS candidate subsets, The method according to claim 3, wherein the first MTS candidate subset is composed of default MTS candidates within the MTS candidate set.
6. The method according to claim 1, wherein the high-level syntax header is a slice header, a picture header, a picture parameter set (PPS), a video parameter set (VPS), an adaptive parameter set (APS), or a sequence parameter set (SPS).
7. A method for video decoding in a video decoder, Determining multiple transform selection (MTS) selection information for a plurality of transform coefficient blocks in a coded video bitstream based on a plurality of previously decoded transform coefficient blocks in a previously decoded region, wherein the MTS selection information indicates (i) threshold information or (ii) at least one of a plurality of MTS candidate subsets for a transform coefficient block among the plurality of transform coefficient blocks, Determining which MTS candidate subset among the plurality of MTS candidate subsets is selected for the transform coefficient block based on the MTS selection information; Inverse-transforming the transform coefficient block based on MTS candidates included in the determined MTS candidate subset A method including.
8. The MTS selection information indicates the threshold information including at least one threshold of the plurality of transform coefficient blocks, The step of determining the MTS selection information includes determining the at least one threshold based on coefficient information of the plurality of previously decoded transform coefficient blocks, The method according to claim 7, wherein the step of determining which MTS candidate subset further includes determining the MTS candidate subset based on the at least one threshold and one of (i) the number of non-zero coefficients in the transform coefficient block or (ii) the position of the last significant coefficient in the scan order within the transform coefficient block.
9. The method according to claim 8, wherein the coefficient information of the plurality of previously decoded transform coefficient blocks indicates (i) an average value of non-zero coefficients of the plurality of previously decoded transform coefficient blocks, or (ii) an average position of the last significant coefficient in the scanning order of the plurality of previously decoded transform coefficient blocks.
10. The MTS selection information indicates the threshold information including at least one threshold of the plurality of transform coefficient blocks, A plurality of coefficient information is associated with the plurality of previously decoded transform coefficient blocks, Each of the plurality of coefficient information corresponds to each type of a plurality of types of block sizes within the previously decoded region, The step of determining the MTS selection information includes the step of determining the at least one threshold based on one coefficient information corresponding to each type of block size to which the transform coefficient block belongs, The method according to claim 7, wherein the step of determining which MTS candidate subset further includes the step of determining the MTS candidate subset based on the at least one threshold and one of (i) the number of non-zero coefficients within the transform coefficient block or (ii) the position of the last significant coefficient in the scanning order within the transform coefficient block.
11. The MTS selection information indicates the plurality of MTS candidate subsets, The step of determining the MTS selection information, for the plurality of transform coefficient blocks, determining MTS candidates and an order of the MTS candidates used to form the plurality of MTS candidate subsets from an MTS candidate set based on statistical information of transform types of the plurality of previously decoded transform coefficient blocks; and determining the plurality of MTS candidate subsets from the MTS candidate set based on the MTS candidates and the order of the MTS candidates; including The method according to claim 7, wherein the step of determining which MTS candidate subset includes the step of determining the MTS candidate subset to be one of the plurality of MTS candidate subsets.
12. The MTS selection information indicates the plurality of MTS candidate subsets, A plurality of statistical information of transform types is associated with the plurality of previously decoded transform coefficient blocks, Each of the plurality of statistical information of transform types corresponds to each type of a plurality of types of block sizes within the previously decoded region, The step of determining the MTS selection information includes: For the transform coefficient block, determining MTS candidates and the order of the MTS candidates to be used to form the plurality of MTS candidate subsets from the MTS candidate set based on one statistical information of a transform type corresponding to the block size of the type to which the transform coefficient block belongs; Determining the plurality of MTS candidate subsets from the MTS candidate set based on the MTS candidates and the order of the MTS candidates; and The method according to claim 7, wherein the step of determining which MTS candidate subset includes the step of determining the MTS candidate subset to be one of the plurality of MTS candidate subsets. **Claim 13** An apparatus for video decoding, comprising a processing circuit configured to perform the method according to any one of claims 1 to 6. **Claim 14** An apparatus for video decoding, comprising a processing circuit configured to perform the method according to any one of claims 7 to 12. **Claim 15** A computer program for causing a computer to execute the method according to any one of claims 1 to 6. **Claim 16** A computer program for causing a computer to execute the method according to any one of claims 7 to 12.
Citation Information
Patent Citations
Image coding method based on multiple transform selection and device therefor
US20210211727A1