Context-adaptive transformation set

A neural network-based transform set selection scheme enhances video coding efficiency by adapting transform sets based on neighboring samples and prediction modes, addressing inefficiencies in existing technologies like AV1 and HEVC.

JP7771272B2Active Publication Date: 2025-11-17TENCENT AMERICA LLC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024091076
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-06-22
Filing Date
2024-06-05
Publication Date
2025-11-17
Estimated Expiration
2041-06-29

AI Technical Summary

Technical Problem

Existing video coding technologies, such as AV1 and HEVC, face challenges in optimizing transform set selection for improved compression efficiency, particularly in handling various prediction modes and neighboring sample information.

Method used

A method and system utilizing a neural network-based transform set selection scheme that selects a transform set based on neighboring reconstructed samples and coded information, including inter prediction mode, to enhance image and video compression.

Benefits of technology

Improves compression efficiency by dynamically adapting transform sets based on neighboring sample information, leading to more effective encoding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007771272000015
    Figure 0007771272000015
  • Figure 0007771272000016
    Figure 0007771272000016
  • Figure 0007771272000017
    Figure 0007771272000017
Patent Text Reader

Abstract

To provide a system and a method for coding and decoding a coding bitstream.SOLUTION: A method includes coding a block of an image from a coding bitstream, the coding step includes the steps of selecting a transform set on the basis of at least one adjacent reconstructed sample from one or more previously decoded adjacent blocks or from a previously decoded image, and inverse-transforming coefficients of the block using a transform from the transform set.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 076,817, filed September 10, 2020, and U.S. Provisional Application No. 63 / 077,381, filed September 11, 2020, the disclosures of which are incorporated herein by reference in their entireties.

[0002] TECHNICAL FIELD Embodiments of the present disclosure relate to advanced video coding techniques, and more particularly to primary transform set and secondary transform set selection schemes. [Background technology]

[0003] AOMedia Video 1 (AV1) is an open video coding format designed for video transmission over the Internet. It was developed as a successor to VP9 by the Alliance for Open Media (AOMedia), a consortium founded in 2015 that includes semiconductor companies, video-on-demand providers, video content producers, software developers, and web browser vendors. Many of the AV1 project's components were sourced from previous research efforts by Alliance members. Individual contributors initiated experimental technology platforms several years ago: Xiph / Mozilla's Daala released its code in 2010, Google's experimental VP9 evolution project VP10 was announced on September 12, 2014, and Cisco's Thor on August 11, 2015. Based on the VP9 codebase, AV1 incorporates additional technologies, some of which were developed in these experimental formats. The first version of the AV1 reference codec, version 0.1.0, was published on April 7, 2016. The Alliance announced the release of the AV1 Bitstream Specification, along with reference, software-based encoders and decoders, on March 28, 2018. The validated version 1.0.0 of the specification was released on June 25, 2018. The "AV1 Bitstream & Decoding Process Specification" was released on January 8, 2019, as the validated version 1.0.0 of Errata 1. The AV1 Bitstream Specification includes reference video codecs. "AV1 Bitstream & Decoding Process Specification" (Version 1.0.0 with Errata 1), The Alliance for Open Media (Alliance for Open Media) (January 8, 2019), is incorporated herein by reference in its entirety.

[0004] The High Efficiency Video Coding (HEVC) standard is being jointly developed by the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) standards organizations. To develop the HEVC standard, these two standards organizations collaborate in a partnership known as the Joint Collaborative Team on Video Coding (JCT-VC). The first edition of the HEVC standard was completed in January 2013 and became an aligned text published by both the ITU-T and ISO / IEC. Additional work has since been organized to extend the standard to support several additional application scenarios, including extended range usage with enhanced precision and color format support, scalable video coding, and 3D / stereo / multiview video coding. The HEVC standard is now MPEG-H Part 2 (ISO / IEC 23008-2) at ISO / IEC and ITU-T Recommendation H.265. The HEVC standard "SERIES H: AUDIOVISUAL AND MULTIMEDIA SYSTEMS, Infrastructure of Audiovision Services - Coding of moving video," ITU-T H.265, International Telecommunication Union (April 2015), specifications of which are incorporated herein by reference.

[0005] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (Version 1), 2014 (Version 2), 2015 (Version 3), and 2016 (Version 4). Since then, they have been studying the potential need for standardization of future video coding technologies with significantly higher compression capabilities than HEVC. In October 2017, they jointly issued a call for proposals for video compression technologies with capabilities exceeding those of HEVC (CfP). By February 15, 2018, 22 CfP responses for standard dynamic range (SDR), 12 for high dynamic range (HDR), and 12 for 360 video categories had been submitted. In April 2018, the 122 MPEG / 10 Joint Video Exploration Team-Joint Video Experts Team (JVET) meeting evaluated all of the received CfP responses. After careful evaluation, the JVET formally launched the standardization of next-generation video coding beyond HEVC, namely, the so-called Versatile Video Coding (VVC). The VVC standard, "Versatile Video Coding (Draft 7)," JVET-P2001-vE, Joint Video Experts Team (October 2019), is incorporated herein by reference in its entirety. Another specification of the VVC standard, "Versatile Video Coding (Draft 10)," JVET-S2001-vE, Joint Video Experts Team (July 2020), is incorporated herein by reference in its entirety. Summary of the Invention

[0006] According to an embodiment, a primary and secondary transform set selection scheme using nearby reconstructed samples is provided. According to an embodiment, a neural network based transform set selection scheme for image and video compression is provided.

[0007] According to one or more embodiments, there is provided a method executed by at least one processor, the method including receiving a coded bitstream and decoding a block of an image from the coded bitstream, the decoding including selecting a transform set based on at least one neighboring reconstructed sample from one or more previously decoded neighboring blocks or from a previously decoded image, and inverse transforming coefficients of the block using a transform from the transform set.

[0008] According to one or more embodiments, the step of selecting the transform set is further based on coded information of the prediction mode.

[0009] According to one embodiment, the coded information is of inter prediction mode.

[0010] According to one embodiment, the step of selecting a transform set includes: selecting a subgroup of transform sets from the group of transform sets based on the first coding information; and selecting a transform set from the subgroup.

[0011] According to one embodiment, the step of selecting a transform set from the subgroup includes a step of selecting a transform set based on second coding information, and the method further includes a step of selecting a transform candidate from the transform set based on an index value signaled in the coded bitstream.

[0012] According to one embodiment, the at least one adjacent reconstructed sample comprises a sample reconstructed from one or more previously decoded adjacent blocks.

[0013] According to one embodiment, selecting a transform set comprises selecting a transform set from a group of transform sets, the group of transform sets comprising only secondary transform kernels.

[0014] According to one embodiment, the second transformation kernel is non-separable.

[0015] According to one embodiment, the step of selecting a transformation set is performed by inputting information of at least one nearby reconstruction sample into a neural network and identifying the transformation set based on an index that is output from the neural network.

[0016] According to one embodiment, the set of transformations are quadratic transformations.

[0017] According to one or more embodiments, a system is provided, comprising: at least one memory configured to store computer program code; and at least one processor configured to access the program code and operate as instructed by the computer program code, the computer program code including: decoding code configured to cause the at least one processor to decode a block of an image from a received coded bitstream, the decoding code including: transform set selection code configured to cause the at least one processor to select a transform set based on at least one neighboring reconstructed sample from one or more previously decoded neighboring blocks or from a previously decoded image; and transform code configured to cause the at least one processor to inverse transform coefficients of the block using a transform from the transform set. Includes.

[0018] According to one embodiment, the transform set is further selected based on the coding information of the prediction mode.

[0019] According to one embodiment, the coding information is of inter prediction mode.

[0020] According to one embodiment, the transform set selection code causes the at least one processor to select a subgroup of transform sets from said group of transform sets based on the first coding information and to select a transform set from the subgroup.

[0021] According to one embodiment, the transform set selection code is configured to cause the at least one processor to select a transform set based on the second coding information, and the decoding code further includes a transform selection code configured to cause the at least one processor to select a transform candidate from the transform set based on an index value signaled in the coded bitstream.

[0022] According to one embodiment, the at least one adjacent reconstructed sample comprises a sample reconstructed from one or more previously decoded adjacent blocks.

[0023] According to one embodiment, the transform set selection code is configured to select a transform set from a group of transform sets, the group of transform sets including only secondary transform kernels.

[0024] According to one embodiment, the second transformation kernel is non-separable.

[0025] According to one embodiment, the transform set selection code is configured to cause at least one processor to input information of at least one adjacent reconstruction sample to a neural network and identify a transform set based on an index output from the neural network.

[0026] According to one or more embodiments, a non-transitory computer-readable medium is provided that stores computer instructions that, when executed by at least one processor, cause the at least one processor to decode a block of an image from a received coding bitstream by: selecting a transform set based on at least one neighboring reconstructed sample from one or more previously decoded neighboring blocks or from a previously decoded image; and inverse transforming coefficients of the block using a transform from the transform set. [Brief explanation of the drawings]

[0027] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Figure 1] FIG. 1 is a diagram illustrating a simplified block diagram of a communication system according to one embodiment. [Figure 2] FIG. 2 is a diagram illustrating a simplified block diagram of a communication system according to one embodiment. [Figure 3] FIG. 3 is a schematic diagram illustrating a simplified block diagram of a decoder according to one embodiment. [Figure 4] FIG. 4 is a schematic diagram illustrating a simplified block diagram of an encoder according to one embodiment. [Figure 5A] FIG. 5A is a diagram showing a first exemplary partition structure of VP9. [Figure 5B] FIG. 5B is a diagram illustrating a second exemplary partition structure of VP9. [Figure 5C] FIG. 5C is a diagram illustrating a third exemplary partition structure of VP9. [Figure 5D] FIG. 5D is a diagram illustrating a fourth exemplary partition structure of VP9. [Figure 6A] FIG. 6A shows a first exemplary partition structure of AV1. [Figure 6B] FIG. 6B shows a second exemplary partition structure of AV1. [Figure 6C] FIG. 6C shows a third exemplary partition structure of AV1. [Figure 6D] FIG. 6D shows a fourth exemplary partition structure of AV1. [Figure 6E] FIG. 6E is a diagram showing a fifth exemplary partition structure of AV1. [Figure 6F] FIG. 6F shows a sixth exemplary partition structure of AV1. [Figure 6G] FIG. 6G shows a seventh exemplary partition structure of AV1. [Figure 6H] FIG. 6H shows an eighth exemplary partition structure of AV1. [Figure 6I] FIG. 6I shows a ninth exemplary partition structure of AV1. [Figure 6J] FIG. 6J is a diagram showing a tenth exemplary partition structure of AV1. [Figure 7] FIG. 7 shows eight nominal angles for AV1. [Figure 8] FIG. 8 is a diagram showing the current block and samples. [Figure 9] FIG. 9 is a diagram illustrating an example of the recursive intra-filtering mode. [Figure 10] FIG. 10 is a diagram showing reference lines adjacent to a coding block unit. [Figure 11] Figure 11 is a table of AV1 hybrid transform kernels and their availability. [Figure 12] FIG. 12 is a diagram of a low frequency non-separable transformation process. [Figure 13] FIG. 13 is an explanatory diagram of a matrix. [Figure 14] FIG. 14 is a diagram for explaining two-dimensional convolution of a kernel and an image. [Figure 15] FIG. 15 is a diagram illustrating max pooling of image patches. [Figure 16A]FIG. 16A illustrates the first intra-decoding process. [Figure 16B] FIG. 16B illustrates a second intra-decoding process. [Figure 17] FIG. 17 illustrates an example of a convolutional neural network filter architecture. [Figure 18] FIG. 18 illustrates an example of a convolutional neural network filter architecture. [Figure 19] FIG. 19 is a diagram illustrating an example of a high-density residual unit structure. [Figure 20] FIG. 20 is a diagram showing the first process. [Figure 21] FIG. 21 is a diagram illustrating the second process. [Figure 22] FIG. 22 is a table of mappings from intra prediction modes to transform set indices. [Figure 23A] FIG. 23A is an explanatory diagram of a first residual pattern according to a comparative example. [Figure 23B] FIG. 23B is an explanatory diagram of a second residual pattern according to a comparative example. [Figure 23C] FIG. 23C is an explanatory diagram of a third residual pattern according to a comparative example. [Figure 23D] FIG. 23D is an explanatory diagram of a fourth residual pattern according to a comparative example. [Figure 24] FIG. 24 is a schematic diagram of a decoder according to one embodiment of the present disclosure. [Figure 25] FIG. 25 is a diagram of a computer system suitable for implementing embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0028] In this disclosure, the term block may be interpreted as a prediction block, a coding block, or a coding unit (CU). The term "block" here may also be used to refer to a transform block.

[0029] In this disclosure, the term "transform set" refers to a group of transform kernel (or candidate) options. A transform set may include one or more transform kernel (or candidate) options. According to an embodiment of the present disclosure, if one or more transform options are available, an index may be signaled to indicate which of the transform options in the transform set applies to the current block.

[0030] In this disclosure, the term "prediction mode set" refers to a group of prediction mode options. A prediction mode set may include multiple prediction mode options. According to an embodiment of the present disclosure, when multiple prediction mode options are available, an index may be further signaled to indicate which one of the prediction mode options in the prediction mode set is applied to the current block to perform prediction.

[0031] In this disclosure, the term "neighboring reconstructed sample sets" refers to samples reconstructed from neighboring previously decoded blocks or groups of reconstructed samples within a previously decoded image.

[0032] In this disclosure, the term "neural network" refers to the general concept of a data processing structure having one or more layers, as described herein with respect to "Deep Learning for Video Coding." According to embodiments of the present disclosure, any neural network may be configured to implement the embodiments.

[0033] FIG. 1 shows a simplified block diagram of a communication system (100) according to one embodiment of the present disclosure. The system (100) may include at least two terminals (110, 120) interconnected via a network (150). For unidirectional data transmission, a first terminal (110) may code video data locally for transmission to the other terminal (120) via the network (150). The second terminal (120) may receive the other terminal's coded video data from the network (150), decode the coded data, and display the recovered video data. Unidirectional data transmission may be common in media delivery applications, etc.

[0034] 1 illustrates a second pair of terminals (130, 140) configured to support bidirectional transmission of coded video, such as may occur during a video conference. For bidirectional transmission of data, each terminal (130, 140) may code video data captured at a local location for transmission to the other terminal over a network (150). Each terminal (130, 140) may also receive coded video data transmitted by the other terminal, decode the coded video data, and display the recovered video data on a local display device.

[0035] In FIG. 1 , the terminals (110-140) may be depicted as servers, personal computers, smartphones, and / or any other type of terminal. For example, the terminals (110-140) may be laptop computers, tablet computers, media players, and / or dedicated videoconferencing devices. The network (150) represents any number of networks that convey coded video data between the terminals (110-140), including, for example, wired and / or wireless communication networks. The communication network (150) may exchange data within circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this description, the architecture and topology of the network (150) are not important to the operation of the present invention, unless otherwise described below.

[0036] 2 illustrates the placement of a video encoder and decoder in a streaming environment as an example of an application of the disclosed subject matter. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, and storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0037] As shown in FIG. 2 , the streaming system (200) may include a capture subsystem (213) that may include a video source (201) and an encoder (203). The video source (201) may be, for example, a digital camera and may be configured to generate an uncompressed video sample stream (202). The uncompressed video sample stream (202) may provide a high data volume when compared to an encoded video bitstream and may be processed by an encoder (203) coupled to the camera (201). The encoder (203) may include hardware, software, or a combination thereof, and may enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video bitstream (204) may include a lower data volume when compared to the sample stream and may be stored on a streaming server (205) for future use. One or more streaming clients (206) may access the streaming server (205) to retrieve a video bitstream (209), which may be a copy of the encoded video bitstream (204).

[0038] In embodiments, the streaming server (205) may also function as a media-aware network element (MANE). For example, the streaming server (205) may be configured to prune the encoded video bitstream (204) to tailor potentially different bitstreams to one or more streaming clients (206). In embodiments, a MANE may be provided separately from the streaming server (205) within the streaming system (200).

[0039] The streaming client (206) may include a video decoder (210) and a display (212). The video decoder (210) may, for example, decode a video bitstream (209), which may be a received copy of the encoded video bitstream (204), and generate a transmitted video sample stream (211) that may be rendered on a display (212) or other rendering device (not shown). In some streaming systems, the video bitstreams (204, 209) may be encoded according to a particular video coding / compression standard. Examples of such standards include, but are not limited to, ITU-T Recommendation H.265. A video coding standard known as Versatile Video Coding (VCC) is under development. Embodiments of the present disclosure may be used in connection with VVC.

[0040] FIG. 3 illustrates an example functional block diagram of a video decoder (210) attached to a display (212) according to one embodiment of the present disclosure.

[0041] The video decoder (210) may include a channel (312), a receiver (310), a buffer memory (315), an entropy decoder / parser (320), a scalar / inverse transform unit (351), an intra prediction unit (352), a motion compensated prediction unit (353), an aggregator (355), a loop filter unit (356), a reference picture memory (357), and a current picture memory (357). In at least one embodiment, the video decoder (210) may include an integrated circuit, a series of integrated circuits, and / or other electronic circuitry. The video decoder (210) may also be implemented partially or completely in software running on one or more CPUs with associated memory.

[0042] In this and other embodiments, the receiver (310) can receive one or more coded video sequences, one of which is to be decoded by the decoder (210), with the decoding of each coded video sequence being independent of the other coded video sequences. The coded video sequences can be received from a channel (312), which can be a hardware / software link to a storage device that stores the encoded video data. The receiver (310) can receive the encoded video data along with other data, such as coded audio data and / or auxiliary data streams, which can be transferred using respective entities (not shown). The receiver (310) can separate the coded video sequences from other data. To combat network jitter, a buffer memory (315) can be coupled between the receiver (310) and the entropy decoder / parser (320) (hereinafter "parser"). If the receiver 310 is receiving data from a store-and-forward device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer 315 may be unused or small. For use on best-effort packet networks such as the Internet, the buffer 315 may be required and may be relatively large and adaptively sized.

[0043] The video decoder (210) may include a parser (320) for reconstructing symbols (321) from the coded video sequence. These symbol categories include, for example, information used to manage the operation of the decoder (210) and potential information for controlling a rendering device, such as a display (212), which may be coupled to the decoder as shown in FIG. 2. The control information for the rendering device(s) may be in the form of a supplemental enhancement information (SEI) message or a video usability information (VUI) parameter set fragment (not shown). The parser (320) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow principles including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (320) may extract a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroups may include groups of pictures (GOPs), images, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), and the like.

[0044] The parser (320) may also extract information from the coded video sequence, such as transform coefficients, quantization parameter values, motion vectors, etc. The parser (320) may perform entropy decoding / parsing operations on the video sequence received from the buffer (315) to generate symbols (321).

[0045] The reconstruction of the symbols (321) can include multiple different units, depending on the type of video image or portion thereof being coded (e.g., inter- and intra-image, inter- and intra-block) and other factors. Which units are included and how can be controlled by subgroup control information parsed from the coded video sequence by the parser (320). The flow of such subgroup control information between the parser (320) and the following units is not shown for clarity.

[0046] In addition to the functional blocks already mentioned, the decoder (210) can be conceptually divided into several functional units, as described below. In a practical implementation operating under commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, a conceptual subdivision into the following functional units is appropriate:

[0047] One unit may be a scalar / inverse transform unit (351), which may receive quantized transform coefficients as symbol(s) (321) from the parser (320) as well as control information including the transform used, block size, quantization coefficients, quantization scaling matrix, etc. The scalar / inverse transform unit (351) may output blocks containing sample values ​​that may be input to an aggregator (355).

[0048] In some cases, the output samples of the scaler / inverse transform unit (351) may relate to intra-coded blocks; i.e., blocks that do not use prediction information from a previously reconstructed image, but can use prediction information from a previously reconstructed portion of the current image. Such prediction information may be provided by an intra-image prediction unit (352). In some cases, the intra-image prediction unit (352) generates blocks of the same size and shape as the block being reconstructed using surrounding already reconstructed information fetched from the (partially reconstructed) current image from a current image memory (358). The aggregator (355) may add, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351).

[0049] In other cases, the output samples of the scalar / inverse transform unit (351) may relate to inter-coding and potentially to a motion-compensated block. In such cases, the motion-compensated prediction unit (353) may access a reference picture memory (357) to fetch samples used for prediction. After motion-compensating the fetched samples according to the symbols (321) associated with the block, these samples may be added by the aggregator (355) to the output of the scalar / inverse transform unit (351) (referred to as residual samples or residual signals in this case) to generate output sample information. The addresses in the reference picture memory (357) from which the motion-compensated prediction unit (353) fetches prediction samples may be controlled by a motion vector. The motion vector may be available to the motion-compensated prediction unit (353) in the form of a symbol (321), which may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolating sample values ​​as fetched from the reference picture memory (357), motion vector prediction mechanisms, etc., if sub-sample accurate motion vectors are used.

[0050] The output samples of the aggregator (355) can be subjected to various loop filtering techniques in the loop filter unit (356). Video compression techniques can include in-loop filtering techniques controlled by parameters contained in the coded video bitstream and made available to the loop filter unit (356) as symbols (321) from the parser (320), but can also be responsive to meta-information obtained during a previous partial decoding (in decode order) of the coded image or coded video sequence, or to previously reconstructed and loop-filtered sample values.

[0051] The output of the loop filter unit (356) can be a sample stream that can be output to a rendering device such as a display (212) or stored in a reference image memory (357) for use in future inter-image prediction.

[0052] Once a coded picture is fully reconstructed, it can be used as a reference picture for future predictions. Once a coded picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (320)), the current reference picture becomes part of the reference picture memory (357), and a new current picture memory can be reallocated before beginning reconstruction of a subsequent coded picture.

[0053] The video decoder (210) may perform decoding operations according to a predetermined video compression technology, which may be documented in a standard such as ITU-T Rec. 265. The coded video sequence may conform to the syntax specified by the video compression technology or standard being used, in the sense that it conforms to the syntax of the video compression technology or standard as specified in the video compression technology document or standard, particularly the profile document therein. To comply with some video compression technologies or standards, the complexity of the coded video sequence may also be within a range defined by the level of the video compression technology or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further constrained by the specification of a hypothetical reference decoder (HRD) and HRD buffer management metadata signaled in the coded video sequence.

[0054] In one embodiment, the receiver (310) can receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence(s). The additional data may be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant images, forward error correction codes, etc.

[0055] FIG. 4 illustrates an example functional block diagram of a video encoder (203) associated with a video source (201) according to one embodiment of the present disclosure.

[0056] The video encoder (203) may include an encoder, for example, a source coder (430), a coding engine (432), a (local) decoder (433), a reference image memory (434), a predictor (435), a transmitter (440), an entropy coder (445), a controller (450), and a channel (460).

[0057] The encoder (203) can receive video samples from a video source (201) (not part of the encoder) that can capture video images to be coded by the encoder (203).

[0058] The video source (201) can provide a source video sequence to be coded by the encoder (203) in the form of a digital video sample stream, which can be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media distribution system, the video source (201) can be a storage device that stores prepared video. In a video conferencing system, the video source (203) can be a camera that captures local video information as a video sequence. The video data can be provided as multiple individual images that, when viewed in sequence, create motion. The images themselves can be organized as a spatial array of pixels, each of which can contain one or more samples, depending on the sampling structure, color space, etc., in use. Those skilled in the art can readily understand the relationship between pixels and samples. The following discussion focuses on samples.

[0059] According to one embodiment, the encoder (203) can code and compress images of a source video sequence into a coded video sequence (443) in real time or under any other time constraint required by the application. Enforcing an appropriate coding rate is one function of the controller (450). The controller (450) can also control and be functionally coupled to other functional units, as described below. Couplings are not shown for clarity. Parameters set by the controller (450) can include rate control-related parameters (picture skip, quantizer, lambda value for rate-distortion optimization techniques, ...), picture size, picture group layout, maximum motion vector search range, etc. Those skilled in the art can readily identify other functions of the controller (450) as they may be relevant to optimizing the video encoder (203) for a particular system design.

[0060] Some video encoders operate in what those skilled in the art would readily recognize as a "coding loop." As an oversimplified explanation, the coding loop consists of the encoding portion of the source coder (430) (responsible for generating symbols based on the input picture to be coded and one or more reference pictures) and a (local) decoder (433) embedded in the encoder (203), which reconstructs the symbols and generates the sample data that a (remote) decoder would also generate if the compression between the symbols and the coded video bitstream were lossless in a particular video compression technology. The reconstructed sample stream is input to a reference picture memory (434). Because decoding the symbol stream yields bit-exact results independent of the decoder location (local or remote), the reference picture memory contents are also bit-exact between the local and remote encoders. In other words, the predictive portion of the encoder "sees" the exact same sample values ​​as the reference picture samples that the decoder "sees" when using prediction during decoding. This basic principle of reference image synchrony (and the resulting drift when synchrony cannot be maintained, for example due to channel errors) is known to those skilled in the art.

[0061] The operation of the "local" decoder (433) may be the same as the operation of the "remote" decoder (210), as already described in detail above in connection with Figure 3. However, because symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder (445) and parser (320) is lossless, the entropy decoding portion of the decoder (210), including the channel (312), receiver (310), buffer (315), and parser (320), may not be fully implemented in the local decoder (433).

[0062] An observation that can be made in this regard is that any decoder technology, except for parsing / entropy decoding, present in the decoder may need to be present in a substantially identical functional form in the corresponding encoder. For this reason, the disclosed subject matter focuses on decoder operation. A description of the encoder technology may be omitted, as it may be the inverse of the decoder technology described overall. Only in certain areas is a more detailed explanation necessary, which is provided below.

[0063] As part of its operation, the source coder (430) can perform motion-compensated predictive coding, which codes an input frame with reference to one or more previously coded frames from the video sequence designated as “reference frames.” In this manner, the coding engine (432) codes differences between pixel blocks of the input frame and pixel blocks of one or more reference frames that can be selected as the predictive reference(s) for the input frame.

[0064] The local video decoder (433) may decode the coded video data of a frame that may be designated as a reference image based on the symbols generated by the source coder (430). The operation of the coding engine (432) may advantageously be a lossy process. If the coded video data can be decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence may typically be a replica of the source video sequence, possibly with some errors. The local video decoder (433) repeats the decoding process performed by the video decoder on the reference frame, which may result in a reconstructed reference frame to be stored in the reference image memory (434). In this way, the encoder (203) can locally store a copy of the reconstructed reference frame that has common content with the reconstructed reference frame that would be obtained by the far-end video decoder (without transmission errors).

[0065] The predictor (435) may perform a prediction search for the coding engine (432). That is, for a new frame to be coded, the predictor (435) may search the reference picture memory (434) for sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture motion vectors, block shapes, etc., that may serve as suitable prediction references for the new picture. The predictor (435) may operate on a sample block-by-sample block basis to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (435), the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (434).

[0066] The controller (450) may manage the coding operations of the video coder (430), including, for example, setting the parameters and subgroup parameters used to encode the video data.

[0067] The outputs of all of the above-mentioned functional units may undergo entropy coding in an entropy coder (445), which converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.

[0068] The transmitter (440) may buffer the coded video sequence(s) created by the entropy coder (445) for transmission over a communication channel (460), which may be a hardware / software link, to a storage device that may store the encoded video data. The transmitter (440) may merge the coded video data from the video coder (430) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (not shown).

[0069] The controller (450) may manage the operation of the encoder (203). During coding, the controller (450) may assign a particular coded picture type to each coded picture, which may affect the coding technique that may be applied to each picture. For example, pictures are often assigned as intra-pictures (I-pictures), predicted pictures (P-pictures), or bidirectionally predicted pictures (B-pictures).

[0070] An intra-picture (I-picture) may be one that can be coded and decoded without using any other frame of the sequence as a source of prediction. Some video codecs allow different types of intra-pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of these variations of I-pictures and their respective uses and characteristics.

[0071] A predicted image (P-image) may be coded and decoded using inter- or intra-prediction, which uses at most one motion vector and reference index to predict the sample values ​​of each block.

[0072] Bi-directionally predicted images (B-pictures) may be coded and decoded using inter- or intra-prediction, which uses up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predicted images may use two or more reference images and associated metadata for the reconstruction of a block.

[0073] A source image is typically spatially divided into multiple sample blocks (e.g., 4x4, 8x8, 4x8, or 16x16 blocks of samples each) and coded block by block. Blocks can be predictively coded with reference to other (already coded) blocks, determined by the coding assignment applied to each image of the block. For example, blocks of an I image can be nonpredictively coded, or they can be predictively coded with reference to already coded blocks of the same image (spatial prediction or intra-prediction). Pixel blocks of a P image can be nonpredictively coded via spatial prediction or via temporal prediction with respect to one previously coded reference image. Blocks of a B image can be nonpredictively coded via spatial prediction with reference to one or two previously coded reference images, or via temporal prediction.

[0074] The video encoder (203) may perform coding operations according to a predetermined video coding technique or standard, such as ITU-T Rec. H.265. In doing so, the video coder (203) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.

[0075] In one embodiment, the transmitter (440) can transmit additional data along with the encoded video. The source coder (430) can include such data as part of the coded video sequence. The additional data can include temporal, spatial, and SNR enhancement layers, as well as other types of redundant data, such as redundant images and slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, etc.

[0076] [VP9 and AV1 coding block partitions]

[0077] Referring to the partition structures (502)-(508) in Figures 5A-D, VP9 uses a 4-way partition tree from the 64x64 level down to the 4x4 level, with some additional restrictions on the 8x8 block. Note that the partitions shown as R in Figure 5D refer to recursion, in that the same partition tree is repeated at lower scales until the lowest 4x4 level is reached.

[0078] See the partition structures (511) to (520) in Figures 6A to 6J. AV1 not only extends the partition tree to a 10-way structure, but also extends the maximum size (called a superblock in VP9 / AV1 terminology) to start at 128x128. Note that this includes 4:1 / 1:4 rectangular partitions, which did not exist in VP9. As shown in Figures 6C to 6F, a partition type with three subpartitions is called a "T-shaped" partition. Rectangular partitions cannot be further subdivided. In addition to the size of the coding block, a coding tree depth can be defined to indicate the division depth from the root node. Specifically, the coding tree depth of the root node, e.g., 128x128, is set to 0, and after further dividing the tree block, the coding tree depth is incremented by 1.

[0079] Instead of enforcing a fixed transform unit size as in VP9, ​​AV1 allows the partitioning of luma coding blocks into transform units of multiple sizes, which can be expressed as recursive partitions up to two levels down. To incorporate AV1's extended coding block partitioning, square, 2:1 / 1:2, and 4:1 / 1:4 size transforms from 4x4 to 64x64 can be supported. For chroma blocks, only the largest possible transform units are allowed.

[0080] [HEVC block partitioning]

[0081] In HEVC, coding tree units (CTUs) can be divided into coding units (CUs) by using a quadtree (QT) structure, denoted as a coding tree, to adapt to various local characteristics. The decision of whether to code an image region using inter-image (temporal) prediction or intra-image (spatial) prediction can be made at the CU level. Each CU can be further divided into one, two, or four prediction units (PUs) according to a PU partition type. The same prediction process is applied within a PU, and related information is transmitted to the decoder on a PU-by-PU basis. After applying the prediction process based on the PU partition type to obtain residual blocks, the CU can be divided into transform units (TUs) according to another quadtree structure, such as the CU's coding tree. One of the key features of the HEVC structure is its multiple partition concept, including CUs, PUs, and TUs. In HEVC, a CU or TU can only have a square shape, while a PU can have a square or rectangular shape for inter-prediction blocks. In HEVC, a coding block can be further divided into four square sub-blocks, and a transform can be performed on each sub-block (i.e., TU). Each TU can be further recursively divided (using quadtree division) into smaller TUs, which are called residual quadtrees (RQTs).

[0082] At image boundaries, HEVC employs implicit quad-tree splitting to maintain the quad-tree split until the block is sized to fit the image boundary.

[0083] [Quadtree with nested multi-type tree coding block structure in VVC]

[0084] In VVC, a quadtree with nested multitype trees using a binary-ternary segmentation structure replaces the concept of multiple partition unit types. That is, VVC does not include a separation of the concepts of CU, PU, ​​and TU, except when necessary for CUs whose size is too large for the maximum transform length, improving flexibility in CU partition shape. In the coding tree structure, CUs can be either square or rectangular. Coding tree units (CTUs) are first partitioned using a quaternary tree (also known as a quadtree) structure. The quaternary tree leaf nodes can then be further partitioned using a multitype tree structure. There are four multitype tree structures: vertical bisection (SPLIT_BT_VER), horizontal bisection (SPLIT_BT_HOR), vertical trisection (SPLIT_TT_VER), and horizontal trisection (SPLIT_TT_HOR). The multitype tree leaf nodes are sometimes called coding units (CUs), and as long as the CUs are not too large relative to the maximum transform length, this segmentation can be used for prediction and transform processing without further division. This means that in most cases, CUs, PUs, and TUs have the same block size in a quadtree with a nested multitype tree coding block structure. An exception occurs when the maximum supported transform length is smaller than the width or height of the color components of the CU. An example of block partitioning is when a CTU is divided into multiple CUs with a quadtree and nested multitype tree coding block structure, and with quadtree partitions and multitype tree partitions. The quadtree with nested multitype tree partitions provides a content adaptive coding tree structure composed of CUs.

[0085] In VVC, the maximum supported luma transform size is 64x64, and the maximum supported chroma transform size is 32x32. If the width or height of a CB is larger than the maximum transform width or height, the CB is automatically split horizontally and / or vertically to meet the transform size constraint in that direction.

[0086] In VTM7, the coding tree method supports the ability for luma and chroma to have separate block tree structures. For P and B slices, the luma and chroma CTBs of one CTU may have to share the same coding tree structure. However, for I slices, luma and chroma can have separate block tree structures. When applying the separate block tree mode, the luma CTB is partitioned into CUs by one coding tree structure, and the chroma CTB is partitioned into chroma CUs by another coding tree structure. This means that a CU in an I slice can contain a coding block for the luma component or a coding block for two chroma components, and a CU in a P slice or B slice can contain coding blocks for all three color components unless the video is monochrome.

[0087] [Directional Intra Prediction in AV1]

[0088] VP9 supports eight directional modes, corresponding to angles from 45 degrees to 207 degrees. AV1 extends the directional intra modes to finer granularity angles to take advantage of the greater spatial redundancy in directional textures. The original eight angles are slightly modified to create nominal angles, designated V_PRED (542), H_PRED (543), D45_PRED (544), D135_PRED (545), D113_PRED (546), D157_PRED (547), D203_PRED (548), and D67_PRED (549), as shown in Figure 7 for the current block (541). For each nominal angle, there are seven finer angles, for a total of 56 directional angles in AV1. The prediction angle is the nominal intra-angle plus an angle delta, which is -3 to 3 times the step size of 3 degrees. In AV1, eight nominal modes and five non-angle-smooth modes are first signaled. Then, if the current mode is an angle mode, an index is further signaled to indicate the angle delta relative to the corresponding nominal angle. To implement directional prediction modes in AV1 in a generic way, all 56 directional intra-prediction modes in AV1 are implemented with a unified directional predictor that projects each pixel to a reference sub-pixel position and interpolates the reference pixel using a two-tap bilinear filter.

[0089] [Non-directional smooth intra predictor in AV1]

[0090] AV1 has five non-directional smooth intra prediction modes: DC, PAETH, SMOOOTH, SMOOTH_V, and SMOOTH_H. For DC prediction, the average of the neighboring samples above and to the left is used as the predictor for the block to be predicted. The PAETH predictor first takes the top, left, and top-left reference samples, and then sets the closest (top + left - left) value as the predictor for the pixel to be predicted. Figure 8 shows the locations of the top sample (554), left sample (556), and top-left sample (558) relative to a pixel (552) in the current block (550). In SMOOTH, SMOOTH_V, and SMOOOTH_H modes, the current block (550) is predicted using quadratic interpolation in the vertical or horizontal direction, or an average in both directions.

[0091] [Recursive filtering-based intra-predictor]

[0092] To capture the attenuated spatial correlation due to edge references, filter intra modes are designed for luma blocks. Five filter intra modes are defined in AV1, each represented by a set of eight 7-tap filters that reflect the correlation between a pixel in a 4x2 patch and its seven neighbors. In other words, the weighting coefficients of the 7-tap filters depend on the position. For example, as shown in Figure 9, an 8x8 block (560) can be divided into 8x42 patches. These patches are denoted as B0, B1, B2, B3, B4, B5, B6, and B7 in Figure 9. For each patch, seven neighbors, denoted R0 through R6, can be used to predict pixels in the current patch. In patch B0, all neighbors may already be reconstructed. However, in other patches, some neighbors may not be reconstructed, and the predicted values ​​of the neighbors are used as references. For example, since not all neighbors in patch B7 are reconstructed, the predicted samples of the neighbors are used instead.

[0093] [Chroma predicted from Luma]

[0094] Chroma from Luma (CfL) is a chroma-only intra predictor that models chroma pixels as a linear function of the co-reconstructed luma pixels. CfL prediction can be expressed in equation (1) as: CfL(α)=α×L AC +DC (formula 1) Here, LAC denotes the AC contribution of the luma component, α denotes a parameter of the linear model, and DC denotes the DC contribution of the chroma component. Specifically, the reconstructed luma pixels are subsampled to the chroma resolution and then averaged to form the AC contribution. Instead of requiring the decoder to calculate scaling parameters to approximate the chroma AC components from the AC contributions, as in some background techniques, AV1 CfL can determine the parameter α based on the original chroma pixels and signal them in the bitstream. This reduces decoder complexity and results in more accurate predictions. The DC contribution of the chroma components can be calculated using an intra-DC mode, which is sufficient for most chroma content and has a mature, fast implementation.

[0095] [Multi-line intra prediction]

[0096] Multi-line intra prediction can use more reference lines for intra prediction, and the encoder determines and signals which reference lines are used to generate the intra predictor. The reference line index can be signaled before the intra prediction mode, and if a non-zero reference line index is signaled, only the most probable mode can be allowed. Figure 10 shows an example of four reference lines (570), each consisting of six segments, i.e., segments A through F, with a reference sample at the top left. Furthermore, segments A and F are packed with the nearest samples from segments B and E, respectively.

[0097] [AV1 primary conversion]

[0098] To support extended coding block partitioning, multiple transform sizes (e.g., ranging from 4 points to 64 points for each dimension) and transform shapes (e.g., square; rectangle with width / height ratios of 2:1 / 1:2 and 4:1 / 1:4) are introduced into AV1.

[0099] The 2D transform process can include the use of hybrid transform kernels (e.g., composed of a different one-dimensional (1D) transform for each dimension of the coding residual block). According to one embodiment, the primary 1D transforms are (a) a 4-point, 8-point, 16-point, 32-point, or 64-point DCT-2; (b) a 4-point, 8-point, or 16-point asymmetric DST (DST-4, DST-7) and their inverse versions; and (c) a 4-point, 8-point, 16-point, or 32-point discriminant transform. The basis functions of the DCT-2 and asymmetric DST used in AV1 are listed in Table 1 below. Table 1 shows the AV1 primary transform basis functions DCT-2, DST-4, and DST-7 for an N-point input. [Table 1]

[0100] The availability of hybrid transform kernels can be based on transform block size and prediction mode. This dependency is listed in table 580 of FIG. 11. Table 580 shows AV1 hybrid transform kernels and their availability based on prediction mode and block size. In table 580, the symbols "→" and "↓" indicate the horizontal and vertical dimensions, respectively, and "R" and "X" indicate the availability and unavailability of the kernel for that block size and prediction mode, respectively.

[0101] For chroma components, the selection of the transform type can be done implicitly. For intra-prediction residuals, the transform type can be selected according to the intra-prediction mode, as specified in Table 2 below. For inter-prediction residuals, the transform type can be selected according to the transform type selection of the co-located luma block. Therefore, for chroma components, there may be no transform type to signal in the bitstream. [Table 2]

[0102] [Secondary transformation in VVC]

[0103] Referring to Figure 12, in VVC, a low-frequency non-separable transform (LFNST), also known as a contractive quadratic transform, can be applied between the forward linear transform (591) and quantization (593) (at the encoder) and between the dequantization (594) and the inverse linear transform (596) (at the decoder) to further decorrelate the linear transform coefficients. For example, a forward LFNST (592) can be applied by the encoder, and an inverse LFNST (595) can be applied by the decoder. Depending on the block size, a 4x4 non-separable transform or an 8x8 non-separable transform can be applied in LFNST. For example, a 4x4 LFNST can be applied to small blocks (i.e., min(width, height) < 8), and an 8x8 LFNST can be applied to larger blocks (i.e., min(width, height) > 4). For a 4x4 forward LFNST and an 8x8 forward LFNST, the forward LFNST (592) can have 16 and 64 input coefficients, respectively. For a 4x4 inverse LFNST and an 8x8 inverse LFNST, the inverse LFNST (595) can have 8 and 16 input coefficients, respectively.

[0104] The application of the non-separable transform used in LFNST is described as follows using the input as an example: To apply a 4x4 LFNST, a 4x4 input block X shown in equation (2) below is first transformed into a vector x as shown in equation (3) below. This can be represented as JPEG0007771272000003.jpg1417 (hereafter referred to as X̂):

number

number

[0105] A non-separable transformation is

number

[0106] A. Reduced Non-Separable Transform

[0107] LFNST can be based on a direct matrix multiplication approach and applies a non-separable transform so that it can be implemented in a single pass without multiple iterations. However, the non-separable transform matrix dimension can be reduced or minimized, minimizing the computational complexity and the memory space for storing the transform coefficients. Therefore, the reduced non-separable transform (RST) method can be used in LFNST. The main idea of the reduced non-separable transform is to map an N-dimensional vector (where N is typically equal to 64 in an 8×8 NSST) to an R-dimensional vector in a different space. Here, N / R (R < N) is the reduction factor. Thus, instead of an N×N matrix, the RST matrix becomes an R×N matrix (600), as shown in FIG. 13.

[0108] The RxN matrix (600) has R rows of the transform, which is an R basis in N-dimensional space. The inverse transform matrix of the RT can be the transpose of its forward transform. For an 8x8 LFNST, a reduction factor of 4 can be applied, and the traditional 64x64 direct matrix, which is the size of a non-separable transform matrix for 8x8, can be reduced to a 16x48 direct matrix. Therefore, a 48x16 inverse RST matrix can be used at the decoder side to generate the core (primary) transform coefficients in the 8x8 upper-left region. The 16x48 matrix is ​​applied instead of the 16x64 with the same transform set configuration, and each matrix can take 48 input data from three 4x4 blocks in the upper-left 8x8 block, except for the lower-right 4x4 block. With the help of the reduced dimension, the memory usage for storing all LFNST matrices can be reduced from 10KB to 8KB with a reasonable performance degradation. To reduce complexity, LFNST can be restricted to apply only when all coefficients outside the first coefficient subgroup are insignificant. Thus, when LFNST is applied, all primary transform coefficients may have to be zero. This allows LFNST index signaling to be conditioned on the last-significant position, thus avoiding the extra coefficient scanning in current LFNST designs that may be required to check significant coefficients only at certain positions. LFNST's worst-case processing (in terms of multiplications per pixel) limits non-separable transforms of 4x4 and 8x8 blocks to 8x16 and 8x48 transforms, respectively. In such cases, the last valid scan position when applying LFNST may have to be less than 8 for other sizes less than 16. For blocks with shapes of 4xN and Nx4 and N>8, the restriction may mean that LFNST is currently applied only once, and only to the top-left 4x4 region. When LFNST is applied, all first-order-only coefficients can be zero; in such a case, the number of operations for the first-order transform can be reduced. From the encoder's perspective, the quantization of coefficients is significantly simplified when the LFNST transform is tested.Rate-distortion optimal quantization can be performed on at most the first 16 coefficients (in scan order), and the remaining coefficients can be forced to be zero.

[0109] B.LFNST conversion selection

[0110] For each transform set used in LFNST, there can be four transform sets and two non-separable transform matrices (kernels). The mapping from intra prediction modes to transform sets can be predefined as shown in Table 3 below. If any of the three CCLM modes (INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the current block (81<=predModeIntra=83), transform set 0 can be selected for the current chroma block. For each transform set, the selected non-separable secondary transform candidate can be further specified by an explicitly signaled LFNST index. The index can be signaled in the bitstream once per intra CU after the transform coefficients. [Table 3]

[0111] C. LFNST Index Signaling and Interaction with Other Tools

[0112] Because LFNST can be restricted to be applicable only when all coefficients outside the first coefficient subgroup are insignificant, LFNST index coding may depend on the position of the last significant coefficient. In addition, the LFNST index may be context coded but may be independent of the intra prediction mode, and only the first bin may be context coded. Furthermore, LFNST can be applied to intra CUs in both intra slices and inter slices, and to both luma and chroma. When dual tree is enabled, LFNST indices for luma and chroma can be signaled separately. For inter slices (dual tree disabled), a single LFNST index is signaled and can be used for both luma and chroma.

[0113] When an intra-subpartition (ISP) mode is selected, the LFNST may be disabled and the RST index may not be signaled. This is because even if the RST is applied to all feasible partition blocks, the performance improvement may be small. Furthermore, disabling the RST for the ISP-predicted residual may reduce encoding complexity. Also, when a matrix-based intra-prediction (MIP) mode is selected, the LFNST may be disabled and the index may not be signaled.

[0114] Considering that CUs larger than 64x64 may be implicitly split (TU tiling) due to the existing maximum transform size limitation (e.g., 64x64), LFNST index search may increase data buffering by four times for a certain number of decode pipeline stages. Therefore, the maximum size allowed for LFNST may be limited to 64x64. According to an embodiment, LFNST may be enabled only for DCT2.

[0115] [AV1 Residual Coding]

[0116] For each transform unit, AV1 coefficient coding may begin with signaling of a skip sign, followed by the transform kernel type and the end-of-block (eob) position when the skip sign is 0. Each coefficient value may then be mapped to multiple level maps and signs.

[0117] After the eob position is coded, the lower level map and the mid level map can be coded in reverse scan order, with the former indicating whether the magnitude of the coefficient is between 0 and 2, and the latter indicating whether the range is between 3 and 14. In the next step, the signs of the coefficients and the residual values ​​of coefficients greater than 14 using an Exp-Golomb code can be coded in forward scan order.

[0118] Regarding the use of context modeling, lower-level map coding can incorporate transform size and direction, as well as up to five neighboring coefficient information, while mid-level map coding can take a similar approach to lower-level amp coding, except that the number of neighboring coefficients is reduced to two. The exponential-Golomb code for the residual level and the sign of the AC coefficient can be coded without a context model, while the sign of the DC coefficient is coded using the dc sign of its neighboring transform unit.

[0119] [Deep Learning for Video Coding]

[0120] Deep learning is a set of learning methods that attempt to model data with complex architectures that combine different nonlinear transformations. The basic brick of deep learning is the neural network, which is connected to form a deep neural network.

[0121] Artificial neural networks are applications that are nonlinear with respect to a parameter θ that relates the inputs x and the output y = f(x,θ). The parameter θ is estimated from training samples. Neural networks can be used for regression or classification. There are several types of neural network architectures: (a) multilayer perceptrons, which are the oldest form of neural network; (b) convolutional neural networks (CNNs), which are particularly well suited for image processing; and (c) recurrent neural networks, which are used for sequential data such as text or time series.

[0122] Deep learning and neural networks can be used in video coding for two main reasons: first, unlike traditional machine learning algorithms, deep learning algorithms scan data and search for features, which eliminates the need for feature engineering; and second, deep learning models generalize well to new data, especially in image-related tasks.

[0123] A.CNN Layer

[0124] The advantages of CNNs compared to multi-layer perceptrons are twofold: neurons in a layer are only connected to a small region before them, which significantly reduces the amount of weights; moreover, CNNs are translation invariant, making them particularly suitable for processing images without losing spatial information. CNNs consist of several types of layers: convolutional layers, pooling layers, and fully connected layers.

[0125] (1) Convolutional layer

[0126] The discrete convolution between two functions f and g is defined as shown in equation (4) below:

number

[0127] For a two-dimensional signal such as an image, the following equation (5) can be considered for a two-dimensional convolution:

number

[0128] Referring to Figure 14, the principle of 2D convolution is to drag a convolution kernel (612) over an image (610). At each location, a convolution is applied between the convolution kernel and the portion of the image currently being processed (611). The convolution kernel is then moved by a number s of pixels, where s is called the stride. Sometimes, to control the size of the output, zero padding is added around the image, which is a margin of size p containing zero values. Assume that a C0 kernel (also called a filter) of size kxk is applied to each image. If the input image has size W, then i ×H i ×C i In the case of (W i is the width, H i is the height, C i is the number of channels, usually C i =3), the output volume is W0 × H0 × C0, where C0 corresponds to the number of kernels, and W0 and H0 have the relationship shown in equations (6) and (7).

number

[0129] The convolution operation can be combined with an activation function φ to add nonlinearity to the network: z(x) = φ(K*x+b), where b is the bias. One example is the Rectified Linear Unit (ReLU) activation function, which performs a max(0,x) operation.

[0130] (2) Pooling layer

[0131] CNNs also have pooling layers that can reduce network dimensionality, also known as subsampling, by taking the average or maximum over patches of an image (average pooling or max pooling). Like convolutional layers, pooling layers operate on small patches of an image with a stride. In one example, referring to FIG. 15, considering a 4×4 input patch (620) where max pooling is performed with a stride s=2, the output dimension of the output (622) is half the horizontal and vertical input dimension. While it is also possible to reduce dimensionality using convolutional layers by taking a stride greater than 1 without zero padding, the advantage of pooling is that it reduces the network's sensitivity to small transformations of the input image.

[0132] (3) Fully connected layer

[0133] After multiple convolutional and pooling layers, a CNN typically ends with several fully connected layers, where the tensors output by the preceding convolutional / pooling layers are transformed into a single vector of values.

[0134] B. Application of CNN to Video Coding

[0135] (1) Loop filtering

[0136] JVET-I0022 provides a convolutional neural network filter (CNNF) for intraframes. The CNNF serves as a loop filter for intraframes, replacing the Joint Exploration Model (JEM) filters, namely the bilateral filter (BF), deblocking filter (DF), and sample adaptive offset (SAO). Figure 16A shows the JEM intradecoding process (630), which includes entropy decoding (631), inverse quantization (632), inverse transform (633), BF (634), DF (635), SAO (636), prediction (637), and adaptive loop filter (ALF) (638). Figure 16B shows the intradecoding process, which includes the CNNF (644) instead of the BF (634), DF (635), and SAO (636). For B and P frames, the filters are kept the same as in JEM7.0.

[0137] Referring to Figures 16B and 17, the CNNF (644) may include two inputs: a reconstruction parameter (652) and a quantization parameter (QP) (654), which may allow a single set of parameters to be used to adapt to reconstructions with different qualities. Both of the two inputs may be normalized for better convergence in the training process. To reduce complexity, a simple CNN with 10 layers may be employed. The CNN may consist of one concatenation layer (656), seven convolutional layers (658A-G), each followed by a ReLU layer, one convolutional layer (660), and one summation layer (662). These layers may be connected one by one to form a network. It may be understood that the above layer parameters may be included in the convolutional layer. By connecting the reconstructed Y, U, and V to the summation layer, the network is normalized to learn the characteristics of the residual between the reconstructed image and the original image. According to one embodiment, simulation results report BD rate savings of -3.57%, -6.17% and -7.06% for luma and both chroma components for JEM-7.0 with AI configuration, and encoding and decoding times of 107% and 12887%, respectively, compared to the anchor.

[0138] JVET-N0254 reports experimental results of a dense residual convolutional neural network based on an in-loop filter (DRNLF). Referring now to FIG. 18, a structural block diagram of an example dense residual network (DRN) (670) is shown. The network structure can include N dense residual units (DRUs) (672A-N), where M represents the number of convolution kernels. For example, N can be set to 4 and M can be set to 32 as a trade-off between computational efficiency and performance. A normalized QP map (674) can be concatenated with the reconstructed frame as input to the DRN (670).

[0139] According to an embodiment, the DRUs (672A-N) may each have the structure (680) shown in FIG. 19. The DRUs may propagate inputs directly to subsequent units via shortcuts. To further reduce computational cost, a 3x3 depthwise separable convolution (DSC) layer may be applied to the DRUs.

[0140] The output of the network can have three channels, corresponding to Y, Cb, and Cr, respectively. The filter can be applied to both intra and inter images. An additional flag can be signaled for each CTU to indicate whether the DRNLF is on or off. Experimental results for one embodiment show BD rates of −1.52%, −2.12%, and −2.73% for the Y, Cb, and Cr components, respectively, in the all-intra configuration, −1.45%, −4.37%, and −4.27% in the random access configuration, and −1.54%, −6.04%, and −5.86% in the low-latency configuration. In this embodiment, the decoding times are 4667%, 7156%, and 9127% in the AI, RA, and LDB configurations.

[0141] (2) Intra prediction

[0142] 20 and 21, diagrams of a first process (690A) and a second process (690B) for intra-prediction modes are shown. Intra-prediction modes can be used to generate intra-image prediction signals on rectangular blocks in future video codecs. These intra-prediction modes perform two main steps: first, extract a set of features from decoded samples; second, use these features to select an affine linear combination of predefined image patterns as a prediction signal; and specific signaling schemes can be used for intra-prediction modes.

[0143] Referring to Figure 20, on a given MxN block (692A), with M < 32 and N < 32, generation of a luma prediction signal pred is performed by processing a set of reference samples r through a neural network. The reference samples r may consist of K rows of size N+K above the block (692A) and K columns of size M on the left. The number K may depend on M and N. For example, K may be set to 2 for all M and N.

[0144] The neural network (696A) can extract a vector of features f tr from the reconstructed samples as follows: r is considered as a vector in a real vector space of dimension d0, where d0 = K * (N + M + K) denotes the number of samples in r. For fixed integral square matrices A1 and A2, each having d0 rows and columns, and fixed integral bias vectors b1 and b2 of dimension d0, first calculate the following equation (8): t1=ρ(A1·r+b1) (Equation 8)

[0145] In equation (8), "·" denotes the usual matrix-vector product. Furthermore, the function ρ is an integer approximation of the ELU function ρ0, where the latter function is defined on a p-dimensional vector v, as shown in equation (9) below.

number

[0146] For a fixed integer d1 with 0≦d1≦d0, there can be a predetermined integral matrix A3, e.g., a predetermined integral bias vector b3 of dimension d1, with rows of d1 and columns of d0 and one or more bias weights (694A), and thus calculate the feature vector f tr as shown in equation (11) below. ftr=ρ(A·t2+b3) (Equation 11)

[0147] The value of d1 depends on M and N. At present, assume d1 = d0.

[0148] Among the feature vectors ftr, the final prediction signal pred is generated using an affine linear map, followed by a standard clipping operation that depends on the bit depth. Therefore, there is a predetermined matrix A4 with M*N rows and d1 columns, and a predetermined bias vector b4 of dimension M*N. Thus, it is calculated as follows in Equation (12). pred = Clip(A4·ftr + b4) (Equation 12)

[0149] Referring to FIG. 21 here, n different intra prediction modes (698B) are used, where n is set to 35 when max(M, N) < 32 and 11 otherwise. Therefore, an index premode having 0 ≦ premode < n is signaled by the encoder and parsed by the decoder, and can be used in the following high culture. One has n = 3 + 2 k where k = 3 when max(M, N) = 32 and k = 5 otherwise. In the first step, an index predIdx having 0 ≦ predIdx < n is signaled using the following code. First, one bin encodes whether predIdx < 3. If predIdx < 3, the second bin encodes whether predIdx = 0, and if predIdx ≠ 0, another bin encodes whether predIdx is 1 or equal. If predIdx ≧ 3, the value of predIdx is signaled in a canonical way using k bins.

[0150] From the index predIdx, the actual index predmode is derived using a fully connected neural network (696B) with one hidden layer, which has two rows of size N + 2 and two columns of size M on the left side of the block (692B), taking the reconstructed sample r’ as input.

[0151] The reconstructed sample r’ can be considered as a vector in a real vector space of dimension 2*(M+N+2). There is a fixed square matrix A1’, having a 2*(M+N+2) matrix and one or more bias weights (694B), like a fixed bias vector b1’ in a real vector space of 2*(M+N+2), and thus calculates t1’ as shown in the following formula (13). t1’ = ρ(A1’·r’ + b1) (Formula 13)

[0152] There can be a matrix A2’ having n rows and 2*(M+N+2) columns, and a fixed bias vector b2’ can exist in a real vector space of dimension n, and calculates lgt as shown in the following formula (14). lgt = A2’·t1’ + b2’ (Formula 14)

[0153] Here, the index predmode is derived as the position of the component of lgt that is the predIdx-th largest. Here, two components (lgt) k and (lgt) l are equal for k≠l, (lgt) k is regarded as larger than (lgt) l and fk < l and (lgt) l is regarded as larger than (lgt) k

[0154] [Multi-Transform Selection]

[0155] In addition to the DCT-II that has been used in HEVC, a multi-transform selection (MTS) scheme is used for residual coding of both inter and intra blocks. The scheme can include a plurality of transforms selected from DCT8 / DST7. According to an embodiment, DST-VII and DCT-VIII can be included. Table 4 shows the transform basis functions of DST / DCT selected for N-point input. [Table 4] ​

[0156] To maintain the orthogonality of the transform matrices, the transform matrices may be quantized more precisely than those in HEVC. To keep the intermediate values ​​of the transformed coefficients within the 16-bit range after horizontal and vertical transforms, all coefficients may need to be 10-bit.

[0157] To control the MTS scheme, separate enable flags can be specified at the SPS level for intra and inter, respectively. When MTS is enabled in SPS, a CU level flag is signaled to indicate whether MTS is applied or not. According to an embodiment, MTS can be applied only to luma. MTS signaling can be skipped if one of the following conditions applies: (1) the position of the last significant coefficient of the luma TB is less than 1 (i.e., DC only), or (2) the last significant coefficient of the luma TB is within the MTS zero-out region.

[0158] If the MTS CU flag is equal to zero, DCT2 can be applied in both directions. However, if the MTS CU flag is equal to one, two other flags can be additionally signaled to indicate the transform type in the horizontal and vertical directions, respectively. Table 5 below shows an example of a transform and signaling mapping table. The transform selection for ISP and implicit MTS can be unified by removing the intra mode and block shape dependency. If the current block is in ISP mode, or if the current block is an intra block and explicit MTS for intra and inter is on, only DST7 can be used for both the horizontal and vertical transform cores. For transform matrix precision, an 8-bit primary transform core can be used. Therefore, the transform cores used in HEVC are all kept the same, including 4-point DCT-2 and DST-7, 8-point, 16-point, and 32-point DCT-2. Other transform cores may also use the 8-bit primary transform core, including a 64-point DCT-2, a 4-point DCT-8, an 8-point, a 16-point, a 32-point DCT-7, and a DCT-8. [Table 5]

[0159] To reduce the complexity of large-sized DST-7 and DCT-8, high-frequency transform coefficients may be zeroed out for DST-7 and DCT-8 blocks whose size (width or height, or both width and height) is equal to 32. Only coefficients in the 16x16 lower frequency region may be retained.

[0160] As in HEVC, the remaining blocks can be coded in transform skip mode. To avoid syntax coding redundancy, the transform skip flag may not be signaled if the CU-level MTS_CU_flag is not equal to zero. According to an embodiment, the implicit MTS transform may be set to DCT2 if LFNST or MIP is activated for the current CU. Also, implicit MTS may be enabled even if MTS is enabled for inter-coding blocks.

[0161] [Non-separable quadratic transformation]

[0162] In JEM, a mode-dependent non-separable quadratic transform (NSST) can be applied between the forward core transform and quantization (at the encoder) and between dequantization and the inverse core transform (at the decoder). To maintain low complexity, NSST is only applied to the low-frequency coefficients after the primary transform. If both the width (W) and height (H) of a transform coefficient block are equal to or greater than 8, an 8x8 non-separable quadratic transform can be applied to the top-left 8x8 region of the transform coefficient block. Otherwise, if either W or H of the transform coefficient block is equal to 4, a 4x4 non-separable quadratic transform is applied, and a 4x4 non-separable transform can be performed on the top-left smallest (8,W) by smallest (8,H) region of the transform coefficient block. The above transform selection rules can be applied to both the luma and chroma components.

[0163] The matrix multiplication implementation of the non-separable transform can be performed as described above in the subsection "Quadratic Transforms in VVC" with respect to equations (2)-(3). According to an embodiment, the non-separable quadratic transform can be realized using direct matrix multiplication.

[0164] [Mode-dependent conversion core selection]

[0165] For both 4x4 and 8x8 block sizes, there may be 35x3 non-separable secondary transforms, where 35 is the number of transform sets specified by the intra prediction mode and 3 is the number of non-separable secondary transform (NSST) candidates for each intra prediction mode. The mapping from intra prediction mode to transform sets may be defined as shown in table 700 shown in Figure 22. The transform sets applied to luma / chroma transform coefficients may be specified by the corresponding luma / chroma intra prediction mode according to table 700. For intra prediction modes greater than 34 (diagonal prediction direction), transform coefficient blocks may be swapped before and after the secondary transform in the encoder / decoder.

[0166] For each transform set, the selected non-separable secondary transform candidate may be further specified by an explicitly signaled CU-level NSST index. After using transform coefficients and truncated unary binarization, the index may be signaled in the bitstream once per intra CU. The truncation value may be 2 for planar or DC modes and 3 for angular intra prediction modes. This NSST index may only be signaled if there are more than one non-zero coefficient in the CU. The default value may be zero if not signaled. A zero value for this syntax element may indicate that no secondary transform is applied to the current CU, and values ​​1-3 may indicate a secondary transform to be applied from the set.

[0167] In JEM, NSST may not be applied to blocks coded in transform skip mode. If an NSST index is signaled for a CU and is not equal to zero, NSST may not be used for component blocks coded in transform skip mode in the CU. If a CU with all component blocks is coded in transform skip mode or the number of non-zero coefficients in non-transform skip mode CB is less than two, an NSST index may not be signaled for the CU.

[0168] [Issues of the shape conversion scheme of the comparative embodiment]

[0169] In comparative embodiments, separable transform schemes are not very efficient at capturing directional texture patterns (e.g., edges at 45 / 135 degrees). Non-separable transform schemes are useful for improving coding efficiency in these scenarios. To reduce computational complexity and memory footprint, non-separable transform schemes are typically conceived as secondary transforms applied on top of the low-frequency coefficients of a primary transform. In existing implementations, the selection of the transform kernel to be used (from a group of both primary / secondary and separable / non-separable transform kernels) is based on prediction mode information. However, prediction mode information alone can only provide a rough representation of the entire space of observed residual patterns for that prediction mode, as shown by displays 710, 720, 730, and 740 in Figures 23A-D. Displays 710, 720, 730, and 740 show the observed residual patterns for the D45 (45°) intra-prediction mode in AV1. Neighboring reconstructed samples can provide additional information for a more efficient representation of these residual patterns.

[0170] For transform schemes with multiple transform kernel candidates, the transform set may need to be identified using coding information available at both the encoder and the decoder. In existing multi-transform schemes such as MTS and NSST, the transform set is selected based on coding prediction mode information, such as intra-prediction mode. However, the prediction mode completely covers all statistics of the prediction residual, and neighboring reconstructed samples can provide additional information for more efficient classification of the prediction residual. A neural network-based method can be applied for efficient classification of the prediction residual, thus providing more efficient transform set selection.

[0171] [Exemplary Aspects of the Present Disclosure]

[0172] The embodiments of the present disclosure may be used separately or in combination in any order. Furthermore, each embodiment (e.g., a method, an encoder, and a decoder) may be implemented by processing circuitry (e.g., one or more processors, or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium.

[0173] Embodiments of the present disclosure may incorporate any number of the aspects described above, and may also incorporate one or more of the aspects described below to solve the above-mentioned problems and / or other problems.

[0174] A. First Aspect

[0175] According to an embodiment, neighboring reconstructed samples may be used to select a transform set. In one or more embodiments, from a group of transform sets, a subgroup of transform sets is selected using coded information such as a prediction mode (e.g., an intra-prediction mode or an inter-prediction mode). In one embodiment, from the selected subgroup of transform sets, one transform set is identified using other coded information such as the type of intra / inter-prediction mode, the block size, the predicted block samples of the current block, and neighboring reconstructed samples of the current block. Finally, a transform candidate for the current block is selected from the identified transform sets using an associated index signaled in the bitstream. In one embodiment, from the selected subgroup of transform sets, a final transform candidate is implicitly identified using other coded information such as the type of intra / inter-prediction mode, the block size, the predicted block samples of the current block, and neighboring reconstructed samples of the current block.

[0176] In one or more embodiments, the adjacent reconstructed sample set can include samples from a previously reconstructed adjacent block. In one embodiment, the adjacent reconstructed sample set can include one or more lines of adjacent reconstructed samples from above and to the left. In one example, the number of lines of adjacent reconstructed samples from above and / or to the left is the same as the maximum number of lines of adjacent reconstructed samples used for intra prediction. In one example, the number of lines of adjacent reconstructed samples from above and / or to the left is the same as the maximum number of lines of adjacent reconstructed samples used for CfL prediction mode. In one embodiment, the adjacent reconstructed sample set can include all samples from the adjacent reconstructed block.

[0177] In one or more embodiments, the group of transform sets includes only primary transform kernels, only secondary transform kernels, or a combination of primary and secondary transform kernels. If the group of transform sets includes only primary transform kernels, the primary transform kernels can be separable, non-separable, use different types of DCT / DST, or use different line graph transforms with different self-loop rates. If the group of transform sets includes only secondary transform kernels, the secondary transform kernels can be non-separable, or use different non-separable line graph transforms with different self-loop rates.

[0178] In one or more embodiments, adjacent reconstructed samples can be processed to derive an index associated with a particular transform set. In one embodiment, adjacent reconstructed samples are input to a transform process, and the transform coefficients are used to identify an index associated with a particular transform set. In one embodiment, adjacent reconstructed samples are input to multiple transform processes, and a cost function is used to evaluate a cost value for each transform process. The cost value is then used to select a transform set index. Exemplary cost values ​​include, but are not limited to, the sum of the magnitudes of the first N (e.g., 1, 2, 3, 4, ..., 16) transform coefficients along a scan order. In one embodiment, a classifier is predefined, and adjacent reconstructed samples are input to the classifier to identify a transform set index.

[0179] B. Second Aspect

[0180] According to an embodiment, a neural network-based transform set selection scheme can be provided, where the inputs of the neural network can include, but are not limited to, predicted block samples of the current block, neighboring reconstructed samples of the current block, and the output can be an index used to identify the transform set.

[0181] In one or more embodiments, a group of transform sets is defined, a subgroup of the transform sets is selected using coded information such as a prediction mode (e.g., intra-prediction mode or inter-prediction mode), and then one transform set of the selected subgroup of transform sets is identified using other coded information such as predicted block samples of the current block, neighboring reconstructed samples of the current block, etc. A candidate transform for the current block is then selected from the identified transform set using an associated index signaled in the bitstream.

[0182] In one or more embodiments, the neighboring reconstructed samples can include one or more lines of neighboring reconstructed samples above and to the left. In one example, the number of lines of neighboring reconstructed samples above and / or to the left is equal to the maximum number of lines of neighboring reconstructed samples used for intra prediction. In one example, the number of lines of neighboring reconstructed samples above and / or to the left is equal to the maximum number of lines of neighboring reconstructed samples used for CfL prediction mode.

[0183] In one or more embodiments, the neighboring reconstructed samples and / or the predictive block samples of the current block are input to the neural network, and the output not only includes an identifier for the transform set but also an identifier for the prediction mode set. In other words, the neural network uses the neighboring reconstructed samples and / or the predictive block samples of the current block to identify a particular combination of transform set and prediction mode.

[0184] In one or more embodiments, a neural network is used to identify a transformation set for a secondary transformation. Alternatively, a neural network is used to identify a transformation set used for a primary transformation. Alternatively, a neural network is used to identify a transformation set used to specify a combination of a secondary transformation and a primary transformation. In one embodiment, the secondary transformation uses a non-separable transformation scheme. In one embodiment, the primary transformation can use different types of DCT / DST. In another embodiment, the primary transformation can use different line graph transformations with different self-looping rates.

[0185] In one or more embodiments, for different block sizes, the neighboring reconstructed samples and / or the predicted block samples of the current block may be further upsampled or downsampled before being used as inputs to the neural network.

[0186] In one or more embodiments, for different internal bit depths, the neighboring reconstructed samples and / or predicted block samples of the current block may be further scaled (or quantized) according to the internal bit depth value before being used as inputs to the neural network.

[0187] In one or more embodiments, the parameters used in the neural network depend on coded information, including, but not limited to: whether the block is intra-coded, the block width and / or block height, the quantization parameter, whether the current image is coded as an intra (key)frame, and the intra-prediction mode.

[0188] According to an embodiment, at least one processor and memory storing computer program instructions may be provided. The computer program instructions, when executed by the at least one processor, may implement an encoder or decoder and perform any number of functions described in this disclosure. For example, with reference to FIG. 24, the at least one processor may implement a decoder (800). The computer program instructions may include, for example, a decoding code (810) that configures the at least one processor to decode blocks of an image from a received coded bitstream (e.g., an encoder). The decoding code (810) may include, for example, a transform set selection code (820), a transform selection code (830), and a transform code (840).

[0189] The transform set selection code (820) may cause the at least one processor to select a transform set according to an embodiment of the present disclosure. For example, the transform set selection code (820) may cause the at least one processor to select a transform set based on at least one adjacent reconstructed sample from one or more previously decoded adjacent blocks or from a previously decoded image. According to an embodiment, the transform set selection code (820) may be configured to cause the at least one processor to select a subgroup of transform sets from a group of transform sets based on the first coding information, and to select a transform set from the subgroup according to an embodiment of the present disclosure.

[0190] The transform selection code (830) can cause at least one processor to select transform candidates from the transform set in accordance with embodiments of the present disclosure. For example, the transform selection code (830) can cause at least one processor to select transform candidates from the transform set based on index values ​​signaled in the coded bitstream in accordance with embodiments of the present disclosure.

[0191] The transform code (840) can cause at least one processor to inverse transform the coefficients of the block using a transform (e.g., a candidate transform) from the transform set according to an embodiment of the present disclosure.

[0192] According to an embodiment, the decoding code 810 may cause a neural network to be used in selecting transform groups, transform sub-groups, transform sets, and / or transforms, or to otherwise perform at least a portion of the decoding, according to an embodiment of the present disclosure. According to an embodiment, the decoder (800) may further include neural network code (850) configured to cause at least one processor to implement a neural network, according to an embodiment of the present disclosure.

[0193] According to an embodiment, the encoder-side process corresponding to the above process may be implemented by an encoding code for encoding an image, as will be understood by those skilled in the art based on the above description.

[0194] The techniques of the embodiments of the present disclosure described above can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, Figure 25 illustrates a computer system (900) suitable for implementing embodiments of the disclosed subject matter.

[0195] Computer software may be coded using any suitable machine code or computer language that may be subject to assembly, compilation, linking, or similar mechanisms to generate code containing instructions that may be executed directly, or via interpretation, microcode execution, etc., by a computer central processing unit (CPU), graphics processing unit (GPU), etc.

[0196] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, internet of things devices, and the like.

[0197] 25 for computer system 900 are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of the computer software implementing embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the exemplary embodiment of computer system 900.

[0198] The computer system 900 may include certain human interface input devices that may respond to input by one or more human users, for example, via tactile input (e.g., keystrokes, flipping, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), or olfactory input (not shown). Human interface devices may also be used to capture certain media that do not necessarily involve direct human conscious input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from still image cameras), and video (e.g., two-dimensional video, three-dimensional video, including stereoscopic images).

[0199] The input human interface devices may include one or more of the following (only one of each is shown): a keyboard (901), a mouse (902), a trackpad (903), a touchscreen (910), a data glove, a joystick (905), a microphone (906), a scanner (907), and a camera (908).

[0200] The computer system (900) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen (910), data gloves, or joystick (905), but may also be haptic feedback devices that do not function as input devices). For example, such devices may include audio output devices (e.g., speakers (909), headphones (not shown)), visual output devices (e.g., screens (910), including CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touchscreen input capabilities and each with or without haptic feedback capabilities—some of which may be capable of outputting two-dimensional visual output or three-dimensional or higher-dimensional output via means such as stereographic output: virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), printers (not shown), and the like.

[0201] The computer system (900) may also include human-accessible storage devices and their accessible media, such as optical media drives (920) including CD / DVD ROM / RW with media (921) such as CD / DVD, USB memory (922), removable head drives or solid state drives (923), conventional magnetic media such as tape, floppy disks (not shown), specialized ROM / ASIC / PLD-based devices such as security dongles, etc.

[0202] Those skilled in the art will also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.

[0203] The computer system 900 may also include interfaces to one or more communications networks. Networks may be, for example, wireless, wired, or optical. Networks may further be local, wide-area, metropolitan, vehicular, and industrial, real-time, delay-tolerant, and the like. Examples of networks include Ethernet, WLAN, cellular networks including GSM, 3G, 4G, 5G, LTE, and the like, cable TV, satellite TV, and terrestrial broadcast TV, and industrial and vehicular networks including CANBus. Certain networks require external network interface adapters connected to specific general-purpose data ports or peripheral buses 949 (e.g., USB ports on the computer system 900); others are generally integrated into the core of the computer system 900 by connecting to a system bus (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system), as described below. Using any of these networks, the computer system 900 can communicate with other entities. Such communications can be unidirectional, receive-only (e.g., broadcast television), unidirectional transmit-only (e.g., CAN bus to a particular CAN bus device), or bidirectional, e.g., to other computer systems using local or wide-area digital networks. This type of communication can include communications with cloud computing environments (955). Specific protocols and protocol stacks can be used for each of these networks and network interfaces, as described above.

[0204] The aforementioned human interface devices, human-accessible storage devices, and network interfaces (954) can be connected to the core (940) of the computer system (900).

[0205] The core (940) may include one or more central processing units (CPUs) (941), graphics processing units (GPUs) (942), specialized programmable processing devices in the form of field programmable gate arrays (FPGAs) (943), hardware accelerators 844 for specific tasks, etc. These devices may be connected via a system bus (948), along with read-only memory (ROM) (945), random access memory (946), and internal mass storage devices such as internal non-user-accessible hard drives, SSDs, etc. (947). In some computer systems, the system bus (948) is accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may connect directly to the core's system bus (948) or via a peripheral bus (949). Peripheral bus architectures include PCI, USB, etc. A graphics adapter (950) may be included in the core (940).

[0206] The CPU (941), GPU (942), FPGA (943), and accelerator (944) may combine to execute specific instructions that may constitute the aforementioned computer code. That computer code may be stored in ROM (945) or RAM (946). Transient data may also be stored in RAM (946), while permanent data may be stored, for example, in an internal mass storage device (947). Cache memory, which may be closely associated with one or more of the CPU (941), GPU (942), mass storage device (947), ROM (945), RAM (946), etc., may be used to enable fast storage and retrieval in any of the memory devices.

[0207] The computer-readable medium can have computer code thereon for performing various computer-implemented operations. The media and computer code can be those specially designed and created for the present disclosure, or they can be of the type well known and available in the art of computer software technology.

[0208] As an example, and not by way of limitation, a computer system having the architecture (900), and specifically the core (940), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be media associated with a user-accessible mass storage device, as described above, as well as specific storage devices of the core (940) that are non-transitory in nature, such as the core-internal mass storage device (947) or ROM (945). Software implementing various embodiments of the present disclosure may be stored in such devices and executed by the core (940). The computer-readable media may include one or more memory devices or chips, depending on particular needs. The software may cause the core (940), and specifically the processor (including a CPU, GPU, FPGA, etc.) therein, to perform certain processes or portions thereof described herein, including defining data structures stored in RAM (946) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator 944), which may operate in place of or in conjunction with software to perform particular processes or portions of particular processes described herein. Reference to software includes logic, and vice versa, where appropriate. Reference to a computer-readable medium may include circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry embodying logic for execution, or both, as appropriate. The present disclosure encompasses any suitable combination of hardware and software.

[0209] While this disclosure describes several non-limiting exemplary embodiments, there are modifications, permutations, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to create numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the invention and thus are within its concept and scope.

Claims

1. 1. A method executed by at least one processor, comprising: encoding a block of an image, The encoding step includes: selecting a set of transformations based on at least one neighboring sample from one or more neighboring blocks or from the image; transforming the coefficients of the block using a transform from the transform set; The step of selecting a transformation set comprises: selecting a subgroup of transform sets from the group of transform sets based on information of an intra-prediction mode or an inter-prediction mode; selecting a set of transformations from said subgroup; selecting a transform set from the subgroup includes selecting the transform set based on a type of the intra prediction mode or the inter prediction mode, a block size, a predicted block sample of the block or information of the at least one neighboring sample; The method further comprises selecting candidate transforms from the set of transforms; the at least one adjacent sample includes samples from the one or more adjacent blocks. method.

2. selecting the transform set from the subgroup comprises selecting the transform set based on the information of a type of the intra prediction mode or the inter prediction mode. The method of claim 1.

3. selecting the transform set from the subgroup comprises selecting the transform set based on the information of a type of the inter prediction mode. The method of claim 2.

4. selecting the transform set from the subgroup comprises selecting the transform set based on the information about the block size. The method of claim 1.

5. selecting the transform set from the subgroup comprises selecting the transform set based on the information of the predictive block samples of the block. The method of claim 1.

6. A method executed by at least one processor, comprising: encoding a block of an image, The encoding step includes: selecting a set of transformations based on at least one neighboring sample from one or more neighboring blocks or from the image; transforming the coefficients of the block using a transform from the transform set; The step of selecting a transformation set comprises: selecting a subgroup of transform sets from the group of transform sets based on information of an intra-prediction mode or an inter-prediction mode; selecting a set of transformations from said subgroup; selecting a transform set from the subgroup includes selecting the transform set based on a type of the intra prediction mode or the inter prediction mode, a block size, a predicted block sample of the block or information of the at least one neighboring sample; The method further comprises selecting candidate transforms from the set of transforms; the group of transform sets includes only secondary transform kernels; method.

7. the quadratic transformation kernel is non-separable; The method of claim 6.

8. selecting the transformation set from the subgroup comprises selecting the transformation set based on the information of the at least one neighboring sample. The method of claim 1.

9. A method executed by at least one processor, comprising: encoding a block of an image, The encoding step includes: selecting a set of transformations based on at least one neighboring sample from one or more neighboring blocks or from the image; transforming the coefficients of the block using a transform from the transform set; The step of selecting a transformation set comprises: selecting a subgroup of transform sets from the group of transform sets based on information of an intra-prediction mode or an inter-prediction mode; selecting a set of transformations from said subgroup; selecting a transform set from the subgroup includes selecting the transform set based on a type of the intra prediction mode or the inter prediction mode, a block size, a predicted block sample of the block or information of the at least one neighboring sample; The method further comprises selecting candidate transforms from the set of transforms; the set of transformations are quadratic transformations; method.

10. at least one memory configured to store computer program code; at least one processor configured to access said computer program code and to operate as instructed by said computer program code; A system comprising: The computer program code The at least one processor is configured to perform the method of any one of claims 1 to 9. system.

11. A computer program product for causing at least one processor to carry out a method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method and apparatus for selecting conversions in video coding and video decoding.

    JP2012516625A

  • Digital image coding method, decoding method, apparatus and related computer program

    JP2018524923A

  • Inseparable quadratic transformation for video coding

    JP2018530245A

  • Transform selection for video coding

    JP2019534624A

  • JPP7500732B