Video decoding method, device, computer device and storage medium

By using adjacent reconstructed samples and neural networks for transform set selection in video encoding and decoding, the problem of coarse transform kernel selection in existing technologies is solved, thus improving the compression efficiency of video decoding.

CN114641996BActive Publication Date: 2026-02-06TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180006246.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-06-22
Filing Date
2021-06-29
Publication Date
2026-02-06
Estimated Expiration
2041-06-29

AI Technical Summary

Technical Problem

In existing video encoding and decoding technologies, when selecting a transform kernel based on predictive mode information, only a rough residual mode representation can be provided, lacking effective additional information, resulting in insufficient compression efficiency.

Method used

By using adjacent reconstructed samples for transformation set selection and combining it with neural networks, a more accurate residual pattern representation is provided.

Benefits of technology

By using adjacent reconstructed samples and neural networks, the accuracy of transform set selection is improved, thereby enhancing the compression efficiency of video decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114641996B_ABST
    Figure CN114641996B_ABST
Patent Text Reader

Abstract

A video decoding method, device, computer device and storage medium are disclosed. The method comprises: decoding a block of a picture from a coded bitstream. The decoding comprises: selecting a set of transforms based on at least one neighboring already reconstructed sample, wherein the at least one neighboring already reconstructed sample is from at least one previously decoded neighboring block or from a previously decoded picture; and performing inverse transform on coefficients of the block using a transform in the set of transforms.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references

[0002] This application claims priority to U.S. Patent Application No. 17 / 354,731, filed June 22, 2021; U.S. Provisional Application No. 63 / 076,817, filed September 10, 2020; and U.S. Provisional Application No. 63 / 077,381, filed September 11, 2020, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to video encoding and decoding technology, and more particularly to video decoding methods, apparatus, computer equipment, and storage media. Background Technology

[0004] AOMedia Video 1 (AV1) is an open video codec format designed for video transmission over the Internet. As the successor to VP9, ​​it was developed by the Open Media Consortium (AOMedia), a consortium founded in 2015 that includes semiconductor companies, video-on-demand providers, video content producers, software development companies, and web browser vendors. Many components of the AV1 project stemmed from previous research by consortium members. Individual contributors began developing experimental technology platforms several years ago: Xiph / Mozilla's Daala released code in 2010, Google's experimental VP9 evolution project VP10 was announced on September 12, 2014, and Cisco's Thor released it on August 11, 2015. Building on the VP9 codebase, AV1 incorporates additional technologies, several of which were developed in the form of these experiments. The first version of the AV1 reference codec (version 0.1.0) was released on April 7, 2016. The Alliance for Open Media released the AV1 Streaming Specification, along with software-based encoders and decoders, for reference on March 28, 2018. On June 25, 2018, a final version 1.0.0 of the specification was released. On January 8, 2019, the "AV1 Streaming and Decoding Process Specification," which is final version 1.0.0 and includes errata table 1, was released. The AV1 Streaming Specification includes reference video codecs. The "AV1 Streaming and Decoding Process Specification" (version 1.0.0, including errata table 1) from the Alliance for Open Media (January 8, 2019) is incorporated herein by reference in its entirety.

[0005] The High Efficiency Video Coding (HEVC) standard was jointly developed by the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Experts Group (MPEG) standardization organizations. To develop the HEVC standard, the two standardization organizations worked together in a partnership known as the Joint Collaboration Team on Video Coding (JCT-VC). The first version of the HEVC standard was completed in January 2013, with a unified text published by both ITU-T and ISO / IEC. Since then, the organizations have attached work to extend the standard to support several additional application scenarios, including extended range usage with support for increased precision and color formats, scalable video coding, and 3-D / stereoscopic / multiview video coding. In ISO / IEC, the HEVC standard became MPEG-H Part 2 (ISO / IEC 23008-2), and in ITU-T it became ITU-T Recommendation H.265. The specification for the HEVC standard, “Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services - Coding of moving video, ITU-T H.265,” International Telecommunication Union (April 2015), is incorporated herein by reference in its entirety.

[0006] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Since then, they have been studying the potential need for standardization of future video coding technology that significantly outperforms HEVC in terms of compression capability. In October 2017, they released a Call for Proposals (CfP) on video compression with performance beyond HEVC. As of February 15, 2018, 22 CfP responses for the standard dynamic range (SDR) category, 12 CfP responses for the high dynamic range (HDR) category, and 12 CfP responses for the 360 video category were submitted. In April 2018, all received CfP responses were evaluated in the 122 MPEG / 10th Joint Video Exploration Team - Joint Video Experts Team (JVET) meeting. Through careful evaluation, the JVET officially launched the standardization of the next generation video coding (i.e., the so-called Versatile Video Coding (VVC)) that outperforms HEVC. The specification for the VVC standard, “Versatile Video Coding (Draft 7),” JVET-P2001-vE, Joint Video Experts Team (October 2019), is incorporated herein by reference in its entirety. Another specification for the VVC standard, “Versatile Video Coding (Draft 10),” JVET-S2001-vE, Joint Video Experts Team (July 2020), is incorporated herein by reference in its entirety.

[0007] In existing implementations, based on prediction mode information, a transform kernel to be used is selected. For example, from a set of transform kernels, a primary transform kernel / secondary transform kernel and a separable transform kernel / non-separable transform kernel are selected. However, in terms of the entire residual mode space observed for a prediction mode, only the information of the prediction mode can provide a rough representation. Therefore, some more effective additional information needs to be used to represent the residual mode. SUMMARY

[0008] Embodiments of the present application relate to a video decoding method and device, computer equipment and a storage medium. According to embodiments of the present application, a set selection scheme of primary transform and secondary transform using neighboring reconstructed samples is provided. According to embodiments of the present application, a neural network-based transform set selection scheme is provided for picture and video compression.

[0009] According to embodiments of the present application, a video decoding method comprises:

[0010] receiving a coded bitstream;

[0011] decoding a block of a picture in the coded bitstream, specifically comprising:

[0012] selecting a transform set based on at least one neighboring reconstructed sample, wherein the at least one neighboring reconstructed sample is from at least one previously decoded neighboring block or from a previously decoded picture; and

[0013] performing inverse transform on coefficients of the block using a transform in the transform set.

[0014] According to embodiments of the present application, a video decoding device comprises:

[0015] a decoding module configured to decode a block of a picture in a received coded bitstream;

[0016] the decoding module comprises:

[0017] a transform set selection module configured to select a transform set based on at least one neighboring reconstructed sample, wherein the at least one neighboring reconstructed sample is from at least one previously decoded neighboring block or from a previously decoded picture;

[0018] a transform module configured to perform inverse transform on coefficients of the block using a transform in the transform set.

[0019] According to embodiments of the present application, a computer equipment is provided, comprising a processor and a memory, the memory stores at least one instruction, the at least one instruction is loaded and executed by the processor to implement the above-mentioned video decoding method.

[0020] According to an embodiment of the present application, a non-transitory computer readable medium having stored thereon computer instructions, which when executed by at least one processor, implement the video decoding method described above.

[0021] According to another aspect of the present application, a computer program product or computer program is also provided, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to cause the computer device to perform the video decoding method described above.

[0022] As can be seen from the above technical solution, in the coding process, the adjacent reconstructed samples are used to select the transform set, which fully takes into account that the adjacent reconstructed samples can provide additional information for more effective classification of the prediction residual, and thus can be more efficiently used to select the transform set. BRIEF DESCRIPTION OF DRAWINGS

[0023] Other features, properties, and various advantages of the disclosed subject matter will become further clear from the following detailed description, as well as from the accompanying drawings, in which:

[0024] Figure 1 A simplified schematic diagram of a communication system according to an embodiment of the present application is shown;

[0025] Figure 2 A simplified schematic diagram of a communication system according to an embodiment of the present application is shown;

[0026] Figure 3 A simplified block diagram of a decoder according to an embodiment of the present application is shown;

[0027] Figure 4 A simplified block diagram of an encoder according to an embodiment of the present application is shown;

[0028] Figure 5A A schematic diagram of a first example partition structure for VP9 is shown;

[0029] Figure 5B A schematic diagram of a second example partition structure for VP9 is shown;

[0030] Figure 5C A schematic diagram of a third example partition structure for VP9 is shown;

[0031] Figure 5D A schematic diagram of a fourth example partition structure for VP9 is shown;

[0032] Figure 6A A schematic diagram of a first example partition structure for AV1 is shown;

[0033] Figure 6BA diagram showing a second example partition structure for AV1;

[0034] Figure 6C A diagram showing a third example partition structure for AV1;

[0035] Figure 6D A diagram showing a fourth example partition structure for AV1;

[0036] Figure 6E A diagram showing a fifth example partition structure for AV1;

[0037] Figure 6F A diagram showing a sixth example partition structure for AV1;

[0038] Figure 6G A diagram showing a seventh example partition structure for AV1;

[0039] Figure 6H A diagram showing an eighth example partition structure for AV1;

[0040] Figure 6I A diagram showing a ninth example partition structure for AV1;

[0041] Figure 6J A diagram showing a tenth example partition structure for AV1;

[0042] Figure 7 A diagram showing eight nominal angles for AV1;

[0043] Figure 8 A diagram showing a current block and samples;

[0044] Figure 9 A diagram showing an example recursive intra filtering mode;

[0045] Figure 10 A diagram showing a reference line adjacent to a coding block unit;

[0046] Figure 11 A diagram showing a table of AV1 hybrid transform kernels and their availability;

[0047] Figure 12 A diagram showing a low frequency non-separable transform process;

[0048] Figure 13 An example of a matrix is shown;

[0049] Figure 14 A diagram showing a two-dimensional convolution that explains a kernel and a picture;

[0050] Figure 15 A diagram showing max pooling of small patches of an image;

[0051] Figure 16A A schematic diagram illustrating a first inter-frame decoding process is shown;

[0052] Figure 16B A schematic diagram illustrating a second inter-frame decoding process is shown;

[0053] Figure 17 A schematic diagram illustrating an example of a convolutional neural network filter structure is shown;

[0054] Figure 18 A schematic diagram illustrating an example of a deep residual network is shown;

[0055] Figure 19 A schematic diagram illustrating an example of a deep residual unit structure is shown;

[0056] Figure 20 A schematic diagram illustrating a first process is shown;

[0057] Figure 21 A schematic diagram illustrating a second process is shown;

[0058] Figure 22 A mapping table from inter-prediction modes to transform set indices is shown;

[0059] Figure 23A An example of a first residual mode according to a comparative example is shown;

[0060] Figure 23B An example of a second residual mode according to a comparative example is shown;

[0061] Figure 23C An example of a third residual mode according to a comparative example is shown;

[0062] Figure 23D An example of a fourth residual mode according to a comparative example is shown;

[0063] Figure 24 A block diagram of a decoder according to embodiments of the present application is shown; and

[0064] Figure 25 A block diagram of a computer device according to embodiments of the present application is shown. DETAILED DESCRIPTION

[0065] In the present application, the term "block" can be understood as a prediction block, a coding block or a coding unit (CU). The term "block" can also be used herein to refer to a transform block.

[0066] In this application, the term "transform set" refers to a set of transform kernel (or candidate) options. A transform set can include one or more transform kernel (or candidate) options. According to embodiments of the present application, when more than one transform option is available, an index can be signaled to indicate which of the transform options in the transform set is applied to the current block.

[0067] In this application, the term "prediction mode set" refers to a set of prediction mode options. A prediction mode set can include one or more prediction mode options. According to embodiments of the present application, when more than one prediction mode option is available, an index can be further signaled to indicate which of the prediction mode options in the prediction mode set is applied to the current block to perform prediction.

[0068] In this application, the term "set of neighboring reconstructed samples" refers to a set of reconstructed samples from a previously decoded neighboring block or a reconstructed sample in a previously decoded picture.

[0069] In this application, the term "neural network" refers to the general concept of a data processing structure with one or more layers, as described in "Deep Learning for Video Coding" herein. According to embodiments of the present application, any neural network can be configured to implement these embodiments.

[0070] Figure 1 is a simplified block diagram of a communication system (100) in accordance with embodiments of the present disclosure. The communication system (100) includes at least two terminal devices (110, 120) connected via a network (150). For unidirectional transmission of data, a first terminal device (110) encodes video data at a local location and transmits the encoded video data to the other terminal device (120) via the network (150). The second terminal device (120) receives the encoded video data from the network (150), decodes the encoded video data, and displays the recovered video data. Unidirectional data transmission can be common in applications such as media services.

[0071] Figure 1 A second pair of terminal devices (130, 140) capable of bidirectional transmission of encoded video is shown, for example, during a video conference. For bidirectional transmission of data, each terminal device (130, 140) encodes video data captured at a local location and transmits the encoded video data to the other terminal device via the network (150). Each terminal device (130, 140) also receives encoded video data transmitted by the other terminal device, decodes the encoded video data, and displays the recovered video data on a local display device.

[0072] In Figure 1In general, terminals (110-140) can be servers, personal computers, and smart phones, and / or other types of terminals. For example, terminals (110-140) are laptops, tablets, media players, and / or dedicated video conferencing equipment. Network (150) represents any number of networks that convey coded video data among terminals (110-140), including for example wireline and / or wireless communication networks. Communication network (150) can exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area and / or wide area networks, and / or the Internet. For the purposes of the present application, the architecture and topology of network (150) can be immaterial to the operation of the disclosed subject matter unless otherwise explained herein.

[0073] As one example of the disclosed subject matter, Figure 2 Placement of video encoders and video decoders in a streaming environment is shown. The disclosed subject matter can be equally applicable to other video enabled applications, including, for example, video conferencing, digital TV, storing of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0074] As Figure 2 As shown, streaming system (200) can include a capture subsystem (213), which can include a video source (201) and an encoder (203). Video source (201), for example an electronic camera, can supply the uncompressed video to encoder (203). Encoder (203) can include hardware, software, or a combination of hardware and software to enable or implement aspects of the disclosed subject matter as described in greater detail below. The encoded video bitstream 204, which includes a lower data volume than the supply video from video source (201), can be stored for later use. At least one streaming client (206) can access the encoded video bitstream 204 from the streaming server (205) to reconstruct the video.

[0075] In embodiments of the present application, streaming server (205) can also functionally operate as a Media Aware Network Element (MANE). For example, streaming server (205) can be configured to prune the encoded video bitstream (204) to adapt to possibly different bitstreams sent to streaming clients (206). In embodiments of the present application, a MANE and streaming server (205) can be provided separately in streaming system (200).

[0076] The streaming client (206) includes a video decoder (210) and a display (212). The video decoder (210), for example, can decode a video bitstream (209), which is an incoming copy of the encoded video bitstream (204), and produce an output video sample stream (211) that can be rendered on the display (212) or another rendering device (not depicted). In some streaming systems, the video bitstreams (204, 209) can be encoded according to certain video coding / compression standards. Examples of such standards include, but are not limited to, ITU-T Recommendation H.265. An ongoing video coding standard, popularly known as Versatile Video Coding (VVC). Embodiments of the present application will be placed in the context of VVC.

[0077] Figure 3 is a block diagram of the video decoder (210) that is attached to the display (212) according to embodiments of the present disclosure.

[0078] The video decoder (210) includes a channel (312), a receiver (310), a buffer memory (315), an entropy decoder / pars er (320), a scaler / inverse transform unit (351), an intra prediction unit (352), a motion compensated prediction unit (353), an aggregator (355), a loop filter (356), a reference picture memory (357), and a current picture memory (358). In at least one embodiment, the video decoder (210) includes integrated circuits, sets of integrated circuits, and / or other electronic circuitry. The video decoder (210) can also be partially or entirely embodied by software running on at least one CPU associated with memory.

[0079] In this and other embodiments, the receiver (310) can receive at least one coded video sequence that is to be decoded by the decoder (210); one coded video sequence at a time, where the decoding of each coded video sequence is independent of the other coded video sequences. The coded video sequences can be received from a channel (312), which can be a hardware / software link into a storage device where the encoded video data is stored. The receiver (310) can receive the encoded video data with other data, for example, coded audio data and / or ancillary data streams that can be forwarded to their respective consuming entities (not depicted). The receiver (310) can separate the coded video sequence from the other data. To protect against network jitter, a buffer memory (315) can be coupled between the receiver (310) and the entropy decoder / parsen (320) (hereinafter simply “parsen (320)”). When the receiver (310) is receiving data from a store / forward device or from an isosychronous network, it can not need the buffer memory (315), or can have a small one, depending on the jitter requirements of the network. To be used on a service packet network like the Internet, the buffer memory (315) can be relatively large and can have an adaptive size.

[0080] The video decoder (210) can include a parser (320) to reconstruct symbols (321) from the coded video sequence. Categories of those symbols include information to manage operation of the decoder (210), and potentially information to control a display (212) or other display device that can be coupled to the decoder, as is known. Figure 2The control information for the display device can be Supplemental Enhancement Information (SEI messages) or Parameter Sets fragments (not depicted) of Video Usability Information (VUI). The parser (320) can parse / entropy-decode the received coded video sequence. The coding of the coded video sequence can be in accordance with a video coding technology or standard, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so forth. The parser (320) can extract from the coded video sequence, a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder, based upon at least one parameter corresponding to the group. The subgroups can include Groups of Pictures (GOPs), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs) and so forth. The parser (320) can also extract from the coded video sequence, information such as transform coefficients, quantizer parameter values, motion vectors and so forth.

[0081] The parser (320) can perform entropy-decoding / parsing operation on the video sequence received from the buffer memory (315), thereby creating symbols (321).

[0082] The reconstruction of the symbols (321) can involve a number of different units, depending on the type of coded video picture, or portion of coded video picture (such as intra picture, inter picture, intra block, inter block), and other factors. Which units are involved, and how, can be controlled by subgroup control information that the parser (320) extracts from the coded video sequence. For the sake of brevity, such subgroup control information flow between the parser (320) and the following number of units is not depicted.

[0083] In addition to the functional blocks already mentioned, the decoder (210) can be conceptually subdivided into a number of functional units as described below. In practical implementations operating under commercial constraints, many of these units interact closely with each other, and, at least partially, are integrated with each other. However, for the purpose of describing the disclosed subject matter, the conceptual subdivision into the functional units below is appropriate.

[0084] One unit is the scaler / inverse transform unit (351). The scaler / inverse transform unit (351) receives quantized transform coefficients as symbols (321) and control information including which transform to use, block size, quantization factor, quantization scaling matrices, etc. from the parser (320). The scaler / inverse transform unit (351) can output a block comprising sample values that can be input into the aggregator (355).

[0085] In some cases, the output samples of the scaler / inverse transform unit (351) can belong to an intra coded block; i.e., a block that is not using predictive information from previously reconstructed pictures, but can use predictive information from previously reconstructed parts of the current picture. Such predictive information can be provided by the intra picture prediction unit (352). In some cases, the intra picture prediction unit (352) generates a surrounding block of the same size and shape as the block being reconstructed using reconstructed information extracted from the current (partially reconstructed) picture 309 from the current picture memory (358). In some cases, the aggregator (355) adds, on a per sample basis, the predictive information generated by the intra prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351).

[0086] In other cases, the output samples of the scaler / inverse transform unit (351) can belong to an inter coded and potentially motion compensated block. In this case, the motion compensated prediction unit (353) can access the reference picture memory (357) to fetch samples for prediction. After motion compensation of the fetched samples according to the symbols (321), these samples can be added by the aggregator (355) to the output of the scaler / inverse transform unit (351) (in this case referred to as residual samples or residual signal) to generate the output sample information. The motion compensated prediction unit (353) can control the fetching of prediction samples from addresses within the reference picture memory (357) by motion vectors, and the motion vectors can be available to the motion compensated prediction unit (353) in the form of the symbols (321), e.g., comprising X, Y, and reference picture component. Motion compensation can also include interpolation of sample values fetched from the reference picture memory (357), motion vector prediction mechanisms, etc. when sub-sample precise motion vectors are used.

[0087] Output samples of the aggregator (355) can be subject to various loop filtering techniques in the loop filter unit (356). Video compression technologies can include in-loop filter technologies that are controlled by parameters included in the coded video sequence (also referred to as coded video bitstream) and made available to the loop filter unit (356) as symbols (321) from the parser (320). However, the video compression technologies can also be responsive to meta-information, obtained during the decoding of previous (in decoding order) parts of the coded picture or coded video sequence, as well as responsive to previously reconstructed and loop-filtered sample values, in other embodiments.

[0088] The output of the loop filter unit (356) can be a stream of samples that can be output to the display (212) 312 and that can be stored into the reference picture memory (357) for potential future use in inter-picture prediction of other pictures.

[0089] Once fully reconstructed, certain coded pictures can be used as reference pictures for future prediction. Once a coded picture corresponding to a current picture is fully reconstructed, and the coded picture is identified (by, for example, the parser (320)) as a reference picture, the current reference picture can become a part of the reference picture memory (357), and a fresh current picture buffer can be reallocated to the newly arriving coded picture.

[0090] The video decoder (210) can perform decoding operations according to a predetermined video compression technology in a standard, such as ITU-T H.265. The coded video sequence can conform to a syntax specified by the video compression technology or standard being used, in the sense that the coded video sequence adheres to the syntax of the video compression technology or standard, and the configuration files recorded in the coded video sequence are within the range of values allowed by the video compression technology or standard. In particular, a configuration file can choose, from all the tools available in the video compression technology or standard, certain tools as the only tools to be used under the profile. Also required for conformance is that the complexity of the coded video sequence is within the bounds set by the level of the video compression technology or standard. In some cases, the limits set by the level include maximum picture size, maximum frame rate, maximum reconstruction sample rate (measured in, for example megasamples per second), maximum reference picture size, and so on. Limits set by the level can be further refined by Hypothetical Reference Decoder (HRD) specifications and metadata for HRD buffer management signaled in the coded video sequence, in some cases.

[0091] In an embodiment, the receiver (310) can receive additional (redundant) data with the encoded video. The additional data can be part of the encoded video sequence. The additional data can be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. Additional data can be in the form of, for example, a temporal, spatial, or signal noise ratio (SNR) enhancement layer, redundant slices, redundant pictures, forward error correction codes, etc.

[0092] Figure 4 is a block diagram of a video encoder (203) associated with a video source (201) in accordance with an embodiment of the present disclosure.

[0093] The video encoder (203) can include, for example, an encoder, an encoding engine (432), a (local) decoder (433), a reference picture memory (434), a predictor (435), a transmitter (440), an entropy encoder (445), a controller (450), and a channel (460) as a source coder (430).

[0094] The video encoder (203) can receive video samples from the video source (201) (that is not a part of the encoder) that can capture video pictures to be coded by the encoder (203).

[0095] The video source (201) can provide the source video sequence to be coded by the encoder (203) in the form of a digital video sample stream, which can be in any suitable bit depth (for example: 8 bit, 10 bit, 12 bit,...), any color space (for example BT.601 Y CrCB, RGB,...), and any suitable sampling structure (for example Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (201) can be a storage device which stores previously prepared video. In a videoconferencing system, the video source (201) can be a camera that captures local image information as a video sequence. The video data can be provided as at least two separate pictures which impart motion when viewed in sequence. The pictures themselves can be organized as a spatial array of pixels, wherein each pixel can comprise at least one sample depending on the sampling structure, color space, etc. used. A person having ordinary skill in the art can readily understand the relationship between pixels and samples. The description below focuses on samples.

[0096] According to an embodiment of the present application, the encoder (203) can code and compress the pictures of the source video sequence into a coded video sequence (443) in real time or under any other time constraints as required by the application. Enforcing appropriate coding speed is one function of controller (450). The controller (450) can also control other functional units as described below and can be functionally coupled to these units. For clarity, functional coupling is not shown in the figures. Parameters set by the controller (450) can include rate control related parameters (picture skip, quantizer, lambda value of rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, and so on. The controller (450) can be used for other suitable functions including video encoder (203) optimizations for a certain system design.

[0097] Some video encoders operate in what those skilled in the art refer to as the "encoding loop." As a simple description, in an embodiment, the encoding loop can include the encoding portion in the source coder (430) (hereafter "source coder") that is responsible for creating symbols based on input pictures to be coded and reference pictures, and the (local) decoder (433) embedded in the encoder (203). The decoder (433) reconstructs the symbols to create the sample data in a similar manner as a (remote) decoder would create sample data from the symbols (since any compression between symbols and coded video bitstream in video compression technologies considered in the present application is lossless). The reconstructed sample stream (sample data) is input to the reference picture buffer (434). Since the decoding of symbol streams produces a bit-exact result independent of decoder location (local or remote), the reference picture buffer contents are also bit exact between the local encoder and the remote encoder. In other words, the prediction portion of the encoder "sees" the same reference picture samples that the decoder will "see" when using the prediction during decoding.

[0098] The operation of the "local" decoder (433) can be the same as the "remote" decoder (210) that has been described in detail above, for example. Figure 3 However, when symbols are available, and the entropy encoder (445) and parser (320) are capable of losslessly encoding / decoding symbols into the coded video sequence, the entropy decoding portion of the decoder (210), including channel (312), receiver (310), buffer (315), and parser (320), can not be fully implemented in the local decoder (433).

[0099] At this point it can be observed that any decoder technology, except for the parsing / entropy decoding present in the decoder, must also be present in the corresponding encoder in substantially the same functional form. For this reason, the present application focuses on the decoder operation. The description of the encoder technology can be simplified because the encoder technology is the inverse of the fully described decoder technology. More detailed description is required only in certain areas and is provided below.

[0100] During operation, in some embodiments, the source coder (430) can perform motion compensated predictive coding, which codes an input picture predictively with reference to at least one previously coded picture, designated as a "reference picture". In this manner, the coding engine (432) codes the difference between the pixels of an input picture and previously coded picture pixels of the reference picture, which can be selected as a prediction reference for the input picture.

[0101] The local video decoder (433) can decode coded video data of pictures that can be designated as reference pictures based on symbols created by the source coder (430). The operations of the coding engine (432) can be lossy processes. When the coded video data can be decoded at a video decoder (not shown), the reconstructed video sequence can typically be a replica of the source video sequence with some errors. Figure 4 The local video decoder (433) replicates decoding processes that can be performed by a video decoder on reference pictures and can cause reconstructed reference pictures to be stored in the reference picture memory (434). In this manner, the video encoder (203) can store copies of reconstructed reference pictures locally that have common content as the reconstructed reference pictures that will be obtained by a far-end video decoder (absent transmission errors).

[0102] The predictor (435) can perform a prediction search for the coding engine (432). That is, for each new picture to be coded, the predictor (435) can search the reference picture memory (434) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, and so on, that can serve as appropriate prediction references for the new pictures. The predictor (435) can operate on a sample block-by-pixel block basis to find appropriate prediction references. In some cases, as determined from a search result obtained by the predictor (435), an input picture can have prediction references drawn from at least two reference pictures stored in the reference picture memory (434).

[0103] The controller (450) can manage coding operations of the source coder (430), including, for example, setting of parameters and subgroup parameters used for encoding the video data.

[0104] The outputs of all the above-described functional units can be entropy encoded in an entropy encoder (445). The entropy encoder, for example, includes Huffman coding, variable length coding, arithmetic coding, and so forth, losslessly compresses the symbols generated by various functional units, and thus converts the symbols generated by various functional units into an encoded video sequence.

[0105] The transmitter (440) can buffer the encoded video sequence(s) as created by the entropy coder (445) to prepare it for transmission via a communication channel (460), which can be a hardware / software link leading to a storage device that will store the encoded video data. The transmitter (440) can merge encoded video data from the video coder (430) with other data to be transmitted, for example, encoded audio data and / or ancillary data streams (sources not shown).

[0106] The controller (450) can manage operation of the encoder (203). During coding, the controller (450) can assign to each coded picture a certain coded picture type, which can affect the coding techniques that can be applied to the corresponding picture. For example, pictures often can be assigned as Intra Pictures (I-Pictures), Predictive Pictures (P-Pictures), or Bidi- rectional Pictures (B-Pictures).

[0107] Intra Pictures (I-Pictures), which can be pictures that can be coded and decoded without using any other picture in the sequence as a source of prediction. Some video codecs allow different types of Intra Pictures, including, for example, Independent Decoder Refresh (IDR) Pictures. A person of ordinary skill in the art understands the variants of I-Pictures and their respective applications and features.

[0108] Predictive Pictures (P-Pictures), which can be pictures that can be coded and decoded using either Intra prediction or Inter prediction that uses at most one motion vector and reference index to predict sample values of each block.

[0109] Bidi- rectional Pictures (B-Pictures), which can be pictures that can be coded and decoded using either Intra prediction or Inter prediction that uses at most two motion vectors and reference indices to predict sample values of each block. Similarly, at least two predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.

[0110] A source picture can generally be spatially subdivided into at least two sample blocks (e.g., 4x4, 8x8, 4x8, or 16x16 samples), and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, determined according to the coding assignment of the respective picture to which the block applies. For example, blocks of an I picture can be non-predictively encoded, or the blocks can be predictively encoded with reference to already encoded blocks of the same picture (spatial or intra prediction). Blocks of a P picture can be predictively encoded with reference to one previously encoded reference picture, either by spatial prediction or by temporal prediction. Blocks of a B picture can be predictively encoded with reference to one or two previously encoded reference pictures, either by spatial prediction or by temporal prediction.

[0111] The video encoder (203) can perform encoding operations according to a predetermined video coding technology or standard, such as ITU-T H.265. In its operation, the video encoder (203) can perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancies in the input video sequence. Accordingly, the coded video data can conform to a syntax specified by the video coding technology or standard being used.

[0112] In an embodiment, the transmitter (440) can transmit coded video data and additional data. The video encoder (430) can include this data as part of the coded video sequence. The additional data can comprise temporal / spatial / SNR (signal noise ratio) enhancement layers, other types of redundant data such as redundant pictures and slices, Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) messages, among others.

[0113] [Encoding block partitioning in VP9 and AV1]

[0114] Reference Figure 5A to Figure 5D VP9 uses a 4-way partitioning tree starting at the 64x64 level down to the 4x4 level with some additional restrictions for 8x8 blocks. Note that Figure 5D The partitioning denoted as R in [VP9] refers to a recursion where the same partitioning tree is repeated at a lower scale until the lowest 4x4 level is reached.

[0115] Reference Figure 6A to Figure 6J AV1 not only extends the partitioning tree to a 10-way structure but also increases the maximum size (called superblock in VP9 / AV1 parlance) from 128x128. Note that this includes the 4:1 / 1:4 rectangular partitions that are not present in VP9. As in [VP9], the partitioning tree is applied recursively to each of the 10 partitions. Note that the 4:1 / 1:4 partitions are not present in VP9. Figure 6C to Figure 6FAs shown, the partition type with 3 sub-partitions is referred to as a "T-shaped" partition. None of the rectangular partitions can be further subdivided. In addition to the coding block size, a coding tree depth can be defined to indicate the depth of partitioning from the root node. Specifically, the coding tree depth of the root node (e.g., 128x128) is set to 0, and the coding tree depth is increased by 1 after the tree block is further partitioned once.

[0116] Instead of enforcing a fixed transform unit size as in VP9, AV1 allows the luma coding block to be partitioned into multiple sizes of transform units, which can be represented as being partitioned by a downward recursion until level 2. To incorporate the extended coding block partitioning of AV1, square, 2:1 / 1:2, and 4:1 / 1:4, transform sizes from 4x4 to 64x64 can be supported. For chroma blocks, only the largest possible transform unit is allowed.

[0117] [Block partitioning in HEVC]

[0118] In HEVC, a coding tree unit (CTU) can be partitioned into coding units (CUs) using a quad-tree (QT) structure denoted as coding tree. Decisions can be made at the CU level whether to use inter-picture (temporal) or intra-picture (spatial) prediction to code the picture region. Each CU can be further partitioned into one, two, or four prediction units (PUs) according to the PU partition type. Within one PU, the same prediction process can be applied, and the relevant information is sent to the decoder on a PU basis. After applying the prediction process based on the PU partition type to obtain a residual block, the CU can be partitioned into transform units (TUs) according to another quad-tree structure (e.g., coding tree of the CU). One of the key features of the HEVC structure is that it has multiple partitioning concepts including CU, PU, and TU. In HEVC, a CU or TU can only have a square shape, while a PU can have a square or rectangular shape for inter-predicted blocks. In HEVC, one coding block can be further partitioned into four square sub-blocks, and transform is performed on each sub-block (i.e., TU). Each TU can be further recursively partitioned (using quad-tree partitioning) into smaller TUs, which is referred to as a residual quad-tree (RQT).

[0119] At picture boundaries, HEVC employs an implicit quad-tree partitioning such that the block will remain quad-tree partitioned until the size fits the picture boundary.

[0120] [Quad-tree with nested multi-type tree coding block structure in VVC]

[0121] In VVC, quaternary tree with nested multi-type tree using binary and ternary splitting structure replaces the concept of multiple partition unit types. That is, VVC does not include the separation of CU, PU, and TU concepts except for some CUs whose size is too large for the maximum transform length, and VVC supports more flexibility in CU partition shape. In the coding tree structure, a CU can have a square shape or a rectangular shape. A coding tree unit (CTU) is first partitioned by a quaternary tree (also called a quad tree) structure. Then, the quaternary tree leaf node can be further partitioned by a multi-type tree structure. There are four splitting types in the multi-type tree structure: vertical binary splitting (SPLIT_BT_VER), horizontal binary splitting (SPLIT_BT_HOR), vertical ternary splitting (SPLIT_TT_VER), and horizontal ternary splitting (SPLIT_TT_HOR). A multi-type tree leaf node can be referred to as a coding unit (CU), which can be used for prediction and transform processing without any further partitioning unless the CU is too large for the maximum transform length. This means that, in most cases, in the quaternary tree with nested multi-type tree coding block structure, the CU, PU, and TU have the same block size. There is one exception when the maximum supported transform length is smaller than the width or height of the color component of the CU. One example of block partitioning is that a CTU is divided into multiple CUs with quaternary tree and nested multi-type tree coding block structures using quaternary tree partitioning and multi-type tree partitioning. The quaternary tree with nested multi-type tree partitioning provides a content adaptive coding tree structure including CUs.

[0122] In VVC, the maximum supported luma transform size is 64x64, and the maximum supported chroma transform size is 32x32. When the width or height of a CB is larger than the maximum transform width or height, the CB can be automatically split in the horizontal and / or vertical direction to meet the transform size limit in that direction.

[0123] In VTM7, the coding tree scheme supports the ability of luma and chroma having separate block tree structures. For P and B slices, the luma CTBs and chroma CTBs in one CTU must share the same coding tree structure. However, for I slices, luma and chroma can have separate block tree structures. When the separate block tree mode is applied, the luma CTBs are partitioned into CUs by one coding tree structure, and the chroma CTBs are partitioned into chroma CUs by another coding tree structure. This means that a CU in an I slice can consist of a coding block of luma component or two coding blocks of chroma components, and a CU in a P or B slice can consist of coding blocks of all three color components, unless the video is monochrome.

[0124] [Directional intra prediction in AV1]

[0125] VP9 supports eight directional modes corresponding to angles from 45 degrees to 207 degrees. To exploit more kinds of spatial redundancy in directional textures, in AV1, the directional intra modes are extended to a set of angles with finer granularity. The original eight angles are slightly changed and set to nominal angles, the eight nominal angles are named V_PRED (542), H_PRED (543), D45_PRED (544), D135_PRED (545), D113_PRED (546), D157_PRED (547), D203_PRED (548), and D67_PRED (549), as shown in the current block (541) in Figure 7 AV1. For each nominal angle, there are seven smaller angles, so AV1 has 56 directional angles in total. The prediction angle is represented by a nominal intra angle plus an increment angle, which is -3~3 times of 3-degree step. In AV1, first, eight nominal modes and five non-angular smoothing modes are signaled. Then, if the current mode is an angular mode, further signal an index to indicate the angle increment relative to the corresponding nominal angle. To implement the directional prediction modes in AV1 via a general way, all 56 directional intra prediction modes in AV1 are implemented using a unified directional predictor, which projects each pixel to a reference sub-pixel position and interpolates the reference pixel by a 2-tap bilinear filter.

[0126] [Non-directional smoothing intra predictor in AV1]

[0127] In AV1, there are five non-directional smoothing intra prediction modes, which are DC, PAETH, SMOOTH, SMOOTH_V, and SMOOTH_H. For DC prediction, the average of the left and top neighboring samples is used as the prediction value for the block to be predicted. For the PAETH predictor, first, the top, left, and top-left reference samples are retrieved, then the value closest to (top + left - top-left) is set as the prediction value for the pixel to be predicted. Figure 8 The positions of the top sample (554), left sample (556), and top-left sample (558) of a pixel (552) in the current block (550) are shown. For the SMOOTH, SMOOTH_V, and SMOOTH_H modes, in the average direction of the vertical or horizontal direction or both directions, the current block (550) is predicted using quadratic interpolation.

[0128] [Intra predictor based on recursive filtering]

[0129] To capture the spatial correlation of the decay on the edge, filter intra modes are designed for luma blocks. Five filter intra modes are defined for AV1, each represented by a set of eight 7-tap filters, reflecting the correlation between the pixels in a 4x2 patch and their 7 neighbors. In other words, the weighting factors of the 7-tap filters are position dependent. For example, an 8x8 block (560) can be divided into 8 4x2 patches, as shown in Figure 9 These patches are indicated as B0, B1, B2, B3, B4, B5, B6 and B7 in Figure 9 For each patch, its 7 neighbors, indicated by R0 to R6, can be used to predict the pixels in the current patch. For patch B0, all neighbors can have been reconstructed. But for other patches, some neighbors can not have been reconstructed, then the predicted value of the immediate neighbor is used as reference. For example, none of the neighbors of patch B7 is reconstructed, so the predicted sample of the neighbor is used as a substitute.

[0130] [Chroma from luma]

[0131] Chroma from luma (CfL) is a chroma-only intra predictor that models the chroma pixels as a linear function of the coincident reconstructed luma pixels. The CfL prediction can be represented as shown in the following equation (1):

[0132] CfL(a) = a x L AC + DC (equation 1)

[0133] where L AC denotes the AC contribution of the luma component, a denotes the parameter of the linear model, and DC denotes the DC contribution of the chroma component. Specifically, the reconstructed luma pixels are sub-sampled to the chroma resolution, and then the mean is subtracted to form the AC contribution. To approximate the chroma AC component from the AC contribution, the decoder does not need to compute the scaling parameter, as in some background techniques, but AV1 CfL can determine the parameter a based on the original chroma pixels, and their are identified in the bitstream. This reduces the decoder complexity, and results in more accurate prediction. For the DC contribution of the chroma component, it can be computed using the intra DC mode, which is sufficient for most chroma content, and has a mature fast implementation.

[0134] [Multi-line intra prediction]

[0135] Multi-line intra prediction can use more reference lines for intra prediction, where the encoder decides and signals which reference line is used to generate the intra prediction value. The reference line index can be signaled before the intra prediction mode, and only the most probable mode can be allowed in case a non-zero reference line index is signaled. In Figure 10In this case, an example of four reference lines (570) is depicted, where each reference line (570) consists of six segments (i.e., segments A to F) and a top-left reference sample. In addition, segments A and F are padded with the nearest samples from segments B and E, respectively.

[0136] [Primary transform in AV1]

[0137] To support extended coding block partitioning, multiple transform sizes (e.g., ranging from 4 points to 64 points per dimension) and transform shapes (e.g., square; rectangular with width / height ratios of 2:1 / 1:2 and 4:1 / 1:4) are introduced into AV1.

[0138] 2D transform processing can involve the use of hybrid transform kernels (e.g., consisting of different one-dimensional (ID) transforms for each dimension of the coded residual block). According to one embodiment, the primary ID transforms are: (a) 4-point, 8-point, 16-point, 32-point, or 64-point DCT-2; (b) 4-point, 8-point, or 16-point asymmetric DST (DST-4, DST-7) and their flipped versions; and (c) 4-point, 8-point, 16-point, or 32-point identity transform. The basis functions for DCT-2 and asymmetric DST used in AV1 are listed in Table 1 below. Table 1 shows the AV1 primary transform basis functions DCT-2, DST-4, and DST-7 for N-point input.

[0139] Table 1: AV1 primary transform basis functions

[0140]

[0141] The availability of the hybrid transform kernels can be based on the transform block size and the prediction mode. This correlation is summarized in Table 580 of Figure 11 Table 580. Table 580 shows the AV1 hybrid transform kernels based on the prediction mode and block size and their availability. In Table 580, the symbols “→” and “↓” denote the horizontal and vertical dimensions, respectively, and “√” and “x” denote the availability and unavailability of the kernel for the block size and prediction mode, respectively.

[0142] For chroma components, the transform type selection can be done in an implicit manner. For intra-predicted residuals, the transform type can be selected according to the intra-prediction mode as specified in Table 2 below. For inter-predicted residuals, the transform type can be selected according to the transform type selection of the collocated luma block. Therefore, for chroma components, there can be no transform type signaling in the bitstream.

[0143] Table 2: Transform type selection for chroma component intra-predicted residuals.

[0144] Intra prediction Vertical transform Horizontal transform DC_PRED DCT DCT V_PRED ADST DCT H_PRED DCT ADST D45_PRED DCT DCT D135_PRED ADST ADST D113_PRED ADST DCT D157_PRED DCT ADST D203_PRED DCT ADST D67_PRED ADST DCT SMOOTH_PRED ADST ADST SMOOTH_V_PRED ADST DCT SMOOTH_H_PRED DCT ADST PAETH_PRED ADST ADST

[0145] [Secondary transform in VVC]

[0146] Reference Figure 12 In VVC, a low-frequency non-separable transform (LFNST), which is referred to as a reduced secondary transform, can be applied between a forward primary transform (591) and quantization (593) (at the encoder) and between dequantization (594) and inverse primary transform (596) (at the decoder side) to further decorrelate the primary transform coefficients. For example, a forward LFNST (592) can be applied by the encoder, while an inverse LFNST (595) can be applied by the decoder. In LFNST, either a 4x4 non-separable transform or an 8x8 non-separable transform can be applied depending on the block size. For example, a 4x4 LFNST can be applied for small blocks (e.g., min(width, height) < 8), while an 8x8 LFNST can be applied for larger blocks (e.g., min(width, height) > 4). For 4x4 forward LFNST and 8x8 forward LFNST, the forward LFNST (592) can have 16 and 64 input coefficients, respectively. For 4x4 inverse LFNST and 8x8 inverse LFNST, the inverse LFNST (595) can have 8 and 16 input coefficients, respectively.

[0147] The non-separable transform used in LFNST is described below using an input as an example. To apply a 4x4 LFNST, a 4x4 input block X shown below in Equation (2) can first be represented as a vector As shown in Equation (3) below:

[0148]

[0149]

[0150] The non-separable transform can be computed as where denotes a vector of transform coefficients, T is a 16x16 transform matrix. The 16x1 vector of coefficients can be subsequently reorganized into a 4x4 block using a scan order (e.g., horizontal, vertical, or diagonal) for the block. The coefficients with smaller indices can be placed in the 4x4 coefficient block with smaller scan indices.

[0151] A. Reduced non-separable transform

[0152] LFNST can apply non-separable transforms based on a direct matrix multiplication method such that it is implemented in a single pass without multiple iterations. However, it is necessary to reduce the non-separable transform matrix dimension to minimize the computational complexity and memory space to store the transform coefficients. Therefore, a reduced non-separable transform (RST) method can be used in LFNST. The main idea of the reduced non-separable transform is to map an N (for 8x8 NSST, N is usually equal to 64) dimensional vector to an R dimensional vector in a different space, where N / R (R < N) is the reduction factor. Therefore, instead of an N x N matrix, the RST matrix becomes an R x N matrix (600) as shown in Figure 13

[0153] ​In the R x N matrix (600), there are R rows of transformations that are R bases of the N-dimensional space. The inverse transformation matrix of the RT can be the transpose of its forward transformation. For an 8 x 8 LFNST, the reduction factor can be 4, and a 64 x 64 direct matrix as a regular 8 x 8 non-separable transform matrix size can be reduced to a 16 x 48 direct matrix. Thus, a 48 x 16 inverse RST matrix can be used at the decoder side to generate the core (primary) transform coefficients in the 8 x 8 top-left region. When a 16 x 48 matrix is applied instead of a 16 x 64 matrix with the same transform set configuration, each matrix can take 48 input data from three 4 x 4 blocks in the top-left 8 x 8 block except for the bottom-right 4 x 4 block. With the reduced dimension, the memory for storing all LFNST matrices can be reduced from 10 Kb to 8 KB, and the performance degradation is reasonable. To reduce the complexity, the LFNST can be restricted to be applied only when all coefficients except the first coefficient sub-group are non-significant. Thus, when the LFNST is applied, all primary-only transform coefficients must be zero. This allows the adjustment of the LFNST index signaling on the last significant position, thus avoiding the extra coefficient scanning in the current LFNST design, which is only needed when checking the significant coefficients at specific positions. The worst case processing (multiplication per pixel) of the LFNST limits the non-separable transform of 4 x 4 and 8 x 8 blocks to 8 x 16 and 8 x 48 transforms, respectively. In these cases, when the LFNST is applied, the last significant scan position must be less than 8 for other sizes smaller than 16. For blocks with shapes of 4 x N and N x 4 with N > 8, the restriction can mean that the LFNST is now applied only once and only to the top-left 4 x 4 region. Since all primary-only coefficients can be zero when the LFNST is applied, the number of operations of the main transform can be reduced in this case. From the encoder perspective, the quantization of the coefficients is significantly simplified when the LFNST transform is tested. For the first 16 coefficients (in scan order), the rate-distortion optimized quantization can be maximized, and the remaining coefficients can be forced to zero.

[0154] B. LFNST transform selection

[0155] For each transform set used in LFNST, there can be four transform sets and two non-separable transform matrices (kernels). The mapping from the intra prediction mode to the transform set can be pre-defined as shown in Table 3 below. If one of the three CCLM modes (INTRA LT CCLM, INTRA T CCLM or INTRA L CCLM) is used for the current block (81 <= IntraPredMode <= 83), transform set 0 can be selected for the current chroma block. For each transform set, the selected non-separable secondary transform candidate can be further specified by an explicitly signaled LFNST index. After the transform coefficients, this index can be signaled once per intra CU in the bitstream.

[0156] Table 3: Transform selection table

[0157]

[0158] C. LFNST index signaling and interaction with other tools

[0159] Since LFNST can be restricted to apply only when all coefficients outside the first coefficient sub-group are non-significant, LFNST index coding can depend on the position of the last significant coefficient. In addition, the LFNST index can be context coded, but can not be dependent on the intra prediction mode, and only the first bin can be context coded. Furthermore, LFNST can be applied to intra CUs in intra and inter slices, as well as for luma and chroma. If dual tree is enabled, the LFNST index for luma and chroma can be signaled separately. For inter slices (dual tree disabled), a single LFNST index can be signaled and used for luma and chroma.

[0160] When the intra sub-partition (ISP) mode is selected, LFNST can be disabled and the RST index can not be signaled, since the performance improvement can be marginal even if RST is applied to each feasible partition block. In addition, disabling RST for ISP predicted residuals can reduce the encoding complexity. When the matrix-based intra prediction (MIP) mode is selected, LFNST can also be disabled and the index can not be signaled.

[0161] Considering that large CUs larger than 64x64 can be implicitly partitioned (TU tiling) due to the existing maximum transform size limit (e.g., 64x64), LFNST index search can increase the data buffering by a factor of four for a certain number of decoding pipeline stages. Therefore, the maximum size allowed for LFNST can be limited to 64x64. According to embodiments, LFNST can be enabled with DCT2 only.

[0162] [Residual coding in AV1]

[0163] For each transform unit, the AV1 coefficient coding can start with a signaled skip sign, followed by a transform kernel type and an end-of-block (eob) position when the skip sign is zero. Each coefficient value can then be mapped to a multi-level mapping and a sign.

[0164] After the eob position is encoded, the lower-level mapping and the mid-level mapping can be encoded in a reverse scan order, the former can indicate whether the coefficient magnitude is between 0 and 2, and the latter can indicate whether the range is between 3 and 14. In the next step, the sign of the coefficient, and the residual value of the coefficient greater than 14 by Exp-Golomb code, can be encoded in a forward scan order.

[0165] For the use of context modeling, the lower-level mapping coding can incorporate the transform size and orientation, and up to five neighboring coefficient information. On the other hand, the mid-level mapping coding can follow a similar approach as the lower-level amp coding, except that the number of neighboring coefficients is as low as 2. The Exp-Golomb code of the residual level and the sign of the AC coefficient can be encoded without any context model, while the sign of the DC coefficient is encoded using the dc sign of its neighboring transform unit.

[0166] [Deep learning for video coding]

[0167] Deep learning is a set of learning methods that model data with complex architectures that incorporate different nonlinear transformations. The basic building block of deep learning is a neural network, which is combined to form a deep neural network.

[0168] An artificial neural network is a nonlinear application with parameters θ related to an entry x and an output y = f(x, θ). The parameters θ are estimated from learning samples. Neural networks can be used for regression or classification. There are several types of neural network architectures: (a) multilayer perceptron, which is the oldest form of neural network; (b) convolutional neural network (CNN), which is particularly suitable for image processing; and (c) recurrent neural network for sequential data such as text or time series.

[0169] Deep learning and neural networks can be used for video coding, mainly for two reasons: first, unlike traditional machine learning algorithms, deep learning algorithms scan data to search for features, thus not requiring feature engineering. Second, deep learning models generalize well to new data, especially in image-related tasks.

[0170] A. CNN layers

[0171] The strength of CNNs is doubled compared to multilayer perceptrons: CNNs have a drastically reduced number of weights because the neurons in a layer are only connected to a small region of the previous layer; moreover, CNNs are translationally invariant, making them particularly suitable for processing images without losing spatial information. CNNs are composed of several layers, namely convolutional layers, pooling layers, and fully connected layers.

[0172] (1) Convolutional layer

[0173] The discrete convolution between two functions f and g can be defined as shown in equation (4) below:

[0174] (f * g)(x) =∑ t f(t)g(x + t) (equation 4)

[0175] For 2-dimensional signals such as images, the following equation (5) can be considered for 2D convolution:

[0176] (K * I)(i,j) =∑ m,n K(m,n)I(i + n,j + m) (equation 5)

[0177] where K is the convolution kernel applied to the 2D signal (or image) I.

[0178] Referring to Figure 14 , the principle of 2D convolution is to drag the convolution kernel (612) over the image (610). At each position, a convolution is applied between the convolution kernel and a portion of the image (611) currently processed. Then, the convolution kernel moves s pixels, where s is called the stride. Sometimes, zero padding is added, which is a margin containing zero values around the image, of size p, in order to control the size of the output. Assuming that a C0kernel (also called a filter) is applied, each size k x k on the image. If the size of the input image is W i x H i x C i (W i represents the width of the channel, H i represents the height of the channel, C i represents the number of channels, usually C i = 3), the volume of the output is W0x H0x C0, where C0corresponds to the number of kernels, W0and H0have the following relationships shown in equations (6) and (7).

[0179]

[0180] The convolution operation can be combined with an activation function in order to add nonlinearity to the network: where b is the bias. One example is the rectified linear unit (ReLU) activation function that performs a max(0, x) operation.

[0181] (2) Pooling layer

[0182] CNNs also have pooling layers, which allow to reduce the network dimension by taking the average or maximum value over small patches of the image (average pooling or max pooling), also called subsampling. Similar to the convolutional layers, the pooling layers act on small patches of the image, using a stride. In one example, referring to Figure 15 , consider a 4x4 input patch (620) performing max pooling with a stride s = 2, the output dimension of the output (622) is half of the input dimension in both horizontal and vertical directions. It is also possible to reduce the dimension of the convolutional layer by taking a stride larger than 1 without zero padding, but the advantage of pooling is that it makes the network less sensitive to small translations of the input image.

[0183] (3) Fully connected layer

[0184] After several convolutional and pooling layers, the CNNs usually end with several fully connected layers. The tensors output by the previous convolutional / pooling layers are transformed into a single vector value.

[0185] B. Application of CNNs in video coding

[0186] (1) In-loop filtering

[0187] In JVET-I0022, a convolutional neural network filter (CNNF) for intra frames is provided. The CNNF works as an in-loop filter for intra frames to replace the filters in the Joint Exploration Model (JEM), i.e., the bi-directional filter (BF), the de-blocking filter (DF), and the sample adaptive offset (SAO). Figure 16A The intra decoding process of JEM (630) is illustrated, which includes entropy decoding (631), inverse quantization (632, InvQ), inverse transform (633), BF (634), DF (635), SAO (636), prediction (637), and adaptive loop filter (ALF) (638). Figure 16B The intra decoding process is shown, including CNNF (644) instead of BF (634), DF (635), and SAO (636). For B and P frames, the filters can remain the same as in JEM 7.0.

[0188] Referring to Figure 16B and Figure 17, the CNNF (644) can include two inputs: the reconstruction parameters (652) and the quantization parameter (QP) (654), which can adapt the reconstruction of different qualities using a single parameter set. To better converge during the training process, both inputs can be normalized. To reduce complexity, a simple CNN of 10 layers can be employed. The CNN can be composed of one concatenation layer (656), seven convolutional layers (658A through 658G), where each convolutional layer is followed by a ReLU layer, one convolutional layer (660), and a summation layer (662). These layers can be connected one after another and form the network. It can be appreciated that the above layer parameters can be included in the convolutional layers. By connecting the reconstructed Y, U, or V to the summation layer, the network is regularized to learn the characteristics of the residual between the reconstructed image and its original image. According to embodiments, simulation results report that the BD rate of the luminance and two chroma components of JEM-7.0 under the AI configuration are saved by -3.57%, -6.17%, and -7.06%, and the encoding and decoding times are 107% and 12887% compared to the anchor, respectively.

[0189] In JVET-N0254, experimental results of a dense residual convolutional neural network based on an in-loop filter (DRNLF) are reported. Referring now to Figure 18 , a structural block diagram of an exemplary dense residual network (DRN) (670) is depicted. The network structure can include N dense residual units (DRUs) (672A through 672N), and M can represent the number of convolutional kernels. For example, N can be set to 4 and M can be set to 32 as a trade-off between computational efficiency and performance. A normalized QP mapping (674) can be concatenated with the reconstructed frame as an input to the DRN (670).

[0190] According to embodiments, the DRUs (672A through 672N) can each have the structure (680) shown in Figure 19 . The DRUs can propagate the input directly to the subsequent units through a shortcut. To further reduce the computational cost, a 3x3 depth-wise separable convolution (DSC) layer can be applied in the DRUs.

[0191] The output of the network can have three channels, which correspond to Y, Cb, Cr, respectively. The filter can be applied to both intra and inter pictures. An additional flag can be signaled for each CTU to indicate whether DRNLF is on / off. Experimental results of the embodiment show that the BD rate on Y, Cb and Cr components are -1.52%, -2.12% and -2.73% respectively in the all-intra configuration, -1.45%, -4.37% and -4.27% in the random access configuration, and -1.54%, -6.04% and -5.86% in the low-delay configuration. In this embodiment, the decoding time is 4667%, 7156% and 9127% in the AI, RA and LD B configurations.

[0192] (2) Intra prediction

[0193] Referring now to Figure 20 and Figure 21 , examples of a first process (690A) and a second process (690B) for intra prediction modes are depicted. Intra prediction modes can be used to generate an intra picture prediction signal on rectangular blocks in future video codecs. These intra prediction modes perform the following two main steps: First, a set of features is extracted from the decoded samples. Second, these features are used to select an affine linear combination of a predetermined set of image patterns as the prediction signal. In addition, a specific signaling scheme can be used for the intra prediction modes.

[0194] Referring to Figure 20 , on a given MxN block (692A) with M < 32 and N < 32, a set of reference samples r is processed via a neural network to generate a luminance prediction signal pred. The reference samples r can consist of K rows of size N+K and K columns of size M to the left of the block (692A). The number K can depend on M and N. For example, K can be set to 2 for all M and N.

[0195] The neural network (696A) can extract a feature vector ftrfrom the reconstructed samples r as follows. If d0= K*(N+M+K) denotes the number of samples of r, then r is viewed as a vector in a real vector space of dimension d0. For fixed integral square matrices A1and A2of size d0and for fixed integral bias vectors b1and b2of dimension d0, the following equation (8) is first computed.

[0196] t1= p(A1·r + b1) (Equation 8)

[0197] In equation (8), “·” denotes the ordinary matrix-vector product. In addition, the function p is an integer approximation of the ELU function p0, where the latter function is defined on a p-dimensional vector v as shown in the following equation (9).

[0198]

[0199] where p0(v) i and v i denote the i-th component of the vector. Similar operations are applied for t1 and t2 are computed as shown in equation (10) below.

[0200] t2 = p(A2 · t1 + b2) (equation 10)

[0201] For a fixed integer d1 with 0 < d1 < d0, one can predefine an integration matrix A3 of size d1 rows and d0 columns with one or more bias weights (694A), such as a predefined integration bias vector b3 of dimension d1, such that a feature vector ftr is computed as shown in equation (11) below.

[0202] ftr = p(A3 · t2 + b3) (equation 11)

[0203] The value of d1 depends on M and N. Now, let d 1= d0.

[0204] Outside the feature vector ftr, an affine linear mapping is used to generate the final prediction signal pred, followed by a standard clipping operation Clip that depends on the bit-depth. Thus, a predefined matrix A4 of size M*N rows and d1 columns and a predefined bias vector b4 of dimension M*N are used to compute pred as shown in equation (12) below:

[0205] pred = Clip(A4 · ftr + b4) (equation 12)

[0206] Now referring to Figure 21 n different intra prediction modes (698B) will be used, where n is set to 35 if max(M,N) < 32, otherwise it is set to 11. Thus, an index predmode with 0 < predmode < n will be signaled by the encoder and parsed by the decoder and can be signaled using the following syntax. With n = 3 + 2 k where k = 3 if max(M,N) = 32, otherwise k = 5. In a first step, the index predIdx with 0 < predIdx < n is signaled using the following code. First, one bin encodes whether predIdx < 3. If predIdx < 3, a second bin encodes whether predIdx = 0 and, if predIdx ≠ 0, another bin encodes whether predIdx is equal to 1 or 2. If predIdx ≥ 3, k bins are used to signal the value of predIdx in a standard way.

[0207] According to the index predldx, the actual index predmode is derived using a neural network (696B) with one hidden layer of fully connected neurons that uses the reconstructed samples r' as input, the reconstructed samples r' being located in the two upper rows (of size N+2) and the two left columns (of size M) of the block (692B).

[0208] The reconstructed samples r' are considered as a vector in a real vector space of dimension 2*(M+N+2). If there is a fixed square matrix A1' of 2*(M+N+2) rows and columns, and there is one or more bias weights (694B), such as a fixed bias vector b1' in a real vector space of dimension 2*(M+N+2), the t1' is computed as shown in equation (13) below.

[0209] t1' = p(A1' · r' + b1') (equation 13)

[0210] If there is a matrix A2' of n rows and 2*(M+N+2) columns, and there is a fixed bias vector b2' in a real vector space of dimension n, the lgt is computed as shown in equation (14) below.

[0211] lgt = A2' · t1' + b2' (equation 14)

[0212] The index predmode is now derived as the position of the predldx-th largest component of lgt. Here, if two components (lgt) k and (lgt) l are equal, if k < l, (lgt) k is considered greater than (lgt) l , otherwise (lgt) l is considered greater than (lgt) k .

[0213] [Multiple Transform Selection]

[0214] In addition to the DCT-II already employed in HEVC, a multiple transform selection (MTS) scheme can be used for residual coding of inter and intra coded blocks. The scheme can include a number of selected transforms from DCT8 / DST7. According to embodiments, DST-VII and DCT-VIII can be included. Table 4 shows the transform basis functions of the selected DST / DCT for N-point input.

[0215] Table 4: Transform basis functions of DCT-II / VIII and DST VII for N-point input

[0216]

[0217] To maintain the orthogonality of the transform matrix, the transform matrix is more quantization-accurate than in HEVC. To keep the intermediate values of the transform coefficients within the 16-bit range, all coefficients are required to have 10 bits after the horizontal transform and after the vertical transform.

[0218] To control the MTS scheme, separate enabling flags can be specified for intra and inter at the SPS level, respectively. When MTS is enabled at the SPS, a CU-level flag can be signaled to indicate whether to apply MTS. According to embodiments, MTS can be applied only to luma. MTS signaling can be skipped when one of the following conditions is applied: (1) the position of the last significant coefficient of the luma TB is less than 1 (i.e., only DC), or (2) the last significant coefficient of the luma TB is located within the MTS zero-output region.

[0219] If the MTS CU flag is equal to zero, DCT2 can be applied in both directions. However, if the MTS CU flag is equal to 1, two other flags can be additionally signaled to indicate the transform type in the horizontal direction and the vertical direction, respectively. Table 5 below shows an exemplary transform and signaling mapping table. The transform selection of ISP and implicit MTS can be unified by removing the intra mode and block shape dependency. If the current block is ISP mode, or if the current block is an intra block and both intra and inter explicit MTS are on, only DST7 can be used for the horizontal transform core and the vertical transform core. When it comes to transform matrix accuracy, 8-bit primary transform cores can be used. Thus, all transform cores used in HEVC can be kept the same, including 4-point DCT-2 and DST-7, 8-point, 16-point, and 32-point DCT-2. Also, other transform cores, including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7, and DCT-8, can use 8-bit primary transform cores.

[0220] Table 5: Transform and signaling mapping table

[0221]

[0222] To reduce the complexity of large-size DST-7 and DCT-8, for DST-7 and DCT-8 blocks with size (width or height, or both width and height) equal to 32, high-frequency transform coefficients can be zeroed out. Only coefficients within the 16x16 low-frequency region can be kept.

[0223] As in HEVC, the residual of a block can be coded in transform skip mode. To avoid the redundancy of syntax coding, the transform skip flag can not be signaled when CU level MT S CU flag is not equal to zero. According to embodiments, the implicit MTS transform can be set to DCT2 when LFNST or MIP is activated for the current CU. Furthermore, implicit MTS can still be enabled when MTS is enabled for inter coded blocks.

[0224] [Non-separable secondary transform]

[0225] In JEM, a pattern dependent non-separable secondary transform (NSST) can be applied between the forward core transform and quantization (at the encoder) and between dequantization and inverse core transform (at the decoder). To keep low complexity, the NSST can be applied only to the low frequency coefficients after the primary transform. If the width (W) and height (H) of the transform coefficient block are both greater than or equal to 8, an 8x8 non-separable secondary transform can be applied to the top-left 8x8 region of the transform coefficient block. Otherwise, if either W or H of the transform coefficient block is equal to 4, a 4x4 non-separable secondary transform can be applied and a 4x4 non-separable transform can be performed on the top-left min(8, W)xmin(8, H) region of the transform coefficient block. The above transform selection rule can be applied to luma and chroma components.

[0226] The matrix multiplication implementation of the non-separable transform can be performed as described in the above “Secondary transform in VVC” sub-section, referring to equations (2) to (3). According to embodiments, the non-separable secondary transform can be implemented using direct matrix multiplication.

[0227] [Pattern dependent transform core selection]

[0228] For 4x4 and 8x8 block sizes, there can be 35x3 non-separable secondary transforms, where 35 is the number of transform sets specified by the intra prediction mode and 3 is the number of non-separable secondary transform (NSST) candidates for each intra prediction mode. The mapping from the intra prediction mode to the transform set is defined as shown in Table 700 in Figure 22 According to Table 700, the transform set applied to the luma / chroma transform coefficients can be specified by the corresponding luma / chroma intra prediction mode. For intra prediction modes greater than 34 (diagonal prediction directions), the transform coefficient block can be transposed at the encoder / decoder, before / after the secondary transform.

[0229] For each set of transforms, the selected non-separable secondary transform candidate can be further specified by an explicitly signaled CU-level NSST index. The index can be signaled once per intra CU after using the transform coefficients and truncated unary binarization. The truncation value can be 2 in the case of planar or DC mode, and 3 in the case of angular intra prediction modes. The NSST index can be signaled only when there is more than one non-zero coefficient in the CU. When not signaled, the default value can be zero. A zero value of this syntax element can indicate that the secondary transform is not applied to the current CU, and values 1 to 3 can indicate which secondary transform from the set is applied.

[0230] In JEM, NSST can not be applied to blocks coded in transform skip mode. When the NSST index is signaled for a CU and the NSST index is not equal to zero, NSST can not be used for blocks in the CU coded in transform skip mode. The NSST index can not be signaled for a CU when the CU is coded in transform skip mode or the number of non-zero coefficients of non-transform skip mode CBs is less than 2.

[0231] [Problems in transform schemes of comparative embodiments]

[0232] In comparative embodiments, the separable transform scheme is not very effective for capturing directional texture patterns (e.g., edges in 45 / 135 degree directions). The non-separable transform scheme can help improve coding efficiency in those cases. To reduce computational complexity and memory footprint, the non-separable transform scheme is typically designed as a secondary transform applied on top of the low frequency coefficients of the primary transform. In existing implementations, the transform kernel to be used (from a set of transform kernels, primary / secondary and separable / non-separable) is selected based on the prediction mode information. But the prediction mode information alone can only provide a coarse representation of the entire residual mode space observed by that prediction mode, as shown in representations 710, 720, 730, and 740. Representations 710, 720, 730, and 740 show residual modes observed in the D45 (45°) intra prediction mode in AV1. The neighboring already reconstructed samples can provide additional information for a more efficient representation of those residual modes. Figure 23A to Figure 23D

[0233] ​For transform schemes with multiple transform kernel candidates, it can be necessary to use coded information available to both the encoder and the decoder to identify the transform set. In existing multiple transform schemes such as MTS and NSST, the transform set is selected based on coding prediction mode information such as intra prediction mode. However, the prediction mode does not cover all the statistics of the prediction residual completely, and neighboring reconstructed samples can provide additional information for more efficient classification of the prediction residual. A neural network based approach can be applied to the efficient classification of the prediction residual and thus provide more efficient transform set selection.

[0234] [Example aspects of embodiments of the present application]

[0235] Embodiments of the present application can be used alone or in any combination in any order. Furthermore, each of the embodiments (e.g., the method, the encoder, and the decoder) can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program that is stored in a non-transitory computer-readable medium.

[0236] Embodiments of the present application can incorporate any number of aspects as described above. Embodiments of the present application can also incorporate one or more aspects described below and address the problems discussed above and / or other problems.

[0237] A. First aspect

[0238] According to embodiments, the neighboring reconstructed samples can be used for selecting the transform set.

[0239] In one or more embodiments, a sub-group of transform sets is selected from a group of transform sets using first coded information, such as coded information of a prediction mode (e.g., intra prediction mode or inter prediction mode). In one embodiment, one transform set is identified from the selected transform set sub-group using second coded information, such as type of intra / inter prediction mode, block size, prediction block samples of the current block, and neighboring reconstructed samples of the current block. It can be seen that the second coded information is different from the first coded information. Finally, a transform candidate for the current block is selected from the identified transform set using an associated index identified in the bitstream. In one embodiment, the final transform candidate is implicitly identified from the selected transform set sub-group using other coded information, such as type of intra / inter prediction mode, block size, prediction block samples of the current block, and neighboring reconstructed samples of the current block.

[0240] In one or more embodiments, the set of neighboring reconstructed samples can include samples from previously reconstructed neighboring blocks. In one embodiment, the set of neighboring reconstructed samples can include one or more lines of top neighboring reconstructed samples and one or more lines of left neighboring reconstructed samples. In one example, the number of lines of top neighboring reconstructed samples and / or left neighboring reconstructed samples is the same as the maximum number of lines of neighboring reconstructed samples used for intra prediction. In one example, the number of lines of top neighboring reconstructed samples and / or left neighboring reconstructed samples is the same as the maximum number of lines of neighboring reconstructed samples used for CfL prediction mode. In one embodiment, the set of neighboring reconstructed samples can include all samples from neighboring reconstructed blocks.

[0241] In one or more embodiments, the set of transform sets includes only primary transform kernels, only secondary transform kernels, or a combination of primary transform kernels and secondary transform kernels. In the case where the set of transform sets includes only primary transform kernels, the primary transform kernels can be separable, can be non-separable, can use different types of DCT / DST, or use different line graph transforms with different self-loop rates. In the case where the set of transform sets includes only secondary transform kernels, the secondary transform kernels can be non-separable, or use different non-separable line graph transforms with different self-loop rates.

[0242] In one or more embodiments, the neighboring reconstructed samples can be processed to derive an index associated with a particular transform set. In one embodiment, the neighboring reconstructed samples are input to a transform process, and the transform coefficients are used to identify the index associated with the particular transform set. In one embodiment, the neighboring reconstructed samples are input to multiple transform processes, and a cost function is used to evaluate a cost value for each transform process. The cost value is then used to select the transform set index. Exemplary cost values include, but are not limited to, the sum of the magnitudes of the first N (e.g., 1, 2, 3, 4, …, 16) transform coefficients along some scan order. In one embodiment, a classifier is predefined, and the neighboring reconstructed samples are input to the classifier to identify the transform set index.

[0243] B. Second Aspect

[0244] According to embodiments, a neural network based transform set selection scheme can be provided. The input to the neural network includes, but is not limited to, the predicted block samples of the current block, the neighboring reconstructed samples of the current block, and the output can be an index used to identify the transform set.

[0245] In one or more embodiments, a set of transform sets is defined, and using coded information such as a prediction mode (e.g., an intra prediction mode or an inter prediction mode), a sub-set of the transform sets is selected, and then using other coded information such as prediction block samples of the current block, neighboring reconstructed samples of the current block, a transform set of the selected sub-set of transform sets is identified. Then, using an associated index identified in the bitstream, a transform candidate for the current block is selected from the identified transform set.

[0246] In one or more embodiments, the neighboring reconstructed samples can include one or more top neighboring reconstructed sample lines and left neighboring reconstructed sample lines. In one example, the number of top neighboring reconstructed sample lines and / or left neighboring reconstructed sample lines is the same as the maximum number of neighboring reconstructed sample lines used for intra prediction. In one example, the number of top neighboring reconstructed sample lines and / or left neighboring reconstructed sample lines is the same as the maximum number of neighboring reconstructed sample lines used for CfL prediction mode.

[0247] In one or more embodiments, the neighboring reconstructed samples and / or prediction block samples of the current block are inputs to a neural network, and the output includes not only an identifier of a transform set, but also an identifier of a prediction mode set. In other words, the neural network uses the neighboring reconstructed samples and / or prediction block samples of the current block to identify certain combinations of transform sets and prediction modes.

[0248] In one or more embodiments, the neural network is used to identify a transform set for a secondary transform. Optionally, the neural network is used to identify a transform set for a primary transform. Optionally, the neural network is used to identify a transform set for a combination of a secondary transform and a primary transform. In one embodiment, the secondary transform uses a non-separable transform scheme. In one embodiment, the primary transform can use different types of DCT / DST. In another embodiment, the primary transform can use different line graph transforms with different self-loop rates.

[0249] In one or more embodiments, for different block sizes, the neighboring reconstructed samples and / or prediction block samples of the current block can be further up-sampled or down-sampled before being used as inputs to the neural network.

[0250] In one or more embodiments, for different internal bit depths, the neighboring reconstructed samples and / or prediction block samples of the current block can be further scaled (or quantized) according to the internal bit depth values before being used as inputs to the neural network.

[0251] In one or more embodiments, parameters used in the neural network depend on coded information, including but not limited to: whether the block is intra coded, block width and / or block height, quantization parameter, whether the current picture is coded as an intra (key) frame, and intra prediction mode.

[0252] According to embodiments, at least one processor and a memory storing computer program instructions can be provided. The computer program instructions, when executed by the at least one processor, can implement an encoder or a decoder and can perform any number of the functions described in the present application. For example, with reference to Figure 24 , the at least one processor can implement a decoder (800). The computer program instructions can include, for example, a decoding code (810) configured to cause the at least one processor to decode a block of a picture from a coded bitstream received (e.g., from an encoder). The decoding code (810) can include, for example, a transform set selection code (820), a transform selection code (830), and a transform code (840).

[0253] The transform set selection code (820) can cause the at least one processor to select a transform set according to embodiments of the present application. For example, the transform set selection code (820) can cause the at least one processor to select a transform set based on at least one neighboring already reconstructed sample, where the at least one neighboring already reconstructed sample is from at least one previously decoded neighboring block or from a previously decoded picture. According to embodiments, the transform set selection code (820) can be configured to cause the at least one processor to select a sub-group of transform sets from a group of transform sets based on first coded information and select a transform set from the sub-group according to embodiments of the present application.

[0254] The transform selection code (830) can cause the at least one processor to select a transform candidate from a transform set according to embodiments of the present application. For example, the transform selection code (830) can cause the at least one processor to select a transform candidate from a transform set based on an index value identified in a coded bitstream according to embodiments of the present application.

[0255] The transform code (840) can cause the at least one processor to perform an inverse transform on coefficients of the block using a transform (e.g., a transform candidate) from a transform set according to embodiments of the present application.

[0256] According to embodiments of the present application, the decoding code 810 can cause a neural network to be used to select a transform group, a transform sub-group, a transform set, and / or a transform, or otherwise perform at least a portion of the decoding. According to embodiments, the decoder (800) can further include a neural network code (850) configured to cause the at least one processor to implement a neural network according to embodiments of the present application.

[0257] According to embodiments, the encoder side of the above described process, can encode the picture by encoding code based on the above description, as can be understood by a person skilled in the art.

[0258] The techniques in the embodiments of the application described above can be implemented as computer software using computer readable instructions and physically stored in at least one computer readable medium. For example, Figure 25 A computer device 2500 is shown, which is suitable for implementing certain embodiments of the disclosed subject matter.

[0259] The computer software can be coded using any suitable machine code or computer language, that can be subject to assembly, compilation, linking, or like mechanisms to create code, that can be executed by at least one computer central processing unit (CPU), Graphics Processing Unit (GPU), or hardware.

[0260] The instructions can be executed using one or more of serial processing, multi -threaded processing, multi -core processing, message passing, and instruction -level parallelism.

[0261] Figure 25 The components shown in the computer device (900) are exemplary and not limiting, and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing embodiments of the present application. Neither should the configuration of components be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the exemplary embodiment of a computer device (900).

[0262] The computer device (900) can include certain human interface input devices. Such a human interface input device can be a keyboard, pointing device, microphone, or the like.

[0263] The human interface input devices can also be used to capture certain media, such as audio (for example, voice inputs), images (for example, taking a picture), or video (for example, taking a video), and the like.

[0264] Computer device (900) can also include certain human interface output devices. Such human interface output devices can be stimulating human senses using, for example, tactile output, sound, light, and smell / taste. Examples of such human interface output devices can include tactile output devices (for example, a display (910), a projector, speakers (909), a haptic feedback device, etc.), an audio output device (for example, speakers (909), headphones (not depicted)), a video output device (for example, a display (910), a projector, etc.), an ophthalmic device (for example, a near-eye out (NEO) device, a heads-up display (HUD) device, a contact lens device, etc.), a printer (not depicted), and the like. Some of these human interface output devices can also be used as human interface input devices when such devices are capable of generating input corresponding to sensed user interaction. For example, motion of a user’s hand can be translated to position values of a pointer on a display and / or values for controlling an application (for example, a game).

[0265] Computer device (900) can also include human accessible storage devices and their associated media and connectors. Examples can include optical media, various types of removable disks, etc. Note that these components can be used in a wide variety of computer devices such as computer kiosks, cell phones, terminals, set-top boxes, network PCs, and many others.

[0266] Those skilled in the art will further appreciate that the term “computer-readable medium” used in connection with the disclosed subject matter does not include a transitory medium.

[0267] The computer device (900) can also include an interface to one or more communication networks. For example, the network(s) can be a wireless network, a wired network, an optical network, etc. The network(s) can further be a local area network, a wide area network, a metropolitan area network, a vehicular area network, and an industrial network, a real-time network, a delay-tolerant network, etc. The network(s) also includes local area networks, wireless local area networks, cellular networks (GSM, 3G, 4G, 5G, LTE, etc.), cable networks, the Internet, bus and ring networks, and so on. Some networks typically require an external network interface adapter that connects to a communication port or peripheral bus (949) of the computer device (900) (e.g., the USB port of a computer device (900)); others are commonly integrated into the core of the computer device (900) by connection to a system bus, as described below (e.g., an Ethernet interface integrated into a PC computer device or a cellular network interface integrated into a smartphone computer device). Using any of these networks, the computer device (900) can communicate with other entities. The communication can be one-way (only reception), e.g., radio, or one-way (only transmission), e.g., a CAN bus to some CAN bus devices, or two-way, e.g., to other computer devices via local or wide area digital networks. The communication includes communication with a cloud computing environment (955). Each of the above networks and network interfaces can use certain protocols and protocol stacks.

[0268] The above human interface devices, human-accessible storage devices, and network interfaces (954) can be attached to the core (940) of the computer device (900).

[0269] The core (940) can include at least one central processing unit (CPU) (941), graphics processing unit (GPU) (942), dedicated programmable processing units such as field programmable gate arrays (FPGAs) (943), hardware accelerators for certain tasks (944), and so on. These devices, along with read-only memory (ROM) (945), random-access memory (RAM) (946), internal mass storage such as internal non-user accessible hard drives, solid-state drives, or the like (947), can be connected through a system bus (948). In some computer devices, the system bus (948) can be accessible in the form of at least one physical plug to allow for expansion by additional CPUs, graphics processors, and the like. The peripheral devices can be attached either directly to the core's system bus (948), or through a peripheral bus (949). Architectures for a peripheral bus include Peripheral Component Interconnect (PCI), USB, and the like.

[0270] CPUs (941), GPUs (942), FPGAs (943), and accelerators (944) can execute certain instructions that, taken either alone or in combination, can implement the computer code. That computer code can be stored in ROM (945) or RAM (946). Transitional data for the CPU (941) can be stored in RAM (946), whereas permanent data can be stored for example, in the internal mass storage (947). Fast storage and retrieval speeds can be achieved using cache memory, which can be closely associated with one or more processor cores (940), the internal mass storage (947), ROM (945), RAM (946), or the like.

[0271] The computer-readable medium can have computer code thereon for performing various computer-implemented operations. The media and computer code can be those specially designed and constructed for the purposes of the present application, or they can be of the kind known or used by those having skill in the computer software art.

[0272] As an example and not by way of limitation, the computing device (900) having architecture, and specifically the core (940), can provide functionality as a processor, including a CPU, GPU, FPGA, accelerator, or the like, that executes software provided in one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage as described above as well as media stored within the core (940), such as carrier waves for carrier-bound implementations, including but not limited to implementations using Java®applets, Virtual Machines, or other platforms. Software provided by one or more tangible computer-readable media can enable the core (940), and specifically the processor(s) therein, to provide various functionality including functions described herein as being performed by a server, a processor, a processor core, or the like. In some embodiments, software can be downloadable to the core (940) from a remote location, such as a server, or software can be downloaded to core (940) from another computer-readable media, such as a removable storage drive or disk. The software, when run by the core (940), can enable the core (940) to implement various aspects of the present application.

[0273] While at least two example embodiments have been described herein, various alterations, permutations and equivalents thereof will be apparent to those skilled in the art without departing from the scope of the application. Accordingly, the embodiments described herein are intended to be illustrative only and are not limiting of the scope of the application.

Claims

1. A video decoding method, characterized in that, include: Receive encoded bit stream; Decoding blocks of images in the encoded bitstream specifically includes: Using the encoded information of the intra-prediction mode, a subgroup of the transform set is selected from a set of transform sets; A transform set is identified from the subgroup using the type of intra-prediction mode, block size, predicted block samples of the current block, and neighboring reconstructed samples of the current block; Using the associated index identified in the encoded bitstream, a transform candidate for the current block is selected from the identified transform set; and, Using the transformation candidate, perform an inverse transformation on the coefficients of the current block.

2. The method according to claim 1, wherein, The set of transformations includes only quadratic transformation kernels.

3. The method according to claim 2, wherein, The quadratic transformation kernel is inseparable.

4. The method according to any one of claims 1 to 3, wherein, The set of transformations is a quadratic transformation.

5. A video decoding device, characterized in that, include: The decoding module is used to decode blocks of images in the received encoded bitstream; The decoding module includes: A transform set selection module is used to select a subgroup of transform sets from a set of transform sets using encoded information of the intra-prediction mode; identify a transform set from the subgroup using the type of the intra-prediction mode, block size, predicted block samples of the current block, and neighboring reconstructed samples of the current block; and select a transform candidate for the current block from the identified transform set using an associated index identified in the encoded bitstream. A transformation module is used to perform an inverse transformation on the coefficients of the current block using the transformation candidate.

6. A computer device, characterized in that, It includes a processor and a memory, wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the method as described in any one of claims 1-4.

7. A non-transitory computer-readable medium, characterized in that, It stores computer instructions that, when executed by at least one processor, implement the method as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Method and apparatus for transform selection in video encoding and decoding

    US20110268183A1