Neural Network Codec with Hybrid Entropy Model and Flexible Quantization

The hybrid entropy model and flexible quantization in neural video codecs enhance compression efficiency by exploiting spatio-temporal correlations and adaptive bit allocation, addressing limitations in existing technologies and achieving superior performance to conventional standards.

JP2025520660APending Publication Date: 2025-07-03MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024575277
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-06-21
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing neural video codecs struggle to effectively exploit spatio-temporal correlations and achieve efficient rate-distortion performance due to limitations in entropy models and fixed quantization steps, leading to suboptimal compression quality and efficiency compared to conventional standards like H.265.

Method used

Incorporation of a hybrid entropy model that utilizes both spatial and temporal dependencies through a latent prior distribution and double spatial prior distribution, combined with a flexible quantization mechanism for adaptive bit allocation, enabling efficient exploitation of frame correlations and dynamic rate adjustment.

Benefits of technology

The hybrid entropy model and flexible quantization improve rate-distortion performance by leveraging temporal and spatial redundancies, allowing for better compression quality and efficiency in neural video codecs, potentially surpassing conventional standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025520660000001_ABST
    Figure 2025520660000001_ABST
Patent Text Reader

Abstract

Innovations in systems, methods, and software for neural image or video codec features are described herein. For example, a neural video encoder can receive a current video frame, encode the current video frame to produce encoded data, and output the encoded data as part of a bitstream. As part of the encoding, the encoder can determine a current latent representation of the current video frame and encode the current latent representation using an entropy model network that includes one or more convolutional layers. As part of encoding the current latent representation, the encoder can estimate statistical characteristics of a quantized version of the current latent representation based at least in part on a previous latent representation of a previous video frame and entropy encode the quantized version of the current latent representation based at least in part on the estimated statistical characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001]

[0001] Engineers use compression (also called source coding or source encoding) to reduce the bitrate of digital video. Compression reduces the cost of storing and transmitting video information by converting the information into a lower bitrate format. Decompression (also called decoding) reconstructs a version of the original information from the compressed format. A “codec” is an encoder / decoder system.

[0002]

[0002] Over the past few decades, various video codec standards have been adopted, including the ITU-T H.261, H.262 (MPEG-2 or ISO / IEC 13818-2), H.263, H.264 (MPEG-4 AVC or ISO / IEC 14496-10), H.265 / HEVC, H.266 / VVC (ISO / IEC 23090-3 or MPEG-I Part 3) standards, the MPEG-1 (ISO / IEC 11172-2) standard and the MPEG-4 Visual (ISO / IEC 14496-2) standard, as well as the SMPTE 421M (VC-1) standard. Such video codec standards typically define options for the syntax of the encoded video bitstream that detail the parameters within the bitstream when certain features are used in encoding and decoding. Often, the video codec standards also specify details about the decoding operations that a video decoder should perform to obtain a conforming result in decoding. In addition to codec standards, various proprietary codec formats define other options for the syntax of the encoded video bitstream and the corresponding decoding operations.

[0003]

[0003] More recently, some codecs use neural networks and other machine learning methods for data compression. For example, neural image codecs have been developed that compress / decompress images using an entropy model neural network (or "entropy model network" or simply "entropy model") designed to predict the probability distribution of a quantized latent representation of an image. Based on a similar concept, neural video codecs have been developed that compress / decompress video frames using an entropy model. Despite the recent success of neural video codecs compared to conventional video compression / decompression technologies, there is room for improvement to enhance compression quality and / or efficiency.

Summary of the Invention

[0004] In summary, innovations in efficient and high-quality codec technology are described herein. Some of the innovations described herein use an improved entropy model for neural codecs that can efficiently utilize both spatial and temporal dependencies between video frames. Other innovations described herein provide a flexible quantization approach in neural codecs. As described more fully below, the innovations described herein include, but are not limited to, incorporating the previous latent representation of the previous video frame (the "latent prior") into the entropy model to exploit correlations between latent representations, incorporating inter-channel, inter-region prediction (the "dual spatial prior") into the entropy model to exploit spatial redundancy in a manner that is easy to parallelize, and incorporating a flexible quantization mechanism that achieves multiple rates in a single neural codec system by dynamic bit allocation and improves rate-distortion ("RD") performance. The innovations described herein may be implemented in neural video codecs and, optionally, neural image codecs. The innovations described herein may be implemented for future codec standards or formats.

[0005] According to one aspect of the innovation described herein, a neural video encoder can receive a current video frame, encode the current video frame to generate encoded data, and output the encoded data as part of a bitstream. As part of the encoding, the encoder can determine a current latent representation of the current video frame and encode the current latent representation using an entropy model network that includes one or more convolutional layers. As part of encoding the current latent representation using the entropy model network, the encoder can estimate statistical properties of a quantized version of the current latent representation based at least in part on a previous latent representation of a previous video frame and entropy encode the quantized version of the current latent representation based at least in part on the estimated statistical properties. In some cases, using the previous latent representation as an input to the entropy model network helps to exploit temporal redundancy to improve the RD performance of the neural video encoder.

[0006]

[0006] The corresponding neural video decoder can receive the encoded data as part of a bitstream, decode the encoded data to reconstruct a current video frame, and output the reconstructed current video frame. As part of the decoding, the decoder can use an entropy model network that includes one or more convolutional layers to reconstruct a current latent representation of the current video frame. As part of reconstructing the current latent representation, the decoder can estimate statistical properties of a quantized version of the current latent representation based at least in part on a previous latent representation of a previous video frame and entropy decode the quantized version of the current latent representation based at least in part on the estimated statistical properties.

[0007] According to another aspect of the innovation described herein, a neural image encoder or neural video encoder can receive a current frame, encode the current frame to generate encoded data, and output the encoded data as part of a bitstream. As part of the encoding, the encoder can determine a current latent representation of the current frame and encode the current latent representation using an entropy model network that includes one or more convolutional layers. Elements of the current latent representation can be logically organized along the channel dimension and two spatial dimensions. As part of encoding the current latent representation, the encoder can divide the elements of the current latent representation into multiple sets of elements of different channel sets along the channel dimension and different spatial position sets along the two spatial dimensions, where each of the multiple sets of elements has a different combination of one of the different channel sets and one of the different spatial position sets. Then, the encoder can estimate the statistical characteristics of the quantized version of the second set of elements among the multiple sets of elements, including at least partially based on the quantized version of the first set of elements among the multiple sets of elements. Further, the encoder can entropy encode the quantized versions of the multiple sets of elements, respectively, based at least in part on the estimated statistical characteristics. In some cases, using inter-set estimation in the entropy model network helps to exploit spatial redundancy (and potentially channel redundancy) to improve the RD performance of the neural encoder.

[0008]

[0008] The corresponding neural image decoder or neural video decoder can receive the encoded data as part of a bitstream, decode the encoded data, reconstruct the current frame, and output the reconstructed current frame. As part of the decoding, the decoder can reconstruct the current latent representation of the current frame using an entropy model network that includes one or more convolutional layers. The elements of the current latent representation can be logically organized along the channel dimension and two spatial dimensions. The elements of the current latent representation are divided into a plurality of sets of elements of different channel sets along the channel dimension and different spatial position sets along the two spatial dimensions. Each of the plurality of sets of elements has a different combination of one of the different channel sets and one of the different spatial position sets. As part of reconstructing the current latent representation, the decoder can estimate the statistical characteristics of the quantized version of the second set of elements among the plurality of sets of elements, including at least partially based on the quantized version of the first set of elements among the plurality of sets of elements. Furthermore, the decoder can entropy-decode the quantized versions of the plurality of sets of elements respectively based at least partially on the estimated statistical characteristics.

[0009] According to another aspect of the innovation described herein, a neural image encoder or a neural video encoder can receive a current frame, encode the current frame to generate encoded data, and output the encoded data as part of a bitstream. As part of the encoding, the encoder can determine a current latent representation of the current frame. The elements of the current latent representation are logically organized along the channel dimension and two spatial dimensions. As part of the encoding, the encoder can quantize the current latent representation at multiple stages using different quantization step (QS) values, and by quantizing, generate a quantized version of the current latent representation. Further, the encoder can entropy encode the quantized version of the current latent representation. In some cases, using multiple stages of quantization helps provide flexibility in using the neural encoder over a range of QS values for different levels of quality and bitrate.

[0010] A corresponding neural image decoder or neural video decoder can receive the encoded data as part of a bitstream, decode the encoded data to reconstruct the current frame, and output the reconstructed current frame. As part of the decoding, the decoder can reconstruct the current latent representation of the current frame. The elements of the current latent representation are logically organized along the channel dimension and two spatial dimensions. As part of reconstructing the current latent representation, the decoder can entropy decode the quantized version of the current latent representation and inverse quantize the quantized version of the current latent representation at multiple stages using different QS values.

[0011]

[0011] The innovations may be implemented as part of a method, as part of a computer system configured to perform the operations for the method, or as part of one or more computer-readable media that store computer-executable instructions for causing a computer system to perform the operations for the method. The various innovations may be used in combination or separately. This Summary is provided to introduce a set of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. These and other objects, features, and advantages of the present invention will become more apparent from the Detailed Description below, which proceeds with reference to the accompanying figures. [Brief description of the drawings]

[0012]

Figure 1

[0012] FIG. 1 illustrates an example computer system on which some described embodiments may be implemented.

Figure 2A

[0013] FIG. 2a illustrates an example network environment in which some described embodiments may be implemented.

Figure 2B

Figure 3

[0014] FIG. 1 illustrates an example neural video encoder in conjunction with which some described embodiments may be implemented, and an example neural video decoder in conjunction with which some described embodiments may be implemented.

Figure 4

[0015] FIG. 1 illustrates an example entropy model network and segmentation, quantization, dequantization, and concatenation operations for a neural codec in conjunction with which some described embodiments may be implemented.

Figure 5A

[0016] FIG. 5A is a diagram showing features of an exemplary entropy model network that receives a previous latent representation as input.

Figure 5B

[0017] FIG. 5B is a diagram showing features of processing a hyper prior of a current latent representation for an exemplary entropy model network.

Figure 6

[0018] FIG. 8 is a diagram showing features of multi-stage quantization and corresponding inverse quantization in a neural codec system in which some of the described embodiments may be implemented in relation.

Figure 7

[0019] FIG. 9 is a set of screenshots showing features of flexible quantization in a neural codec system.

Figure 8A

[0020] FIG. 8A is a diagram showing an exemplary network structure of a contextual encoder for motion vector information in an exemplary neural video codec system.

Figure 8B

Figure 9A

[0021] FIG. 9A is a diagram showing an exemplary network structure of a contextual encoder for sample value information in an exemplary neural video codec system.

Figure 9B

Figure 10A

[0022] FIG. 10A is a diagram showing an exemplary network structure of a frame generator in an exemplary neural video codec system.

Figure 10B

Figure 11

[0023] It is a diagram showing an exemplary network structure of an entropy model network that uses set - to - set estimation with the previous latent representation as input in an exemplary neural video codec system.

Figure 12

[0024] It is a diagram showing an exemplary network structure of a temporal content encoder.

Figure 13A

[0025] Figure 13A is a diagram showing an exemplary network structure of a hyper - prior distribution encoder for the current latent representation in an exemplary neural video codec system.

Figure 13B

Figure 14A

[0026] Figure 14A is a diagram showing an exemplary network structure of a contextual encoder for the latent representation in an exemplary neural image codec system.

Figure 14B

Figure 15A

[0027] Figure 15A is a flowchart showing a generalized technique for encoding that implements one or more of the innovations described herein.

Figure 15B

Figure 16A

[0028] Figure 16A is a flowchart showing a generalized technique for determining the current latent representation in some exemplary embodiments.

Figure 16B

Figure 17A

[0029] FIG. 17A is a flowchart showing a generalized technique for using a previous latent representation as an input to an entropy model network during encoding in some exemplary embodiments.

Figure 17B

Figure 18A

[0030] FIG. 18A is a flowchart showing a generalized technique for using set - to - set estimation in an entropy model network during encoding in some exemplary embodiments.

Figure 18B

Figure 19A

[0031] FIG. 19A is a flowchart showing a generalized technique for multi - stage quantization in some exemplary embodiments.

Figure 19B

[0013]

[0032] The detailed description presents innovations in efficient and high-quality codec technology that uses an improved entropy model capable of efficiently exploiting both the spatial and temporal dependencies between frames and uses flexible quantization. As more fully described below, the innovations described herein include, but are not limited to, incorporating a latent prior distribution (e.g., a latent representation prior to sample value information or motion vector information) into the entropy model to exploit the correlation between latent representations and thereby improve the RD performance of a neural codec system; exploiting the spatial redundancy between sets of elements in a manner that is amenable to parallelization and thereby improving the RD performance of a neural codec system by incorporating a double spatial prior distribution (e.g., a pipeline that divides elements of a latent representation into multiple sets of elements for inter-set prediction / estimation) into the entropy model; implementing multiple rates in a single neural codec system through dynamic bit allocation and improving the RD performance by incorporating a flexible quantization mechanism. The innovations described herein may be implemented for future video codec standards or formats.

[0014]

[0033] In the examples described herein, the same reference numerals in different figures indicate the same components, modules, or operations. Depending on the context, a given component or module may receive different types of information as input and / or generate different types of information as output or be processed in different ways.

[0015]

[0034] More broadly, there can be various alternatives to the examples described herein. For example, some of the methods described herein can be changed by changing the order of acts of the described method, splitting, repeating, or omitting certain method acts. The various aspects of the disclosed technology can be used in combination or separately. Different embodiments use one or more of the innovations described. Some of the innovations described herein address one or more of the problems pointed out in the background art. Generally, a given technology / tool does not solve all such problems.

[0016] I. Overview of Neural Video Codec

[0035] In recent years, the development of neural image codec technology has been seen. Most neural image codec technologies focus on designing an entropy model for predicting the probability distribution of the quantized latent representation of an image, for example, by using a factorized model, a hyperprior distribution, an auto-regressive prior, a mixture of Gaussian models, a transformer-based model, and the like. Benefiting from these continuously improved entropy models, the compression ratio of neural image codecs has been shown to outperform more conventional image codec technologies such as the intra coding of H.266. Triggered by the success of neural image codec technology, recently, neural video codec technology has been attracting increasing attention.

[0017]

[0036] Most existing research on neural video codecs can be broadly classified into three categories, namely, residual coding-based solutions, conditional coding-based solutions, and 3D autoencoder-based solutions. Residual coding techniques are derived from conventional hybrid video codec architectures. In particular, when encoding the current frame, first, motion-compensated prediction is generated, and then the residual of that prediction with respect to the current frame is encoded. Regarding conditional coding-based solutions, the temporal frame or feature set of the previous frame serves as a condition for encoding the current frame. When compared with residual coding, conditional coding has been shown to have lower or equal entropy bounds. 3D autoencoder-based solutions are a natural extension of neural image codec technology by expanding the dimensions of the input. However, 3D autoencoder-based solutions can be associated with increased encoding latency and can significantly increase memory costs. Generally, most of these existing studies focus on how to generate the latent representation of video frames by exploring different data flows or network structures. Regarding the entropy model, most of these existing methods directly use off-the-shelf solutions (such as super-prior distributions, autoregressive prior distributions, etc.) borrowed from neural image codec technology to encode the latent representation of the current frame. In the design of the entropy model for neural video codec technology, spatio-temporal correlations have not been fully explored. As a result, the RD performance of previous neural video codec technologies has been limited and has been shown to be only slightly better than H.265 encoding.

[0018] II. Overview of an Improved Neural Video Codec with a Hybrid Entropy Model and Flexible Quantization

[0037] The technology described herein improves neural video codecs by incorporating a hybrid entropy model that can efficiently utilize both spatial and temporal correlations between video frames and / or within video frames. Some aspects of the technology described herein can also be used for neural image codecs.

[0019]

[0038] According to one aspect of the disclosed technology, the previous latent representation of the previous video frame (hereinafter also referred to as the "latent prior distribution") is included in the entropy model. Using the latent prior distribution can help utilize the temporal correlation of latent representations between video frames. As will be more fully described below, the quantized latent representation of the previous video frame can be used to predict the distribution of the quantized latent representation of the current video frame. A cascaded training strategy forms a propagation chain of latent representations. Thus, an implicit connection between the latent representation of the current video frame and the latent representation of the long-term reference frame can be established. Such a connection can help the neural codec further utilize the temporal redundancy between latent representations.

[0020]

[0039] According to another aspect of the disclosed technology, for a neural video codec or a neural image codec, in order to utilize the spatial redundancy within a frame, the entropy model includes the characteristics of a double spatial prior distribution. Most existing neural codecs rely on an "autoregressive prior distribution" to utilize spatial correlation. However, the autoregressive prior distribution is a serialized solution and follows a strict scanning order. As a result, neural codecs based on the autoregressive prior distribution are difficult to parallelize and tend to have a very slow speed. In contrast, the double spatial prior distribution described herein is a two-step encoding solution based on a much more time-efficient, improved checkerboard context model. Previously, He et al. presented the checkerboard context model in "Checkerboard context model for efficient learned image compression" in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 14771 - 14780, 2021. There, all channels follow the same encoding order (for example, elements at even positions are always encoded first and then used as context for encoding elements at odd positions). Such an approach may not be able to efficiently handle the content of a specific video because encoding even positions first may have worse RD performance than encoding odd positions first. In contrast, as will be more fully explained below, the double spatial prior distribution first encodes half of the latent representations of elements at both odd and even positions, and then encodes the other half of the latent representations, introducing a mechanism that can benefit from the context from elements at all (both odd and even) positions. Additionally, the correlation between multiple channels of the latent representation can also be utilized during the two-step encoding.Without introducing extra coding dependencies, the double-space prior distribution approach expands the scope or dimension of the spatial context and utilizes the channel context. As a result, more accurate predictions regarding the probability distribution of the quantized latent representation can be achieved.

[0021]

[0040] According to a further aspect of the present disclosure, for a neural video codec or a neural image codec, the entropy model is configured to support an adaptive quantization mechanism. With respect to neural codecs, one challenge is how to achieve smooth rate adaptation with a single model of a trained neural codec. In conventional (non-neural) codecs, smooth rate adaptation can be achieved by adjusting quantization parameters. However, conventional neural codecs lack such an ability and generally use a fixed quantization step ("QS"). To achieve different rates, such conventional neural codecs need to be retrained, which can increase the burden of model training and model memory. In contrast, the adaptive quantization mechanism operating with the improved entropy model described herein enables quantization at multiple granularity levels. For example, as described more fully below, the overall (aggregate) QS can be determined at three different granularities. First, a global QS value can be set by the user for a particular target rate. Then, since different channels may contain information of different importance, the global QS can be multiplied by channel-wise (or per-channel) QS values. And the product of the global QS value and the channel-wise QS values can be further multiplied by the spatial channel-wise (or per-region) QS values generated by the entropy model. Such an adaptive quantization mechanism can help the neural codec handle various types of content and achieve accurate rate adaptation at each position of the global QS. Further, the adaptive quantization mechanism can train the entropy model to learn the QS (especially the spatial channel-wise / per-region QS values), which leads to not only smooth rate adaptation in a single model for different global QS values but also an improvement in RD performance.This is because an adaptive quantization mechanism enables the entropy model to learn to allocate more bits to more important content (through QS values per spatial channel unit / region) that is essential for the reconstruction of current and subsequent video frames. This type of content-adaptive quantization mechanism allows dynamic bit allocation to enhance the final compression ratio.

[0022] III. Exemplary Computer System

[0041] FIG. 1 shows a generalized example of a suitable computer system (100) in which some of the innovations described may be implemented. Since the innovations may be implemented in a variety of general-purpose or special-purpose computer systems, computer system (100) is not intended to suggest any limitation as to the scope of use or functionality.

[0023]

[0042] Referring to FIG. 1, computer system (100) includes one or more processing devices (110, 115) and memories (120, 125). The processing devices (110, 115) execute computer-executable instructions. The processing device can be a general-purpose central processing unit (“CPU”), a processor within an application-specific integrated circuit (“ASIC”), or any other type of processor. In a multiprocessing system, multiple processing devices execute computer-executable instructions to enhance processing power. For example, FIG. 1 shows a CPU (110) and a graphics processing unit or auxiliary processing device (115). The tangible memories (120, 125) can be volatile memories (e.g., registers, caches, RAM) accessible by the processing device, non-volatile memories (e.g., ROM, EEPROM, flash memory, etc.), or some combination of the two. The memories (120, 125) store software (180) that implements one or more innovations related to a hybrid entropy model and / or a neural codec with flexible quantization in the form of computer-executable instructions suitable for execution by the processing device.

[0024]

[0043] A computer system may have additional features. For example, a computer system (100) includes storage (140), one or more input devices (150), one or more output devices (160), and one or more communication connections (170). An interconnect mechanism (not shown), such as a bus, a controller, or a network, interconnects the components of the computer system (100). Generally, operating system software (not shown) provides an operating environment for other software executed in the computer system (100) and coordinates the activities of the components of the computer system (100).

[0025]

[0044] Tangible storage (140) may be removable or non-removable and includes magnetic media such as magnetic disks, magnetic tapes or cassettes, optical media such as CD-ROMs or DVDs, or any other medium that can be used to store information and can be accessed within the computer system (100). Storage (140) stores instructions for software (180) that implements one or more innovations related to a hybrid entropy model and / or a neural codec with flexible quantization.

[0026]

[0045] Input device (150) may be a keyboard, a mouse, a pen, or a touch input device such as a trackball, a voice input device, a scanning device, or another device that provides input to the computer system (100). For video, input device (150) may be a camera, a video card, a screen capture module, a TV tuner card, or a similar device that accepts video input in analog or digital form, or a CD-ROM or CD-RW that reads video input into the computer system (100). Output device (160) may be a display, a printer, a speaker, a CD writer, or another device that provides output from the computer system (100).

[0027]

[0046] The communication connection (170) enables communication via a communication medium with another computing entity. The communication medium carries information such as computer-executable instructions, audio or video input or output, or other data within a modulated data signal. A modulated data signal is a signal that sets or changes one or more of the characteristics of the signal in such a way as to encode information in the signal. By way of example and not limitation, the communication medium may use electricity, light, RF, or other carriers.

[0028]

[0047] The innovation may be described in the broad context of computer-readable media. A computer-readable media is any available tangible media accessible within a computing environment. By way of example and not limitation, in the computer system (100), the computer-readable media includes memories (120, 125), storage (140), and combinations thereof. Thus, the computer-readable media can be, for example, volatile memory, non-volatile memory, optical media, or magnetic media. As used herein, the term computer-readable media does not include transitory signals or propagating carriers.

[0029]

[0048] The innovation may be described in the broad context of computer-executable instructions, such as computer-executable instructions included in program modules, being executed on a target physical or virtual processor in a computer system. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The functions of the program modules may be combined as desired or divided among the program modules in various embodiments. The computer-executable instructions of the program modules may be executed within a local or distributed computer system.

[0030]

[0049] The terms "system" and "device" are used interchangeably herein. Unless the context clearly indicates otherwise, neither term implies any limitation as to the type of computer system or computing device. In general, a computer system or computing device can be local or distributed and can include any combination of dedicated hardware and / or general-purpose hardware and software that implements the functions described herein.

[0031]

[0050] The disclosed methods can also be implemented using dedicated computing hardware configured to execute any of the disclosed methods. For example, the disclosed methods can be implemented by an integrated circuit (e.g., an ASIC such as an ASIC digital signal processor ("DSP"), a graphics processing unit ("GPU"), or a programmable logic device ("PLD") such as a field programmable gate array ("FPGA")) specially designed or configured to perform any of the disclosed methods.

[0032]

[0051] For purposes of presentation, the detailed description uses terms such as "select" and "determine" to describe computer operations in a computer system. These terms are high-level abstractions of operations performed by a computer and should not be confused with acts performed by a human. The actual computer operations corresponding to these terms will vary depending on the implementation.

[0033] IV. Exemplary Network Environment

[0052] Figures 2A and 2B show an exemplary network environment (201, 202) including a video encoder (220) and a video decoder (270). The encoder (220) and the decoder (270) are connected via a network (250) using an appropriate communication protocol. The network (250) can include the Internet or another computer network.

[0034]

[0053] In the network environment (201) shown in FIG. 2A, each real-time communication ("RTC") tool (210) includes both an encoder (220) and a decoder (270) for two-way communication. A given encoder (220) can output the encoded data as part of a bitstream, and the corresponding decoder (270) receives the encoded data from the encoder (220). Two-way communication can be part of a videoconference, a video phone call, or other two- or multi-party communication scenarios. The network environment (201) of FIG. 2A includes two real-time communication tools (210), but the network environment (201) can alternatively include three or more real-time communication tools (210) participating in multi-party communication.

[0035]

[0054] The real-time communication tool (210) manages the encoding by the encoder (220). FIG. 3 shows an exemplary encoder (340) that can be included in the real-time communication tool (210). The real-time communication tool (210) also manages the decoding by the decoder (270). FIG. 3 also shows an exemplary decoder (350) that can be included in the real-time communication tool (210).

[0036]

[0055] In the network environment (202) shown in FIG. 2B, the encoding tool (212) includes an encoder (220) that encodes video for distribution to a plurality of playback tools (214) including a decoder (270). Unidirectional communication may be provided for video surveillance systems, web camera surveillance systems, remote desktop conference presentations or sharing, wireless screen casting, cloud computing or gaming, or other scenarios where video is encoded and sent from one location to one or more other locations. The network environment (202) of FIG. 2B includes two playback tools (214), but the network environment (202) may include more or fewer playback tools (214). Generally, the playback tool (214) communicates with the encoding tool (212) to determine the video stream for the playback tool (214) to receive. The playback tool (214) receives the stream, buffers the received encoded data for an appropriate period, and starts decoding and playback.

[0037]

[0056] FIG. 3 shows an exemplary encoder (340) that may be included in the encoding tool (212). The encoding tool (212) may also include server-side controller logic for managing connections with one or more playback tools (214). The playback tool (214) may include client-side controller logic for managing connections with the encoding tool (212). FIG. 3 also shows an exemplary decoder (350) that may be included in the playback tool (214).

[0038] V. Exemplary Improved Neural Video Codec System

[0057] FIG. 3 shows an exemplary neural video codec system (300) in which several described embodiments may be implemented in relation. The neural video codec system (300) includes a neural video encoder (340) configured to encode a video frame into encoded data using a hybrid entropy model. The neural video encoder (340) may be an embodiment of the encoder (220) depicted in FIGS. 2A-2B. The neural video codec system (300) also includes a neural video decoder (350) configured to reconstruct a video frame from the encoded data using a hybrid entropy model. The neural video decoder (350) may be an embodiment of the decoder (270) depicted in FIGS. 2A-2B. As shown, the neural video encoder (340) may include the neural video decoder (350). In a particular example, the neural video decoder (350) can be a stand-alone system. For example, when a computer system includes only the decoder, the neural video decoder (350) is separate. An exemplary hybrid entropy model is shown in more detail in FIG. 4.

[0039]

[0058] A neural video codec system (300), or a part of a neural video codec system (300) such as a neural video encoder (340) and / or a neural video decoder (350), can be implemented as part of an operating system module, as part of an application library, as part of a stand-alone application, or using dedicated hardware. Overall, the neural video encoder (340) receives a sequence of source video frames from a video source (e.g., a camera, a tuner card, a storage medium, a screen capture module, or other digital video source) and generates encoded data as an output to an output channel (338). The encoded data output to the output channel (338) may include content encoded using one or more of the innovations described herein. When separate, the neural video decoder (350) receives the encoded data from the output channel (338) and generates a reconstructed video frame (320) as an output for an output destination (e.g., a video display device, a storage medium, etc.). The received encoded data may include content encoded using one or more of the innovations described herein. As used herein, the term "frame" generally refers to source image data, encoded image data, or reconstructed image data.

[0040]

[0059] The neural video encoder (340) receives the current video frame (302), encodes the current video frame (302) to generate encoded data, and outputs the encoded data as part of a bitstream supplied to the output channel (338). As part of the encoding, the neural video encoder (340) optionally uses one or more features of the hybrid entropy models described herein. As shown, the neural video encoder (340) also includes at least some components of the neural video decoder (350) for inverse quantization, context decoding, frame generation, buffering, temporal context mining, and motion vector decoding in the reconstruction loop. The neural video decoder (350) can receive the encoded data as part of a bitstream, decode the encoded data to reconstruct the current video frame, and output the reconstructed current video frame (320). As part of the decoding, the neural video decoder (350) optionally uses one or more features of the hybrid entropy models described herein. In FIG. 3, the current video frame (302) is represented by x t where t is the frame index, and the reconstructed current video frame (320) is

[0041]

Number

[0042]

[0060] As described herein, the neural video encoder (340) can be configured to generate temporal context parameters related to the current video frame, and as described below, perform contextual encoding and contextual decoding on the current video frame to reconstruct the current video frame.

[0043] A. Exemplary Generation of Temporal Context Parameters

[0061] Several modules including a motion estimator (326), a motion vector (“MV”) encoder (328), an MV decoder (330), a temporal context mining network (324), and a frame and feature buffer (322) are involved in the generation of the temporal context parameters.

[0044]

[0062] The current video frame x t , and the reconstructed previous video frame (retrieved from the frame and feature buffer (322))

[0045]

Number

[0046]

[0063] The generated set v of MV values t is compressed by the MV encoder (328) and then the reconstructed set of MV values

[0047]

Number

[0048]

Number

[0049]

[0064] The reconstructed set of MV values

[0050]

Number

[0051]

[0065] The temporal context mining network (324) includes one or more convolutional layers and is configured to explore or capture temporal correlations present in the video frames. An exemplary temporal context mining network (324) including an exemplary network structure is described in more detail by Sheng et al., "Temporal Context Mining for Learned Video Compression", arXiv preprint arXiv:2111.13850, 2021 (hereinafter "Sheng 2021"). Generally, the temporal context mining network (324) outputs one or more sets of temporal context parameters at different scales, for example,

[0052]

Number

[0053]

Number

[0054]

Number

[0055]

Number

[0056]

Number

[0057]

Number

[0058]

[0066] One or more generated temporal context parameter sets (e.g.,

[0059]

Number

[0060] B. Exemplary Context Encoding and Decoding

[0067] Conditioned by a multi-scale temporal context parameter set (e.g.,

[0061]

Number

[0062]

[0068] To achieve bitrate savings, the current latent SV representation y tis the quantized version sent to an arithmetic encoder ("AE" 308) that generates a bitstream containing the encoded data of the current latent SV representation before being sent

[0063]

Number

[0064]

Number

[0065]

Number

[0066]

Number

[0067]

Number

[0068]

[0069] As described herein, the AE (308) and AD (312) work in conjunction with an entropy model network (310) to provide entropy encoding and entropy decoding respectively. As further described below, the entropy model network (310) is the quantized version of the current latent SV representation

[0069]

Number

[0070]

Number

[0071]

[0070] Furthermore, the entropy model network (310) can generate QS values (represented by qs sc ) for each of a plurality of spatial regions of the current latent SV representation. The generated qs sc values can be supplied to the quantizer (306) and the inverse quantizer (314). The quantizer (306) and the inverse quantizer (314) can further receive a global QS value (represented by 336, qs global ) and QS values for each of a plurality of channels (represented by 334, qs ch ) for different channels. Based on qs global , qs ch , and qs sc , the quantizer (306) and the inverse quantizer (314) can perform quantization and inverse quantization of multiple granularities, respectively, as described in more detail below.

[0072]

[0071] Still conditioned on the temporal context parameter set (e.g.,

[0073]

Number

[0074]

Number

[0075]

Number

[0076]

[0072] In this encoding / decoding process, by the entropy model network (310)

[0077]

Number

[0078] C. Exemplary reconstruction of the current video frame

[0073] As shown in FIG. 3, the frame generator (318) uses the estimated current feature parameter set

[0079]

Number

[0080] [Number] is configured to generate (320). Generally,

[0081] [Number] represents a reconstructed version of something. For example, in FIG. 3,

[0082] [Number] represents a reconstructed set of feature parameters from the contextual decoder (316). The reconstructed video frame

[0083] [Number] (320) is stored in a frame and feature buffer (322) similar to a decoder picture buffer (“DPB”) and can be used by a motion estimator (326) to generate a set of MV values for subsequent video frames. Further, a frame generator (318) also generates the current set of feature parameters F t The current set of feature parameters F t is also stored in the frame and feature buffer (322) and can be used by a temporal context mining network (324) to generate a set of temporal context parameters for subsequent video frames. In certain cases, the frame generator (318) includes one or more convolutional layers. An exemplary network structure of the frame generator (318) is further described below with reference to FIG. 10A. Alternatively, the frame generator (318) can be implemented using a different network structure.

[0084] D. Exemplary Variations of the Neural Video Codec System

[0074] Depending on the implementation and the type of compression / decompression desired, the modules of the neural video codec system (300) can be added, omitted, split into multiple modules, combined with other modules, and / or replaced with similar modules. The relationships shown between the modules within the neural video encoder (340) and the neural video decoder (350) each show the overall flow of information in the neural video encoder (340) and the neural video decoder (350), and other relationships are not shown for simplicity. Generally, a given module of the neural video codec system (300) can be implemented by software executable on a CPU, by software controlling dedicated hardware (e.g., graphics hardware for video acceleration), or by dedicated hardware (e.g., in an ASIC).

[0085] VI. Exemplary Hybrid Entropy Model

[0075] As described above, both the AE (308) and the AD (312) work in conjunction with the entropy model network (310) to respectively provide entropy encoding and entropy decoding. For that purpose, both entropy encoding and entropy decoding use the probability mass function ("PMF") of the quantized version

[0086]

Number

[0087]

Number

[0088]

Number

[0089]

Number

[0090]

Number

[0091]

Number

[0092]

Number

[0093]

[0076] Figure 4 is for reducing cross-entropy

[0094]

Number

[0095]

[0077] In the depicted example, the Hybrid Entropy Model Network (440) has an input unit (406), a first fusion unit (408), a first statistics / parameter estimator (410), a second fusion unit (412), and a second statistics estimator (414). Each of the first statistics / parameter estimator (410) and the second statistics estimator (414) may include one or more convolutional layers. An exemplary network structure of the Hybrid Entropy Model Network (440) is shown in FIG. 11 and further described below. Alternatively, the Hybrid Entropy Model Network (440) can be implemented using a different network structure.

[0096] A. Latent Prior Distribution

[0078] To improve the estimation of statistics by the entropy model network, the temporal correlation of latent representations between video frames can be utilized. For MV encoding / decoding, the latent prior distribution can be the previous latent MV representation. For SV encoding / decoding, the latent prior distribution can be the previous latent SV representation. Referring to FIG. 4, as described above,

[0097] [Mathematics] The elements of are logically organized in three dimensions, including one channel dimension and two spatial dimensions.

[0098] [Mathematics] (where i, j, and k are the height, width, and channel index, respectively)

[0099] [Mathematics] are represented as elements of. Theoretically, each

[0100] [Mathematics] is the decoded latent representation of the previous video frame

[0101] [Mathematics] with the corresponding elements of

[0102] [Mathematics] may be correlated. Also, each

[0103] [Mathematics] is may be correlated with the elements of the corresponding spatial location of any of the temporal context parameter sets (e.g.,

[0104] [Mathematics] ), and the temporal context parameter set is the previous feature parameter set F related to the previous video frame t-1is determined based in part on and a set of reconstructed MV values of the current video frame

[0105]

Number

[0106]

[0079] FIG. 5A shows

[0107]

Number

[0108]

Number

[0109]

Number

[0110] [Number] (i.e., the prior distribution) can be received. Before being fused by the fusion unit (530), the hyper-prior distribution information

[0111] [Number] is first decoded by the hyper-prior distribution decoder (510) to generate the decoded hyper-prior distribution parameters. An exemplary network structure of the hyper-prior distribution decoder for the latent SV representation is described below with reference to FIG. 13B. Alternatively, the hyper-prior distribution decoder can be implemented using a different network structure. Further, before being fused by the fusion unit (530), the temporal context parameter set

[0112] [Number] is first encoded by the temporal context encoder (520) to generate the temporal context prior. The temporal context encoder (520) may include one or more convolutional layers. An exemplary network structure of the temporal context encoder is described below with reference to FIG. 12. Alternatively, the temporal context encoder can be implemented using a different network structure. FIG. 11 shows an exemplary network structure of the fusion unit (530) (the "Prior Fusion" layer on the left in FIG. 11). Alternatively, the fusion unit (530) can be implemented using a different network structure.

[0113]

[0080] Generally, during encoding, the hyper-prior distribution parameter z t is the current latent SV representation y tis derived from, quantized, entropy-encoded, and output as part of the encoded data. During decoding, or as part of reconstruction during encoding, the hyper-prior distribution parameters

[0114]

Number

[0115]

Number

[0116]

Number

[0117]

Number

[0118]

[0081] In FIG. 5A, only one set of temporal context parameters

[0119]

Number

[0120]

Number

[0121]

Number

[0122]

[0082] As depicted in FIG. 3, the latent prior distribution

[0123]

Number

[0124]

Number

[0125]

Number

[0126]

Number

[0127]

Number

[0128]

Number

[0129]

Number

[0130]

[0083] Furthermore, in some implementations, a cascaded training strategy can be adopted so that the gradient can be backpropagated to multiple video frames. Further details regarding such a cascaded training strategy in different contexts are described in Chan et al., "BasicVSR: The search for essential components in video super-resolution and beyond" in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 4947 - 4956, 2021 and Sheng 2021. Under such a training strategy, a propagation chain of latent representations can be formed. As a result, a connection between the latent representation of the current video frame and the latent representation of the long-term reference video frame is also established. Such a connection can be very helpful in extracting the correlation between the latent representations of multiple video frames, and thus,

[0131]

Number

[0132]

[0084] During decoding, based on the same input as during encoding, the hybrid entropy model network (440) executes operations in the same order to determine the statistical characteristics of each set of elements used in entropy decoding and the QS values for each spatial region (and each channel) used in inverse quantization.

[0133] Figures 4, 5A, and 5B show a hybrid entropy model network for encoding / decoding of latent SV representations. The hybrid entropy model network can also be used for encoding / decoding of latent MV representations. For encoding / decoding of MV information, the input to the hybrid entropy model network is the hyper-prior distribution parameter

[0134] [Number] relating to the current latent MV representation of the current video frame, and the decoded previous latent MV representation

[0135] [Number] (i.e., the latent prior distribution) of the previous video frame. The MV values

[0136] [Number] of the current video frame, from which the above-described set of temporal context parameters is derived, are not input to the hybrid entropy model network for MV encoding / decoding. Before being fused by the fusion unit, the hyper-prior distribution information

[0137] [Number] relating to the latent MV representation is first decoded by a hyper-prior decoder to generate the decoded hyper-prior distribution parameter relating to the MV information. The latent buffer can store the decoded previous latent MV representation

[0138] [Number] (i.e., the latent prior distribution) of the previous video frame.

[0139] B. Double-Space Prior Distribution

[0086] Potential prior distribution to enrich the input

[0140]

Number

[0141]

[0087] For example, to improve the operation efficiency, the features of the double - space prior distribution can be implemented in a two - stage estimation process based on the divided checkerboard context model, as shown in FIG. 4. For simplicity, the arithmetic encoder and arithmetic decoder (e.g., 308 and 312 in FIG. 3) are omitted from the two - stage estimation in FIG. 4. More generally, the features of the double - space prior distribution can be implemented in a multi - stage estimation process with three or more stages. In the corresponding decoding, the features of the double - space prior distribution are implemented in a two - stage estimation process (more generally, a multi - stage estimation process).

[0142]

[0088] FIG. 4 shows the encoding / decoding of the current potential SV representation using the features of the double - space prior distribution. As described above, the elements of the current potential SV representation y t are logically organized in three dimensions including two spatial dimensions and one channel dimension. As shown in FIG. 4, the current potential SV representation y t (402) can be split by a splitter (404) into two blocks (422, 424) of elements along the channel dimension. Each of the two blocks (422, 424) has the same spatial dimensions as y t but only has a part of the channels. For example, assuming that the total number of channels of y t is C, the first block (422) has the lower half of the channels (e.g., y t,k<C / 2) can include elements of, and the second block (424) can include elements of the upper half channel (e.g., y t,k≧C / 2 ) can include elements of.

[0143]

[0089] As shown in FIG. 4, the two blocks (422, 424) of elements can be quantized by a quantizer (416) to generate quantized versions of the two blocks (e.g.,

[0144]

Number

[0145]

[0090] Each quantized block (e.g.,

[0146]

Number

[0147]

Number

[0148]

Number

[0149]

[0091] During the first stage of the dual-space prior distribution estimation process, the first statistical / parameter estimator (410) can be configured to encode the elements of the first even set (426) while setting the elements of the first odd set (430) to zero. Further, the first statistical / parameter estimator (410) can be configured to encode the elements of the second odd set (428) while setting the elements of the second even set (432) to zero. Encoding the elements of the first even set (426) and encoding the elements of the second odd set (428) can be performed simultaneously or substantially simultaneously (e.g., by parallel computing).

[0150]

[0092] During the estimation process of the first stage, the first statistical / parameter estimator (410) can estimate the statistical characteristics (e.g., μ idx and σ idx , where idx = {t, (i + j) % 2 == 0, k < C / 2}) and the statistical characteristics (e.g., μ idx and σ idx , where idx = {t, (i + j) % 2 == 1, k ≥ C / 2}) of the elements of the second odd set (428). During the estimation process of the first stage, the first statistical / parameter estimator (410) can also determine the QS parameter

[0151]

Number

[0152]

[0093] FIG. 11 shows an exemplary network structure of the first statistical / parameter estimator (410) (the “parameter estimation” structure on the left in FIG. 11). Alternatively, the first statistical / parameter estimator (410) may be implemented using a different network structure.

[0153]

[0094] After the first step of encoding, the quantized first even set (426) and the second odd set (428) can be fused together by the second fusion unit (412), and the second fusion unit (412) further receives at least a part of the output channels from the first statistical / parameter estimator (410) as inputs, and further generates a context for the second stage of estimation. FIG. 11 shows an exemplary network structure of the second fusion unit (412) (the “prior distribution fusion” layer on the right in FIG. 11). Alternatively, the second fusion unit (412) may be implemented using a different network structure.

[0154]

[0095] During the second stage of the double-space prior distribution estimation process, the second statistical estimator (414) may be configured to encode the elements of the first odd set (430) and the elements of the second even set (432). Similarly, encoding the elements of the first odd set (430) and encoding the elements of the second even set (432) may be performed simultaneously or substantially simultaneously (e.g., by parallel computing).

[0155]

[0096] FIG. 11 shows an exemplary network structure of the second statistical estimator (414) (the “parameter estimation” structure on the right in FIG. 11). Alternatively, the second statistical estimator (414) may be implemented using a different network structure.

[0156]

[0097] During the second stage of the estimation process, the second statistical estimator (414) determines the statistical characteristics (e.g., μ idx and σ idx, where idx = {t, (i + j) % 2 == 1, k < C / 2}, and the statistical characteristics (e.g., μ idx and σ idx , where idx = {t, (i + j) % 2 == 0, k ≧ C / 2}) can be estimated. Since the second fusion unit (412) fuses at least a part of the results of the first statistical estimator (410), the estimation process of the second stage can benefit from the context from all spatial positions. As a result, the estimation of the statistical characteristics by the second statistical estimator (414) can be more accurate by leveraging the estimation results obtained by the first statistical estimator (410).

[0157]

[0098] Entropy coding for the first even set (426), the second odd set (428), the first odd set (430), and the second even set (432) can occur simultaneously using the respective statistical characteristics of the different sets of elements.

[0158]

[0099] During decoding, based on the same input, the entropy model network (440) executes operations in the same order to determine the statistical characteristics of each set of elements and the QS values for each spatial region (and for each channel). When reconstructing the current latent SV representation during decoding, the statistical characteristics estimated by the first statistical / parameter estimator (410) can be used for entropy decoding of the elements of the first even set (426) and the second odd set (428), and the statistical characteristics estimated by the second statistical estimator (414) can be used for entropy decoding of the elements of the first odd set (430) and the second even set (432). During decoding, and as part of the reconstruction loop during encoding, the decoded elements of the first even set (426), the second odd set (428), the first odd set (430), and the second even set (432) are the reconstructed first block (434, e.g.,

[0159]

Number

[0160]

Number

[0161]

Number

[0162] Compared with the conventional checkerboard context model described in He et al., "Checkerboard context model for efficient learned image compression" in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 14771-14780, 2021, the segmented checkerboard context model described herein expands the scope of spatial context by channel splitting. In other words, the segmentation of the elements of the quantized version of the latent representation is not only in two spatial dimensions (e.g., based on the odd or even positions of the elements), but also along the channel dimension. As a result, a more accurate estimation of statistical characteristics can be achieved. Further, for both the first and second estimation stages, the first quantized block and the second quantized block can be added and sent to an arithmetic encoder (e.g., 308 of FIG. 3). Thus, the characteristics of the double spatial prior distribution described herein do not introduce any additional coding delay when compared to the conventional checkerboard model described by He et al.

[0163]

[0101] Importantly, the characteristics of the double spatial prior distribution described herein can mine the correlation between channels. For example, the first even set (426) quantized for the first stage of the estimation process can also be used as a condition for encoding the second even set (432) during the second stage of the estimation process. Similarly, the second odd set (428) quantized for the first stage of the estimation process can also be used as a condition for encoding the first odd set (430) during the second stage of the estimation process. As a result, the double spatial prior distribution, by more efficiently utilizing the correlation between spatial positions and channel dimensions,

[0164]

Number

[0165]

[0102] As depicted in Figure 4,

[0166]

Number

[0167]

Number

[0168]

[0103] Depending on the situation,

[0169]

Number

[0170]

Number

[0171]

[0104] More generally, the current potential SV representation y t (and the resulting quantized version

[0172]

Number

[0173]

[0105] Based at least in part on the quantized version of the first set of elements among the plurality of sets of elements, the statistical characteristics of the quantized version of the second set of elements among the plurality of sets of elements can be estimated. For example, in FIG. 4, the statistical characteristics of the first odd set (430) and / or the second even set (432) can be estimated based at least in part on the first even set (426) and / or the second odd set (428). When the estimation process is divided into multiple stages, as described above, the (quantized) results of the encoding from the previous stage can be fused to provide the input for the subsequent estimation stage.

[0174]

[0106] FIGS. 4 and 11 show an entropy model network for encoding / decoding the current latent SV representation using the characteristics of the double-space prior distribution. The entropy model network using the characteristics of the double-space prior distribution can also be used for encoding / decoding the current latent MV representation. For encoding / decoding MV information, the input to the entropy model network includes hyper-prior distribution parameters

[0175]

Number

[0176]

Number

[0177]

[0107] During decoding, based on the same input, the entropy model network performs operations in the same order to determine the statistical characteristics of each set of elements and the QS values for each region (and for each channel) of the latent MV representation space. When reconstructing the current latent MV representation during decoding, the statistical characteristics estimated by the first statistical / parameter estimator can be used for entropy decoding of the elements of the first stage set, and the statistical characteristics estimated by the second statistical estimator can be used for entropy decoding of the elements of the second stage set. During decoding and as part of the reconstruction loop during encoding, the decoded elements of each set of elements can be inverse quantized by an inverse quantizer using the QS values for each spatial region (and for each channel) from the entropy model network to generate a reconstructed set of MV values. As further described below, the inverse quantizer can be configured to enable inverse quantization at multiple granularity levels. Finally, the reconstructed set of MV values is the reconstructed current latent MV representation

[0178]

Number

[0179]

[0108] As described above, an entropy model network implementing a multi-stage estimation process can be a hybrid entropy model network that accepts a latent prior distribution as an input. Alternatively, an entropy model network implementing a multi-stage estimation process can operate without accepting a latent prior distribution as an input.

[0180] C. Quantization of Multiple Granularities

[0109] Conventional neural codecs cannot handle rate adaptation with a single model of the trained neural codec. To achieve different rates, the entropy models of conventional neural codecs need to be retrained by adjusting their weights according to the RD criterion for different rates. Such an approach can significantly increase the training cost of the neural codec and the memory burden of the model. To achieve a wide range of rates in a single model of the trained neural codec, an adaptive quantization mechanism can be integrated with an entropy model network (e.g., a hybrid entropy model network (440) that accepts a latent prior distribution as an input and implements the characteristics of a double-space prior distribution). As described below, the adaptive quantization mechanism can be used when encoding / decoding the current latent SV representation or the current latent MV representation. Similarly, the adaptive quantization mechanism can be used with a hybrid entropy model network that accepts a latent prior distribution as an input or an entropy model network that does not accept a latent prior distribution as an input.

[0181]

[0110] As shown in FIG. 4, the current latent SV representation y t (402) includes three sets of quantization step (「QS」) values, i.e., a global QS value (qs for adjusting the bit rate and the overall qualityglobal The quantization step values per channel (qs ch represented by, which can also be referred to as the "quantization step value per channel"), and the quantization step values per region (qs sc represented by, which can also be referred to as the "quantization step value per spatial channel") for different spatial regions of the current latent SC representation can be used by the quantizer (416) for quantization. As described herein, different spatial regions are associated with different positions or areas / blocks (determined, for example, by height, width, and channel indices i, j, and k) of the current latent SV representation, and the QS values per region are channel specific (however, alternatively, the QS values per region can be channel independent). For example, two different spatial regions of the same channel (having, for example, different i or j indices) can have two different QS values per region, and two same spatial regions of two different channels (having, for example, the same i and j indices) can also have different QS values per region.

[0182]

[0111] Global QS value qs global can be a fixed value predefined by the user as part of the overall setting of quality and bitrate. The global QS value qs global can be the same parameter for the current latent SV representation and the current latent MV representation, or the current latent SV representation and the current latent MV representation can have different global QS values.

[0183]

[0112] Channel-specific QS value qs for different channels chcan be configured as part of a neural video codec system (e.g., 300) based on the importance of each channel and can be learned during the model training process. The QS value for each channel can be applied to a specific channel. The per-channel QS values for different channels can be the same or different. For example, any of the per-channel QS values for the lower half channels (

[0184]

Number

[0185]

Number

[0186]

[0113] The per-region QS values qs sc for different spatial regions can be generated by the above-described hybrid entropy model network (440), or other entropy model networks. The current latent SV representation and the current latent MV representation can have the per-region QS values qs sc . In some exemplary implementations, the per-region QS values qs sc are generated by the entropy model network from the same input during encoding and decoding, so the per-region QS values qs sc are not encoded and are not sent within the encoded data. For example, in one particular example, the per-region QS values (qs sc ) for different spatial regions can be generated by the first statistical estimator (410).

[0187]

[0114] In FIG. 4, the hybrid entropy model network (440) generates the per-region QS values qs sc for all sets of elements of the current latent representation. In some other implementations, it is possible for the per-region QS values (qs sc ) for some spatial regions to be generated by the first statistical / parameter estimator (410), and the per-region QS values (qs sc) can be generated by a second statistical estimator (414), for example, based on the encoded elements of a first odd set (430) and a second even set (432). In this case, the second statistical estimator (414) outputs the QS value (qs sc ) for each region regarding the first odd set (430) and the second even set (432) in addition to the statistical characteristics of the first odd set (430) and the second even set (432). Alternatively, the QS value qs sc for each region regarding all sets of elements of the current latent representation can be generated by the second statistical estimator (414).

[0188]

[0115] During decoding and as part of the reconstruction loop during encoding, the corresponding inverse quantization is applied. Generally, the inverse quantizer (418) can use the same global QS value qs global , the QS value qs for each of a plurality of channels ch , and the QS value qs for each of a plurality of regions sc . The global QS value qs global can be sent to the decoder together with other encoded data in the bitstream. Since only a single number is sent for each frame or video (or two values are sent if different values are used for MV information and SV information), the overhead of sending qs global can be ignored.

[0189]

[0116] FIG. 6 shows an exemplary implementation of quantization and inverse quantization with multiple granularities. As shown, quantization can be performed in three consecutive stages. In the first quantization stage (610), the current latent SV representation y t is first quantized using the global QS value qs global . In the second quantization stage (620), the output of the first quantization stage (610) is quantized using the QS value qs for each channel chis further quantized using (e.g., each channel is quantized by a QS value specific to that channel). In a third quantization stage (630), the output of the second quantization stage (620) is the region-wise QS value qs sc is further quantized using (e.g., each element is quantized by a QS value specific to the spatial location and channel of that element). The output of the third quantization stage (630) is the final quantized version of the current latent SV representation

[0190]

Number

[0191]

[0117] Similarly, inverse quantization can be performed in three consecutive stages, but in the reverse order. As shown, the quantized version of the current latent SV representation

[0192]

Number

[0193]

Number

[0194]

Number

[0195]

Number

[0196]

Number

[0197]

[0118] FIG. 6 shows a particular order for applying a global QS value qs global , per-channel QS values qs ch , and per-region QS values qs sc during quantization and inverse quantization. Alternatively, the global QS value qs global , per-channel QS values qs ch , and per-region QS values qs sc can be applied in a different order during quantization and inverse quantization.

[0198]

[0119] Since the global QS value is a single value applied to all spatial positions of a given latent representation and the elements of all channels, qs global can bring about a coarse quantization effect for controlling the target rate. Since different channels carry information with different importance, the QS value for each channel can scale or adjust the quantization steps in different channels. Furthermore, different spatial positions for each channel can also have different characteristics due to various image or video contents. Therefore, the QS value for each region can be used for a more precise adjustment of the quantization step size for each position of each channel.

[0199]

[0120] As described above, the QS value qs sc for each region is generated by an entropy model network (e.g., a hybrid entropy model network (440)). Therefore, for each image or video frame, qs sc is dynamically changed to adapt to the content of the image or video. Such content adaptation not only helps to achieve smooth bitrate adjustment, but also can improve the final rate-distortion performance through content-adaptive bit allocation. In particular, more important information that is essential for reconstruction and / or is referred to by the encoding of subsequent video frames is assigned a smaller quantization value, and vice versa.

[0200]

[0121] An example of visualization is shown in FIG. 7. In this example, the upper left panel shows the input video frame. The QS value qs sc for each region generated by the hybrid entropy model is shown in the lower left panel. The upper right panel shows the quantized latent representation of the input video frame without using the QS value qs sc for each region. The lower right panel shows the QS value qs scShows the quantized latent representation of the input video frame for use. In this example, the hybrid entropy model learns that moving players are more important and generates per-region QS values for these regions. In contrast, background regions (e.g., indicated by horizontal and vertical lines in the lower right corner of the panel) are associated with per-region QS values for larger regions, thus resulting in a significant bitrate savings. For example, the bitrate is 0.065 bits per pixel (BPP) without using per-region QS values, but is reduced to 0.056 BPP when using per-region QS values, thus resulting in a reduction of approximately 13.8% in BPP with similar image quality.

[0201]

[0122] In many of the foregoing examples, each spatial location (or each combination or spatial location and channel) has its own per-region QS value. Alternatively, the per-region QS value can be shared among multiple spatially adjacent locations, e.g., with respect to a block or window.

[0202]

[0123] In many of the foregoing examples, there are three stages of quantization and three corresponding stages of inverse quantization. Alternatively, it is possible to have fewer or more stages. For example, there are two stages of quantization and two corresponding stages of inverse quantization, with a global QS value applied in one stage and per-region, per-channel QS values applied in another stage. Or, as another example, there are four stages of quantization and four corresponding stages of inverse quantization, with a global QS value applied in one stage, per-channel QS values applied in another stage, and hierarchical QS values applied in the remaining stages for different levels of spatial granularity.

[0203] VII. Exemplary Neural Network Structure

[0124] A convolutional neural network ("CNN") is used in some components of the neural video codec system (300) and neural image codec system described herein. Generally, a CNN includes one or more convolutional layers. A convolutional layer includes a set of filters (also called kernels), and the parameters of those filters can be learned through a training process. A convolutional layer uses kernels to compute a convolution operation on the input values (e.g., sample values, MV values for the first layer, or outputs from previous layers for subsequent layers) of an input image or video frame in order to extract basic features embedded in the image or video frame. The size of the kernel is usually smaller than the input image or video frame. Each kernel is convolved with the image or video frame to create an activation map (also called a "feature map") made up of neurons. The output volume of a convolutional layer is obtained by stacking the activation maps of all the kernels along the depth dimension (an example of the channel dimension). In addition to convolutional layers, some CNNs may also include one or more sub-pixel convolutional layers, one or more pooling layers, and / or one or more rectified linear unit ("ReLU") correction layers. A sub-pixel convolutional layer performs a standard convolution operation followed by a pixel-shuffling operation. Placed between two convolutional layers, a pooling layer receives multiple activation maps and applies a pooling operation to each of the activation maps to reduce the spatial dimension while preserving important characteristics of the activation maps. An ReLU correction layer acts as an activation function by replacing all negative values received as input with zero.

[0204]

[0125] This section describes an exemplary network structure of selected components of the neural video codec system (300) and neural image codec system. For the convolutional layers and sub-pixel convolutional layers depicted in the example, the notation (K,C in ,C out, (S) indicates the kernel size, the number of input channels, the number of output channels, and the stride respectively. Generally, the stride is a kernel parameter that modifies the amount of movement on an image or video frame.

[0205]

[0126] Current video frame x t And an exemplary network structure of a contextual encoder (e.g., 304) and a contextual decoder (e.g., 316) for the potential SV representation are shown in FIGS. 9A and 9B respectively. In the depicted example, the input to the contextual encoder is the current video frame x (having three input channels corresponding to three color components for each spatial location of the frame) t And a multi-scale temporal context parameter set having a total of 64 channels for each of the multi-scale temporal context parameter sets of different spatial resolutions (original, 2x downsampled, 4x downsampled)

[0206]

Number

[0207]

Number

[0208]

Number

[0209]

Number

[0210]

[0127] An exemplary network structure of the hybrid entropy model network (e.g., 440) is shown in FIG. 11. The inputs to the hybrid entropy model network include the decoded hyper-prior distribution parameters with 192 channels, the temporal context prior distribution with 192 channels, and the latent prior distribution with 96 channels (i.e., the decoded previous latent SV representation of the previous video frame

[0211]

Number

[0212]

[0128] The decoded hyper-prior distribution parameters can be generated by the hyper-prior distribution decoder (510) of FIG. 5A by decoding the hyper-prior distribution generated previously from the current latent representation y t as described above with reference to FIG. 5A. The temporal context prior distribution can be generated by the temporal context encoder (520). In particular, the input to the temporal context encoder is the temporal context parameter set

[0213]

Number

[0214]

Number

[0215]

[0129] Referring back to FIG. 11, the hybrid entropy model network is configured to perform the estimation process in two stages. The first stage has one convolutional layer acting as a first fusion unit (e.g., 408) for fusing prior distributions, followed by two convolutional layers and two leaky ReLU layers working together as a first statistical / parameter estimator (e.g., 410). In the first stage of estimation, the hybrid entropy model network not only estimates the means and scale parameters of the probability distributions of two parts of the quantized latent current latent SV representation, i.e.,

[0216]

Number

[0217]

Number

[0218]

[0130] Exemplary network structures of the MV context encoder (e.g., of the MV encoder 328) and the MV context decoder (e.g., of the MV decoder 330) are shown in FIGS. 8A and 8B, respectively. The input to the MV context encoder is the current set v of MV values with two channels for the horizontal MV component and the vertical MV component, respectively t and the output of the MV context encoder is the current latent MV representation of the current MV values, represented by mv_y t with a 16x downsampled spatial resolution and 64 channels. The MV context decoder generally follows the reverse structure. As shown, the MV context encoder has one convolutional layer and the MV context decoder has one sub-pixel convolutional layer. Both the MV context encoder and the MV context decoder include a plurality of residual blocks including the downsample residual block of the MV context encoder and the upsample residual block of the MV context decoder. Further details regarding these residual blocks can be found in Sheng 2021.

[0219]

[0131] In some exemplary implementations, the MV encoder (328) and the MV decoder (330) include an entropy model network similar to one of the entropy model networks described above with respect to the SV information. For example, the MV encoder (328) and the MV decoder (330) may have a network structure similar to the network structure shown in FIG. 11, but include a hybrid entropy model network that does not have a temporal context prior distribution (based on the temporal context parameter set) as an input. Instead, the input includes a hyper-prior distribution of the current latent MV representation and a prior distribution (previous latent MV representation). The number of input channels is different for the first prior distribution fusion layer (e.g., 192 channels less), and accordingly, the number of output channels and input channels of each layer of the network structure can be reduced, but the overall composition can be the same.

[0220]

[0132] In some exemplary implementations, similar to what was described above with respect to the SV information, the MV encoder (328) includes a hyper-prior encoder and a hyper-prior decoder, and the MV decoder (330) includes a hyper-prior decoder. For example, the hyper-prior encoder and hyper-prior decoder for the current latent MV representation can have a network structure similar to the hyper-prior encoder and hyper-prior decoder shown in FIGS. 13A and 13B.

[0221]

[0133] An exemplary network structure of a frame generator (e.g., 318) is shown in FIG. 10A. In the depicted example, the frame generator has a W-Net-based structure. Further details regarding W-Net-based structures in different contexts are described in Xia and Kulis, "W-net: A deep model for fully unsupervised image segmentation", arXiv preprint arXiv:1711.08506 (2017). Such a network design can effectively expand the receptive field of the model with an acceptable complexity and improve the model's generation ability. As shown, the input to the frame generator is a high-resolution current feature parameter set with 32 channels (output from the contextual decoder)

[0222]

Number

[0223]

Number

[0224]

Number

[0225]

[0134] As described above, the hybrid entropy model network described herein supports content-adaptive quantization that enables handling multiple rates with a single model. Certain aspects of the entropy model network can be used for image encoding / decoding and video encoding / decoding. Thus, a neural image codec system that supports such capabilities can be implemented for intra-frame encoding / decoding. Exemplary network structures of a neural image context encoder and a neural image context decoder are shown in FIGS. 14A-14B, respectively. As shown, the neural image context encoder includes a convolutional layer and a plurality of residual blocks including downsample residual blocks. The neural image context decoder includes one convolutional layer, one sub-pixel convolutional layer, one U-Net, and a plurality of residual blocks including upsample residual blocks. (It is possible to have a network structure similar to that depicted in FIG. 10A) The U-Net is incorporated into the neural image context decoder to improve the generation ability of the neural image codec system.

[0226]

[0135] In some examples, the same quantization / inverse quantization of multiple granularities described above is, for example, the intra-latent SV representation intra_y of the current video frame x t ​t to generate a quantized version of, and a version of the intra-SV representation

[0227]

Number

[0228]

[0136] In some examples, a similar entropy model network can be used to determine the statistical characteristics of a quantized version of the intra-latent SV representation intra_y t and generate a QS value. For example, the only difference can be the input to the entropy model. For a neural image codec, the input to the entropy model network can include only the corresponding hyper-prior distribution of the intra-latent SV representation, without a latent prior distribution and a temporal context prior distribution (since there is no frame before encoding / decoding).

[0229] VIII. Exemplary Methods of Neural Encoding and Neural Decoding

[0137] This section describes exemplary methods of neural encoding and neural decoding. The methods described herein can be executed by computer-executable instructions stored on one or more computer-readable media (e.g., storage or other tangible media) or stored on one or more computer-readable storage devices (e.g., causing a computing system to execute the method). Such methods can be executed in software, firmware, hardware, or combinations thereof. Such methods can be executed, at least in part, by a computing system (e.g., one or more computing devices). The actions shown can also be described from another perspective while still practicing the technology. For example, "receive" can also be described as "send" from a different perspective.

[0230]

[0138] Figures 15A and 15B are flowcharts showing overall methods 1500 and 1550 for neural encoding and neural decoding, respectively, and can be implemented, for example, by the video codec system (300) of FIG. 3, or another neural video codec system or neural image codec system.

[0231]

[0139] As shown in FIG. 15A, the neural encoding method 1500 begins at 1510 where the current frame is received by the neural encoder. In a particular example, the neural encoder can be a neural video encoder (e.g., 340) configured to encode a video frame (e.g., 302) as depicted in FIG. 3. In a particular example, the neural encoder can be a neural image encoder configured to encode a single image. For example, the neural encoder can be a motion vector encoder (328) configured to encode a set of MV values v t Or, as another example, the neural encoder can be configured to encode sample values. At 1520, the neural encoder can encode the current frame to generate encoded data. And at 1530, the neural encoder can output the encoded data as part of a bitstream.

[0232]

[0140] As shown in FIG. 15B, the neural decoding method 1550 begins at 1560 where encoded data as part of a bitstream can be received by the neural decoder. In a particular example, the neural decoder can be a neural video decoder (e.g., 350) configured to generate a decoded video frame (e.g., 320) as depicted in FIG. 3. In a particular example, the neural decoder can be a neural image decoder configured to decode a single image. For example, the neural decoder can be a decoded set of MV values

[0233] [Numerical] It can be a motion vector decoder (330) configured to generate. Alternatively, as another example, the neural decoder can be configured to decode sample values. At 1570, the neural decoder can decode the encoded data to reconstruct the current frame. And at 1580, the neural decoder can output the reconstructed current frame.

[0234]

[0141] FIG. 16A shows an exemplary method 1600 for encoding a current frame that can be used in combination with the techniques shown in FIGS. 17A, 18A, and / or 19A as described below. FIG. 16B shows an exemplary method 1650 for decoding encoded data that can be used in combination with the techniques shown in FIGS. 17B, 18B, and / or 19B as described below.

[0235]

[0142] As shown in FIG. 16A, at 1610, the neural encoder can determine the current latent representation of the current frame (e.g., y t , mv_y t ). And at 1620, the neural encoder can encode the current latent representation using an entropy model network (e.g., 310) including one or more convolutional layers.

[0236]

[0143] As shown in FIG. 16B, at 1660, the neural decoder can reconstruct the current latent representation of the current frame using an entropy model network (e.g., 310) including one or more convolutional layers. At 1670, the neural decoder uses a contextual decoder (e.g., 316) including one or more convolutional layers to obtain the current set of feature parameters of the current frame from the current latent representation (e.g.,

[0237] [Number] ) can be estimated. And at 1680, the neural decoder can reconstruct the current frame from the estimated current feature parameter set.

[0238]

[0144] FIG. 17A shows an exemplary method 1700 for encoding a current latent representation by using a latent prior distribution as an input to an entropy model network. Correspondingly, FIG. 17B shows an exemplary method 1750 for reconstructing the current latent representation of the current frame by using a latent prior distribution as an input to an entropy model network. The neural encoder implementing method 1700 is a neural video encoder, and the neural decoder implementing method 1750 is a neural video decoder.

[0239]

[0145] As shown in FIG. 17A, at 1710, the neural encoder (as part of the entropy model network) estimates the statistical characteristics (e.g., mean value, scale value of the probability distribution function) of the quantized version of the current latent representation (e.g.,

[0240] [Number] ) based at least in part on the previous latent representation of the previous video frame (e.g.,

[0241] [Number] ). And at 1720, the neural encoder can entropy-encode the quantized version of the current latent representation based at least in part on the estimated statistical characteristics.

[0242]

[0146] For example, the current latent representation is the current latent SV representation of the current video frame, and the previous latent representation is the previous latent SV representation of the previous video frame. In this case, when the neural encoder determines the current latent representation, the neural encoder uses a contextual encoder including one or more convolutional layers to determine the current latent SV representation.

[0243]

[0147] Alternatively, as another example, the current latent representation is the current MV representation of the current video frame, and the previous latent representation is the previous latent MV representation of the previous video frame. In this case, when the neural encoder determines the current latent representation, the neural encoder uses motion estimation to determine the MV value of the current video frame with respect to the previous video frame, and uses an MV contextual encoder to determine the current latent MV representation from the MV value.

[0244]

[0148] The neural encoder can quantize the current latent representation, thereby generating a quantized version of the current latent representation. In doing so, the neural encoder can apply at least some QS values (such as QS values for each region) determined using an entropy model network based at least in part on the previous latent representation.

[0245]

[0149] As shown in FIG. 17B, at 1760, the neural decoder (using an entropy model network) determines the statistical characteristics (such as the mean value, scale value of the probability distribution function) of the quantized version of the current latent representation (for example,

[0246]

Number

[0247]

Number

[0248]

[0150] For example, the current latent representation is the current latent SV representation of the current video frame, and the previous latent representation is the previous latent SV representation of the previous video frame. In this case, the neural decoder can use a context decoder to estimate the current set of feature parameters of the current video frame from the current latent SV representation, reconstruct the current video frame from the estimated current set of feature parameters, and output the reconstructed current video frame.

[0249]

[0151] Alternatively, as another example, the current latent representation is the current latent MV representation of the current video frame, and the previous latent representation is the previous latent MV representation of the previous video frame. In this case, the neural decoder can use an MV context decoder to determine the MV values of the current video frame from the current latent MV representation.

[0250]

[0152] The neural decoder can inverse-quantize the quantized version of the current latent representation. In doing so, the neural decoder can apply at least some QS values (such as QS values for each region) determined using an entropy model network based at least in part on the previous latent representation.

[0251]

[0153] During encoding or decoding, the estimation of statistical characteristics using an entropy model network can also be at least partially based on other inputs such as the hyper-prior distribution parameters of the current video frame (generated from the current latent representation using a hyper-prior distribution encoder) and / or the set of temporal context parameters of the current video frame (generated from the previous feature parameter set of the previous video frame and the MV value of the current video frame using a temporal context mining network).

[0252]

[0154] FIG. 18A shows an exemplary method 1800 for encoding a current latent representation using a double-space prior distribution in an entropy model network. Correspondingly, FIG. 18B shows an exemplary method 1850 for reconstructing a current latent representation using a double-space prior distribution in an entropy model network. The neural encoder implementing method 1800 can be either a neural video encoder or a neural image encoder. The neural decoder implementing method 1850 can be either a neural video decoder or a neural image decoder.

[0253]

[0155] As shown in FIG. 18A, at 1810, the neural encoder can (prior to the estimation by the entropy model network) divide the elements of the current latent representation (e.g., y t , mv_y t ) into multiple sets of elements of different channel sets along the channel dimension and different spatial position sets along the two spatial dimensions. Each of the multiple sets of elements has a different combination of one of the different channel sets and one of the different spatial position sets. At 1820, the neural encoder can (as part of the entropy model network) the quantized versions of the multiple sets of elements (e.g.,

[0254]

Number

[0255]

[0156] For example, the current latent representation is the current latent SV representation of the current frame. In this case, when the neural encoder determines the current latent representation, the neural encoder uses the context encoder to determine the current latent SV representation.

[0256]

[0157] Alternatively, as another example, the current latent representation is the current latent MV representation of the current frame. In this case, when the neural encoder determines the current latent representation, the neural encoder uses motion estimation to determine the MV value of the current frame with respect to the previous frame, and uses the MV context encoder to determine the current latent MV representation from the MV value.

[0257]

[0158] The neural encoder can quantize the current latent representation, thereby generating quantized versions of a plurality of sets of elements of the current latent representation. In doing so, the neural encoder can apply at least some QS values (such as QS values for each region) determined using the entropy model network.

[0258]

[0159] As shown in FIG. 18B, at 1860, the neural decoder is a quantized version of the current latent representation (e.g.,

[0259]

Number

[0260]

Number

[0261]

[0160] For example, the current latent representation is the current latent SV representation of the current frame. In this case, the neural decoder can use a contextual decoder to estimate the current set of feature parameters of the current frame from the current latent SV representation, reconstruct the current frame from the estimated current set of feature parameters, and output the reconstructed current frame.

[0262]

[0161] Alternatively, as another example, the current latent representation is the current latent MV representation of the current frame. In this case, the neural decoder can use an MV context decoder to determine the MV value of the current frame from the current latent MV representation.

[0263]

[0162] The neural decoder can inverse-quantize quantized versions of multiple sets of elements of the current latent representation. In so doing, the neural decoder can apply at least some QS values (such as QS values for each region) determined using an entropy model network.

[0264]

[0163] The number of sets of elements depends on the implementation. In some exemplary implementations, due to the characteristics of the double-space prior distribution, the multiple sets of elements include: (a) a first set of elements having elements of a first channel set among different channel sets and a first spatial position set among different spatial position sets; (b) a second set of elements having elements of the first channel set and a second spatial position set among different spatial position sets; (c) a third set of elements having elements of a second channel set among different channel sets and elements of the second spatial position set; and (d) a fourth set of elements having elements of the second channel set and elements of the first spatial position set. For example, the first channel set includes the lower half of the channels, and the second channel set includes the upper half of the channels. Alternatively, the first channel set includes the even channels, and the second channel set includes the odd channels. The first spatial position set can include even positions, while the second spatial position set can include odd positions, or vice versa. Alternatively, the elements of the current latent representation are divided into more or fewer sets of elements.

[0265] [

[0164] ]FIG. 19A shows an exemplary method 1900 for performing quantization of multiple granularities. Correspondingly, FIG. 19B shows an exemplary method 1950 for performing inverse quantization of multiple granularities. The neural encoder that implements method 1900 can be either a neural video encoder or a neural image encoder. The neural decoder that implements method 1950 can be either a neural video decoder or a neural image decoder.

[0266] [

[0165] ]As shown in FIG. 19A, at 1910, the neural encoder can determine the current latent representation of the current frame (e.g., y t , mv_y t ). The elements of the current latent representation are logically organized along the channel dimension and two spatial dimensions. At 1920, the neural encoder can quantize the current latent representation in multiple stages using different quantization step values (e.g., qs global , qs ch , and qs sc ), thereby generating a quantized version of the current latent representation (e.g.,

[0267] [[Number]] ). Then, at 1930, the neural encoder can entropy encode the quantized version of the current latent representation.

[0268] [

[0166] ]For example, the current latent representation is the current latent SV representation of the current frame. In this case, when the neural encoder determines the current latent representation, the neural encoder uses a contextual encoder to determine the current latent SV representation.

[0269]

[0167] Alternatively, as another example, the current latent representation is the current latent MV representation of the current frame. In this case, when the neural encoder determines the current latent representation, the neural encoder uses motion estimation to determine the MV value of the current frame relative to the previous frame, and uses the MV context encoder to determine the current latent MV representation from the MV value.

[0270]

[0168] The neural encoder can use an entropy model network to estimate the statistical characteristics of the quantized version of the current latent representation, and the entropy encoding can use such statistical characteristics. The neural encoder can also use an entropy model network to determine at least a part of the QS value.

[0271]

[0169] As shown in FIG. 19B, in 1960, the neural decoder can receive a quantized version of the current latent representation (e.g.,

[0272]

Number

[0273]

[0170] For example, the current latent representation is the current latent SV representation of the current frame. In this case, the neural decoder can use a context decoder to estimate the current feature parameter set of the current frame from the current latent SV representation, reconstruct the current frame from the estimated current feature parameter set, and output the reconstructed current frame.

[0274]

[0171] Alternatively, as another example, the current latent representation is the current latent MV representation of the current frame. In this case, the neural decoder can use an MV context decoder to determine the MV value of the current frame from the current latent MV representation.

[0275]

[0172] The neural decoder can use an entropy model network to estimate the statistical characteristics of the quantized version of the current latent representation, and the entropy decoder can use such statistical characteristics. The neural decoder can also use an entropy model network to determine at least a part of the QS value.

[0276]

[0173] During encoding or decoding, the different QS values used for quantization (encoding) or inverse quantization (encoding / reconstruction or decoding) may include a global QS value for adjusting the bit rate and overall quality. The encoded data may include one or more syntax elements indicating a global QS value that is allowed to vary within a range. The different QS values may also include per-channel QS values for different channels of the current latent representation. The per-channel QS values may be predefined. Alternatively, the per-channel QS values may be allowed to change over time, in which case the encoded data may include syntax elements indicating the per-channel QS values. The different QS values may also include per-region QS values for different spatial regions of the current latent representation. The different spatial regions may be associated with different positions or areas of the current latent representation. The per-region QS values may be channel-specific or channel-independent. Alternatively, the different QS values may include other and / or additional QS values.

[0277]

[0174] In some exemplary implementations, for quantization, the multiple steps include: a first step of quantizing each element of the current latent representation using a global QS value; a second step of quantizing each element of the current latent representation of different channels using per-channel QS values; and a third step of quantizing each element of the current representation of different spatial regions for different channels using per-region QS values. For inverse quantization, the multiple steps include: a first step of inverse quantizing each element of the current representation of different spatial regions for different channels using per-region QS values; a second step of inverse quantizing each element of the current latent representation of different channels using per-channel QS values; and a third step of inverse quantizing each element of the current latent representation using a global QS value. Alternatively, the steps of quantization and inverse quantization may be performed in a different order.

[0278] IX. Exemplary Experimental Results

[0175] To evaluate the performance of the neural video codec technology described in this specification, experimental studies were conducted.

[0279]

[0176] For the training of the neural video codec in the experimental scenario, the training data is obtained from Vimeo-90k described in Xue et al., "Video Enhancement with Task-Oriented Flow", International Journal of Computer Vision (IJCV) 127,8:1106~1125, 2019. The videos are randomly cropped into 256x256 patches. The tests use the same test sequences described in Sheng 2021. All sequences are widely used in conventional video codecs and neural video codecs including HEVC classes B, C, D, E, and RGB. Additionally, 1080p videos from the UVG and MCL-JCV datasets are also tested. The UVG dataset is described in Mercat et al., "UVG dataset: 50 / 120fps 4K sequences for video codec analysis and development", Proceedings of the 11th ACM Multimedia Systems Conference, 297~302, 2020. The MCL-JCV dataset is described in Wang et al., "MCL-JCV: a JND-based H. 264 / AVC video quality assessment dataset", 2016 IEEE International Conference on Image Processing (ICIP), IEEE, 1509~1513 (2016).

[0280]

[0177] Regarding the training of neural video codecs in an experimental scenario, 96 frames are tested for each video. To approach a realistic scenario, the intra period is set to 32. The training uses a low-latency encoding setting, similar to most existing research. The compression ratio is measured by the BD-Rate described in Bjontegaard, "Calculation of average PSNR differences between RD curves", VCEG-M33 (2001), where negative values indicate bitrate savings and positive values indicate an increase in bitrate. In addition to the x265 encoder using the veryslow preset, the benchmarks for comparison also included HM-16.20 and VTM-13.2, which represent the optimal encoders of the H.265 standard and H.266 standard, respectively. For HM and VTM, the configuration with the highest compression ratio is used. The experimental results are compared with existing state-of-the-art neural video codecs including DVC_Pro, MLVC, RLVC, DCVC, and Sheng 2021. The DVC_Pro codec is described in Lu et al., "An end-to-end learning framework for video compression", IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(10):3292~3308 (2020). The MLVC codec is described in Lin et al., "M-LVC: multiple frames prediction for learned video compression", Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2020.The RLVC codec is described in Yang et al., "Learning for Video Compression with Recurrent Auto-Encoder and Recurrent Probability Model", IEEE Journal of Selected Topics in Signal Processing 15(2):388 - 401, 2021. The DCVC codec is described in Li et al., "Deep contextual video compression", Advances in Neural Information Processing Systems 34(2021).

[0281]

[0178] Regarding the training of the neural video codec, for the current latent representation of the MV value v t the entropy model and quantization for the current latent representation are in accordance with the entropy model and quantization method of the current latent SV representation y t The important difference is the input to the entropy model. In the encoding of the current latent representation of the MV value v t the input is the corresponding hyper prior distribution and latent prior distribution, that is, the quantized latent MV representation of the MV value from the previous frame. Since the generation of the temporal context depends on the decoded MV value, there is no temporal context prior distribution. Also, the neural image codec is trained to support the ability to have multiple rates with a single trained model for intra coding, as described above with reference to FIGS. 14A - 14B.

[0282]

[0179] During training, the loss function includes distortion and rate, i.e., loss = λ·D + R, where D refers to the distortion between the input current video frame and the reconstructed current video frame. For different visual goals, the distortion can be, for example, L2 loss or MS-SSIM, and MS-SSIM is described in Wang et al., "Multiscale structural similarity for image quality assessment," The Thirty-Seventh Asilomar Conference on Signals, Systems & Computers, Vol. 2, IEEE, 1398 - 1402, 2003. R represents the bits used to encode the quantized latent SV representation

[0283]

Number

[0284]

[0180] Regarding the training of the neural video codec in the experimental scenario, the training generally uses the multi-stage training method described in Sheng 2021. Further, to support multiple rates with a single model, the experiment uses different λ values in different optimization steps. To simplify the training, four λ values (85, 170, 380, 840) are used. Four random qs global values are set and learned through the RD loss at their respective corresponding λ values. Although only four λ values are used in training, it is noted that by adjusting the qs global during testing, the model can still achieve a wide rate range.

[0285]

[0181] The result of the training is a neural codec in which parameters are set for convolutional layers and other layers of the network structure of each component of the neural encoder and neural decoder. In some exemplary implementations, a QS value for each channel is also defined.

[0286]

[0182] The experimental results show that the neural video codec technology described herein has improved performance compared to existing neural video codec technologies, including the encoder of the latest conventional H.266 standard. For example, the experimental results show that the neural video codec technology described herein achieves bitrate savings of 67.4% and 57.1% respectively compared to DVCPro and DCVC for the UVG dataset. Further, the neural video codec technology described herein achieves an average bitrate savings of 4.7% compared to VTM. This represents the first neural video codec to outperform VTM using the highest compression ratio configuration. In particular, the neural video codec technology described herein exhibits better performance for 1080p videos (HEVC B, HEVC RGB, UVG, MCL-JCV). These results demonstrate the effectiveness of the hybrid entropy model for exploiting correlations from volumed data. When targeted at MS-SSIM, the experimental results also show that the neural video codec technology described herein leads to significant improvements, such as a 46.4% bitrate savings compared to VTM.

[0287]

[0183] As described above, the neural video codec technology described herein can achieve multiple rates with a single model. The global QS value qs global can be flexibly adjusted during testing. The global QS value qs global serves a similar role to the quantization parameter in conventional video codecs. During training, qsglobal can be derived by the RD loss. In the experiment, by interpolating between the maximum and minimum of the learned qs global values, 30 qs global values are manually generated. The experimental results confirm that a single model can achieve fine-grained rate control without any outliers. In contrast, previous methods such as DCVC and Sheng 2021 require different models for each rate point.

[0288]

[0184] The complexity of the model can be compared in terms of model size, MAC (multiply-accumulate), peak feature usage, encoding time, and decoding time. The experiment measures the numerical values using 1080p video frames as input. For the encoding / decoding time, the time on a V100 GPU including the time to write to and read from the bitstream is measured. Since the neural video codec technology described herein supports multi-rate with a single model, it significantly reduces the training and memory burden of the model. In particular, the neural video codec technology described herein brings a significant reduction in encoding / decoding time compared to DCVC, which uses an autoregressive prior distribution model that is difficult to parallelize.

[0289] V. Features

[0185] Different embodiments may include one or more of the features of the invention shown in the following table of features.

[0290]

Table 1

[0291]

Table 2

[0292]

Table 3

[0293]

Table 4

[0294]

Table 5

[0295]

Table 6

[0296]

Table 7

[0297]

Table 8

[0298]

Table 9

[0299]

Table 10

[0300]

Table 11

[0301]

Table 12

[0302]

Table 13

[0303]

Table 14

[0186] Considering the many possible embodiments to which the principles of the disclosed invention may be applied, it should be recognized that the embodiments shown are only preferred examples of the invention and should not be regarded as limiting the scope of the invention. Rather, the scope of the invention is defined by the following claims. Accordingly, we claim as our invention all that falls within the scope and spirit of these claims.

Claims

1. In a computer system implementing a neural video encoder, receiving a current video frame; encoding the current video frame to generate encoded data, wherein the step of encoding the current video frame comprises: determining a current latent representation of the current video frame; encoding the current latent representation using an entropy model network comprising one or more convolutional layers; and the step of encoding the current latent representation using the entropy model network comprises: estimating statistical characteristics of a quantized version of the current latent representation based at least in part on a previous latent representation of a previous video frame; entropy encoding the quantized version of the current latent representation based at least in part on the estimated statistical characteristics; and outputting the encoded data as part of a bitstream. A method comprising:

2. The method of claim 1, further comprising: quantizing the current latent representation to thereby generate the quantized version of the current latent representation. A method comprising:

3. The method of claim 2, wherein the step of encoding the current latent representation using the entropy model network comprises: determining at least some quantization step (QS) values for the current latent representation based at least in part on the previous latent representation, wherein the step of quantizing uses the at least some QS values. A method further comprising:

4. The method of claim 1, wherein the current latent representation is a current latent sample value (「SV」) representation of the current video frame, the previous latent representation is a previous latent SV representation of the previous video frame, and the step of determining the current latent representation comprises determining the current latent SV representation using a contextual encoder comprising one or more convolutional layers. A method comprising:

5. The method according to claim 1, wherein the current latent representation is the current latent motion vector ( "MV") representation of the current video frame, the previous latent representation is the previous latent MV representation of the previous video frame, and the step of determining the current latent representation comprises: using motion estimation to determine the MV value of the current video frame relative to the previous video frame; determining the current latent MV representation from the MV value using an MV context encoder comprising one or more convolutional layers; A method comprising:

6. In a computer system implementing a neural video decoder, receiving encoded data as part of a bitstream; decoding the encoded data to reconstruct a current video frame, the step of decoding the encoded data reconstructing a current latent representation of the current video frame using an entropy model network comprising one or more convolutional layers; comprising, the step of reconstructing the current latent representation comprises: estimating statistical characteristics of a quantized version of the current latent representation based at least in part on a previous latent representation of a previous video frame; entropy decoding the quantized version of the current latent representation based at least in part on the estimated statistical characteristics; and steps; A method comprising:

7. The method according to claim 6, further comprising: inverse quantizing the quantized version of the current latent representation; A method further comprising:

8. The method according to claim 7, wherein the step of reconstructing the current latent representation using the entropy model network comprises: determining at least some quantization step (QS) values for the current latent representation based at least in part on the previous latent representation, the step of inverse quantizing using the at least some QS values; A method further comprising:

9. The method according to claim 6, wherein the current latent representation is the current latent sample value ( "SV") representation of the current video frame, the previous latent representation is the previous latent SV representation of the previous video frame, and the method comprises: estimating, using a contextual decoder comprising one or more convolutional layers, a current set of feature parameters of the current video frame from the current latent SV representation; reconstructing the current video frame from the estimated current set of feature parameters; outputting the reconstructed current video frame; and a method further comprising.

10. The method according to claim 6, wherein the current latent representation is a current latent motion vector ("MV") representation of the current video frame, the previous latent representation is a previous latent MV representation of the previous video frame, and the method further comprises determining MV values of the current video frame from the current latent MV representation using an MV contextual decoder comprising one or more convolutional layers.

11. The method according to any one of claims 1 to 10, wherein the step of estimating the statistical characteristics of the quantized version of the current latent representation is at least partially based also on hyper-prior distribution parameters of the current video frame, and the hyper-prior distribution parameters are generated from the current latent representation using a hyper-prior distribution encoder comprising one or more convolutional layers.

12. The method according to any one of claims 1 to 4 and 6 to 9, wherein the step of estimating the statistical characteristics of the quantized version of the current latent representation is at least partially based also on one or more temporal context parameter sets of the current video frame, and the one or more temporal context parameter sets are generated from the previous set of feature parameters of the previous video frame and the motion vector ("MV") values of the current video frame using a temporal context mining network comprising one or more convolutional layers.

13. The method according to any one of claims 1 to 10, wherein the statistical characteristics include one or more mean values and one or more scale parameters of a probability distribution function of the quantized version of the current latent representation.

14. The method according to any one of claims 1 to 13, wherein the elements of the current latent representation are logically organized along a channel dimension and two spatial dimensions.

15. The method according to claim 14, wherein the step of estimating the statistical characteristics of the quantized version of the current latent representation comprises: dividing the elements of the current latent representation into a plurality of sets of elements of different channel sets along the channel dimension and different spatial position sets along the two spatial dimensions, each of the plurality of sets of elements having a different combination of one of the different channel sets and one of the different spatial position sets; estimating the statistical characteristics of the quantized version of a second set of elements among the plurality of sets of elements, at least partially based on the quantized version of a first set of elements among the plurality of sets of elements; and a method.

16. The method according to claim 15, wherein the plurality of sets of elements comprise: a first set of elements having elements of a first channel set among the different channel sets and a first spatial position set among the different spatial position sets; a second set of elements having elements of the first channel set and a second spatial position set among the different spatial position sets; a third set of elements having elements of a second channel set among the different channel sets and elements of the second spatial position set; and a fourth set of elements having elements of the second channel set and elements of the first spatial position set. and a method.

17. The method according to claim 16, wherein the step of estimating the statistical characteristics of the quantized version of the current latent representation comprises: estimating the statistical characteristics of the quantized version of the first set of elements; estimating the statistical characteristics of the quantized version of the third set of elements; and fusing the quantized version of the first set of elements and the quantized version of the third set of elements with other inputs. estimating statistical characteristics of the quantized version of the fourth set of elements using the result of the fusion; comprising; the step of estimating the statistical characteristics of the quantized version of the second set of elements also uses the result of the fusion; method.

18. The method according to claim 1, quantizing the current latent representation at each of the plurality of stages using different quantization step ("QS") values at the plurality of stages, thereby generating the quantized version of the current latent representation; the method further comprising.

19. The method according to claim 6, inverse quantizing the quantized version of the current latent representation at each of the plurality of stages using different quantization step ("QS") values at the plurality of stages; the method further comprising.

20. The method according to claim 18 or 19, wherein the different QS values are a global QS value for adjusting bitrate and overall quality, a per-channel QS value for different channels of the current latent representation, a per-region QS value for different spatial regions of the current latent representation, wherein the different spatial regions are associated with different positions or areas of the current latent representation, and the per-region QS values for the different regions are per-channel specific or channel-independent; the method comprising.

21. In a computer system implementing a neural image encoder or a neural video encoder, receiving a current frame; encoding the current frame to generate encoded data, wherein the step of encoding the current frame comprises determining a current latent representation of the current frame, wherein the elements of the current latent representation are logically organized along a channel dimension and two spatial dimensions; encoding the current latent representation using an entropy model network comprising one or more convolutional layers; and the step of encoding the current latent representation using the entropy model network comprises Dividing the elements of the current latent representation into a plurality of sets of elements of different channel sets along the channel dimension and different spatial position sets along the two spatial dimensions, wherein each of the plurality of sets of elements has a different combination of one of the different channel sets and one of the different spatial position sets; Estimating the statistical characteristics of the quantized versions of the plurality of sets of elements, respectively, including estimating the statistical characteristics of the quantized version of a second set of elements among the plurality of sets of elements, at least partially based on the quantized version of a first set of elements among the plurality of sets of elements; Entropy encoding the quantized versions of the plurality of sets of elements, respectively, based at least in part on the estimated statistical characteristics; Including steps; Outputting the encoded data as part of a bitstream; A method including.

22. The method according to claim 21, wherein: The step of quantizing the current latent representation, thereby generating the quantized versions of the plurality of sets of elements of the current latent representation; A method further comprising.

23. The method according to claim 22, wherein the step of encoding the current latent representation using the entropy model network: Determining at least some quantization step (QS) values for the current latent representation, wherein the step of quantizing uses the at least some QS values; A method further comprising.

24. The method according to claim 21, wherein the current latent representation is the current latent sample value ( "SV") representation of the current frame, and the step of determining the current latent representation includes determining the current latent SV representation using a contextual encoder including one or more convolutional layers;

25. The method according to claim 21, wherein the current latent representation is the current latent motion vector ( "MV") representation of the current frame, and the step of determining the current latent representation: A step of using motion estimation to determine an MV value of the current frame with respect to a previous frame; A step of determining the current latent MV representation from the MV value using an MV context encoder including one or more convolutional layers A method comprising.

26. In a computer system implementing a neural image decoder or a neural video decoder, A step of receiving encoded data as part of a bitstream; A step of decoding the encoded data to reconstruct a current frame, wherein the step of decoding the encoded data is A step of reconstructing a current latent representation of the current frame using an entropy model network including one or more convolutional layers, wherein elements of the current latent representation are logically organized along a channel dimension and two spatial dimensions, and the elements of the current latent representation are divided into a plurality of sets of elements of different channel sets along the channel dimension and different spatial position sets along the two spatial dimensions, and each of the plurality of sets of elements has a different combination of one of the different channel sets and one of the different spatial position sets, the step Including steps and A method including, wherein the step of reconstructing the current latent representation is A step of respectively estimating statistical characteristics of quantized versions of the plurality of sets of elements, including a step of estimating statistical characteristics of a quantized version of a second set of elements among the plurality of sets of elements based at least in part on the quantized version of a first set of elements among the plurality of sets of elements; A step of entropy decoding the quantized versions of the plurality of sets of elements respectively based at least in part on the estimated statistical characteristics; A method including.

27. The method according to claim 26, wherein A step of inverse quantizing the quantized versions of the plurality of sets of elements of the current latent representation A method further including.

28. The method according to claim 27, wherein the step of reconstructing the current latent representation using the entropy model network is Determining at least some quantization step (QS) values for the current latent representation, wherein the step of inverse quantization uses the at least some QS values A method further comprising. **Claim 29** The method according to claim 26, wherein the current latent representation is the current latent sample value ("SV") representation of the current frame, and the method comprises Estimating a current set of feature parameters of the current frame from the current latent SV representation using a contextual decoder including one or more convolutional layers; Reconstructing the current frame from the estimated current set of feature parameters; Outputting the reconstructed current frame A method further comprising. **Claim 30** The method according to claim 26, wherein the current latent representation is the current latent motion vector ("MV") representation of the current frame, and the method further comprises determining MV values of the current frame from the current latent MV representation using an MV contextual decoder including one or more convolutional layers. **Claim 31** The method according to any one of claims 21 to 30, wherein the statistical characteristics each include one or more mean values and one or more scale parameters of the probability distribution function of each of the quantized versions of the plurality of sets of elements. **Claim 32** The method according to any one of claims 21 to 30, wherein the plurality of sets of elements comprises The first set of elements, having elements of the first channel set of the different channel sets and the first spatial position set of the different spatial position sets; The second set of elements, having elements of the first channel set and the second spatial position set of the different spatial position sets; The third set of elements, having elements of the second channel set of the different channel sets and the second spatial position set; The fourth set of elements, having elements of the second channel set and the first spatial position set A method comprising. **Claim 33** The method according to claim 32, wherein the step of respectively estimating the statistical characteristics of the quantized versions of the plurality of sets of elements comprises: estimating the statistical characteristics of the quantized version of the first set of elements; estimating the statistical characteristics of the quantized version of the third set of elements; fusing the quantized version of the first set of elements and the quantized version of the third set of elements with other inputs; using the result of the fusion to estimate the statistical characteristics of the quantized version of the fourth set of elements; and the step of estimating the statistical characteristics of the quantized version of the second set of elements also uses the result of the fusion. Method.

34. The method according to claim 32, wherein the first channel set includes the first half of the channels; the second channel set includes the second half of the channels; the first set of spatial positions includes one of even and odd positions; the second set of spatial positions includes the other of the even and odd positions. Method.

35. The method according to claim 32, wherein the entropy model network comprises: a first fusion stage for fusing inputs; a first estimation stage for using the output from the first fusion stage to estimate the statistical characteristics of the quantized version of the first set of elements and the statistical characteristics of the quantized version of the third set of elements; a second fusion stage for fusing the quantized version of the first set of elements and the quantized version of the third set of elements with other inputs, wherein the other inputs include some output from the first estimation stage; a second estimation stage for using the output from the second fusion stage to estimate the statistical characteristics of the quantized version of the second set of elements and the statistical characteristics of the quantized version of the fourth set of elements. Method.

36. The method according to any one of claims 21 to 30, wherein the step of estimating the statistical characteristics of the quantized versions of the plurality of sets of elements comprises: The hyper-prior distribution parameters of the current frame, which are hyper-prior distribution parameters generated from the current latent representation using a hyper-prior distribution encoder including one or more convolutional layers, and when the current frame is the current video frame, the previous latent representation of the previous video frame and a method that is also at least partially based thereon. **Claim 37** The method according to any one of claims 21 to 24 and 26 to 29, wherein the step of estimating the statistical characteristics of the quantized version of the plurality of sets of elements is The hyper-prior distribution parameters of the current frame, which are hyper-prior distribution parameters generated from the current latent representation using a hyper-prior distribution encoder including one or more convolutional layers, and when the current frame is the current video frame, the previous latent representation of the previous video frame, and one or more temporal context parameter sets of the current video frame, which are one or more temporal context parameter sets generated from the previous feature parameter set of the previous video frame and the motion vector ("MV") values of the current video frame using a temporal context mining network including one or more convolutional layers and a method that is also at least partially based thereon. **Claim 38** The method according to claim 21, wherein the step of quantizing the current latent representation at each of the plurality of stages using different quantization step ("QS") values, thereby generating the quantized version of the current latent representation is further included. **Claim 39** The method according to claim 26, wherein the step of inverse quantizing the quantized version of the current latent representation at each of the plurality of stages using different quantization step ("QS") values is further included. **Claim 40** The method according to claim 38 or 39, wherein the different QS values are a global QS value for adjusting the bitrate and overall quality, and a per-channel QS value for different channels of the current latent representation, and QS values for a plurality of regions for different spatial regions of the current latent representation, wherein the different spatial regions are associated with different positions or regions of the current latent representation, and the QS values for the different regions are either channel-specific or channel-independent, and the QS values for the plurality of regions A method comprising. **Claim 41** In a computer system implementing a neural image encoder or a neural video encoder, Receiving a current frame; Encoding the current frame to generate encoded data, wherein the step of encoding the current frame comprises: Determining a current latent representation of the current frame, wherein the elements of the current latent representation are logically organized along a channel dimension and two spatial dimensions; Quantizing the current latent representation at each of the plurality of stages using different quantization step ("QS") values at the plurality of stages, thereby generating a quantized version of the current latent representation; Entropy encoding the quantized version of the current latent representation; Steps comprising; Outputting the encoded data as part of a bitstream; A method comprising. **Claim 42** The method according to claim 41, wherein the step of encoding the current frame comprises: Estimating statistical characteristics of the quantized version of the current latent representation using an entropy model network comprising one or more convolutional layers, wherein the step of entropy encoding is at least partially based on the estimated statistical characteristics; The method further comprising. **Claim 43** The method according to claim 42, wherein the step of encoding the current frame comprises: Using the entropy model network to determine at least a portion of the QS values; The method further comprising. **Claim 44** The method according to claim 41, wherein the current latent representation is a current latent sample value ("SV") representation of the current frame, and the step of determining the current latent representation comprises determining the current latent SV representation using a contextual encoder comprising one or more convolutional layers.

45. The method according to claim 41, wherein the current latent representation is the current latent motion vector ( "MV") representation of the current frame, and the step of determining the current latent representation comprises: using motion estimation to determine the MV value of the current frame relative to the previous frame; using an MV context encoder including one or more convolutional layers to determine the current latent MV representation from the MV value and a method comprising:

46. In a computer system implementing an image decoder or a video decoder, receiving the encoded data as part of a bitstream; decoding the encoded data to reconstruct the current frame, the step of decoding the encoded data comprising: reconstructing the current latent representation of the current frame, the elements of the current latent representation being logically organized along a channel dimension and two spatial dimensions; and a method comprising: the step of reconstructing the current latent representation comprising: entropy decoding a quantized version of the current latent representation; inverse quantizing the quantized version of the current latent representation at each of the plurality of stages using different quantization step ( "QS") values at the plurality of stages and a method comprising:

47. The method according to claim 46, wherein the step of decoding the current frame comprises: using an entropy model network including one or more convolutional layers to estimate statistical characteristics of the quantized version of the current latent representation, the step of entropy decoding being at least partially based on the estimated statistical characteristics; and a method further comprising:

48. The method according to claim 47, wherein the step of decoding the current frame comprises: using the entropy model network to determine at least a part of the QS values and a method further comprising:

49. The method according to claim 46, wherein the current latent representation is the current latent sample value ( "SV") representation of the current frame, and the method comprises: Estimating, using a contextual decoder including one or more convolutional layers, a current set of feature parameters of the current frame from the current latent SV representation; Reconstructing the current frame from the estimated current set of feature parameters; Outputting the reconstructed current frame; A method further comprising.

50. The method according to claim 46, wherein the current latent representation is a current latent motion vector ("MV") representation of the current frame, and the method further comprises determining MV values of the current frame from the current latent MV representation using an MV contextual decoder including one or more convolutional layers.

51. The method according to any one of claims 41 to 50, wherein the different QS values are A global QS value for adjusting the bit rate and overall quality, Per-channel QS values for different channels of the current latent representation, Per-region QS values for different spatial regions of the current latent representation, wherein the different spatial regions are associated with different positions or areas of the current latent representation, and the per-region QS values for the different regions are per-channel specific or channel-independent; A method comprising.

52. The method according to any one of claims 41 to 45, wherein the plurality of steps are A first step including quantizing each element of the current latent representation using a global QS value among the different QS values; A second step including quantizing each element of the current latent representation of different channels using per-channel QS values among the different QS values; A third step including quantizing each element of the current representation of different spatial regions for different channels using per-region QS values among the different QS values; A method comprising.

53. The method according to any one of claims 46 to 50, wherein the plurality of steps are A first step including inverse quantizing each element of the current representation of different spatial regions for different channels using per-region QS values among the different QS values; A second stage including the step of inverse quantizing each element of the current latent representation of the different channels using QS values for each of a plurality of channels among the different QS values; A third stage including the step of inverse quantizing each element of the current latent representation using a global QS value among the different QS values A method comprising.

54. The method according to any one of claims 41 to 50, wherein the different QS values include a global QS value, the encoded data includes one or more syntax elements indicating the global QS value, and the global QS value is allowed to vary within a range.

55. The method according to any one of claims 41 to 50, wherein the different QS values include QS values for each of a plurality of channels, the QS values for each of the plurality of channels are predefined or the encoded data includes a syntax element indicating the QS values for each of the plurality of channels. A method.

56. The method according to any one of claims 41 to 50, wherein statistical characteristics of the quantized version of the current latent representation are hyper-prior distribution parameters of the current frame, hyper-prior distribution parameters generated from the current latent representation using a hyper-prior distribution encoder including one or more convolutional layers, and when the current frame is the current video frame, the previous latent representation of the previous video frame A method further comprising the step of estimating at least partially based on.

57. The method according to any one of claims 41 to 44 and 46 to 49, wherein statistical characteristics of the quantized version of the current latent representation are hyper-prior distribution parameters of the current frame, hyper-prior distribution parameters generated from the current latent representation using a hyper-prior distribution encoder including one or more convolutional layers, and when the current frame is the current video frame, the previous latent representation of the previous video frame, and One or more temporal context parameter sets of the current video frame, using a temporal context mining network including one or more convolutional layers, from the previous feature parameter set of the previous video frame and the motion vector ("MV") values of the current video frame One or more temporal context parameter sets generated and A method further comprising the step of estimating at least partially based on. **Claim 58** The method according to any one of claims 41 to 50, wherein the step of estimating the statistical characteristics of the quantized version of the current latent representation, Dividing the elements of the current latent representation into a plurality of sets of elements of different channel sets along the channel dimension and different spatial position sets along the two spatial dimensions, each of the plurality of sets of elements having a different combination of one of the different channel sets and one of the different spatial position sets, the step of, Estimating the statistical characteristics of the quantized version of the second set of elements among the plurality of sets of elements, at least partially based on the quantized version of the first set of elements among the plurality of sets of elements A method further comprising the steps of including. **Claim 59** The method according to claim 58, wherein the plurality of sets of elements are The first set of elements, the first set of elements having elements of a first channel set among the different channel sets and a first spatial position set among the different spatial position sets, The second set of elements, the second set of elements having elements of the first channel set and a second spatial position set among the different spatial position sets, A third set of elements, the third set of elements having elements of a second channel set among the different channel sets and elements of the second spatial position set, A fourth set of elements, the fourth set of elements having elements of the second channel set and a first spatial position set Including, method. **Claim 60** The method according to claim 59, wherein the step of estimating the statistical characteristics of the quantized version of the current latent representation is estimating statistical characteristics of the quantized version of the first set of elements; estimating statistical characteristics of the quantized version of the third set of elements; fusing the quantized version of the first set of elements and the quantized version of the third set of elements with other inputs; using the result of the fusion to estimate statistical characteristics of the quantized version of the fourth set of elements; comprising; the step of estimating the statistical characteristics of the quantized version of the second set of elements also uses the result of the fusion; method. **Claim 61** One or more non-transitory computer-readable media storing computer-executable instructions for causing a computer system to perform the operations of the method according to any one of claims 1 to 60 when programmed by the computer-executable instructions. **Claim 62** A computer system configured to perform the operations of the method according to any one of claims 1 to 5, 11 to 25, 30 to 45, and 51 to 60, a frame buffer configured to store the current frame or the current video frame; a video encoder configured to perform the encoding step; an encoded data buffer configured to store the encoded data for output; a computer system comprising. **Claim 63** A computer system configured to perform the operations of the method according to any one of claims 6 to 20, 26 to 40, and 46 to 60, an encoded data buffer configured to store the encoded data; a video decoder configured to perform the decoding step; a frame buffer configured to store the reconstructed current video frame for output; a computer system comprising.