Method and device for content-adaptive online education in neural image compression

A neural network-based video encoding and decoding method addresses inefficiencies in existing video coding technologies by adapting to content-specific characteristics, improving compression efficiency and reducing redundancy.

KR102996814B1Active Publication Date: 2026-07-29TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2022-04-29
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in achieving efficient compression ratios while maintaining acceptable image quality, particularly in applications requiring high bandwidth and storage efficiency, such as consumer streaming and television distribution, due to limitations in intra-prediction and motion compensation techniques.

Method used

Implementing a neural network-based approach for video encoding and decoding, where a pre-trained neural network is updated using neural network update information within the coded bitstream to enhance intra-prediction and motion compensation, allowing for more efficient compression and reconstruction of video data.

Benefits of technology

The neural network-based approach significantly improves compression efficiency by adapting to content-specific characteristics, reducing redundancy and enhancing the compression ratio without compromising image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure R1020227041731_ABST
    Figure R1020227041731_ABST
Patent Text Reader

Abstract

The aspects of the present disclosure provide a method, an apparatus for video decoding, and a computer-readable non-transient storage medium. The apparatus may include a processing circuit. The processing circuit is configured to decode neural network update information within a coded bitstream for a neural network in the video decoder. The neural network is composed of pre-trained parameters. The neural network update information indicates an alternative parameter corresponding to the encoded image to be reconstructed and to the pre-trained parameter among the pre-trained parameters. The processing circuit is configured to update the neural network in the video decoder based on the alternative parameter. The processing circuit is configured to decode the encoded image based on the updated neural network for the encoded image. The processing circuit is configured to update the neural network of the video decoder based on replacement parameters. The processing circuit is configured to decode the encoded image based on the updated neural network for the encoded image.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] Integration by reference

[0002] The present application claims the benefit of priority to U.S. Provisional Application No. 63 / 182,396, "Content-adaptive Online Training in Neural Image Compression," filed April 30, 2021, and the benefit of priority to U.S. Patent Application No. 17 / 729,994, "METHOD AND APPARATUS FOR CONTENT-ADAPTIVE ONLINE TRAINING IN NEURAL IMAGE COMPRESSION," filed April 26, 2022. The disclosures of these prior applications are incorporated into the supplementary application by reference in their entirety.

[0003] The present disclosure describes embodiments generally related to video coding. Background Technology

[0004] The background description provided in this specification is generally intended to provide the context of the present disclosure. To the extent described in this background section, the works of the currently named inventors and aspects of description that may not be recognized as prior art at the time of filing are not recognized as prior art of the present disclosure, either expressly or impliedly.

[0005] Video coding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital images and / or videos may contain a series of pictures, each having, for example, a spatial dimension of 1920×1080 luminance samples (also called luminance samples) and associated chrominance samples (also called chroma samples). The series of pictures may have, for example, 60 pictures per second or a fixed or variable picture rate of 60 Hz (informally also known as the frame rate). Uncompressed images and / or videos have specific bitrate requirements. For example, 1080p60 4:2:0 video with 8 bits per sample (1920×1080 luminance sample resolution at a 60 Hz frame rate) requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires more than 600 GB of storage space.

[0006] One objective of video coding and decoding may be to reduce the redundancy of input image and / or video signals through compression. Compression can help reduce the aforementioned bandwidth and / or storage requirements by more than double digits in some cases. Both lossless compression and lossy compression, as well as combinations thereof, can be used. Lossless compression refers to a technique that can reconstruct an exact copy of the original signal from the compressed original signal. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signals is small enough to make the reconstructed signal useful for the intended application. For video, lossy compression is widely employed. The amount of acceptable distortion varies depending on the application. For example, users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio may reflect the following: the higher the acceptable / tolerated distortion, the higher the compression ratio may be. Although the description in this specification uses video encoding / decoding as an exemplary example, the same technology may be applied to image encoding / decoding in a similar manner without departing from the spirit of this disclosure.

[0007] Video encoders and decoders can utilize techniques from a wide range of categories, including, for example, motion compensation, transform, quantization, and entropy coding.

[0008] Video codec technologies may include a technique known as intra coding. In intra coding, sample values ​​are represented without reference to samples of a previously reconstructed reference picture or other data. In some video codecs, a picture is spatially subdivided into sample blocks. If all blocks of samples are coded in intra mode, that picture may be an intra picture. Derivations such as intra pictures and independent decoder refresh pictures can be used to reset the decoder state and thus serve as the first picture or still image of the coded video bitstream and video session. Samples in an intra block may be exposed to transformation, and transformation coefficients may be quantized before entropy coding. Intra prediction may be a technique that minimizes sample values ​​in the pre-transform domain. In some cases, the smaller the DC value after conversion and the smaller the AC coefficient, the fewer bits are required at a given quantization step size to represent the block after entropy coding.

[0009] For example, traditional intra-coding, such as that known from MPEG-2 generative coding techniques, does not use intra-prediction. However, some new video compression techniques include methods that attempt to use, for example, surrounding sample data and / or metadata acquired during the encoding and / or decoding of spatially neighboring and preceding data blocks in the decoding order. These techniques will henceforth be referred to as "intra-prediction" techniques. In at least some cases, intra-prediction uses only reference data from the current picture being reconstructed, rather than the reference picture.

[0010] There may be various forms of intra-prediction. If one or more of these techniques can be used in a given video coding technique, the technique in use may be coded in an intra-prediction mode. In some cases, a mode may have submodes and / or parameters, which may be coded individually or included in a mode codeword. Which codeword is used for a given combination of mode, submode, and / or parameter can affect the coding efficiency gain through intra-prediction, and thus may also affect the entropy coding technique used to convert the codeword into a bitstream.

[0011] Intra prediction of specific modes was introduced with H.264, improved in H.265, and further enhanced in modern coding techniques such as the Joint Exploration Model (JEM), Universal Video Coding (VVC), and Benchmark Set (BMS). A predictor block can be formed using neighbor samples belonging to already available samples. The sample values ​​of neighbor samples are copied to the predictor block according to the direction. References to the direction in use can be encoded in the bitstream or predicted themselves.

[0012] Referring to FIG. 1a, a subset of nine known predictor directions from the 33 possible predictor directions of H.265 (corresponding to the 33 angle modes of the 35 intra modes) is shown in the lower right. The point where the arrows converge (101) indicates the predicted sample. The arrows indicate the direction in which the sample is predicted. For example, arrow (102) indicates that the sample (101) is predicted to the upper right at an angle of 45 degrees from the horizontal from one sample or sample. Similarly, arrow (103) indicates that the sample (101) is predicted to the lower left of the sample (101) at an angle of 22.5 degrees from the horizontal from one sample or sample.

[0013] Referring still to FIG. 1a, a square block (104) of 4×4 samples (indicated by a thick dashed line) is shown in the upper left corner. The square block (104) contains 16 samples, each labeled "S", a position in the Y dimension (e.g., column index), and a position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (from the top) and the first sample in the X dimension (from the left). Similarly, sample S44 is the fourth sample of the block (104) in both the Y and X dimensions. Since the block size is 4×4 samples, S44 is located in the lower right corner. Additional reference samples following a similar numbering scheme are shown. The reference samples are labeled R, a Y position (e.g., row index) and an X position (column index) relative to the block (104). In both H.264 and H.265, the predicted sample is adjacent to the block being rebuilt; therefore, there is no need to use negative values.

[0014] Intra-picture prediction can be performed by copying reference sample values ​​from appropriate neighbor samples by the signaled prediction direction. For example, assume that the coded video bitstream includes signaling indicating a prediction direction corresponding to the arrow (102) for this block—that is, samples are predicted from one prediction sample or samples located at the top right at a 45-degree angle from the horizontal. In this case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.

[0015] In some cases, to calculate reference samples, especially when the direction cannot be divided equally into 45 degrees, the values ​​of multiple reference samples may be combined, for example, through interpolation.

[0016] As video coding technology has advanced, the number of possible directions has increased. H.264 (2003) could represent nine different directions. H.265 (2013) increased this to 33, and at the time of its release, JEM / VVC / BMS could support up to 65 directions. Experiments were conducted to identify the most probable direction, and specific techniques of entropy coding were used to represent such high-probability directions with a small number of bits, thereby accommodating a specific penalty for low-probability directions. Additionally, the direction itself can be predicted from neighboring directions used in already decoded neighboring blocks.

[0017] FIG. 1b illustrates a schematic diagram (110) showing 65 intra-predicted directions according to JEM to explain the number of predicted directions increasing over time.

[0018] The mapping of intra-predicted direction bits within a coded video bitstream representing direction can vary depending on the video coding technique; for example, it ranges from a simple direct mapping of the predicted direction to complex adaptive schemes involving intra-predicted modes, codewords, and the most likely mode, as well as similar techniques. However, in all cases, there may be a specific direction in the video content that is statistically less likely to occur than other specific directions. Since the goal of video compression is to reduce redundancy, in well-functioning video coding techniques, less likely directions are represented by a larger number of bits than more likely directions.

[0019] Motion compensation may be a lossy compression technique and may relate to a technique used to predict a newly reconstructed picture or part of a picture after a block of sample data from a previously reconstructed picture or part thereof (reference picture) has been spatially moved in a direction indicated by a motion vector (hereinafter MV). In some cases, the reference picture may be identical to the picture currently being reconstructed. The MV may be two or three dimensions in X and Y, with a third being an indication of the reference picture in use (the latter may indirectly be a time dimension).

[0020] In some video compression techniques, an MV applicable to a certain area of ​​sample data can be predicted from other MVs, for example, from MVs associated with other areas of sample data that are spatially adjacent to the area being reconstructed and precede it in the decoding order. This significantly reduces the amount of data required to code the MV, thereby eliminating redundancy and increasing the compression ratio. For example, when coding an input video signal derived from a camera (referred to as natural video), MV prediction can work effectively because there is a statistical probability that an area larger than the area where a single MV can be applied moves in a similar direction; thus, in some cases, it can be predicted using similar motion vectors derived from MVs in neighboring areas. As a result, the MV found for a given area is similar or identical to the MV predicted from surrounding MVs, which can be represented by fewer bits than would have been used if the MV had been coded directly after entropy coding. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., MV) derived from the original signal (i.e., sample stream). In other cases, the MV prediction itself may be lost, for example, because rounding error occurs when calculating the predictor from several surrounding MVs.

[0021] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). Among the many MV prediction mechanisms provided by H.265, the technique described here is hereinafter referred to as "spatial merge".

[0022] Referring to FIG. 2, the current block (201) may include samples discovered by an encoder during the motion search process so as to be predictable from a previous block of the same size that is spatially shifted. Instead of directly coding the MV, the MV may be derived from metadata associated with one or more reference pictures, for example from the most recent reference picture (in decoding order), using an MV associated with one of five neighboring samples denoted as A0, A1 and B0, B1, B2 (202–206, respectively). In H.265, the MV prediction may use a predictor from the same reference picture that is being used by the neighboring block.

[0023] The aspects of the present disclosure provide a method and apparatus for video encoding and decoding. In some examples, the apparatus for video decoding includes a processing circuit. The processing circuit is configured to decode neural network update information within a coded bitstream for a neural network in a video decoder. The neural network is composed of pre-trained parameters. The neural network update information indicates a replacement parameter corresponding to an encoded image to be reconstructed and corresponding to a pre-trained parameter among the pre-trained parameters. The processing circuit is configured to update the neural network in the video decoder based on the replacement parameter. The processing circuit is configured to decode the encoded image based on the updated neural network for the encoded image.

[0024] In one embodiment, the neural network update information further directs one or more alternative parameters for one or more remaining neural networks in the video decoder. The processing circuit is configured to update the one or more remaining neural networks based on the one or more alternative parameters.

[0025] In one embodiment, the coded bitstream further directs one or more encoded bits used to determine a context model for decoding the encoded image. The video decoder includes a main decoder network, a context model network, an entropy parameter network, and a hyperdecoder network. The neural network is one of the main decoder network, the context model network, the entropy parameter network, and the hyperdecoder network. The processing circuit is configured to decode the one or more encoded bits using the hyperdecoder network. The processing circuit may determine a context model using the context model network and the entropy parameter network based on one or more decoded bits of the encoded image and quantized potentials available to the context model network. The processing circuit may decode the encoded image using the main decoder network and the context model.

[0026] In one embodiment, the pre-trained parameter is a pre-trained bias term.

[0027] In one embodiment, the pre-trained parameter is a pre-trained weighting coefficient.

[0028] In one embodiment, the neural network update information indicates a plurality of alternative parameters corresponding to a plurality of pre-trained parameters among the pre-trained parameters for the neural network. The plurality of pre-trained parameters includes the pre-trained parameters, and the plurality of pre-trained parameters include one or more pre-trained bias terms and one or more pre-trained weighting coefficients. The processing circuit is configured to update the neural network in the video decoder based on the plurality of alternative parameters including the alternative parameters.

[0029] In one embodiment, the neural network update information indicates the difference between the alternative parameter and the pre-trained parameter. The processing circuit may determine the alternative parameter based on the sum of the difference and the pre-trained parameter.

[0030] In one embodiment, the processing circuit can decode another encoded image from the coded bitstream based on the updated neural network.

[0031] The aspects of the present disclosure also provide a non-transient computer-readable storage medium for storing a program, said program executable to perform a method for video encoding and decoding by at least one processor. Brief explanation of the drawing

[0032] Additional features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. Figure 1a is a schematic diagram of an exemplary subset of intra-prediction modes. Figure 1b is a diagram of an exemplary intra-predicted direction. FIG. 2 illustrates a current block (201) and surrounding samples according to one embodiment. FIG. 3 is a schematic diagram of a simplified block diagram of a communication system (300) according to one embodiment. FIG. 4 is a schematic diagram of a simplified block diagram of a communication system (400) according to one embodiment. FIG. 5 is a schematic diagram of a simplified block diagram of a decoder according to one embodiment. FIG. 6 is a schematic diagram of a simplified block diagram of an encoder according to one embodiment. FIG. 7 illustrates a block diagram of an encoder according to another embodiment. FIG. 8 illustrates a block diagram of a decoder according to another embodiment. FIG. 9 illustrates an exemplary NIC framework according to one embodiment of the present disclosure. FIG. 10 illustrates an exemplary convolutional neural network (CNN) of a main encoder network according to one embodiment of the present disclosure. FIG. 11 illustrates an exemplary CNN of a main decoder network according to one embodiment of the present disclosure. FIG. 12 shows an exemplary CNN of a hyperencoder according to one embodiment of the present disclosure. FIG. 13 illustrates an exemplary CNN of a hyper decoder according to one embodiment of the present disclosure. FIG. 14 illustrates an exemplary CNN of a context model network according to one embodiment of the present disclosure. FIG. 15 illustrates an exemplary CNN of an entropy parameter network according to one embodiment of the present disclosure. FIG. 16a illustrates an exemplary video encoder according to one embodiment of the present disclosure. FIG. 16b illustrates an exemplary video decoder according to one embodiment of the present disclosure. FIG. 17 illustrates an exemplary video encoder according to one embodiment of the present disclosure. FIG. 18 illustrates an exemplary video decoder according to one embodiment of the present disclosure. FIG. 19 illustrates a flowchart schematically describing a process according to one embodiment of the present disclosure. FIG. 20 illustrates a flowchart schematically describing a process according to one embodiment of the present disclosure. FIG. 21 is a schematic diagram of a computer system according to one embodiment. Specific details for implementing the invention

[0033] FIG. 3 shows a simplified block diagram of a communication system (300) according to one embodiment of the present disclosure. The communication system (300) includes a plurality of terminal devices capable of communicating with each other, for example, through a network (350). For example, the communication system (300) includes a first pair of terminal devices (310, 320) interconnected through the network (350). In the example of FIG. 3, the first pair of terminal devices (310, 320) performs unidirectional transmission of data. For example, the terminal device (310) may encode video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to another terminal device (320) through the network (350). The encoded video data may be transmitted in the form of one or more coded video bit streams. The terminal device (320) receives coded video data from the network (350), decodes the coded video data to restore a video picture, and can display the video picture according to the restored video data. Unidirectional data transmission may be common in media serving applications, etc.

[0034] In another example, the communication system (300) includes a second pair of terminal devices (330, 340) that perform bidirectional transmission of coded video data, which may occur, for example, during a video conference. For bidirectional transmission of data, in one example, each terminal device of the pair of terminal devices (330, 340) may code video data (e.g., a stream of video pictures captured by the terminal device) for transmission to another terminal device of the pair of terminal devices (330, 340) via a network (350). Each terminal device of the pair of terminal devices (330, 340) may also receive coded video data transmitted by another terminal device of the pair of terminal devices (330, 340), decode the coded video data to restore the video, and display the video picture on an accessible display device according to the restored video data.

[0035] In the example of FIG. 3, the terminal devices (310, 320, 330, 340) may be exemplified as servers, personal computers, and smartphones, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure find applications using laptop computers, tablet computers, media panels, and / or dedicated video conferencing equipment. The network (350) represents any number of networks that transmit coded video data between terminal devices (310, 320, 330, 340), including, for example, wired and / or wireless communication networks. The communication network (350) may exchange data over circuit-switched and / or packet-switched channels. Representative networks include communication networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of the network (350) may not be important to the operation of the present disclosure unless described below.

[0036] FIG. 4 illustrates the placement of a video encoder and a video decoder in a streaming environment as an example of an application for the disclosed subject. The disclosed subject may be equally applicable to other video-enabled applications, such as storing compressed video on digital media including, for example, video conferencing, digital TV, CD, DVD, memory stick, etc.

[0037] A streaming system may include a video source (401) and a capture subsystem (413) which may include, for example, a digital camera, to generate, for example, a stream (402) of uncompressed video pictures. In one example, the stream (402) of video pictures includes samples captured by a digital camera. The stream (402) of video pictures, indicated by a bold line to emphasize a high data volume compared to encoded video data (404) (or coded video bitstream), may be processed by an electronic device (420) comprising a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof that can enable or implement aspects of the disclosed subject matter as described in more detail below. The encoded video data (404) (or encoded video bitstream (404)) is indicated by a thin line to emphasize a lower data volume compared to the stream (402) of the video picture and can be stored in a streaming server (405) for later use. One or more streaming client subsystems, such as the client subsystems (406, 408) in FIG. 4, can access the streaming server (405) to retrieve copies (407, 409) of the encoded video data (404). The client subsystem (406) may include a video decoder (410) in, for example, an electronic device (430). The video decoder (410) decodes the incoming copy (407) of the encoded video data and generates an outgoing stream of a video picture (411) that can be rendered on a display (412) (e.g., a display screen) or another rendering device (not shown). In some streaming systems, encoded video data (404, 407, 409) (e.g., video bitstream) may be encoded according to a specific video coding / compression standard.Examples of such standards include ITU-T Recommendation H.265. For instance, a video coding standard under development is informally known as VVC (Versatile Video Coding). The disclosed topics may be used in the context of VVC.

[0038] Note that the electronic device (420, 430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).

[0039] FIG. 5 illustrates a block diagram of a video decoder (510) according to one embodiment of the present disclosure. The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used instead of the video decoder (510) in the example of FIG. 3.

[0040] A receiver (531) may receive one or more coded video sequences to be decoded by a video decoder (510); in the same or different embodiments, at one coded video sequence at a time, the decoding of each coded video sequence is independent of other coded video sequences. The coded video sequences may be received from a channel (501) which may be a hardware / software link to a storage device that stores the encoded video data. The receiver (531) may receive the encoded video data along with other data, e.g., encoded audio data and / or auxiliary data streams, which may be forwarded to each user entity (not shown). The receiver (531) may separate the coded video sequences from the other data. To prevent network jitter, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter "parser (520)"). In certain applications, the buffer memory (515) is part of the video decoder (510). In other cases, it may be outside the video decoder (510) (not shown). In yet another case, the buffer memory (not shown) may be outside the video decoder (510), for example, to prevent network jitter, and there may be an additional buffer memory (515) inside the video decoder (510), for example, to handle playout timing. When the receiver (531) is receiving data from a storage / forwarding device or an isosynchronous network that has sufficient bandwidth and controllability, the buffer memory (515) may not be needed or may be small.For use in a best-effort packet network such as the Internet, the buffer memory (515) may be required, may be relatively large, may have an adaptive size, and may be implemented at least partially in a similar element (not shown) outside of the operating system or video decoder (510).

[0041] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from a coded video sequence. Categories of these symbols include information used to manage the operation of the video decoder (510) and potential information for controlling a rendering device (512) (e.g., a display picture) that is not an integrated part of the electronic device (530) but can be connected to the electronic device (530) as illustrated in FIG. 5. The control information for the rendering device(s) may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (520) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow video coding techniques or standards and may follow various principles including variable-length coding, Huffman coding, and arithmetic coding with or without context sensitivity. The parser (520) may extract a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to a group. The subgroups may include a Group of Picture (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), etc. The parser (520) may also extract transform coefficients, quantizer parameter values, motion vectors, etc. from the coded video sequence information.

[0042] The parser (520) can generate symbols (521) by performing entropy decoding / parsing operations on a video sequence received from the buffer memory (515).

[0043] The reconstruction of the symbol (521) may include several different units depending on the type of the coded video picture or a part thereof (e.g., inter- and intra-pictures, inter- and intra-blocks) and other factors. Which units are involved and how they are involved may be controlled by subgroup control information parsed from the coded video sequence by the parser (520). The flow of such subgroup control information between the parser (520) and the number of units below it is not shown for clarity.

[0044] Beyond the previously mentioned functional blocks, the video decoder (510) can be conceptually subdivided into a number of functional units as described below. In actual implementations operating under commercial constraints, many of these units may interact closely with one another and be integrated with one another, at least partially. However, to explain the subject of disclosure, it is appropriate to conceptually subdivide it into the functional units below.

[0045] The first unit is a scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives quantized transform coefficients as well as control information, including transforms for use as symbol(s) (521) from the parser (520), block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (551) can output a block containing sample values ​​that can be input to an aggregator (555).

[0046] In some cases, the output sample of the scaler / inverse conversion unit (551) may be related to an intra-coded block, that is, a block that does not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed part of the current picture. Such prediction information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) uses already reconstructed surrounding information drawn from the current picture buffer (558) to generate a block of the same size and shape as the block being reconstructed. The current picture buffer (558) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. The aggregator (555) adds, in some cases, the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse conversion unit (551) on a sample basis.

[0047] In other cases, the output samples of the scaler / inverse unit (551) may belong to an intercoded and potentially motion-compensated block. In this case, the motion-compensated prediction unit (553) may access the reference picture memory (557) to retrieve samples used for prediction. After motion-compensating the samples retrieved according to the symbols (521) belonging to the block, these samples may be added to the output of the scaler / inverse unit (551) (in this case, referred to as residual samples or residual signals) by the aggregator (555) to generate output sample information. The address in the reference picture memory (557) from which the motion-compensated prediction unit (553) retrieves the prediction samples may be controlled by the MV and may be used by the motion-compensated unit (553) in the form of a symbol (521) that may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​drawn from reference picture memory (557) when the accurate motion vector of a subsample is used, a motion vector prediction mechanism, etc.

[0048] The output sample of the aggregator (555) may be subject to various loop filtering techniques in the loop filter unit (556). The video compression technique may include an in-loop filter technique that is controlled by parameters contained in the coded video sequence (also referred to as the coded video bit stream) and made available to the loop filter unit (556) as a symbol (521) from the parser (520), but may also respond to meta-information obtained while decoding a previous (in decoding order) part of the coded picture or coded video sequence, as well as respond to previously reconstructed and loop-filtered sample values.

[0049] The output of the loop filter unit (556) may be a sample stream that can be output to the render device (512) as well as stored in the reference picture memory (557) for use in future inter-picture prediction.

[0050] Once a specific coded picture has been completely reconstructed, it can later be used as a reference picture for prediction. For example, when a coded picture corresponding to the current picture is completely reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and the new current picture buffer can be reallocated before starting the reconstruction of the next coded picture.

[0051] A video decoder (510) can perform decoding operations according to a video compression technology predetermined in a standard, such as ITU-T Rec. H.265. The coded video sequence may follow the syntax specified by the video compression technology or standard used in that it complies with both the syntax of the video compression technology or standard and the profile documented in the video compression technology. In particular, the profile may select a specific tool as the only tool available in that profile among all tools available in the video compression technology or standard. Additionally, for compliance to be required, the complexity of the coded video sequence may be within the range defined by the level of the video compression technology or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further limited by HRD specifications and metadata for managing the virtual reference decoder (HRD) buffer signaled in the coded video sequence.

[0052] In one embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence(s). The additional data may be used by the video decoder (510) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, a temporal layer, a spatial layer or an SNR enhancement layer, a redundant slice, a redundant picture, a forward error correction code, etc.

[0053] FIG. 6 illustrates a block diagram of a video encoder (603) according to one embodiment of the present disclosure. The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmitting circuit). The video encoder (603) may be used instead of the video encoder (403) in the example of FIG. 4.

[0054] The video encoder (603) can receive video samples from a video source (601) (not part of the electronic device (620) in the example of FIG. 6) capable of capturing video image(s) to be encoded by the video encoder (603). In another example, the video source (601) is part of the electronic device (620).

[0055] A video source (601) may provide a source video sequence to be coded by a video encoder (603) in the form of a digital video sample stream, which may have any appropriate bit depth (e.g., 8-bit, 10-bit, 12-bit,…), any color space (e.g., BT.601 Y CrCB, RGB,…), and any appropriate sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (601) may be a storage device that stores pre-prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that impart motion when viewed sequentially. The picture itself may consist of a spatial array of pixels, and each pixel may contain one or more samples depending on the sampling structure, color space, etc. being used. A person skilled in the art will easily understand the relationship between pixels and samples. The following description focuses on the sample.

[0056] According to one embodiment, the video encoder (603) can code and compress a picture of a source video sequence into a coded video sequence (643) in real time or under any other time constraint required by the application. Enforcing an appropriate coding rate is one of the functions of the controller (650). In some embodiments, the controller (650) controls other function units and is functionally coupled to other function units as described below. For clarity, the coupling is not shown. Parameters set by the controller (650) may include rate control-related parameters (picture skip, quantizer, lambda value of rate distortion optimization technique,…), picture size, picture group (GOP) layout, maximum motion vector search range, etc. The controller (650) may be configured to have other appropriate functions belonging to the video encoder (603) optimized for a specific system design.

[0057] In some embodiments, the video encoder (603) is configured to operate in a coding loop. For the sake of oversimplification, in one example, the coding loop may include a source coder (630) (e.g., responsible for generating symbol and reference picture(s), such as a symbol stream, based on the input picture to be coded), and a (local) decoder (633) built into the video encoder (603). The decoder (633) reconstructs the symbols to generate sample data in the same way that the (remote) decoder also generates (since any compression between the symbols and the coded video bitstream is lossless in the video compression techniques considered in the disclosed subject). The reconstructed sample stream (sample data) is input into the reference picture memory (634). Since the decoding of the symbol stream leads to a bit-exact result regardless of the decoder location (local or remote), the contents of the reference picture memory (634) are also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "recognizes" as reference picture samples exactly the same sample values ​​that the decoder "sees" when using prediction during decoding. The basic principle of reference picture synchronicity (and, for example, resulting drift, where simultaneity cannot be maintained due to channel error) is also used in some related techniques.

[0058] The operation of the “local” decoder (633) may be the same as the operation of the “remote” decoder, such as the video decoder (510), which has already been described in detail in relation to FIG. 5. Briefly referring also to FIG. 5, since symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder (645) and parser (520) may be lossless, the entropy decoding part of the video decoder (510) may not be fully implemented in the local decoder (633), including the buffer memory (515) and parser (520).

[0059] In one embodiment, the decoder technology, excluding the parsing / entropy decoding present in the decoder, exists in the corresponding encoder in the same or substantially the same functional form. Accordingly, the subject matter of the disclosure focuses on the decoder operation. The description of the encoder technology may be omitted as it is the opposite of the comprehensively described decoder technology. A more detailed description in specific areas is provided below.

[0060] In operation, in some examples, the source coder (630) may perform motion-compensated predictive coding to predictively code an input picture by referencing one or more previously coded pictures from a video sequence designated as a "reference picture." In this way, the coding engine (632) codes the difference between a pixel block of the input picture and a pixel block of the reference picture(s) that can be selected as predictive reference(s) for the input picture.

[0061] The local video decoder (633) can decode the coded video data of a picture that can be designated as a reference picture based on symbols generated by the source coder (630). The operation of the coding engine (632) may advantageously be a lossy process. When the coded video data can be decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence may generally be a replica of the source video sequence with some errors. The local video decoder (633) can replicate the decoding process that can be performed by the video decoder on the reference picture and allow the reconstructed reference picture to be stored in the reference picture cache (634). In this way, the video encoder (603) can locally store a copy of the reconstructed reference picture having common content as the reconstructed reference picture to be acquired by the far-end video decoder (no transmission errors).

[0062] The predictor (635) can perform a predictive search for the coding engine (632). That is, for a new picture to be coded, the predictor (635) can search the reference picture memory (634) for sample data (which is a candidate reference pixel block) that can serve as a suitable predictive reference for the new picture, or for specific metadata such as a reference picture motion vector, block shape, etc. The predictor (635) can operate on a sample block-by-pixel block basis to find a suitable predictive reference. In some cases, as determined by the search result obtained by the predictor (635), the input picture may have a predictive reference drawn from a number of reference pictures stored in the reference picture memory (634).

[0063] The controller (650) can manage the coding operation of the source coder (630), including the settings of parameters and subgroup parameters used to encode video data, for example.

[0064] The output of all the aforementioned function units may be subject to entropy coding in the entropy coder (645). The entropy coder (645) converts the symbols generated by the various function units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable-length coding, and arithmetic coding.

[0065] The transmitter (640) may buffer the coded video sequence(s) generated by the entropy coder (645) and prepare for transmission through a communication channel (660), which may be a hardware / software link to a storage device for storing the encoded video data. The transmitter (640) may merge the coded video data from the video coder (603) with other data to be transmitted, e.g., coded audio data and / or an auxiliary data stream (source not shown).

[0066] The controller (650) can manage the operation of the video encoder (603). During coding, the controller (650) can assign a specific coded picture type to each coded picture, which can affect the coding technique that can be applied to each picture. For example, a picture can often be designated as one of the following picture types:

[0067] An Intra Picture (I Picture) may be one that can be coded and decoded without using any other picture within the sequence as a prediction source. Some video codecs allow different types of Intra Pictures, including, for example, Independent Decoder Refresh Picture ("IDR"). Those skilled in the art are aware of these variations of I Pictures and their respective applications and characteristics.

[0068] The predictive picture (P picture) may be coded and decoded using intra-prediction or inter-prediction, which uses at most one motion vector and reference index to predict the sample values ​​of each block.

[0069] A bidirectionally predictive picture (B-picture) may be coded and decoded using intra-prediction or inter-prediction, which utilize up to two motion vectors and reference indices to predict sample values ​​for each block. Similarly, a multiple-predictive picture may use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0070] A source picture is typically subdivided spatially into multiple sample blocks (e.g., 4×4, 8×8, 4×8, or 16×16 sample blocks) and can be coded block by block. Blocks can be predictively coded by referencing other (already coded) blocks as determined by the coding assignment applied to each picture in the block. For example, blocks in picture I can be non-predictively coded or predictively coded by referencing already coded blocks (spatial prediction or intra prediction) of the same picture. Pixel blocks in picture P can be predictively coded via spatial prediction or temporal prediction by referencing one previously coded reference picture. Blocks in picture B can be predictively coded via spatial prediction or temporal prediction by referencing one or two previously coded reference pictures.

[0071] The video encoder (603) may perform coding operations according to a predetermined video coding technique or standard, such as ITU-T Rec. H.265. In such operations, the video encoder (603) may perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Thus, the coded video data may follow the syntax specified by the video coding technique or standard used.

[0072] In one embodiment, the transmitter (640) may transmit additional data along with the encoded video. The source coder (630) may include this data as part of the encoded video sequence. The additional data may include other forms of redundant data such as time / space / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0073] Video can be captured as a temporal sequence of multiple source pictures (video pictures). Intra-picture prediction (often abbreviated as intra prediction) utilizes spatial correlations within a given picture, while inter-picture prediction utilizes (temporal or other) correlations between pictures. In one example, a specific picture currently being encoded / decoded, referred to as the current picture, is partitioned into blocks. When a block within the current picture is similar to a reference block in a reference picture that was previously encoded and is still buffered in the video, the block within the current picture can be encoded as a vector referred to as a motion vector. The motion vector points to the reference block within the reference picture and, if multiple reference pictures are in use, may have three dimensions that identify the reference picture.

[0074] In some embodiments, a bi-prediction technique may be used for inter-picture prediction. According to the bi-prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture, in which the decoding order of the current picture in the video is both preceding (however, in display order, they may be past and future, respectively). A block of the current picture may be coded by a first motion vector pointing to a first reference block of the first reference picture and a second motion vector pointing to a second reference block of the second reference picture. A block may be predicted by a combination of the first reference block and the second reference block.

[0075] In addition, coding efficiency can be improved by using merge mode technology in inter-picture prediction.

[0076] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed on a block basis. For example, according to the HEVC standard, pictures within a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and the CTUs within the pictures have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Typically, a CTU includes three coding tree blocks (CTBs), one luminance CTB and two chroma CTBs. Each CTU may be a tetrad tree that is repeatedly divided into one or more coding units (CUs). For example, a 64×64 pixel CTU may be divided into one 64×64 pixel CU, four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine a prediction type for the CU, such as an inter-prediction type or an intra-prediction type. The CU is divided into one or more prediction units (PUs) based on temporal and / or spatial predictability. Typically, each PU includes a luminal prediction block (PB) and two chroma PBs. In one embodiment, the prediction operation in coding (encoding / decoding) is performed on a unit of prediction blocks. If a luminal prediction block is used as an example of a prediction block, the prediction block includes a matrix of values ​​(e.g., luminal values) for pixels, such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0077] FIG. 7 illustrates a video encoder (703) according to another embodiment of the present disclosure. The video encoder (703) is configured to receive a processing block (e.g., a prediction block) of sample values ​​within a current video picture in a video picture sequence and to encode the processing block into a coded picture that is part of a coded video sequence. In one example, the video encoder (703) is used instead of the video encoder (403) in the example of FIG. 4.

[0078] In the HEVC example, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as a prediction block of 8×8 samples. The video encoder (703) determines whether the processing block is best coded using an intra-mode, inter-mode, or positive prediction mode, for example, using rate distortion optimization. When the processing block is coded in intra-mode, the video encoder (703) may use an intra-prediction technique to encode the processing block into a coded picture; when the processing block is coded in inter-mode or positive prediction mode, the video encoder (703) may use an inter-prediction technique or a positive prediction technique, respectively, to encode the processing block into a coded picture. In a particular video coding technique, the merge mode may be an inter-picture prediction submode in which the motion vector is derived from one or more motion vector predictors without the benefit of the coded motion vector component outside the predictor. In a particular other video coding technique, there may be a motion vector component applicable to the subject block. In one example, the video encoder (703) includes other components such as a mode determination module (not shown) for determining the mode of the processing block.

[0079] In the example of FIG. 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), and a general controller (721), and includes an entropy encoder (725) combined together as shown in FIG. 7.

[0080] The inter-encoder (730) receives a sample of the current block (e.g., processing block), compares the block with one or more reference blocks within a reference picture (e.g., blocks within a previous picture and a subsequent picture), generates inter-prediction information (e.g., description of redundant information according to inter-encoding techniques, motion vectors, merge mode information), and calculates an inter-prediction result (e.g., prediction block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture that is decoded based on encoded video information.

[0081] The intra-encoder (722) receives a sample of the current block (e.g., processing block), and in some cases, compares the block with a block already coded in the same picture, generates quantized coefficients after conversion, and in some cases also generates intra-prediction information (e.g., intra-prediction direction information according to one or more intra-encoding techniques). In one example, the intra-encoder (722) also calculates an intra-prediction result (e.g., prediction block) based on the intra-prediction information and reference block of the same picture.

[0082] A general controller (721) is configured to determine general control data and to control other components of the video encoder (703) based on the general control data. In one example, the general controller (721) determines the mode of the block and provides a control signal to the switch (726) according to the mode. For example, if the mode is an intra mode, the general controller (721) controls the switch (726) to select an intra mode result for use in the residual calculator (723) and controls the entropy encoder (725) to select intra prediction information and include the intra prediction information in the bit stream; if the mode is an inter mode, the general controller (721) controls the switch (726) to select an inter prediction result for use in the residual calculator (723) and controls the entropy encoder (725) to select inter prediction information and include the inter prediction information in the bit stream.

[0083] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra-encoder (722) or the inter-encoder (730). The residual encoder (724) is configured to operate based on the residual data to encode the residual data and generate transformation coefficients. In one example, the residual encoder (724) is configured to convert the residual data from the spatial domain to the frequency domain and generate transformation coefficients. The transformation coefficients then undergo quantization processing to obtain quantized transformation coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transformation and generate decoded residual data. The decoded residual data can be appropriately used by the intra-encoder (722) and the inter-encoder (730). For example, the inter encoder (730) can generate a decoded block based on decoded residual data and inter prediction information, and the intra encoder (722) can generate a decoded block based on decoded residual data and intra prediction information. The decoded block is appropriately processed to generate a decoded picture, and the decoded picture can be buffered in a memory circuit (not shown) and, in some examples, used as a reference picture.

[0084] The entropy encoder (725) is configured to format the bit stream to include the encoded block. The entropy encoder (725) is configured to include various information according to an appropriate standard, such as the HEVC standard. In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other appropriate information in the bit stream. According to the disclosed subject invention, it is noted that when coding a block in an inter mode or a merged submode of both prediction modes, there is no residual information.

[0085] FIG. 8 illustrates a drawing of a video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive a coded picture that is part of a coded video sequence and to decode the coded picture to generate a reconstructed picture. In one example, the video decoder (810) is used instead of the video decoder (410) in the example of FIG. 4.

[0086] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872) combined together as shown in FIG. 8.

[0087] The entropy decoder (871) may be configured to reconstruct, from the coded picture, specific symbols representing the syntax elements constituting the coded picture. These symbols may include, for example, the mode in which the block is coded (e.g., intra mode, inter mode, positive prediction mode, merge submode, or the latter two in other submodes), prediction information (e.g., intra prediction information or inter prediction information) capable of identifying specific samples or metadata used for prediction by the intra decoder (872) or the inter decoder (880), respectively, and residual information in the form of, for example, quantized transformation coefficients. In one example, if the prediction mode is an inter mode or positive prediction mode, the inter prediction information is provided to the inter decoder (880); if the prediction type is an intra prediction type, the intra prediction information is provided to the intra decoder (872). The residual information may be dequantized and provided to the residual decoder (873).

[0088] The inter decoder (880) is configured to receive inter prediction information and generate an inter prediction result based on the inter prediction information.

[0089] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.

[0090] The residual decoder (873) is configured to perform inverse quantization to extract inverse quantized transform coefficients and process the inverse quantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require specific control information (including quantizer parameters (QP)), and that information may be provided by the entropy decoder (871) (the data path is not shown as it may contain only a small amount of control information).

[0091] The reconstruction module (874) is configured to combine the residuals and prediction results output by the residual decoder (873) in the spatial domain (which may be output by the inter-prediction module or the intra-prediction module) to form a reconstructed block, the reconstructed block may be part of a reconstructed picture, and the reconstructed picture may also be part of a reconstructed video. Note that other appropriate actions, such as deblocking, may be performed to improve visual quality.

[0092] It should be noted that the video encoder (403, 603, 703) and video decoder (410, 510, 810) may be implemented using any suitable technology. In one embodiment, the video encoder (403, 603, 703) and video decoder (410, 510, 810) may be implemented using one or more integrated circuits. In another embodiment, the video encoder (403, 603, 603) and video decoder (410, 510, 810) may be implemented using one or more processors that execute software instructions.

[0093] The present disclosure describes video coding techniques related to neural image compression techniques and / or neural video compression techniques, such as artificial intelligence (AI)-based neural image compression (NIC). Aspects of the present disclosure include content-adaptive online training in NICs, such as NIC methods for end-to-end (E2E) optimized image coding frameworks based on neural networks. Neural networks (NN) may include artificial neural networks (ANN), such as deep neural networks (DNN), convolutional neural networks (CNN), etc.

[0094] In one embodiment, the related hybrid video codec is difficult to optimize as a whole. For example, in a hybrid video codec, improvement of a single module (e.g., an encoder) may not lead to coding gain in overall performance. In an NN-based video coding framework, different modules can be optimized together from input to output to improve an end objective (e.g., rate-distortion performance, such as the rate-distortion loss L described in this disclosure) by performing a learning process or a training process (e.g., a machine learning process), thus becoming an E2E optimized NIC.

[0095] An exemplary NIC framework or system can be described as follows. The NIC framework uses an input image x as input to a neural network encoder (e.g., an encoder based on a neural network such as a DNN) to obtain, for example, a compressed representation (e.g., a compact representation) that can be compacted for storage and transmission purposes. It can calculate. A neural network decoder (e.g., a decoder based on a neural network such as a DNN) is a compressed representation Output image (also called reconstructed image) using as input It can be reconstructed. In various embodiments, the input image x and the reconstructed image is in the spatial domain and is a compressed representation It is in a domain different from the spatial domain. In some examples, the compressed representation It is quantized and entropy-coded.

[0096] In some examples, the NIC framework can use a variational autoencoder (VAE) architecture. In a VAE architecture, the neural network encoder can directly use the entire input image x as input to the neural network encoder. The entire input image x is a compressed representation To compute, it can pass through a set of neural network layers acting as black boxes. Compressed representation is the output of the neural network encoder. The neural network decoder is the full compressed representation It can take as input. Compressed expression is a reconstructed image To calculate, it can pass through a different set of neural network layers acting as different black boxes. Rate-distortion (RD) loss L ( ) is a reconstructed image distortion loss A compact expression with a trade-off hyperparameter λ and ) It can be optimized to achieve a trade-off between the bit consumption R of.

[0097] L ( ) = ) + R ( ) Equation 1

[0098] Neural networks (e.g., ANNs) can learn to perform tasks from examples without task-specific programming. An ANN may consist of connected nodes or artificial neurons. Connections between nodes can transmit signals from a first node to a second node (e.g., a receiving node), and the signals may be modified by weights, which can be indicated by weighting coefficients for the connections. The receiving node can generate an output signal by processing signals from the node(s) transmitting the signal(s) to the receiving node (i.e., input signal(s) for the receiving node) and then applying a function to the input signals. The function may be a linear function. In one example, the output signal is a weighted summation of the input signals. In another example, the output signal is further modified by a bias, which can be indicated by a bias term, and thus the output signal is the sum of the bias and the weighted summation of the input signals. The function may include, for example, a weighted sum or a non-linear operation on the weighted sum of the input signal(s) and the bias. The output signal may be transmitted to node(s) (downstream node(s)) connected to the receiving node. The ANN may be represented or configured by parameters (e.g., weights of the connection and / or bias). The weights and / or bias may be obtained by training the ANN, for example, where the weights and / or bias can be iteratively adjusted. A trained ANN composed of determined weights and / or determined biases may be used to perform a task.

[0099] Nodes in an ANN can be organized into an appropriate architecture. In various embodiments, nodes in an ANN consist of layers including an input layer that receives input signal(s) for the ANN and an output layer that outputs output signal(s) from the ANN. In one embodiment, the ANN further includes layer(s), such as hidden layer(s), between the input layer and the output layer. Different types of transformations can be performed for each input of different layers. Signals can travel from the input layer to the output layer.

[0100] An ANN having multiple layers between an input layer and an output layer can be called a DNN. In one embodiment, the DNN is a feedforward network in which data flows from the input layer to the output layer without looping back. In one example, the DNN is a fully connected network in which each node of one layer is connected to all nodes of the next layer. In one embodiment, the DNN is a recurrent neural network (RNN) in which data can flow in any direction. In one embodiment, the DNN is a CNN.

[0101] A CNN may include an input layer, an output layer, and hidden layer(s) between the input layer and the output layer. The hidden layer(s) may include convolutional layer(s) that perform convolutions, such as two-dimensional (2D) convolution (e.g., used in an encoder). In one embodiment, the 2D convolution performed in the convolutional layer is between a convolutional kernel (also called a filter or channel, such as a 5 x 5 matrix) and an input signal to the convolutional layer (e.g., a 2D matrix, such as a 2D image, or a 256 x 256 matrix). In various examples, the dimension of the convolutional kernel (e.g., 5 x 5) is smaller than the dimension of the input signal (e.g., 256 x 256). Therefore, the part covered by the convolution kernel in the input signal (e.g., a 256 x 256 matrix) (e.g., a 5 x 5 area) is smaller than the area of ​​the input signal (e.g., a 256 x 256 area), and thus can be called the receptive field at each node of the next layer.

[0102] During convolution, the inner product of the convolution kernel and the corresponding recept field in the input signal is calculated. Therefore, since each element of the convolution kernel is a weight applied to the corresponding sample in the recept field, the convolution kernel contains weights. For example, a convolution kernel represented by a 5 x 5 matrix has 25 weights. In some examples, a bias is applied to the output signal of the convolution layer, and the output signal is based on the sum of the inner product and the bias.

[0103] Since a convolutional kernel can shift along the input signal (e.g., a 2D matrix) by a size called the stride, the convolutional operation generates a feature map or activation map (e.g., another 2D matrix), which contributes to the input of the next layer in a CNN. For example, the input signal is a 2D image with 256 x 256 samples, and the stride is 2 samples (e.g., a stride of 2). When the stride is 2, the convolutional kernel shifts by 2 samples in the X direction (e.g., horizontal direction) and / or the Y direction (e.g., vertical direction).

[0104] Multiple convolutional kernels can be applied to the input signal in the same convolutional layer to generate multiple feature maps, each of which can represent a specific feature of the input signal. Generally, a convolutional layer with N channels (i.e., N convolutional kernels), each convolutional kernel with M x M samples, and a stride S can be designated as Conv: MxM cN sS. For example, each convolutional layer with 192 channels, a convolutional kernel with 5 x 5 samples, and a stride of 2 is designated as Conv: 5x5 c192 s2. The hidden layer(s) may include inverse convolutional layer(s) that perform inverse convolution, such as 2D inverse convolution (e.g., used in a decoder). Inverse convolution is the inverse of convolution. An inverse convolutional layer with 192 channels, each inverse convolutional kernel with 5 x 5 samples, and a stride of 2 is specified as DeConv: 5x5 c192 s2.

[0105] In various embodiments, CNNs have the following advantages. The number of trainable parameters (i.e., parameters to be trained) in a CNN can be much smaller than the number of trainable parameters in a DNN such as a feedforward DNN. In a CNN, a relatively large number of nodes can share the same filter (e.g., the same weight) and the same bias (if bias is used), and thus a single bias and single weight vector can be used across all recept fields sharing the same filter, which can reduce the memory footprint. For example, for an input signal with 100 x 100 samples, a convolutional layer with a convolutional kernel of 5 x 5 samples has 25 trainable parameters (e.g., weights). If bias is used, one channel uses 26 trainable parameters (e.g., 25 weights and 1 bias). If the convolutional layer has N channels, the total number of trainable parameters is 26 x N. On the other hand, for a fully connected layer in a DNN, 100 x 100 (i.e., 10,000) weights are used for each node in the next layer. If there are L nodes in the next layer, the total number of learnable parameters is 10,000 x L.

[0106] CNNs may include one or more other layer(s), such as pooling layer(s), fully connected layer(s) that can connect every node in one layer to every node in another layer, and normalization layer(s). Layers in a CNN can be arranged in an appropriate order and any appropriate architecture (e.g., feedforward architecture, recurrent architecture). For example, other layer(s), such as pooling layer(s), fully connected layer(s), and normalization layer(s), follow the convolutional layer.

[0107] Pooling layers can be used to reduce the dimensionality of data by combining the outputs from multiple nodes of one layer into a single node of the next layer. Pooling operations for a pooling layer that takes a feature map as input are described below. This description can be appropriately adapted to other input signals. The feature map can be divided into sub-regions (e.g., rectangular sub-regions), and the features in each sub-region can be independently downsampled (or pooled) into a single value by taking, for example, the average value in average pooling or the maximum value in maximum pooling.

[0108] The pooling layer can perform pooling such as local pooling, global pooling, max pooling, and mean pooling. Pooling is a form of non-linear downsampling. Local pooling combines a small number of nodes in a feature map (e.g., a local cluster of nodes, such as 2 x 2 nodes). Global pooling can combine all nodes in a feature map, for example.

[0109] Pooling layers can reduce the size of representations, thereby reducing the number of parameters, memory footprint, and computational load in CNNs. In one example, a pooling layer is inserted between consecutive convolutional layers in a CNN. In another example, an activation function, such as a rectified linear unit (ReLU) layer, follows the pooling layer. For example, the pooling layer is omitted between consecutive convolutional layers in a CNN.

[0110] The normalization layer can be ReLU, leaky ReLU, generalized divisive normalization (GDN), or inverse GDN. ReLU can remove negative values ​​from input signals, such as feature maps, by applying a non-saturating activation function that sets negative values ​​to zero. Leaky ReLU can have a small gradient (e.g., 0.01) for negative values ​​instead of a flat gradient (e.g., 0). Therefore, if the value x is greater than 0, the output of leaky ReLU is x. Otherwise, the output of leaky ReLU is x multiplied by the small gradient (e.g., 0.01). In one example, the gradient is determined before training and is therefore not learned during training.

[0111] FIG. 9 illustrates an exemplary NIC framework (900) (e.g., NIC system) according to one embodiment of the present disclosure. The NIC framework (900) may be based on a neural network such as a DNN and / or a CNN. The NIC framework (900) may be used to compress (e.g., encode) an image and to decompress (e.g., decode or reconstruct) the compressed image (e.g., encoded image). The NIC framework (900) may include two sub-neural networks, namely a first sub-NN (951) and a second sub-NN (952) implemented using a neural network.

[0112] The first sub-NN (951) may be similar to an autoencoder and is a compressed image of the input image x Generates and compresses the image The image reconstructed by extracting it It can be trained to acquire. The first sub-NN (951) may include a plurality of components (or modules), such as a main encoder neural network (or main encoder network) (911), a quantizer (912), an entropy encoder (913), an entropy decoder (914), and a main decoder neural network (or main encoder network) (915). Referring to FIG. 9, the main encoder network (911) can generate a latent value or latent representation y from an input image x (e.g., an image to be compressed or encoded). For example, the main encoder network (911) is implemented using a CNN. The relationship between the latent representation y and the input image x can be described using Equation 2.

[0113] Equation (2)

[0114] Here, the parameter represents parameters such as weights and biases used in the convolution kernels in the main encoder network (911) (if biases are used in the main encoder network (911)).

[0115] The potential expression y is a potential value quantized using a quantizer (912). It can be quantized to generate. Quantized potential is, for example, a compressed representation of an input image compressed image (e.g., encoded image) To generate, it can be compressed using lossless compression by an entropy encoder (913). The entropy encoder (913) can use entropy coding techniques such as Huffman coding, arithmetic coding, etc. In one example, the entropy encoder (913) uses arithmetic encoding and is an arithmetic encoder. In one example, the encoded image (931) is transmitted as a coded bitstream.

[0116] The encoded image (931) can be decompressed (e.g., entropy decoding) by an entropy decoder (914) to produce an output. The entropy decoder (914) may use entropy coding techniques such as Huffman coding, arithmetic coding, etc., corresponding to the entropy encoding techniques used in the entropy encoder (913). In one example, the entropy decoder (914) uses arithmetic decoding and is an arithmetic decoder. For example, lossless compression is used in the entropy encoder (913) and lossless decompression is used in the entropy decoder (914), and noise such as that caused by the transmission of the encoded image (931) can be omitted, so that the output from the entropy decoder (914) is a quantized potential value am.

[0117] The main decoder network (915) is a quantized potential An image reconstructed by decoding It can generate. For example, the main decoder network (915) is implemented using a CNN. Reconstructed image (i.e., the output of the main decoder network (915)) and the quantized potential The relationship between (i.e., the input of the main decoder network (915)) can be described using Equation 3.

[0118] Equation 3

[0119] Here, the parameter represents parameters such as weights and biases (if biases are used in the main decoder network (915)) used in the convolutional kernels in the main decoder network (915). Thus, the first sub-NN (951) compresses (e.g., encodes) the input image x to obtain an encoded image (931) and decompresses (e.g., decodes) the encoded image to reconstruct an image. You can obtain the reconstructed image It may differ from the input image x due to the quantization loss introduced by the quantizer (912).

[0120] The second sub-NN (952) is a quantized potential used for entropy coding An entropy model (e.g., a prior probabilistic model) can be learned for it. Thus, the entropy model may be a conditional entropy model, such as a Gaussian scale model (GSM) or a Gaussian mixture model (GMM) that depends on the input image x. The second sub-NN (952) may include a context model NN (916), an entropy parameter NN (917), a hyper-encoder (921), a quantizer (922), an entropy encoder (923), an entropy decoder (924), and a hyper-decoder (925). The entropy model used in the context model NN (916) is a potential value (e.g., a quantized potential value It may be an autoregressive model for ). In one example, a hyperencoder (921), a quantizer (922), an entropy encoder (923), an entropy decoder (924), and a hyperdecoder (925) form a hyperneural network (e.g., a hyperprior NN). The hyperneural network may represent information useful for correcting context-based predictions. Data from the context model NN (916) and the hyperneural network may be combined by an entropy parameter NN (917). The entropy parameter NN (917) may generate parameters such as mean and scale parameters for an entropy model, such as a conditional Gaussian entropy model (e.g., GMM).

[0121] Referring to FIG. 9, at the encoder side, the quantized latent from the quantizer (912) is supplied to the context model NN (916). On the decoder side, the quantized potential from the entropy decoder (914) is fed to the context model NN (916). The context model NN (916) can be implemented using a neural network such as a CNN. The context model NN (916) is the quantized potential available to the context model NN (916). In context Output based on Can create. Context may include a previously quantized potential on the encoder side or a previously entropy-decoded quantized potential on the decoder side. Output of the context model NN (916). and input (e.g., The relationship between ) can be described using Equation 4.

[0122] Equation 4

[0123] Here, the parameter represents parameters such as weights and biases used in the convolutional kernel of the context model NN (916) (where biases are used in the context model NN (916)).

[0124] Output from context model NN(916) and output from hyperdecoder (925) is supplied to the entropy parameter NN (917) and output Generates. The entropy parameter NN (917) can be implemented using a neural network such as a CNN. Output of the entropy parameter NN (917). and input (e.g., eg, and The relationship between ) can be described using Equation 5.

[0125] Equation 5

[0126] Here, the parameter represents parameters such as weights and biases (if biases are used in the entropy parameter NN (917)) used in the convolution kernel in the entropy parameter NN (917). Output of the entropy parameter NN (917). It can be used to determine the entropy model (e.g., conditioning), and thus the conditional entropy model is, for example, the output from the hyperdecoder (925). It can depend on the input image x through. In an example, the output It includes parameters such as mean and scale parameters used to condition an entropy model (e.g., GMM). Referring to FIG. 9, an entropy model (e.g., conditional entropy model) can be used by an entropy encoder (913) and an entropy decoder (914) in entropy coding and entropy decoding, respectively.

[0127] The second sub-NN (952) can be described as follows. A potential value y can be fed to a hyper-encoder (921) to generate a hyper-potential value z. For example, the hyper-encoder (921) is implemented using a neural network such as a CNN. The relationship between the hyper-potential value z and the potential value y can be described using Equation 6. It may be as follows.

[0128] Equation 6

[0129] Here, the parameter represents parameters such as weights and biases used in the convolution kernel in the hyperencoder (921) (where biases are used in the hyperencoder (921)).

[0130] The hyperpotential z is quantized by a quantizer (922) and the quantized potential It generates a quantized potential value z^, which can be compressed using lossless compression by, for example, an entropy encoder (923) to generate side information such as an encoded bit (932) from a hyper-neural network. The entropy encoder (923) can use entropy coding techniques such as Huffman coding, arithmetic coding, etc. In one example, the entropy encoder (923) uses arithmetic encoding and is an arithmetic encoder. In one example, side information such as an encoded bit (932) can be transmitted as a coded bitstream together with, for example, an encoded image (931).

[0131] Additional information, such as encoded bits (932), can be decompressed (e.g., entropy decoding) by an entropy decoder (924) to produce an output. The entropy decoder (924) may use entropy coding techniques such as Huffman coding, arithmetic coding, etc. In one example, the entropy decoder (924) uses arithmetic decoding and is an arithmetic decoder. For example, lossless compression is used in the entropy encoder (923), lossless decompression is used in the entropy decoder (924), noise, etc., caused by the transmission of additional information can be omitted, and the output from the entropy decoder (924) is a quantized potential value It can be. The hyperdecoder (925) is a quantized potential Decode and output Can generate. Output c and quantized potential The relationship between them can be described using Equation 7.

[0132] Equation 7

[0133] Here, the parameter represents parameters such as weights and biases used in the convolution kernel in the hyperdecoder (925) (if biases are used in the hyperdecoder (925)).

[0134] As described above, the compressed or encoded bit (932) can be added to the coded bitstream as additional information, which enables the trophy decoder (914) to use a conditional entropy model. Thus, the entropy model can be image-dependent and spatially adaptive, and therefore can be more accurate than a fixed entropy model.

[0135] The NIC framework (900) may be appropriately adapted, for example, to omit one or more components illustrated in FIG. 9, modify one or more components illustrated in FIG. 9, and / or include one or more components not illustrated in FIG. 9. In one example, a NIC framework using a fixed entropy model includes a first sub-NN (951) and does not include a second sub-NN (952). In one example, the NIC framework includes components of the NIC framework (900) excluding an entropy encoder (923) and an entropy decoder (924).

[0136] In one embodiment, one or more components of the NIC framework (900) illustrated in FIG. 9 are implemented using neural network(s) such as CNN(s). Each NN-based component in the NIC framework (e.g., NIC framework (900)) (e.g., main encoder network (911), main decoder network (915), context model NN (916), entropy parameter NN (917), hyper-encoder (921) or decoder (925) of the hyper-NIC framework) may include any suitable architecture (e.g., having any suitable combination of layers), any suitable type of parameter (e.g., weights, biases, combinations of weights and biases, and biases), and any suitable number of parameters.

[0137] In one embodiment, the main encoder network (911), the main decoder network (915), the context model NN (916), the entropy parameter NN (917), the hyperencoder (921), and the hyperdecoder (925) are implemented using their respective CNNs.

[0138] FIG. 10 illustrates an exemplary CNN of a main encoder network (911) according to one embodiment of the present invention. For example, the main encoder network (911) includes four sets of layers, each set of layers including a convolutional layer 5x5 c192 s2 followed by a GDN layer. One or more layers illustrated in FIG. 10 may be modified and / or omitted. Additional layer(s) may be added to the main encoder network (911).

[0139] FIG. 11 illustrates an exemplary CNN of a main decoder network (915) according to one embodiment of the present disclosure. For example, the main decoder network (915) includes three sets of layers, each set of layers including an inverse convolutional layer 5x5 c192 s2 followed by an IGDN layer. Additionally, an inverse convolutional layer 5x5 c3 s2 and an IGDN layer follow the three sets of layers. One or more layers illustrated in FIG. 11 may be modified and / or omitted. Additional layer(s) may be added to the main decoder network (915).

[0140] FIG. 12 illustrates an exemplary CNN of a hyperencoder (921) according to one embodiment of the present disclosure. For example, the hyperencoder (921) includes a convolutional layer 3x3 c192 s1 followed by a Leaky ReLU, a convolutional layer 5x5 c192 s2 followed by a Leaky ReLU, and a convolutional layer 5x5 c192 s2. One or more layers illustrated in FIG. 12 may be modified and / or omitted. Additional layer(s) may be added to the hyperencoder (921).

[0141] FIG. 13 illustrates an exemplary CNN of a hyperdecoder (925) according to one embodiment of the present disclosure. For example, the hyperdecoder (925) includes an inverse convolutional layer 5x5 c192 s2 followed by a Leaky ReLU, an inverse convolutional layer 5x5 c288 s2 followed by a Leaky ReLU, and an inverse convolutional layer 3x3 c384 s1. One or more layers illustrated in FIG. 13 may be modified and / or omitted. Additional layer(s) may be added to the hyperencoder (925).

[0142] FIG. 14 illustrates an exemplary CNN of a context model NN (916) according to one embodiment of the present disclosure. For example, since the context model NN (916) includes a masked convolution 5x5 c384 s1 for context prediction, the context of Equation 4 It includes a restricted context (e.g., a 5x5 convolutional kernel). The convolutional layer of FIG. 14 can be modified. Additional layer(s) can be added to the context model NN (916).

[0143] FIG. 15 illustrates an exemplary CNN of an entropy parameter NN (917) according to one embodiment of the present disclosure. For example, the entropy parameter NN (917) includes a convolutional layer 1x1 c640 s1 followed by a Leaky ReLU, a convolutional layer 1x1 c512 s1 followed by a Leaky ReLU, and a convolutional layer 1x1 c384 s1. One or more layers illustrated in FIG. 15 may be modified and / or omitted. Additional layer(s) may be added to the entropy parameter NN (917).

[0144] The NIC framework (900) may be implemented using a CNN as described with reference to FIGS. 10 through 15. The NIC framework (900) may be appropriately adapted so that one or more components of the NIC framework (900) (e.g., (911), (915), (916), (917), (921), and / or (925)) are implemented using any appropriate type of neural network (e.g., a CNN or a non-CNN-based neural network). One or more other components of the NIC framework (900) may be implemented using neural network(s).

[0145] A NIC framework (900) including a neural network (e.g., CNN) can be trained to learn the parameters used in the neural network. For example, when using a CNN, the weights and biases used in the convolutional kernels of the main encoder network (911) (where biases are used in the main encoder network (911)), the weights and biases used in the convolutional kernels of the main decoder network (915) (where biases are used in the main decoder network (915)), the weights and biases used in the convolutional kernels of the hyper-encoder (921) (where biases are used in the hyper-encoder (921)), the weights and biases used in the convolutional kernels of the hyper-decoder (925) (where biases are used in the hyper-decoder (925)), the weights and biases used in the convolutional kernel(s) of the context model NN (916) (where biases are used in the context model NN (916)), and the weights and biases used in the convolutional kernels of the entropy parameter NN (917) (where biases are used in the entropy parameter NN (917)), - The parameters represented by can each be learned during the training process.

[0146] In one example, referring to FIG. 10, the main encoder network (911) includes four convolutional layers, each convolutional layer having a 5x5 and 192-channel convolutional kernel. Thus, the number of weights used in the convolutional kernels of the main encoder network (911) is 19200 (i.e., 4x5x5x192). The parameters used in the main encoder network (911) include 19200 weights and optional biases. Additional parameter(s) may be included if biases and / or additional NN(s) are used in the main encoder network (911).

[0147] Referring to FIG. 9, the NIC framework (900) includes at least one component or module built into the neural network(s). The at least one component may include at least one of a main encoder network (911), a main decoder network (915), a hyper-encoder (921), a hyper-decoder (925), a context model NN (916), and an entropy parameter NN (917). At least one component may be trained individually. In one example, the training process is used to learn parameters for each component individually. At least one component may be trained together as a group. In one example, the training process is used to learn parameters for a subset of at least one component together. In one example, the training process is used to learn parameters for all of at least one component, so it is called E2E optimization.

[0148] In the training process for one or more components in the NIC framework (900), the weights (or weighting coefficients) of one or more components may be initialized. In one example, the weights are initialized based on a pre-trained corresponding neural network model (e.g., DNN model, CNN model). In one example, the weights are initialized by setting the weights to random numbers.

[0149] For example, after the weights are initialized, a training image set may be used to train one or more components. The training image set may include any suitable images of any appropriate size(s). In some examples, the training image set includes raw images, natural images, and computer-generated images.

[0150] In some examples, the training image set includes residual images containing residual data in the spatial domain. The residual data can be computed by a residual calculator (e.g., residual calculator (723)). In some examples, the training images within the training image set (e.g., raw images and / or residual images containing residual data) can be divided into blocks of appropriate size, and the blocks and / or images can be used to train a neural network in the NIC framework. Thus, a neural network can be trained in the NIC framework using raw images, residual images, blocks derived from raw images and / or blocks derived from residual images.

[0151] For brevity, the training process below is explained using training images as examples. The explanation can be appropriately adapted to the training blocks. Training images of the training image set t It can generate a compressed representation (e.g., encoded information, e.g., as a bitstream) through the encoding process of FIG. 9. The encoded information is an image reconstructed through the decoding process described in FIG. 9. It can calculate and reconstruct.

[0152] In the case of the NIC framework (900), two competing targets, e.g., reconstruction quality and bit consumption, are balanced. Quality loss function (e.g., distortion or distortion loss) ) is reconstruction (e.g., reconstructed image) It can be used to indicate reconstruction quality, such as the difference between the original image (e.g., training image t). The rate (or rate loss) R can be used to indicate the bit consumption of the compressed representation. In one example, the rate loss R includes additional information, for instance, used to determine the context model.

[0153] In the case of neural image compression, a differentiable approximation of quantization can be used for E2E optimization. In various examples, during the training process of neural network-based image compression, noise injection is used to simulate quantization, and thus quantization is not performed by a quantizer (e.g., quantizer (912)) but is simulated by noise injection. Thus, training with noise injection can approximate the quantization error in various ways. Since a bits per pixel (BPP) estimator can simulate an entropy coder, entropy coding can be simulated by an entropy encoder (e.g., (913)) and an entropy decoder (e.g., (914)). During the training process, the rate loss R in the loss function L shown in FIG. 1 can be estimated, for example, based on noise injection and the BPP estimator. Generally, a higher rate R may result in lower distortion D, and a lower R may result in higher distortion D. Therefore, in Equation 1, the trade-off hyperparameter λ is the joint RD loss L It can be used to optimize, where L, the sum of λD and R, can be optimized. The training process is the combined RD loss L This can be used to adjust the parameters of one or more components (e.g., (911), (915)) in the NIC framework (900) so that they are minimized or optimized.

[0154] Since various models can be used to determine the distortion loss D and the rate loss R, the combined RD loss L can be determined from Equation 1. In one example, the distortion loss ) is expressed as the peak signal-to-noise ratio (PSNR), a metric based on mean squared error, multiscale structural similarity (MS-SSIM) quality index, and a weighted combination of PSNR and MS-SSIM.

[0155] In one example, the target of the learning process is to train an encoding neural network (e.g., encoding DNN), such as a video encoder to be used on the encoder side, and a decoding neural network (e.g., decoding DNN), such as a video decoder to be used on the decoder side. In one example, referring to FIG. 9, the encoding neural network may include a main encoder network (911), a hyper-encoder (921), a hyper-decoder (925), a context model NN (916), and an entropy parameter NN (917). The decoding neural network may include a main decoder network (915), a hyper-decoder (925), a context model NN (916), and an entropy parameter NN (917). The video encoder and / or video decoder may include other component(s) based on and / or not based on the NN(s).

[0156] The NIC framework (e.g., NIC framework (900)) can be trained in an E2E manner. In one example, the encoding neural network and the decoding neural network are updated together in the training process based on the backpropagated gradient in an E2E manner.

[0157] After the neural network parameters of the NIC framework (900) are trained, one or more components of the NIC framework (900) can be used to encode and / or decode an image. In one embodiment, on the encoder side, a video encoder is configured to encode an input image x into an encoded image (931) to be transmitted as a bitstream. The video encoder may include several components of the NIC framework (900). In one embodiment, on the decoder side, a corresponding video decoder converts the encoded image (931) in the bitstream into a reconstructed image It is configured to decode. The video decoder may include several components of the NIC framework (900).

[0158] In one example, the video encoder includes all components of the NIC framework (900), for example, when content-adaptive online training is employed.

[0159] FIG. 16a illustrates an exemplary video encoder (1600A) according to one embodiment of the present disclosure. The video encoder (1600A) includes a main encoder network (911), a quantizer (912), an entropy encoder (913), and a second sub-NN (952) as described with reference to FIG. 9, and a detailed description is omitted for brevity. FIG. 16b illustrates an exemplary video decoder (1600B) according to one embodiment of the present disclosure. The video decoder (1600B) may correspond to the video encoder (1600A). The video decoder (1600B) may include a main decoder network (915), an entropy decoder (914), a context model NN (916), an entropy parameter NN (917), an entropy decoder (924), and a hyper-decoder (925). Referring to FIG. 6a and FIG. 16b, at the encoder side, a video encoder (1600A) can generate an encoded image (931) and an encoded bit (932) to be transmitted in a bitstream. At the decoder side, a video decoder (1600B) can receive and decode the encoded image (931) and the encoded bit (932).

[0160] FIGS. 17 and FIGS. 18 respectively illustrate an exemplary video encoder (1700) and a corresponding video decoder (1800) according to one embodiment of the present disclosure. Referring to FIG. 17, the encoder (1700) includes a main encoder network (911), a quantizer (912), and an entropy encoder (913). Examples of the main encoder network (911), the quantizer (912), and the entropy encoder (913) are described with reference to FIG. 9. Referring to FIG. 18, the video decoder (1800) includes a main decoder network (915) and an entropy decoder (914). Examples of the main decoder network (915) and the entropy decoder (914) are described with reference to FIG. 9. Referring to FIG. 9. Referring to FIGS. 17 and 18, a video encoder (1700) can generate an encoded image (931) to be transmitted as a bitstream. A video decoder (1800) can receive and decode the encoded image (931).

[0161] As described above, a NIC framework (900) including a video encoder and a video decoder may be trained based on images and / or blocks within a training image set. In some examples, one or more images to be compressed (e.g., encoded) and / or transmitted have properties that are significantly different from the training image set. Therefore, encoding and decoding one or more images using a video encoder and a video decoder, respectively, trained based on the training image set results in a relatively poor RD loss L (e.g., relatively large distortion and / or a relatively large bit rate). Accordingly, aspects of the present disclosure describe a content-adaptive online training method for a NIC.

[0162] To distinguish between a training process based on a training image set and a content-adaptive online training process based on compression (e.g., encoding) and / or one or more images to be transmitted, the NIC framework (900), video encoder, and video decoder trained by the training image set are respectively referred to as the pre-trained NIC framework (900), pre-trained video encoder, and pre-trained video decoder. The parameters in the pre-trained NIC framework (900), pre-trained video encoder, or pre-trained video decoder are respectively referred to as NIC pre-trained parameters, encoder pre-trained parameters, and decoder pre-trained parameters. In one example, the NIC pre-trained parameters include encoder pre-trained parameters and decoder pre-trained parameters. In one example, the encoder pre-trained parameters and decoder pre-trained parameters do not overlap if none of the encoder pre-trained parameters are included in the decoder pre-trained parameters. For example, the encoder pre-trained parameters in (1700) (e.g., pre-trained parameters in the main encoder network (911)) and the decoder pre-trained parameters in (1800) (e.g., pre-trained parameters in the main decoder network (915)) do not overlap. In one example, the encoder pre-trained parameters and the decoder pre-trained parameters overlap if at least one of the encoder pre-trained parameters is included in the decoder pre-trained parameters. For example, the encoder pre-trained parameters in (1600A) (e.g., pre-trained parameters in the context model NN (916)) and the decoder pre-trained parameters in (1600B) (e.g., pre-trained parameters in the context model NN (916)) overlap. The NIC pre-trained parameters may be obtained based on blocks and / or images of the training image set.

[0163] The content-adaptive online training process can be described as a fine-tuning process and is described below. One or more of the NIC pre-trained parameters of the pre-trained NIC framework (900) may be further trained (e.g., fine-tuned) based on one or more images to be encoded and / or transmitted, wherein one or more images may differ from the training image set. One or more pre-trained parameters used for the NIC pre-trained parameters may be fine-tuned by optimizing the combined RD loss L based on one or more images. One or more pre-trained parameters fine-tuned by one or more images are referred to as one or more substitute parameters or one or more fine-tuning parameters. In one embodiment, after one or more of the NIC pre-trained parameters are fine-tuned (e.g., substituted) by one or more substitute parameters, neural network update information is encoded into a bitstream to indicate one or more substitute parameters or a subset of one or more substitute parameters. In one example, the NIC framework (900) is updated (or fine-tuned) so that one or more pre-trained parameters are each substituted by one or more substitute parameters.

[0164] In the first scenario, one or more pre-trained parameters include a first subset of one or more pre-trained parameters and a second subset of one or more pre-trained parameters. One or more alternative parameters include a first subset of one or more alternative parameters and a second subset of one or more alternative parameters.

[0165] A first subset of one or more pre-trained parameters is used in a pre-trained video encoder and is replaced, for example, with a first subset of one or more surrogate parameters during the training process. Thus, the pre-trained video encoder is updated with a video encoder updated by the training process. The neural network update information may indicate a second subset of one or more surrogate parameters to replace a second subset of one or more surrogate parameters. One or more images may be encoded using the updated video encoder and transmitted as a bitstream along with the neural network update information.

[0166] At the decoder side, a second set of one or more pre-trained parameters is used in a pre-trained video decoder. In one embodiment, the pre-trained video decoder receives and decodes neural network update information to determine a second subset of one or more replacement parameters. The pre-trained video decoder is updated to an updated video decoder when the second subset of one or more pre-trained parameters in the pre-trained video decoder is replaced by a second subset of one or more replacement parameters. One or more encoded images can be decoded using the updated video decoder.

[0167] FIGS. 16a and 16b illustrate an example of a first scenario. For example, one or more pre-trained parameters include N1 pre-trained parameters in a pre-trained context model NN (916) and N2 pre-trained parameters in a pre-trained main decoder network (915). Thus, a first subset of one or more pre-trained parameters includes N1 pre-trained parameters, and a second subset of one or more pre-trained parameters is identical to one or more pre-trained parameters. Thus, the N1 pre-trained parameters in the pre-trained context model NN (916) can be replaced with N1 corresponding replacement parameters so that the pre-trained video encoder (1600A) can be updated to the updated video encoder (1600A). The pre-trained context model NN (916) is also updated to the updated context model NN (916). At the decoder side, N1 pre-trained parameters can be replaced with N1 corresponding replacement parameters and N2 pre-trained parameters can be replaced with N2 corresponding replacement parameters, so that the pre-trained context model NN (916) is updated to become the updated context model NN (916) and the pre-trained main decoder network (915) is updated to become the updated main decoder network (915). Thus, the pre-trained video decoder (1600B) can be updated to the updated video decoder (1600B).

[0168] In the second scenario, none of the one or more pre-trained parameters are used in the pre-trained video encoder on the encoder side. Rather, one or more pre-trained parameters are used in the pre-trained video decoder on the decoder side. Therefore, the pre-trained video encoder is not updated and remains a pre-trained video encoder even after the training process. In one embodiment, the neural network update information represents one or more alternative parameters. One or more images can be encoded using the pre-trained video encoder and transmitted as a bitstream along with the neural network update information.

[0169] On the decoder side, a pre-trained video decoder can receive and decode neural network update information to determine one or more surrogate parameters. The pre-trained video decoder is updated to an updated video decoder when one or more pre-trained parameters in the pre-trained video decoder are replaced by one or more surrogate parameters. One or more encoded images can be decoded using the updated video decoder.

[0170] FIGS. 16a and FIGS. 16b illustrate an example of a second scenario. For example, one or more pre-trained parameters include N2 pre-trained parameters in a pre-trained main decoder network (915). Thus, none of the one or more pre-trained parameters are used in a pre-trained video encoder on the encoder side (e.g., pre-trained video encoder (1600A)). Thus, the pre-trained video encoder (1600A) remains a pre-trained video encoder after the training process. On the decoder side, the N2 pre-trained parameters can be replaced with N2 corresponding replacement parameters that update the pre-trained main decoder network (915) to an updated main decoder network (915). Thus, the pre-trained video decoder (1600B) can be updated to an updated video decoder (1600B).

[0171] In the third scenario, one or more pre-trained parameters are used in a pre-trained video encoder and are replaced, for example, with one or more alternative parameters during the training process. Thus, the pre-trained video encoder is updated with a video encoder updated by the training process. One or more images can be encoded using the updated video encoder and transmitted as a bitstream. Neural network update information is not encoded in the bitstream. On the decoder side, the pre-trained video decoder is not updated and remains a pre-trained video decoder. One or more encoded images can be decoded using the pre-trained video decoder.

[0172] FIGS. 16a and FIGS. 16b illustrate an example of a third scenario. For example, one or more pre-trained parameters are in a pre-trained main encoder network (911). Thus, one or more pre-trained parameters in the pre-trained main encoder network (911) can be replaced with one or more alternative parameters so that the pre-trained video encoder (1600A) can be updated to the updated video encoder (1600A). The pre-trained main encoder network (911) is also updated to the updated main encoder network (911). On the decoder side, the pre-trained video decoder (1600B) is not updated.

[0173] In various examples, as described in the first, second, and third scenarios, video decoding can be performed by pre-trained decoders with different functions, including decoders that have the ability to update pre-trained parameters and decoders that do not.

[0174] In one example, compression performance can be improved by coding one or more images with an updated video encoder and / or an updated video decoder compared to coding one or more images with a pre-trained video encoder and a pre-trained video decoder. Thus, a content-adaptive online training method can be used to adapt a pre-trained NIC framework (e.g., pre-trained NIC framework (900)) to target image content (e.g., one or more images to be transmitted), and thus fine-tune the pre-trained NIC framework. Thus, the video encoder on the encoder side and / or the video decoder on the decoder side can be updated.

[0175] The content-adaptive online learning method can be used as a preprocessing step (e.g., a pre-encoding step) to improve the compression performance of a pre-trained E2E NIC compression method.

[0176] In one embodiment, one or more images include a single input image, and the fine-tuning process is performed on the single input image. The NIC framework (900) is trained and updated (e.g., fine-tuned) based on the single input image. An updated video encoder on the encoder side and / or an updated video decoder on the decoder side may be used to code the single input image and optionally another input image. Neural network update information may be encoded as a bitstream together with the encoded single input image.

[0177] In one embodiment, one or more images include multiple input images, and the fine-tuning process is performed with multiple input images. The NIC framework (900) is trained and updated (e.g., fine-tuned) based on multiple input images. An updated video encoder on the encoder side and / or an updated decoder on the decoder side may be used to code multiple input images and optionally other input images. Neural network update information may be encoded into a bitstream along with the encoded multiple input images.

[0178] Rate loss R can be increased by signaling neural network update information in the bitstream. If one or more images include a single input image, the neural network update information is signaled for each encoded image, and a first increase in rate loss R is used to indicate an increase in rate loss R due to signaling of neural network update information per image. If one or more images include multiple input images, the neural network update information is signaled and shared for multiple input images, and a second increase in rate loss R is used to indicate an increase in rate loss R due to signaling of neural network update information per image. Because the neural network update information is shared by multiple input images, the second increase in rate loss R may be smaller than the first increase in rate loss R. Therefore, in some examples, it may be advantageous to fine-tune the IC framework using multiple input images.

[0179] In one embodiment, one or more pre-trained parameters to be updated are in one component of the pre-trained NIC framework (900). Accordingly, that one component of the pre-trained NIC framework (900) is updated based on one or more alternative parameters, and other components of the pre-trained NIC framework (900) are not updated.

[0180] One of the above components may be a pre-trained context model NN (916), a pre-trained entropy parameter NN (917), a pre-trained main encoder network (911), a pre-trained main decoder network (915), a pre-trained hyper-encoder (921) or a pre-trained hyper-decoder (925). The pre-trained video encoder and / or pre-trained video decoder may be updated depending on which of the components of the pre-trained NIC framework (900) is updated.

[0181] In one example, one or more pre-trained parameters to be updated are in a pre-trained context model NN (916), and thus the pre-trained context model NN (916) is updated and the remaining components (911), (915), (921), (917) and (925) are not updated. In one example, a pre-trained video encoder on the encoder side and a pre-trained video decoder on the decoder side contain a pre-trained context model NN (916), and thus both the pre-trained video encoder and the pre-trained video decoder are updated.

[0182] In one example, one or more pre-trained parameters to be updated are in a pre-trained hyperdecoder (925), and thus the pre-trained hyperdecoder (925) is updated and the remaining components (911), (915), (916), (917) and (921) are not updated. Thus, the pre-trained video encoder is not updated and the pre-trained video decoder is updated.

[0183] In one embodiment, one or more pre-trained parameters to be updated are in a plurality of components of the pre-trained NIC framework (900). Thus, a plurality of components of the pre-trained NIC framework (900) are updated based on one or more alternative parameters. In one example, a plurality of components of the pre-trained NIC framework (900) include all components composed of neural networks (e.g., DNN, CNN). In one example, a plurality of components of the pre-trained NIC framework (900) include a pre-trained main encoder network (911), a pre-trained main decoder network (915), a pre-trained context model NN (916), a pre-trained entropy parameter NN (917), a pre-trained hyper-encoder (921), and a pre-trained hyper-decoder (925), which are CNN-based components.

[0184] As described above, in one example, one or more pre-trained parameters to be updated are in a pre-trained video encoder of a pre-trained NIC framework (900). In one example, one or more pre-trained parameters to be updated are in a pre-trained video decoder of a NIC framework (900). In one example, one or more pre-trained parameters to be updated are in a pre-trained video encoder and a pre-trained video decoder of a pre-trained NIC framework (900).

[0185] The NIC framework (900) may be based on a neural network, and, for example, one or more components of the NIC framework (900) may include a neural network such as a CNN, a DNN, etc. As previously mentioned, the neural network may be specified by various types of parameters such as weights, biases, etc. Each neural network-based component in the NIC framework (900) (e.g., context model NN (916), entropy parameter NN (917), main encoder network (911), main decoder network (915), hyperencoder (921) or hyperdecoder (925)) may be composed of appropriate parameters such as weights, biases, or a combination of weights and biases. When a CNN(s) are used, the weights may include elements of the convolutional kernel. One or more types of parameters may be used to specify the neural network. In one embodiment, one or more pre-trained parameters to be updated are bias term(s), and only the bias term(s) are replaced by one or more alternative parameters. In one embodiment, one or more pre-trained parameters to be updated are weights, and only the weights are replaced by one or more replacement parameters. In one embodiment, one or more pre-trained parameters to be updated include weight and bias term(s), and all pre-trained parameters including weight and bias term(s) are replaced by one or more replacement parameters. In one embodiment, other parameters may be used to specify the neural network, and the other parameters may be fine-tuned.

[0186] The fine-tuning process may include multiple epochs (e.g., iterations) in which one or more pre-trained parameters are updated in an iterative fine-tuning process. The fine-tuning process may be stopped when the training loss has flattened or is about to flatten. In one example, the fine-tuning process is stopped when the training loss (e.g., RD ​​loss L) is less than a first threshold. In another example, the fine-tuning process is stopped when the difference between two consecutive training losses is less than a second threshold.

[0187] Two hyperparameters (e.g., step size and maximum number of steps) can be used in the fine-tuning process along with a loss function (e.g., RD ​​loss L). The maximum number of iterations can be used as a threshold for the maximum number of iterations to terminate the fine-tuning process. In one example, the fine-tuning process is stopped when the number of iterations reaches the maximum number of iterations.

[0188] The step size may represent the learning rate of an online training process (e.g., an online fine-tuning process). The step size may be used for a gradient descent algorithm or for backpropagation calculations performed in the fine-tuning process. The step size may be determined using any appropriate method. In one embodiment, different step sizes are used for images with different types of content to achieve optimal results. Different types may refer to different variances. In one example, the step size is determined based on the variance of the images used to update the NIC framework. For example, the step size of an image with high variance is larger than the step size of an image with low variance if the high variance is greater than the low variance.

[0189] In one embodiment, a first step size may be used to execute a specific number of iterations (e.g., 100). Then, a second step size (e.g., adding or subtracting a size increment to the first step size) may be used to execute a specific number of iterations. The step size to be used may be determined by comparing the results of the first step size and the second step size. Two or more step sizes may be tested to determine the optimal step size.

[0190] The step size may vary during the fine-tuning process. The step size may have an initial value at the start of the fine-tuning process, and the initial value may be reduced (e.g., by half) at later stages of the fine-tuning process, for example, after a certain number of iterations, to achieve fine-tuning. The step size or learning rate may vary by the scheduler during iterative online training. The scheduler may include parameter tuning methods used to adjust the step size. The scheduler may determine a value for the step size so that the step size can increase, decrease, or remain constant over multiple intervals. In one example, the learning rate is changed by the scheduler at each step. A single scheduler or multiple different schedulers may be used for different images. Therefore, multiple sets of surrogate parameter(s) may be generated based on multiple schedulers, and one of the multiple sets of surrogate parameter(s) with better compression performance (e.g., smaller RD loss) may be selected.

[0191] At the end of the fine-tuning process, one or more updated parameters may be calculated for each of one or more alternative parameters. In one embodiment, one or more updated parameters are calculated as the difference between one or more alternative parameters and one or more pre-trained parameters. In one embodiment, one or more updated parameters are each of one or more alternative parameters.

[0192] In one embodiment, one or more updated parameters may be generated from one or more alternative parameters using, for example, a specific linear or non-linear transformation, and the one or more updated parameters are representative parameter(s) generated based on the one or more alternative parameters. The one or more alternative parameters are transformed into one or more updated parameters for better compression.

[0193] A first subset of one or more updated parameters corresponds to a first subset of one or more replacement parameters, and a second subset of one or more updated parameters corresponds to a second subset of one or more replacement parameters.

[0194] In one example, one or more updated parameters may be compressed using, for example, the LZMA2 algorithm, a variation of the Lempel-Ziv-Markov chain algorithm (LZMA), the bzip2 algorithm, etc. In one example, compression is omitted for one or more updated parameters. In some embodiments, one or more updated parameters or a second subset of one or more updated parameters may be encoded as a bitstream as neural network update information, and the neural network update information indicates one or more replacement parameters or a second subset of one or more replacement parameters.

[0195] After the fine-tuning process, in some examples, a pre-trained video encoder on the encoder side may be updated or fine-tuned based on (i) a first subset of one or more alternate parameters or (ii) one or more alternate parameters. An input image (e.g., one of one or more images used in the fine-tuning process) may be encoded into a bitstream using the updated video encoder. Thus, the bitstream contains both the encoded image and neural network update information.

[0196] Where applicable, in one example, neural network update information is decoded (e.g., decompressed) by a pre-trained video decoder to obtain one or more updated parameters or a second subset of one or more updated parameters. In one example, one or more alternative parameters or a second subset of one or more alternative parameters may be obtained based on the relationship between one or more updated parameters and the aforementioned one or more alternative parameters. The pre-trained video decoder may be fine-tuned, and the decoded updated video may be used to decode an encoded image as described above.

[0197] The NIC framework may include any type of neural network and may use any neural network-based image compression method, such as a context hyperprior encoder-decoder framework (e.g., the NIC framework shown in FIG. 9), a scale hyperprior encoder-decoder framework, a Gaussian Mixture Likelihoods framework and a variation of the Gaussian Mixture Likelihoods framework, an RNN-based recursive compression method and a variation of the RNN-based recursive compression method.

[0198] Compared to related E2E image compression methods, the content-adaptive online training method and apparatus of the present disclosure may have the following advantages. The adaptive online training mechanism is utilized to improve NIC coding efficiency. Using a flexible and general framework allows for the accommodation of various types of pre-trained frameworks and quality metrics. For example, specific pre-trained parameters of various types of pre-trained frameworks can be replaced using online training with the images to be encoded and transmitted.

[0199] FIG. 19 illustrates a flowchart schematically describing a process (1900) according to an embodiment of the present disclosure. The process (1900) may be used to encode images such as raw images or residual images. In various embodiments, the process (1900) is executed by a processing circuit, such as a processing circuit of a terminal device (310, 320, 330, 340), a processing circuit performing the function of a video encoder (1600A), or a processing circuit performing the function of a video encoder (1700). In one example, the processing circuit performs a combination of the functions of (i) one of a video encoder (403), a video encoder (603), and a video encoder (703), and (ii) one of a video encoder (1600A) and a video encoder (1700). In some embodiments, the process (1900) is implemented as a software instruction, and thus, when the processing circuit executes the software instruction, the processing circuit performs the process (1900). The process begins at (S1901). In one example, the NIC framework is based on a neural network. In one example, the NIC framework is the NIC framework (900) described with reference to FIG. 9. The NIC framework may be based on a CNN, such as the one described with reference to FIG. 10 through FIG. 15. The video encoder (e.g., (1600A) or (1700)) and the corresponding video decoder (e.g., (1600B) or (1800)) may include a number of components in the NIC framework as described above. The neural network-based NIC framework is pre-trained, and thus the video encoder and video decoder are pre-trained. The process (1900) proceeds to step (S1910).

[0200] In (S1910), a fine-tuning process for the NIC framework is performed based on one or more images (or input image(s)). The input image(s) may be any suitable image(s) having any suitable size(s). In some examples, the input image(s) include raw image(s), natural image(s), computer-generated image(s), etc., in the spatial domain.

[0201] In some examples, the input image(s) contain residual data in the spatial domain calculated, for example, by a residual calculator (e.g., residual calculator (723)). Components of various devices can be appropriately combined to achieve (S1910), for example, with reference to FIGS. 7 and FIGS. 9, and the residual data from the residual calculator is combined with the image and supplied to the main encoder network (911) of the NIC framework.

[0202] In one or more pre-trained neural networks of a NIC framework (e.g., a pre-trained NIC framework), one or more parameters (e.g., one or more pre-trained parameters) may be updated to become one or more replacement parameters as described above. In one embodiment, one or more parameters in one or more neural networks are updated, for example, during the training process described in (S1910) at each step.

[0203] In one embodiment, at least one neural network in a video encoder (e.g., a pre-trained video encoder) is composed of a first subset of one or more pre-trained parameters, and thus at least one neural network in the video encoder can be updated based on a corresponding first subset of one or more alternative parameters. In one example, the first subset of one or more alternative parameters includes all of one or more alternative parameters. In one example, at least one neural network in the video encoder is updated when one or more pre-trained parameters are each replaced by the first subset of one or more alternative parameters. In one example, at least one neural network in the video encoder is updated iteratively in a fine-tuning process. In one example, none of one or more pre-trained parameters are included in the video encoder, and thus the video encoder is not updated and remains a pre-trained video encoder.

[0204] In (S1920), one or more images may be encoded using a video encoder having at least one updated neural network. In one example, one or more images is encoded after at least one neural network in the video encoder has been updated.

[0205] Step (S1920) can be appropriately adapted. For example, if none of the one or more alternative parameters are included in at least one neural network in the video encoder, the video encoder is not updated, and thus one or more images can be encoded using a pre-trained video encoder (e.g., a video encoder including at least one pre-trained neural network).

[0206] In (S1930), neural network update information indicating a second subset of one or more alternative parameters may be encoded into a bitstream. In one example, the second subset of one or more alternative parameters is used to update at least one neural network in a video decoder on the decoder side. Step (S1930) may be omitted, and, for example, if the second subset of one or more alternative parameters does not contain parameters and no neural network update information is signaled in the bitstream, no neural network in the video decoder is updated.

[0207] In (S1940), a bitstream containing one or more encoded images and neural network update information may be transmitted. Step (S1940) may be appropriately adapted. For example, if step (S1930) is omitted, the bitstream does not contain neural network update information. The process (1900) proceeds to (S1999) and terminates.

[0208] The process (1900) can be appropriately adapted to various scenarios, and the steps of the process (1900) can be adjusted accordingly. One or more steps of the process (1900) may be adapted, omitted, repeated, and / or combined. Any appropriate order may be used to implement the process (1900). Additional step(s) may be added. For example, in addition to encoding one of the one or more images, other image(s), such as the remaining(s) of the one or more images, are encoded in (S1920) and transmitted in (S1940).

[0209] In some examples of the process (1900), one or more images are encoded by an updated video encoder and transmitted as a bitstream. Since the fine-tuning process is based on one or more images, the fine-tuning process is based on the context to be encoded and is therefore context-based.

[0210] In some examples, neural network update information further indicates which parameter(s) are a second subset of one or more pre-trained parameters (or a corresponding second subset of one or more alternative parameters), so that the corresponding pre-trained parameter(s) in the video decoder can be updated. The neural network update information may indicate component information (e.g., (915)), layer information (e.g., layer 4 DeConv: 5x5 c3 s2), channel information (e.g., second channel), a second subset of one or more pre-trained parameters, etc., which is a subset of one or more pre-trained parameters. Thus, referring to FIG. 11, the second subset of one or more alternative parameters includes a convolutional kernel of the second channel, which is DeConv: 5x5 c3 s2 in the main decoder network (915). Thus, the convolutional kernel of the second channel, which is DeConv: 5x5 c3 s2 in the pre-trained main decoder network (915), is updated. In some examples, component information (e.g., (915)), layer information (e.g., layer 4 DeConv:5x5 c3 s2), channel information (e.g., channel 2), a second subset of one or more pre-trained parameters, etc. are predetermined and stored in a pre-trained video decoder and are therefore not signaled.

[0211] FIG. 20 illustrates a flowchart schematically illustrating a process (2000) according to one embodiment of the present disclosure. The process (2000) may be used to reconstruct an encoded image. In various embodiments, the process (2000) is executed by a processing circuit, such as a processing circuit of a terminal device (310, 320, 330, 340), a processing circuit that performs the function of a video decoder (1600B), or a processing circuit that performs the function of a video decoder (1800). In one example, the processing circuit performs a combination of the functions of (i) one of a video decoder (410), a video decoder (510), and a video decoder (810), and (ii) one of a video decoder (1600B) or a video decoder (1800). In some embodiments, the process (2000) is implemented as a software instruction, and thus, when the processing circuit executes the software instruction, the processing circuit performs the process (2000). The process starts at (S2001). In one example, the NIC framework is based on a neural network. In one example, the NIC framework is the NIC framework (900) described with reference to FIG. 9. The NIC framework may be based on a CNN such as the one described with reference to FIG. 10 through 15. A video decoder (e.g., (1600B) or (1800)) may include a number of components in the NIC framework as described above. A NIC framework based on a neural network may be pre-trained. A video decoder may be pre-trained with pre-trained parameters. The process (2000) proceeds to (S2010).

[0212] In (S2010), neural network update information within the coded bitstream can be decoded. The neural network update information may be information about the neural network in the video decoder. The neural network may be composed of pre-trained parameters. The neural network update information may correspond to the encoded image to be reconstructed and may indicate alternative parameters corresponding to the pre-trained parameters among the pre-trained parameters.

[0213] In one example, the pre-trained parameter is the pre-trained bias term.

[0214] In one example, the pre-trained parameters are pre-trained weighting coefficients.

[0215] In one embodiment, the video decoder includes a plurality of neural networks. The plurality of neural networks includes the neural networks described above. Neural network update information may indicate update information for one or more remaining neural networks among the plurality of neural networks. For example, the neural network update information further indicates one or more alternative parameters for one or more remaining neural networks among the plurality of neural networks. One or more alternative parameters correspond to one or more pre-trained parameters for each of the one or more remaining neural networks. In one example, each pre-trained parameter and one or more pre-trained parameters are each pre-trained bias term. In one example, each pre-trained parameter and one or more pre-trained parameters are each pre-trained weighting coefficient. In one example, the pre-trained parameter and one or more pre-trained parameters include one or more pre-trained bias terms and one or more pre-trained weighting coefficients in the plurality of neural networks.

[0216] In an example, neural network update information directs update information for a subset of multiple neural networks, while the remaining subset of multiple neural networks is not updated.

[0217] In one example, the video decoder is the video decoder (1800) shown in FIG. 18. The neural network is the main decoder network (915).

[0218] In one example, the video decoder is the video decoder (1600B) illustrated in FIG. 16. A number of neural networks of the video decoder include a main decoder network (915), a context model NN (916), an entropy parameter NN (917), and a hyperdecoder (925). The neural network is one of the main decoder network (915), the context model NN (916), the entropy parameter NN (917), and the hyperdecoder (925). For example, the neural network is the context model NN (916). Neural network update information directs one or more alternative parameters for one or more of the remaining neural networks of the video decoder (1600B) (e.g., the main decoder network (915), the entropy parameter NN (917), and / or the hyperdecoder (925)).

[0219] In one example, neural network update information indicates multiple surrogate parameters corresponding to multiple pre-trained parameters among the pre-trained parameters for the neural network. The multiple pre-trained parameters include the aforementioned pre-trained parameters. The multiple pre-trained parameters include one or more pre-trained bias terms and one or more pre-trained weighting coefficients.

[0220] In (S2020), alternative parameters can be determined based on neural network update information. In one embodiment, the updated parameters are obtained from neural network update information. In one example, the updated parameters can be obtained from neural network update information by decompression. In one example, the neural network update information indicates that the updated parameters are the difference between the alternative parameters and the pre-trained parameters, and the alternative parameters can be calculated based on the sum of the updated parameters and the pre-trained parameters. In one embodiment, the alternative parameters are determined to be the updated parameters. In one embodiment, the updated parameters are representative parameters generated based on the alternative parameters on the encoder side (e.g., using a linear or non-linear transformation), and the alternative parameters are obtained based on the representative parameters.

[0221] In (S2030), the neural network in the video decoder may be updated (or fine-tuned) based on alternative parameters, for example, by replacing a pre-trained parameter with an alternative parameter in the neural network. If the video decoder includes multiple neural networks and the neural network update information directs update information (e.g., additional alternative parameter(s)) for the multiple neural networks, the multiple neural networks may be updated. For example, the neural network update information further includes one or more alternative parameters for one or more remaining neural networks in the video decoder, and one or more remaining neural networks may be updated based on one or more alternative parameters.

[0222] In (S2040), the encoded image within the bitstream may be decoded by an updated video decoder based on, for example, an updated neural network. The output image generated in (S2040) may be any suitable image(s) having any appropriate size(s). In some examples, the output image includes a reconstructed raw image in the spatial domain, a natural image, a computer-generated image, etc.

[0223] In some examples, the output image of the video decoder contains residual data in the spatial domain, and thus additional processing may be used to generate a reconstructed image based on the output image. For example, the reconstruction module (874) is configured to combine the residual data in the spatial domain and the prediction result (output by the inter or intra prediction module) to form a reconstructed block that may be part of the reconstructed image. Additional appropriate actions, such as deblocking operations, may be performed to improve visual quality. Various components of the device may be appropriately combined to achieve (S2040), and, for example as illustrated in FIGS. 8 and 9, residual data and the corresponding prediction result from the main decoder network (915) of the video decoder are supplied to the reconstruction module (874) to generate a reconstructed image.

[0224] In one example, the bitstream further includes one or more encoded bits used to determine a context model for decoding the encoded image. The video decoder may include a main decoder network (e.g., (911)), a context model network (e.g., (916)), an entropy parameter network (e.g., (917)), and a hyperdecoder network (e.g., (925)). The neural network is one of the main decoder network, the context model network, the entropy parameter NN, and the hyperdecoder network. One or more encoded bits may be decoded using the hyperdecoder network. The entropy model (e.g., context model) may be determined using the context model network and the entropy parameter network based on the decoded bits of the encoded image available to the context model network and the quantized potentials and decoded bits. The encoded image may be decoded using the main decoder network and the entropy model.

[0225] The process (2000) proceeds to (S2099) and terminates.

[0226] The process (2000) can be appropriately adapted to various scenarios, and the steps of the process (2000) can be adjusted accordingly. One or more steps of the process (2000) may be adapted, omitted, repeated, and / or combined. Any appropriate order may be used to implement the process (2000). Additional step(s) may be added.

[0227] For example, in (S2040), one or more additional encoded images from the encoded bitstream are decoded based on an updated neural network. Thus, the encoded image and the one or more additional encoded images can share the same neural network update information.

[0228] The embodiments of the present disclosure may be used individually or combined in any order. Additionally, each method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a computer-readable non-universal medium.

[0229] The present disclosure makes no limitations on the methods used in encoders, such as neural network-based encoders, or decoders, such as neural network-based decoders. The neural network(s) used in encoders, decoders, etc., may be any suitable type of neural network(s), such as DNNs, CNNs, etc.

[0230] Accordingly, the content-adaptive online training method of the present invention can accommodate different types of NIC frameworks, such as different types of encoding DNN, decoding DNN, encoding CNN, decoding CNN, etc.

[0231] The technology described above may be implemented as computer software that uses computer-readable instructions and can be physically stored on one or more computer-readable media. For example, FIG. 21 illustrates a computer system (2100) suitable for implementing a specific embodiment of the disclosed invention.

[0232] Computer software may be coded using any suitable machine code or computer language capable of generating code containing instructions that can be executed directly, or through interpretation, micro-code execution, etc., by a computer central processing unit (CPU), graphics processing unit (GPU), etc., via assembly, compilation, linking, or similar mechanisms.

[0233] The instruction can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0234] The components of the computer system (2100) illustrated in FIG. 21 are essentially exemplary and are not intended to imply any limitation to the scope of use or function of the computer software implementing the embodiments of the present disclosure. The configuration of the components should not be interpreted as having any dependency or requirement related to any one or combination of the components shown in the exemplary embodiments of the computer system (2100).

[0235] The computer system (2100) may include a specific human interface input device. This human interface input device may respond to input by one or more human users, for example, tactile input (e.g., keystroke, swip, data glove movement), audio input (e.g., voice, clap), visual input (e.g., gesture), and olfactory input (not shown). The human interface device may also be used to capture specific media that are not necessarily directly related to conscious input by a person, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, pictures acquired from a still image camera), and videos (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0236] The input human interface device may include one or more of a keyboard (2101), a mouse (2102), a trackpad (2103), a touch screen (2110), a data glove (not shown), a joystick (2105), a microphone (2106), a scanner (2107), and a camera (2108) (each only one is shown).

[0237] The computer system (2100) may include a specific human interface output device. Such human interface output device may stimulate the senses of one or more human users, for example, through tactile output, sound, light and smell / taste. These human interface output devices may include haptic output devices (e.g., touch screen (2110), data gloves (not shown), or haptic feedback via a joystick (2105), but there may also be haptic feedback devices that do not serve as input devices), audio output devices (e.g., speakers (2109), headphones (not shown)), visual output devices (e.g., screens (2110) including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input functions and with or without haptic feedback functions—some of which may produce two-dimensional visual output or three-dimensional or higher output through means such as stereographic output, virtual-reality glasses (not shown), holographic display and smoke tank (not shown)—), and printers (not shown).

[0238] The computer system (2100) may also include human-accessible storage devices and associated media, such as optical media including a CD / DVD ROM RW (2120) having a CD / DVD media (2121), a thumb drive (2122), a removable hard drive or solid-state drive (2123), legacy magnetic media such as tape and floppy disk (not shown), and specialized ROM / ASIC / PLD-based devices such as a security dongle (not shown).

[0239] Those skilled in the art should also understand that the term “computer-readable medium” as used in connection with the subject matter disclosed herein does not include a transmitting medium, a carrier wave, or other transient signals.

[0240] The computer system (2100) may also include an interface (2154) to one or more communication networks (2155). The network may be, for example, a wireless, wired, or optical network. The network may also be a local, wide-area, metropolitan, automotive and industrial, real-time, latency-tolerant, etc. Examples of networks include cellular networks including Ethernet, wireless LAN, GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, automotive and industrial networks including CANBus, etc. A specific network generally requires an external network interface adapter attached to a specific general-purpose data port or peripheral bus (2149) (e.g., a USB port of the computer system (2100)); others are generally integrated into the core of the computer system (2100) by attaching to a system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (2100) can communicate with other entities. This communication may be unidirectional, receive-only (e.g., TV broadcasting), unidirectional transmit-only (e.g., from a CANbus to a specific CANbus device), or bidirectional (e.g., using a local or wide-area digital network to another computer system). Specific protocols and protocol stacks may be used for each of the networks and network interfaces as described above.

[0241] The aforementioned human interface device, human-accessible storage device, and network interface can be attached to the core (2140) of the computer system (2100).

[0242] The core (2140) may include one or more central processing units (CPUs) (2141), graphics processing units (GPUs) (2142), specialized programmable processing units in the form of a field programmable gate area (FPGA) (2143), hardware accelerators (2144) for specific tasks, graphics adapters (2150), etc. Along with read-only memory (ROM) (2145), random access memory (2146), and internal mass storage devices such as internal hard drives, SSDs, etc. (2147) that are inaccessible to the user, these devices may be connected via a system bus (2148). In some computer systems, the system bus (2148) may be accessible in the form of one or more physical plugs that enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be connected directly to the core's system bus (2148) or via a peripheral bus (2149). For example, the screen (2110) can be connected to a graphics adapter (2150). Architectures for peripheral buses include PCI, USB, etc.

[0243] The CPU (2141), GPU (2142), FPGA (2143), and accelerator (2144) can be combined to execute specific instructions that can construct the aforementioned computer code. The computer code may be stored in ROM (2145) or RAM (2146). Transitional data may also be stored in RAM (2146), while permanent data may be stored, for example, in an internal mass storage device (2147). Fast storage and retrieval of any of the memory elements may be made possible through the use of a cache memory that may be closely associated with one or more CPUs (2141), GPUs (2142), mass storage devices (2147), ROM (2145), RAM (2146), etc.

[0244] A computer-readable medium may have computer code for performing various computer-implemented operations. The medium and the computer code may be specially designed and constructed for the purposes of this disclosure, or they may be of a type well known and available to those skilled in the art of computer software.

[0245] As an example, not limited to, the architecture (2100), specifically the core (2140), may provide functionality as a result of processor(s) (including CPU, GPU, FPGA, accelerator, etc.) executing software implemented on one or more types of computer-readable media. Such computer-readable media may be media associated with a mass storage device accessible to a user as described above, as well as specific storage devices of the core (2140) of a non-transient nature, such as a mass storage device (2147) or ROM (2145) within the core. Software implementing various embodiments of the present disclosure may be stored on such devices and executed by the core (2140). The computer-readable media may include one or more memory elements or chips as needed. Software may enable the core (2140) and, in particular, the internal processor (including a CPU, GPU, FPGA, etc.) to execute a specific process or a specific part of a specific process described herein, including defining data structures stored in RAM (2146) and modifying such data structures according to a process defined by the software. Additionally or alternatively, the computer system may provide a function that is otherwise implemented in a circuit (e.g., accelerator (2144)) as a result of logic hardwired, which can operate instead of or with the software to execute a specific process or a specific part of a specific process described herein. References to software may include logic, and where appropriate, vice versa. References to computer-readable media may include circuits that store software for execution (e.g., integrated circuit (IC)), circuits that implement logic for execution, or both, where appropriate. The present disclosure includes any suitable combination of hardware and software.

[0246] 부록 A: 약어

[0247] JEM: joint exploration model

[0248] VVC: versatile video coding

[0249] BMS: benchmark set

[0250] MV: Motion Vector

[0251] HEVC: High Efficiency Video Coding

[0252] SEI: Supplementary Enhancement Information

[0253] VUI: Video Usability Information

[0254] GOPs: Groups of Pictures

[0255] TUs: Transform Units,

[0256] PUs: Prediction Units

[0257] CTUs: Coding Tree Units

[0258] CTBs: Coding Tree Blocks

[0259] PBs: Prediction Blocks

[0260] HRD: Hypothetical Reference Decoder

[0261] SNR: Signal Noise Ratio

[0262] CPUs: Central Processing Units

[0263] GPUs: Graphics Processing Units

[0264] CRT: Cathode Ray Tube

[0265] LCD: Liquid-Crystal Display

[0266] OLED: Organic Light-Emitting Diode

[0267] CD: Compact Disc

[0268] DVD: Digital Video Disc

[0269] ROM: Read-Only Memory

[0270] RAM: Random Access Memory

[0271] ASIC: Application-Specific Integrated Circuit

[0272] PLD: Programmable Logic Device

[0273] LAN: Local Area Network

[0274] GSM: Global System for Mobile communications

[0275] LTE: Long-Term Evolution

[0276] CANBus: Controller Area Network Bus

[0277] USB: Universal Serial Bus

[0278] PCI: Peripheral Component Interconnect

[0279] FPGA: Field Programmable Gate Areas

[0280] SSD: solid-state drive

[0281] IC: Integrated Circuit

[0282] CU: Coding Unit

[0283] NIC: Neural Image Compression

[0284] RD: Rate-Distortion

[0285] E2E: End to End

[0286] ANN: Artificial Neural Network

[0287] DNN: Deep Neural Network

[0288] CNN: Convolution Neural Network

[0289] Although the present disclosure describes some exemplary embodiments, there are modifications, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, those skilled in the art will understand that numerous systems and methods can be devised that implement the principles of the present disclosure and thus fall within the spirit and scope of the present disclosure, even though they are not expressly illustrated or described herein.

Claims

Claim 1 A method for video decoding in a video decoder, comprising: decoding neural network update information within a coded bitstream for a neural network in the video decoder, wherein the neural network is one of a main decoder network, a context model network, an entropy parameter network, and a hyperdecoder network and is composed of pre-trained parameters, and the neural network update information corresponding to an encoded image is reconstructed and indicates an alternative parameter corresponding to a pre-trained parameter among the pre-trained parameters; decoding one or more encoded bits using the hyperdecoder network, wherein the coded bitstream further indicates one or more encoded bits used to determine a context model for decoding the encoded image; determining the context model using the context model network and the entropy parameter network based on a quantized latent of the encoded image available to the context model network and the one or more decoded bits; and updating the neural network in the video decoder based on the alternative parameter. A method comprising the step of decoding the encoded image based on the updated neural network using the main decoder network and the context model. Claim 2 The method of claim 1, wherein the neural network update information further indicates one or more alternative parameters for one or more remaining neural networks in the video decoder, and the method further comprises the step of updating one or more remaining neural networks based on the one or more alternative parameters. Claim 3 delete Claim 4 A method according to claim 1, wherein the pre-trained parameter is a pre-trained bias term. Claim 5 In claim 1, the method wherein the pre-trained parameter is a pre-trained weighting coefficient. Claim 6 A method according to claim 1, wherein the neural network update information indicates a plurality of alternative parameters corresponding to a plurality of pre-trained parameters among pre-trained parameters for the neural network, wherein the plurality of pre-trained parameters include the pre-trained parameters, and the plurality of pre-trained parameters include one or more pre-trained bias terms and one or more pre-trained weighting coefficients, and the updating step includes the step of updating the neural network in the video decoder based on the plurality of alternative parameters including the alternative parameters. Claim 7 A method according to claim 1, wherein the neural network update information indicates the difference between the alternative parameter and the pre-trained parameter, and the method further comprises the step of determining the alternative parameter according to the sum of the difference and the pre-trained parameter. Claim 8 A method according to claim 1, further comprising the step of decoding another encoded image from the coded bitstream based on the updated neural network. Claim 9 A device for video decoding comprising a processing circuit, wherein the processing circuit is configured to perform the method of any one of claims 1, 2, or 4 through 8. Claim 10 A computer-readable non-transient storage medium for storing a program, wherein the program is executable by at least one processor to perform the method of any one of claims 1, 2, or 4 through 8. Claim 11 delete Claim 12 delete Claim 13 delete Claim 14 delete Claim 15 delete Claim 16 delete Claim 17 delete Claim 18 delete Claim 19 delete Claim 20 delete