Method and apparatus for video decoding

By updating neural network parameters in the video decoder, the problem of low encoding efficiency caused by the difference in image characteristics and training image sets in the prior art is solved, and a better video decoding effect is achieved.

CN115735359BActive Publication Date: 2025-05-06TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202280003936.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-04-26
Filing Date
2022-04-29
Publication Date
2025-05-06
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

Existing video encoding techniques lead to relatively poor losses, such as rate distortion loss, which affects encoding efficiency when processing images that are significantly different from the training image set.

Method used

By using neural network update information in the video decoder, the neural network in the video decoder is updated based on replacement parameters, thereby optimizing or minimizing losses.

Benefits of technology

Improves the efficiency of video decoding and optimizes losses, especially when the image characteristics are different from the training image set.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115735359B_ABST
    Figure CN115735359B_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure provide methods, devices, computer apparatuses, and non-transitory computer-readable storage media for video decoding. The method for video decoding includes: decoding neural network update information for a neural network in a video decoder in a coded bitstream, the neural network being configured with pre-trained parameters, the neural network update information corresponding to a coded image to be reconstructed and indicating replacement parameters corresponding to pre-trained parameters in the pre-trained parameters; updating the neural network in the video decoder based on the replacement parameters; and decoding the coded image based on the updated neural network for the coded image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of priority to U.S. Patent Application No. 17 / 729,994, filed on April 26, 2022, "METHOD AND APPARATUS FOR CONTENT-ADAPTIVE ONLINE TRAINING IN NEURAL IMAGE COMPRESSION," which claims the benefit of priority to U.S. Provisional Application No. 63 / 182,396, filed on April 30, 2021, "CONTENT-ADAPTIVE ONLINE TRAINING IN NEURAL IMAGE COMPRESSION." The disclosures of the prior applications are hereby incorporated by reference in their entirety. Technical Field

[0003] The present application generally relates to video coding and decoding technology, and more particularly to a method and device for video decoding. Background Art

[0004] The purpose of the background description provided herein is to generally present the context of the present disclosure. To the extent the work is described in this background section, the work of the presently named inventors and aspects of the description that may not otherwise be described as prior art at the time of filing are neither explicitly nor implicitly admitted as prior art to the present disclosure.

[0005] Video encoding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital images and / or videos can include a series of pictures, each picture having spatial dimensions of, for example, 1920×1080 luminance samples and associated chrominance samples. This series of pictures can have a fixed or variable picture rate (also informally referred to as a frame rate) of, for example, 60 pictures per second or 60Hz. Uncompressed images and / or videos have specific bit rate requirements. For example, 1080p60 4:2:0 video (1920×1080 luminance sample resolution at 60Hz frame rate) with 8 bits per sample requires a bandwidth of nearly 1.5Gbit / s. One hour of such video requires more than 600 gigabytes of storage space.

[0006] One purpose of video encoding and decoding can be to reduce redundancy in input images and / or video signals by compression. Compression can help reduce the bandwidth and / or storage space requirements mentioned above, in some cases by two orders of magnitude or more. Both lossless compression and lossy compression, as well as combinations thereof, can be used. Lossless compression refers to a technique that can reconstruct an exact copy of the original signal based on the compressed original signal. When lossy compression is used, the reconstructed signal may be different from the original signal, but the distortion between the original signal and the reconstructed signal is small enough to enable the reconstructed signal to be used for the intended application. In the case of video, lossy compression is widely used. The amount of distortion tolerated depends on the application; for example, users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio can reflect that a higher allowable / tolerable distortion can produce a higher compression ratio. Although the description herein uses video encoding / decoding as an illustrative example, the same technique can be applied to image encoding / decoding in a similar manner without departing from the spirit of the present disclosure.

[0007] Video encoders and decoders may utilize techniques from several broad categories including, for example, motion compensation, transforms, quantization, and entropy coding.

[0008] Video codec techniques may include techniques known as intra-frame coding. In intra-frame coding, sample values ​​are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, pictures are spatially subdivided into blocks of samples. When all sample blocks are encoded in intra-frame mode, the picture may be an intra-frame picture. Intra-frame pictures and their derivatives (e.g., independent decoder refresh pictures) may be used to reset decoder states, and may therefore be used as the first picture in a coded video bitstream and video session, or as a still image. Samples of intra-frame blocks may be subjected to transformation, and transform coefficients may be quantized before entropy coding. Intra-frame prediction may be a technique for minimizing sample values ​​in a pre-transform domain. In some cases, the smaller the DC value after transformation and the smaller the AC coefficient, the fewer bits are required to represent the block after entropy coding at a given quantization step size.

[0009] Conventional intra-frame coding, such as that known from, for example, MPEG-2 generation coding techniques, does not use intra-frame prediction. However, some newer video compression techniques include techniques that attempt to use intra-frame prediction based on surrounding sample data and / or metadata obtained, for example, during encoding and / or decoding of spatially adjacent and preceding data blocks in decoding order. Such techniques are hereinafter referred to as "intra-frame prediction" techniques. Note that, in at least some cases, intra-frame prediction uses reference data only from the current picture under reconstruction, and not reference data from reference pictures.

[0010] There can be many different forms of intra-frame prediction. When more than one such technique can be used in a given video coding technique, the techniques used can be encoded in intra-frame prediction mode. In some cases, a mode can have sub-modes and / or parameters, and these sub-modes and / or parameters can be encoded separately or included in the mode codeword. Which codeword is used for a given mode, sub-mode and / or parameter combination can affect the coding efficiency gain through intra-frame prediction, and therefore the entropy coding technique used to convert the codeword into a bitstream can also affect the coding efficiency gain through intra-frame prediction.

[0011] Some modes of intra prediction were introduced with H.264, refined in H.265, and further refined in newer coding techniques such as the joint exploration model (JEM), versatile video coding (VVC), and the benchmark set (BMS). Using neighboring sample values ​​belonging to already available samples, a predictor block can be formed. Sample values ​​of neighboring samples are copied into the predictor block according to the direction. A reference to the used direction can be encoded in the bitstream or can be predicted itself.

[0012] Neural image compression technology and / or neural video compression technology, such as artificial intelligence (AI)-based neural image compression (NIC), has begun to be applied to video coding and decoding technology.

[0013] The NIC framework including the video encoder and the video decoder can be trained based on images and / or blocks in the set of training images. In the case where one or more images to be compressed (e.g., encoded) and / or transmitted have significantly different characteristics from the set of training images, encoding and decoding the one or more images using the video encoder and the video decoder trained based on the set of training images, respectively, may result in relatively poor losses, such as rate-distortion (RD) losses (e.g., relatively large distortion and / or relatively large bit rate). Summary of the invention

[0014] Aspects of the present disclosure provide methods and apparatus for video encoding and decoding.

[0015] In some examples, a method for video decoding includes: decoding neural network update information for a neural network in a video decoder in a coded bitstream, the neural network being configured with pre-trained parameters, the neural network update information corresponding to a coded image to be reconstructed and indicating replacement parameters corresponding to pre-trained parameters in the pre-trained parameters; updating the neural network in the video decoder based on the replacement parameters; and decoding the coded image based on the updated neural network for the coded image.

[0016] In some examples, a device for video decoding includes a neural network update information decoding module, an update module, and an image decoding module. The neural network update information decoding module is used to decode the neural network update information for the neural network in the video decoder in the encoded bitstream. The neural network is configured with pre-trained parameters. The neural network update information corresponds to the encoded image to be reconstructed and indicates replacement parameters corresponding to the pre-trained parameters in the pre-trained parameters. The update module is used to update the neural network in the video decoder based on the replacement parameters. The image decoding module is used to decode the encoded image based on the updated neural network for the encoded image.

[0017] In an embodiment, the neural network update information further indicates one or more replacement parameters for one or more remaining neural networks in the video decoder. The update module is used to update the one or more remaining neural networks based on the one or more replacement parameters.

[0018] In an embodiment, the coded bitstream also indicates one or more coded bits used to determine a context model for decoding the coded image. The video decoder includes a main decoder network, a context model network, an entropy parameter network, and a super decoder network. The neural network is one of the main decoder network, the context model network, the entropy parameter network, and the super decoder network. The image decoding module is used to decode the one or more coded bits using the super decoder network, determine the context model using the context model network and the entropy parameter network based on the quantized potential and one or more decoded bits of the coded image available to the context model network, and decode the coded image using the main decoder network and the context model.

[0019] In the example, the pre-trained parameter is a pre-trained bias term.

[0020] In the example, the pre-trained parameters are pre-trained weight coefficients.

[0021] In an example, the neural network update information indicates a plurality of replacement parameters corresponding to a plurality of pre-trained parameters in the pre-trained parameters for the neural network. The plurality of pre-trained parameters include pre-trained parameters, and the plurality of pre-trained parameters include one or more pre-trained bias terms and one or more pre-trained weight coefficients. The update module may update the neural network in the video decoder based on the plurality of replacement parameters including the replacement parameters.

[0022] In an embodiment, the neural network update information indicates a difference between the replacement parameter and the pre-trained parameter. The device also includes a determination module, which can determine the replacement parameter based on the difference and the sum of the pre-trained parameter.

[0023] In an embodiment, the image decoding module may decode additional encoded images in the encoded bitstream based on the updated neural network.

[0024] Aspects of the present disclosure also provide a computer device. The computer device includes a processor and a memory. The memory is used to store program code and transmit the program code to the processor. The processor is used to execute a method for video decoding according to instructions in the program code.

[0025] Aspects of the present disclosure also provide a non-transitory computer-readable storage medium storing a program, the program being executable by at least one processor to perform a method for video decoding.

[0026] According to the method and device for video decoding provided by the present disclosure, neural network update information for a neural network in a video decoder in a coded bitstream is decoded, the neural network is configured with pre-trained parameters, the neural network update information corresponds to a coded image to be reconstructed and indicates replacement parameters corresponding to the pre-trained parameters in the pre-trained parameters, the neural network in the video decoder is updated based on the replacement parameters, and the coded image is decoded based on the updated neural network for the coded image, which can optimize or minimize the loss. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0028] Figure 1A is a schematic illustration of an exemplary subset of intra prediction modes.

[0029] Figure 1B is a diagram of exemplary intra prediction directions.

[0030] FIG. 2 illustrates a current block and surrounding samples according to an embodiment.

[0031] Figure 3 is a schematic illustration of a simplified block diagram of a communication system according to an embodiment.

[0032] Figure 4 is a schematic illustration of a simplified block diagram of a communication system according to an embodiment.

[0033] Figure 5 is a schematic illustration of a simplified block diagram of a decoder according to an embodiment.

[0034] Figure 6 is a schematic illustration of a simplified block diagram of an encoder according to an embodiment.

[0035] Figure 7 A block diagram of an encoder according to another embodiment is shown.

[0036] Figure 8 A block diagram of a decoder according to another embodiment is shown.

[0037] Fig. 9 An exemplary NIC framework according to an embodiment of the present disclosure is shown.

[0038] Fig.10 An exemplary convolution neural network (CNN) of a main encoder network according to an embodiment of the present disclosure is shown.

[0039] Fig.11 An exemplary CNN of the main decoder network according to an embodiment of the present disclosure is shown.

[0040] Fig.12 An exemplary CNN of a superencoder according to an embodiment of the present disclosure is shown.

[0041] Fig.13 An exemplary CNN of a superdecoder according to an embodiment of the present disclosure is shown.

[0042] Fig.14 An exemplary CNN of a context model network according to an embodiment of the present disclosure is shown.

[0043] Fig.15 An exemplary CNN of an entropy parameter network according to an embodiment of the present disclosure is shown.

[0044] Fig.16A An exemplary video encoder according to an embodiment of the present disclosure is shown.

[0045] Fig. 16B An exemplary video decoder according to an embodiment of the present disclosure is shown.

[0046] Fig.17 An exemplary video encoder according to an embodiment of the present disclosure is shown.

[0047] Fig.18 An exemplary video decoder according to an embodiment of the present disclosure is shown.

[0048] Fig.19 A flowchart outlining a process according to an implementation of the present disclosure is shown.

[0049] Fig. 20 A flowchart outlining a process according to an implementation of the present disclosure is shown.

[0050] Fig.21is a schematic illustration of a computer system according to an embodiment. DETAILED DESCRIPTION

[0051] Figure 1A is a schematic illustration of an exemplary subset of intra prediction modes. Figure 1A , depicted at the bottom right is a subset of nine predictor directions known from the 33 possible predictor directions of H.265 (corresponding to the 33 angular modes of the 35 intra modes). The point where the arrows intersect (101) represents the sample being predicted. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted based on one or more samples at a 45 degree angle to the horizontal line at the top right. Similarly, arrow (103) indicates that sample (101) is predicted based on one or more samples at a 22.5 degree angle to the horizontal line at the bottom left of sample (101).

[0052] Still refer to Figure 1A , depicted at the top left is a square block (104) of 4×4 samples (indicated by bold dashed lines). The square block (104) includes 16 samples, each of which is labeled with "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (from the top) and the first sample in the X dimension (from the left). Similarly, sample S44 is the fourth sample in the block (104) in both the Y dimension and the X dimension. Since the size of the block is 4×4 samples, S44 is at the bottom right. Also shown are reference samples that follow a similar numbering scheme. The reference samples are labeled with R, their Y position (e.g., row index) and X position (column index) relative to the block (104). In both H.264 and H.265, the prediction samples are adjacent to the block under reconstruction; therefore, there is no need to use negative values.

[0053] Intra-picture prediction can work by copying reference sample values ​​from appropriate neighboring samples along a signaled prediction direction. For example, assume that the coded video bitstream includes signaling that indicates, for this block, a prediction direction consistent with arrow (102) - that is, the sample is predicted based on one or more prediction samples to the upper right at a 45 degree angle to the horizontal. In this case, samples S41, S32, S23, and S14 are predicted based on the same reference sample R05. Sample S44 is then predicted based on reference sample R08.

[0054] In some cases, the values ​​of multiple reference samples may be combined, for example, by interpolation, in order to compute a reference sample; in particular, when the directions are not evenly divisible at 45 degrees.

[0055] As video coding technology has evolved, the number of possible directions has also increased. In H.264 (2003), nine different directions could be represented. In H.265 (2013), this increased to 33, and JEM / VVC / BMS can support up to 65 directions when published. Experiments have been conducted to identify the most likely directions, and certain techniques in entropy coding are used to represent these possible directions with a small number of bits, at the expense of fewer possible directions. In addition, the direction itself can sometimes be predicted based on neighboring directions used in adjacent decoded blocks.

[0056] Figure 1B A schematic diagram (110) is shown depicting 65 intra prediction directions according to JEM to illustrate that the number of prediction directions increases over time.

[0057] The mapping of intra-prediction direction bits representing directions in the coded video bitstream can vary according to different video coding techniques; and the mapping can range, for example, from simple direct mapping of prediction directions to intra-prediction modes, codewords, complex adaptive schemes involving most probable modes, and similar techniques. However, in all cases, there may be certain directions that are statistically less likely to occur in the video content than certain other directions. Since the goal of video compression is to reduce redundancy, in a well-functioning video coding technique, those less probable directions will be represented by a larger number of bits than more probable directions.

[0058] Motion compensation may be a lossy compression technique and may involve a technique in which a block of sample data from a previously reconstructed picture or portion thereof (reference picture) is used to predict a reconstructed picture or picture portion after being spatially shifted in a direction indicated by a motion vector (hereinafter MV). In some cases, the reference picture may be the same as the picture under current reconstruction. The MV may have two dimensions, X and Y, or three dimensions, the third dimension being an indication of the reference picture in use (indirectly, the third dimension may be a temporal dimension).

[0059] In some video compression techniques, an MV applicable to a particular region of sample data can be predicted based on other MVs, for example, based on an MV related to another region of sample data that is spatially adjacent to the region under reconstruction and precedes the MV in decoding order. The above prediction can significantly reduce the amount of data required to encode the MV, thereby eliminating redundancy and increasing compression. MV prediction can work effectively, for example, because when encoding an input video signal derived from a camera (called natural video), there is a statistical probability that a larger region than the region to which a single MV is applicable moves in a similar direction, and therefore in some cases similar motion vectors derived from MVs of adjacent regions can be used to predict the larger region. This makes the MV obtained for a given region similar or identical to the MV predicted from the surrounding MVs, and in turn, the MV can be represented after entropy coding with a smaller number of bits than would be used if the MV was encoded directly. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., MV) derived from the original signal (i.e., sample stream). In other cases, MV prediction itself can be lossy, for example due to rounding errors when calculating a predictor based on several surrounding MVs.

[0060] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T H.265 Recommendation, "High Efficiency Video Coding", December 2016). Among the multiple MV prediction mechanisms provided by H.265, described here is a technique referred to as "spatial merging" hereinafter.

[0061] 2, the current block (201) includes samples obtained by the encoder during the motion search process that can be predicted from a previous block of the same size that has been spatially shifted. Instead of encoding the MV directly, the MV associated with any of the five surrounding samples represented by A0, A1 and B0, B1, B2 (corresponding to 202 to 206, respectively) can be used to derive the MV from the reference picture associated with one or more reference pictures, such as the most recent (in decoding order). In H.265, MV prediction can use a predictor from the same reference picture being used by a neighboring block.

[0062] Figure 3 A simplified block diagram of a communication system (300) according to an embodiment of the present disclosure is shown. The communication system (300) includes a plurality of terminal devices that can communicate with each other via, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). Figure 3In the example of , a first pair of terminal devices (310) and (320) perform unidirectional transmission of data. For example, the terminal device (310) can encode video data (e.g., a video picture stream captured by the terminal device (310)) for transmission to another terminal device (320) via a network (350). The encoded video data can be transmitted in the form of one or more encoded video bitstreams. The terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to recover the video picture, and display the video picture based on the recovered video data. Unidirectional data transmission may be common in media service applications, etc.

[0063] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data, which bidirectional transmission may occur, for example, during a video conference. For the bidirectional transmission of data, in the example, each of the terminal devices (330) and (340) can encode video data (e.g., a video picture stream captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) via the network (350). Each of the terminal devices (330) and (340) can also receive the encoded video data transmitted by the other of the terminal devices (330) and (340), and can decode the encoded video data to restore the video picture, and can display the video picture at an accessible display device based on the restored video data.

[0064] exist Figure 3 In the example of , terminal devices (310), (320), (330) and (340) can be shown as servers, personal computers and smart phones, but the principles of the present disclosure may not be so limited. Implementations of the present disclosure are applicable to laptop computers, tablet computers, media players and / or dedicated video conferencing equipment. Network (350) represents any number of networks that transmit encoded video data between terminal devices (310), (320), (330) and (340), including, for example, wired (wired) and / or wireless communication networks. Communication network (350) can exchange data in circuit switching channels and / or packet switching channels. Representative networks include telecommunication networks, local area networks, wide area networks and / or the Internet. For the purposes of this discussion, unless otherwise described herein below, the architecture and topology of network (350) may be irrelevant to the operation of the present disclosure.

[0065] As examples of applications of the disclosed subject matter, Figure 4The placement of a video encoder and a video decoder in a streaming environment is shown. The disclosed subject matter may be equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc., and the like.

[0066] The streaming system may include a capture subsystem (413) that may include a video source (401), such as a digital camera, that creates, for example, an uncompressed video picture stream (402). In an example, the video picture stream (402) includes samples captured by the digital camera. The video picture stream (402) is depicted as a thick line to emphasize the high amount of data when compared to the encoded video data (404) (or encoded video bitstream), which may be processed by an electronic device (420) coupled to the video source (401) and including a video encoder (403). The video encoder (403) may include hardware, software, or a combination thereof to implement or implement aspects of the disclosed subject matter as described in more detail below. The encoded video data (404) (or the encoded video bitstream (404)) is depicted as a thin line to emphasize the lower amount of data when compared to the video picture stream (402), and the encoded video data (404) (or the encoded video bitstream (404)) can be stored on the streaming server (405) for future use. One or more streaming client subsystems, such as Figure 4 The client subsystems (406) and (408) in the streaming server (405) can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes the incoming copy (407) of the encoded video data and creates an outgoing video picture stream (411) that can be presented on a display (412) (e.g., a display screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (404), (407) and (409) (e.g., a video bitstream) can be encoded according to certain video encoding / compression standards. Examples of these standards include ITU-T Recommendation H.265. In the example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter can be used in the context of VVC.

[0067] Note that the electronic devices (420) and (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).

[0068] Figure 5 A block diagram of a video decoder (510) according to an embodiment of the present disclosure is shown. The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., receiving circuitry). The video decoder (510) may replace Figure 4 The video decoder (410) in the example is used.

[0069] A receiver (531) may receive one or more encoded video sequences to be decoded by a video decoder (510); in the same or another embodiment, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the other encoded video sequences. The encoded video sequence may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (531) may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective consuming entities (not depicted). The receiver (531) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (515) may be coupled between the receiver (531) and an entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) is part of the video decoder (510). In other applications, the buffer memory (515) may be external to the video decoder (510) (not depicted). In still other applications, there may be a buffer memory (not depicted) external to the video decoder (510) to, for example, prevent network jitter, and in addition there may be another buffer memory (515) internal to the video decoder (510) to, for example, handle playout timing. When the receiver (531) receives data from a store / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (515) may not be needed, or the buffer memory (515) may be small. For use over a best effort packet network such as the Internet, a buffer memory (515) may be needed, which may be relatively large and may advantageously have an adaptive size, and may be implemented at least in part in an operating system or similar element (not depicted) external to the video decoder (510).

[0070] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the encoded video sequence. The categories of these symbols include: information for managing the operation of the video decoder (510), and possibly information for controlling a rendering device such as a rendering device (512) (e.g., a display screen), which is not part of the electronic device (530) but can be coupled to the electronic device (530), such as Figure 5 . The control information of the rendering device may be in the form of a Supplemental Enhancement Information (SEI message) or a Video Usability Information (VUI) parameter set fragment (not depicted). The parser (520) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may comply with a video coding technique or standard and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) may extract a subgroup parameter set for at least one subgroup of the pixel subgroups in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroup may include: a group of pictures (Group of Pictures, GOP), a picture, a tile, a slice, a macroblock, a coding unit (Coding Unit, CU), a block, a transform unit (Transform Unit, TU), a prediction unit (Prediction Unit, PU), etc. The parser (520) may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the coded video sequence.

[0071] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to create symbols (521).

[0072] The reconstruction of the symbol (521) may involve a number of different units depending on the type of coded video picture or portion thereof (e.g., inter- and intra-pictures, inter- and intra-blocks) and other factors. Which units are involved and how they are involved may be controlled by subgroup control information parsed from the coded video sequence by the parser (520). For clarity, such subgroup control information flow between the parser (520) and the following multiple units is not depicted.

[0073] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually subdivided into a number of functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the following functional units.

[0074] The first unit is a sealer / inverse transform unit (551). The sealer / inverse transform unit (551) receives quantized transform coefficients and control information (including which transform to use, block size, quantization factor, quantization scaling matrix, etc.) from the parser (520) as (one or more) symbols (521). The sealer / inverse transform unit (551) can output blocks including sample values, which can be input into the aggregator (555).

[0075] In some cases, the output samples of the scaler / inverse transform (551) may belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates a block of the same size and shape as the block being reconstructed using surrounding reconstructed information obtained from a current picture buffer (558). For example, the current picture buffer (558) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (555) adds the prediction information already generated by the intra-prediction unit (552) to the output sample information as provided by the scaler / inverse transform unit (551) on a per-sample basis.

[0076] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to an inter-frame coded and possibly motion compensated block. In such a case, the motion compensated prediction unit (553) may access the reference picture memory (557) to obtain samples for prediction. After motion compensation of the obtained samples according to the symbols (521) belonging to the block, these samples may be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (in this case referred to as residual samples or residual signal), thereby generating output sample information. The address within the reference picture memory (557) from which the motion compensated prediction unit (553) obtains the predicted samples may be controlled by a motion vector, which may be obtained by the motion compensated prediction unit (553) in the form of a symbol (521), which may have, for example, an X component, a Y component and a reference picture component. Motion compensation may also include interpolation of sample values ​​obtained from the reference picture memory (557) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.

[0077] The output samples of the aggregator (555) may be subjected to various loop filtering techniques in a loop filter unit (556). The video compression techniques may include in-loop filter techniques controlled by parameters included in the coded video sequence (also referred to as a coded video bitstream), which parameters may be obtained by the loop filter unit (556) as symbols (521) from the parser (520), but the in-loop filter techniques may also be responsive to meta-information obtained during decoding of a previous portion (in decoding order) of the coded picture or coded video sequence, and to previously reconstructed and loop filtered sample values.

[0078] The output of the loop filter unit (556) may be a sample stream that may be output to a rendering device (512) and stored in a reference picture memory (557) for use in future inter-picture prediction.

[0079] Once fully reconstructed, certain coded pictures can be used as reference pictures for use in future predictions. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557) and a new current picture buffer can be reallocated before starting to reconstruct a subsequent coded picture.

[0080] The video decoder (510) may perform decoding operations according to a predetermined video compression technology in a standard such as ITU-T H.265 Recommendation. The coded video sequence may conform to the syntax specified by the video compression technology or standard used in the sense that the coded video sequence follows both the syntax of the video compression technology or standard and the profile recorded in the video compression technology or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technology or standard as tools available only under the profile. For compliance, it is also required that the complexity of the coded video sequence is within the limits defined by the level of the video compression technology or standard. In some cases, the level limits the maximum picture size, the maximum frame rate, the maximum reconstruction sample rate (measured in, for example, millions of samples per second), the maximum reference picture size, etc. In some cases, the limits set by the level may be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the coded video sequence.

[0081] In an embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence (one or more). The additional data may be used by the video decoder (510) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0082] Figure 6 A block diagram of a video encoder (603) according to an embodiment of the present disclosure is shown. The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit system). The video encoder (603) can replace Figure 4 The video encoder (403) in the example is used.

[0083] The video encoder (603) can be used to obtain the video source (601) (not Figure 6 In an example of an electronic device (620) receiving video samples, a video source (601) can capture (one or more) video images to be encoded by a video encoder (603). In another example, the video source (601) is a part of the electronic device (620).

[0084] The video source (601) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (603), the digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (601) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (601) may be a camera device that captures local image information as a video sequence. The video data may be provided as a plurality of separate pictures that are given motion when viewed sequentially. The picture itself may be organized as a spatial pixel array, wherein each pixel may include one or more samples, depending on the sampling structure, color space, etc. used. The relationship between pixels and samples may be easily understood by those skilled in the art. The following description focuses on samples.

[0085] According to an embodiment, the video encoder (603) can encode and compress the pictures of the source video sequence into a coded video sequence (643) in real time or under any other time constraints required by the application. Implementing the appropriate encoding speed is a function of the controller (650). In some embodiments, the controller (650) controls other functional units as described below and is functionally coupled to the other functional units. The coupling is not depicted for simplicity. The parameters set by the controller (650) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization technology...), picture size, picture group (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other suitable functions that belong to the video encoder (603) optimized for a specific system design.

[0086] In some embodiments, the video encoder (603) is configured to operate in an encoding loop. As an extremely simplified description, in an example, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and (one or more) reference pictures) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to the way the (remote) decoder creates sample data (because in the video compression techniques considered in the disclosed subject matter, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input to a reference picture memory (634). Since the decoding of the symbol stream produces bit-accurate results that are independent of the decoder location (local or remote), the contents of the reference picture memory (634) are also bit-accurate between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" sample values ​​that are exactly the same as the sample values ​​that the decoder "sees" when using prediction during decoding as reference picture samples. This basic principle of reference picture synchronization (and the resulting offset if synchronization cannot be maintained, e.g. due to channel errors) is also used in some related techniques.

[0087] The operation of the "local" decoder (633) can be combined with the "remote" decoder such as has been described above. Figure 5 The operation of the video decoder (510) described in detail is the same. However, in addition, briefly refer to Figure 5 , since the symbols are available and the encoding of the symbols into a coded video sequence by the entropy encoder (645) and the decoding of the symbols by the parser (520) can be lossless, the entropy decoding portion of the video decoder (510) including the buffer memory (515) and the parser (520) may not be fully implemented in the local decoder (633).

[0088] In an embodiment, any decoder technology other than the parsing / entropy decoding present in the decoder is present in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosed subject matter focuses on the decoder operation. The description of the encoder technology can be simplified because the encoder technology is opposite to the decoder technology described comprehensively. In some aspects, a more detailed description is provided below.

[0089] In some examples, during operation, the source encoder (630) may perform motion compensated predictive coding, which predictively encodes an input picture with reference to one or more previously encoded pictures from a video sequence designated as “reference pictures.” In this manner, the encoding engine (632) encodes the differences between pixel blocks of an input picture and pixel blocks of reference picture(s) that may be selected as prediction reference(s) for the input picture.

[0090] The local video decoder (633) can decode the encoded video data of the picture that can be designated as the reference picture based on the symbol created by the source encoder (630). The operation of the encoding engine (632) can advantageously be a lossy process. When the encoded video data can be decoded at the video decoder ( Figure 6 When the video encoder (603) is decoded at a remote location (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process that may be performed by the video decoder on the reference picture and may cause the reconstructed reference picture to be stored in the reference picture cache (634). In this way, the video encoder (603) may locally store a copy of the reconstructed reference picture that has common content (absent transmission errors) with the reconstructed reference picture that will be obtained by the remote video decoder.

[0091] The predictor (635) may perform a prediction search for the encoding engine (632). That is, for a new picture to be encoded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture motion vectors, block shapes, etc., that may be used as suitable prediction references for the new picture. The predictor (635) may operate on a sample-block-by-pixel-block basis to find a suitable prediction reference. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory (634).

[0092] The controller (650) may manage encoding operations of the source encoder (630), including, for example, setting parameters and sub-group parameters for encoding video data.

[0093] The outputs of all the above-mentioned functional units may be subjected to entropy coding in the entropy encoder (645). The entropy encoder (645) converts the symbols generated by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0094] The transmitter (640) can buffer the encoded video sequence(s) created by the entropy encoder (645) in preparation for transmission via a communication channel (660), which can be a hardware / software link to a storage device storing the encoded video data. The transmitter (640) can combine the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0095] The controller (650) may manage the operation of the video encoder (603). During encoding, the controller (650) may assign a certain encoding picture type to each encoded picture, which may affect the encoding techniques that may be applied to the corresponding picture. For example, a picture may generally be assigned one of the following picture types:

[0096] An intra picture (I picture) may be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures. Those skilled in the art are aware of those variations of I pictures and their corresponding applications and features.

[0097] A predictive picture (P picture) may be a picture that can be encoded and decoded using inter prediction or intra prediction to predict sample values ​​of each block using at most one motion vector and a reference index.

[0098] Bidirectional predictive pictures (B pictures), which can be pictures that can be encoded and decoded using inter-prediction or intra-prediction that uses up to two motion vectors and reference indices to predict sample values ​​for each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstruction of a single block.

[0099] The source picture may typically be spatially subdivided into a number of blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples, respectively), and coded block by block. These blocks may be predictively coded with reference to other (coded) blocks, which are determined by the coding allocation applied to the corresponding picture of the block. For example, blocks of an I picture may be non-predictively coded, or may be predictively coded (spatial prediction or intra-prediction) with reference to coded blocks of the same picture. Pixel blocks of a P picture may be predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B picture may be predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.

[0100] The video encoder (603) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T H.265 Recommendation. In its operation, the video encoder (603) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.

[0101] In an embodiment, the transmitter (640) may transmit additional data along with the encoded video. The source encoder (630) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0102] Video can be captured as multiple source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often simplified as intra-prediction) exploits spatial correlation in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In an example, a particular picture being encoded / decoded (which is referred to as the current picture) is divided into blocks. In the case where a block in the current picture is similar to a reference block in a reference picture that was previously encoded and buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case of using multiple reference pictures, the motion vector may have a third dimension that identifies the reference picture.

[0103] In some embodiments, bidirectional prediction techniques may be used for inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, for example, a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be in the past and future, respectively, in display order). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be predicted by a combination of the first reference block and the second reference block.

[0104] In addition, merge mode techniques can be used in inter-picture prediction to improve coding efficiency.

[0105] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard, a picture in a video picture sequence is divided into coding tree units (CTUs) for compression, and the CTUs in the picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. In general, a CTU includes three coding tree blocks (CTBs), namely a luminance CTB and two chrominance CTBs. Each CTU can be recursively divided into one or more coding units (CUs) in a quadtree. For example, a 64×64 pixel CTU can be divided into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In the example, each CU is analyzed to determine a prediction type for the CU, such as an inter-prediction type or an intra-prediction type. Depending on temporal and / or spatial predictability, the CU is divided into one or more prediction units (PUs). Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In an embodiment, the prediction operation in the codec (encoding / decoding) is performed in units of prediction blocks. Using the luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values ​​(e.g., luma values) such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0106] Figure 7 A diagram of a video encoder (703) according to another embodiment of the present disclosure is shown. The video encoder (703) is configured to receive a processed block (e.g., a prediction block) of sample values ​​within a current video picture in a sequence of video pictures, and encode the processed block into an encoded picture that is part of an encoded video sequence. In an example, the video encoder (703) replaces Figure 4 The video encoder (403) in the example is used.

[0107] In the HEVC example, the video encoder (703) receives a matrix of sample values ​​of a processing block, such as a prediction block of 8×8 samples. The video encoder (703) uses, for example, rate-distortion optimization to determine whether to best encode the processing block using intra mode, inter mode, or bidirectional prediction mode. In the case where the processing block is to be encoded in intra mode, the video encoder (703) may encode the processing block into a coded picture using intra prediction techniques; and in the case where the processing block is to be encoded in inter mode or bidirectional prediction mode, the video encoder (703) may encode the processing block into a coded picture using inter prediction or bidirectional prediction techniques, respectively. In some video coding techniques, the merge mode may be an inter-picture prediction submode, where the motion vector is derived from the predictor without the aid of an encoded motion vector component external to one or more motion vector predictors. In some other video coding techniques, there may be a motion vector component applicable to the object block. In the example, the video encoder (703) includes other components, such as a mode decision module (not shown) that determines the mode of the processing block.

[0108] exist Figure 7 In the example of FIG. 7 , the video encoder ( 703 ) includes Figure 7 An inter-frame encoder (730), an intra-frame encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), an overall controller (721), and an entropy encoder (725) are shown coupled together.

[0109] The inter-frame encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a subsequent picture), generate inter-frame prediction information (e.g., description of redundant information, motion vectors, merge mode information according to an inter-frame coding technique), and calculate an inter-frame prediction result (e.g., a prediction block) based on the inter-frame prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture that is decoded based on the encoded video information.

[0110] The intra encoder (722) is configured to: receive samples of a current block (e.g., a processing block); compare the block with an already encoded block in the same picture in some cases; generate quantization coefficients after transformation; and in some cases also generate intra prediction information (e.g., generate intra prediction direction information according to one or more intra coding techniques). In an example, the intra encoder (722) also calculates an intra prediction result (e.g., a prediction block) based on the intra prediction information and a reference block in the same picture.

[0111] The overall controller (721) is configured to determine overall control data and control other components of the video encoder (703) based on the overall control data. In an example, the overall controller (721) determines a mode of a block and provides a control signal to a switch (726) based on the mode. For example, when the mode is an intra-frame mode, the overall controller (721) controls the switch (726) to select an intra-frame mode result for use by a residual calculator (723), and controls the entropy encoder (725) to select intra-frame prediction information and include the intra-frame prediction information in a bitstream; and when the mode is an inter-frame mode, the overall controller (721) controls the switch (726) to select an inter-frame prediction result for use by a residual calculator (723), and controls the entropy encoder (725) to select inter-frame prediction information and include the inter-frame prediction information in a bitstream.

[0112] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) is configured to operate based on the residual data to encode the residual data to generate a transform coefficient. In an example, the residual encoder (724) is configured to convert the residual data from the spatial domain to the frequency domain and generate a transform coefficient. Then, the transform coefficient is subjected to quantization to obtain a quantized transform coefficient. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be used appropriately by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and the intra prediction information. In some examples, the decoded blocks are processed appropriately to generate decoded pictures, and these decoded pictures can be buffered in memory circuits (not shown) and used as reference pictures.

[0113] The entropy encoder (725) is configured to format the bitstream to include the coded blocks. The entropy encoder (725) is configured to include various information according to a suitable standard, such as the HEVC standard. In an example, the entropy encoder (725) is configured to include overall control data, selected prediction information (e.g., intra-frame prediction information or inter-frame prediction information), residual information, and other suitable information in the bitstream. Note that according to the disclosed subject matter, when the block is encoded in the merge sub-mode of the inter-frame mode or the bidirectional prediction mode, there is no residual information.

[0114] Figure 8A diagram of a video decoder (810) according to another embodiment of the present disclosure is shown. The video decoder (810) is configured to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In an example, the video decoder (810) replaces Figure 4 The video decoder (410) in the example uses.

[0115] exist Figure 8 In the example of FIG. 8 , the video decoder ( 810 ) includes Figure 8 An entropy decoder (871), an inter-frame decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-frame decoder (872) are shown coupled together.

[0116] The entropy decoder (871) can be configured to reconstruct certain symbols from the coded picture, which represent the syntax elements constituting the coded picture. Such symbols may include, for example, a mode for encoding a block (e.g., an intra-frame mode, an inter-frame mode, a bidirectional prediction mode, a merged sub-mode of the latter two, or another sub-mode), prediction information (e.g., intra-frame prediction information or inter-frame prediction information) that can identify certain samples or metadata for prediction by an intra-frame decoder (872) or an inter-frame decoder (880), for example, residual information in the form of quantized transform coefficients, etc. In an example, when the prediction mode is an inter-frame mode or a bidirectional prediction mode, the inter-frame prediction information is provided to the inter-frame decoder (880); and when the prediction type is an intra-frame prediction type, the intra-frame prediction information is provided to the intra-frame decoder (872). The residual information can be subjected to inverse quantization and provided to the residual decoder (873).

[0117] The inter-frame decoder (880) is configured to receive the inter-frame prediction information and generate an inter-frame prediction result based on the inter-frame prediction information.

[0118] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.

[0119] The residual decoder (873) is configured to perform inverse quantization to extract dequantized transform coefficients, and process the dequantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (to include quantizer parameters (Quantizer Parameter, QP)), and the information may be provided by the entropy decoder (871) (data path is not depicted because this may only be a small amount of control information).

[0120] The reconstruction module (874) is configured to combine the residual output by the residual decoder (873) with the prediction result (output by the inter-frame prediction module or the intra-frame prediction module as appropriate) in the spatial domain to form a reconstructed block, which can be part of a reconstructed picture, which can be part of a reconstructed video. Note that other suitable operations such as deblocking operations can be performed to improve visual quality.

[0121] Note that the video encoders (403), (603) and (703) and the video decoders (410), (510) and (810) may be implemented using any suitable technology. In an embodiment, the video encoders (403), (603) and (703) and the video decoders (410), (510) and (810) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603) and (603) and the video decoders (410), (510) and (810) may be implemented using one or more processors that execute software instructions.

[0122] The present disclosure describes video coding techniques related to neural image compression techniques and / or neural video compression techniques, such as neural image compression (NIC) based on artificial intelligence (AI). Aspects of the present disclosure include content adaptive online training in NIC, such as NIC methods for image coding frameworks based on end-to-end (E2E) optimization of neural networks. Neural networks (NNs) may include artificial neural networks (ANNs), such as deep neural networks (DNNs), convolutional neural networks (CNNs), etc.

[0123] In an embodiment, the related hybrid video codec is difficult to optimize as a whole. For example, the improvement of a single module (e.g., encoder) in the hybrid video codec may not result in a coding gain in overall performance. In the NN-based video coding framework, different modules can be jointly optimized from input to output to improve the final goal (e.g., rate-distortion performance, such as the rate-distortion loss L described in the disclosure) by performing a learning process or a training process (e.g., a machine learning process), thereby producing an end-to-end optimized NIC.

[0124] An exemplary NIC framework or system can be described as follows. The NIC framework can use an input image x as an input to a neural network encoder (e.g., an encoder based on a neural network such as a DNN) to compute a compressed representation (e.g., a compact representation) that can be compacted, for example, for storage and transmission purposes. A neural network decoder (e.g., a decoder based on a neural network such as a DNN) can use a compressed representation As input to reconstruct the output image (also called reconstructed image) In various embodiments, the input image x and the reconstructed image In the spatial domain, and the compressed representation In a domain different from the spatial domain. In some examples, the compressed representation quantized and entropy coded.

[0125] In some examples, the NIC framework can use a variational autoencoder (VAE) structure. In the VAE structure, the neural network encoder can directly take the entire input image x as input to the neural network encoder. The entire input image x can be passed through a set of neural network layers that act as a black box to compute a compressed representation . Compressed representation is the output of the neural network encoder. The neural network decoder can convert the entire compressed representation As input. Compressed representation The reconstructed image can be computed through another set of neural network layers, which act as another black box . Can optimize rate distortion (RD) loss To achieve image reconstruction Distortion loss with a compact representation with a trade-off hyperparameter λ A trade-off between the bit consumption R.

[0126]

[0127] A neural network (e.g., ANN) can learn to perform tasks from examples without being programmed for a specific task. ANN can be configured with connected nodes or artificial neurons. The connection between the nodes can transmit a signal from a first node to a second node (e.g., a receiving node), and the signal can be modified by a weight, which can be indicated by a weight coefficient of the connection. The receiving node can process a signal from a node that transmits the signal to the receiving node (i.e., an input signal of the receiving node) and then generate an output signal by applying a function to the input signal. The function can be a linear function. In the example, the output signal is a weighted sum of the input signals. In the example, the output signal is also modified by a bias that can be indicated by a bias term, so that the output signal is the sum of the weighted sum of the bias and the input signal. The function can include nonlinear operations, such as a weighted sum of the input signal or a sum of a bias and a weighted sum. The output signal can be sent to a node (downstream node) connected to the receiving node. ANN can be represented or configured by parameters (e.g., weights and / or biases of the connection). Weights and / or biases can be obtained by training the ANN using examples that can iteratively adjust weights and / or biases. A trained ANN configured with determined weights and / or determined biases may be used to perform a task.

[0128] The nodes in the ANN can be organized in any suitable architecture. In various embodiments, the nodes in the ANN are organized into layers, including an input layer that receives input signals to the ANN and an output layer that outputs output signals from the ANN. In an embodiment, the ANN also includes layers between the input layer and the output layer, such as hidden layers. Different layers can perform different types of transformations on corresponding inputs of different layers. Signals can be transmitted from the input layer to the output layer.

[0129] An ANN with multiple layers between the input layer and the output layer may be referred to as a DNN. In an embodiment, the DNN is a feed-forward network in which data flows from the input layer to the output layer without loopbacks. In an example, the DNN is a fully connected network in which each node in one layer is connected to all nodes in the next layer. In an embodiment, the DNN is a recurrent neural network (RNN) in which data can flow in any direction. In an embodiment, the DNN is a CNN.

[0130] A CNN may include an input layer, an output layer, and a hidden layer between the input layer and the output layer. The hidden layer may include a convolution layer (e.g., used in an encoder) that performs convolutions such as two-dimensional (2D) convolutions. In an embodiment, the 2D convolution performed in the convolution layer is between a convolution kernel (also referred to as a filter or channel, such as a 5×5 matrix) and an input signal to the convolution layer (e.g., a 2D matrix, such as a 2D image, a 256×256 matrix). In various examples, the dimension of the convolution kernel (e.g., 5×5) is smaller than the dimension of the input signal (e.g., 256×256). Therefore, the portion of the input signal (e.g., a 256×256 matrix) covered by the convolution kernel (e.g., a 5×5 area) is smaller than the area of ​​the input signal (e.g., a 256×256 area), and can therefore be referred to as a receptive field in a corresponding node of the next layer.

[0131] During convolution, the dot product of the convolution kernel and the corresponding receptive field in the input signal is calculated. Therefore, each element of the convolution kernel is a weight applied to the corresponding sample in the receptive field, and the convolution kernel includes weights. For example, a convolution kernel represented by a 5×5 matrix has 25 weights. In some examples, a bias is applied to the output signal of the convolution layer, and the output signal is based on the sum of the dot product and the bias.

[0132] The convolution kernel can be moved along the input signal (e.g., a 2D matrix) by an amount called the stride, so that the convolution operation generates a feature map or activation map (e.g., another 2D matrix), which in turn contributes to the input of the next layer in the CNN. For example, the input signal is a 2D image with 256×256 samples, and the stride is 2 samples (e.g., a stride of 2). For a stride of 2, the convolution kernel is moved by 2 samples along the X direction (e.g., horizontal direction) and / or the Y direction (e.g., vertical direction).

[0133] Multiple convolution kernels can be applied to the input signal in the same convolution layer to generate multiple feature maps, respectively, where each feature map can represent a specific feature of the input signal. In general, a convolution layer has N channels (i.e., N convolution kernels), each convolution kernel has M×M samples, and a stride S can be specified as Conv:MxM cN sS. For example, a convolution layer has 192 channels, each convolution kernel has 5×5 samples, and a stride of 2 is specified as Conv:5x5 c192 s2. The hidden layer can include a deconvolution layer (e.g., used in a decoder) that performs deconvolution (e.g., 2D deconvolution). Deconvolution is the inverse of convolution. A deconvolution layer has 192 channels, each deconvolution kernel has 5×5 samples, and a stride of 2 is specified as DeConv:5x5 c192 s2.

[0134] In various embodiments, CNN has the following benefits. Many learnable parameters (i.e., parameters to be trained) in CNN can be significantly smaller than many learnable parameters in DNN (e.g., feedforward DNN). In CNN, a relatively large number of nodes can share the same filter (e.g., the same weight) and the same bias (if bias is used), so memory usage can be reduced because a single vector of biases and weights can be used between all receptive domains sharing the same filter. For example, for an input signal with 100×100 samples, a convolutional layer with a convolution kernel of 5×5 samples has 25 learnable parameters (e.g., weights). If bias is used, one channel uses 26 learnable parameters (e.g., 25 weights and one bias). If the convolutional layer has N channels, the total learnable parameters are 26×N. On the other hand, for a fully connected layer in a DNN, 100×100 (i.e., 10,000) weights are used for each node in the next layer. If the next layer has L nodes, the total learnable parameters are 10,000×L.

[0135] CNN may also include one or more other layers, such as pooling layers, fully connected layers that can connect each node in one layer to each node in another layer, normalization layers, etc. The layers in a CNN may be arranged in any suitable order and in any suitable architecture (e.g., feed-forward architecture, recurrent architecture). In the example, the convolutional layer is followed by other layers, such as pooling layers, fully connected layers, normalization layers, etc.

[0136] A pooling layer can be used to reduce the dimensionality of data by combining the outputs from multiple nodes in one layer into a single node in the next layer. The following describes the pooling operation of a pooling layer with a feature map as input. The description can be appropriately applied to other input signals. The feature map can be divided into sub-regions (e.g., rectangular sub-regions), and the features in each sub-region can be independently downsampled (or pooled) to a single value, for example by taking the average in average pooling or the maximum in max pooling.

[0137] The pooling layer can perform pooling, such as local pooling, global pooling, maximum pooling, average pooling, etc. Pooling is a form of nonlinear downsampling. Local pooling combines a small number of nodes in the feature map (e.g., a local node cluster, such as 2×2 nodes). Global pooling can combine, for example, all nodes of the feature map.

[0138] Pooling layers can reduce the size of representations, thereby reducing the number of parameters, memory usage, and computation in CNNs. In the example, pooling layers are inserted between consecutive convolutional layers in a CNN. In the example, the pooling layer is followed by an activation function, such as a rectified linear unit (ReLU) layer. In the example, pooling layers are omitted between consecutive convolutional layers in a CNN.

[0139] The normalization layer can be a ReLU, a leaky ReLU, a generalized division normalization (GDN), an inverse GDN (IGDN), etc. The ReLU can apply a non-saturating activation function to remove negative values ​​from an input signal (e.g., a feature map) by setting negative values ​​to zero. For negative values, the leaky ReLU can have a small slope (e.g., 0.01) instead of a flat slope (e.g., 0). Therefore, if the value x is greater than 0, the output from the leaky ReLU is x. Otherwise, the output from the leaky ReLU is the value x multiplied by a small slope (e.g., 0.01). In the example, the slope is determined before training and is therefore not learned during training.

[0140] Fig. 9 An exemplary NIC framework (900) (e.g., NIC system) according to an embodiment of the present disclosure is shown. The NIC framework (900) can be based on a neural network, such as a DNN and / or a CNN. The NIC framework (900) can be used to compress (e.g., encode) an image and decompress (e.g., decode or reconstruct) a compressed image (e.g., an encoded image). The NIC framework (900) can include two sub-neural networks implemented using a neural network, a first sub-NN (951) and a second sub-NN (952).

[0141] The first sub-NN (951) can be similar to an autoencoder and can be trained to generate a compressed image of the input image x And for the compressed image Decompress to get the reconstructed image The first sub-NN (951) may include a plurality of components (or modules), such as a main encoder neural network (or main encoder network) (911), a quantizer (912), an entropy encoder (913), an entropy decoder (914), and a main decoder neural network (or main decoder network) (915). Fig. 9 , the main encoder network (911) can generate a potential or latent representation y from an input image x (e.g., an image to be compressed or encoded). In an example, a CNN is used to implement the main encoder network (911). The relationship between the potential representation y and the input image x can be described using Equation 2.

[0142] y=f1(x;θ1) Formula 2

[0143] Here, the parameter θ1 represents the following parameters: such as the weights used in the convolution kernel in the main encoder network (911) and the bias (if a bias is used in the main encoder network (911)).

[0144] The latent representation y may be quantized using a quantizer (912) to generate a quantized latent representation For example, the quantized latent Compress to generate a compressed representation of the input image x A compressed image (eg, an encoded image) (931). The entropy encoder (913) may use an entropy encoding technique such as Huffman encoding, arithmetic encoding, etc. In an example, the entropy encoder (913) uses arithmetic encoding and is an arithmetic encoder. In an example, the encoded image (931) is transmitted in an encoded bitstream.

[0145] The coded image (931) can be decompressed (e.g., entropy decoded) by an entropy decoder (914) to generate an output. The entropy decoder (914) can use an entropy coding technique such as Huffman coding, arithmetic coding, etc. corresponding to the entropy coding technique used in the entropy encoder (913). In an example, the entropy decoder (914) uses arithmetic decoding and is an arithmetic decoder. In an example, lossless compression is used in the entropy encoder (913), lossless decompression is used in the entropy decoder (914), and noise such as that generated by the transmission of the coded image (931) can be ignored, and the output from the entropy decoder (914) is a quantized potential .

[0146] The main decoder network (915) can decode the quantized latent Decode to generate the reconstructed image In an example, the main decoder network (915) is implemented using a CNN. Reconstructing the image (i.e., the output of the main decoder network (915)) and the quantized potential The relationship between (i.e., the input of the main decoder network (915)) can be described using Equation 3.

[0147]

[0148] Wherein, parameter θ2 represents the following parameters: such as weights used in the convolution kernel in the main decoder network (915) and bias (if bias is used in the main decoder network (915)). Therefore, the first sub-NN (951) can compress (e.g., encode) the input image x to obtain an encoded image (931), and decompress (e.g., decode) the encoded image (931) to obtain a reconstructed image Due to the quantization loss introduced by the quantizer (912), the reconstructed image may be different from the input image x.

[0149] The second sub-NN (952) can be used for entropy coding of the quantized potential The entropy model (e.g., a prior probability model) is learned. Therefore, the entropy model can be a conditional entropy model that depends on the input image x, such as a Gaussian mixture model (GMM), a Gaussian scale model (GSM). The second sub-NN (952) can include a context model NN (916), an entropy parameter NN (917), a super encoder (921), a quantizer (922), an entropy encoder (923), an entropy decoder (924), and a super decoder (925). The entropy model used in the context model NN (916) can be about the potential (e.g., the quantized potential ) is an autoregressive model. In an example, a super encoder (921), a quantizer (922), an entropy encoder (923), an entropy decoder (924), and a super decoder (925) form a super neural network (e.g., a super prior NN). The super neural network can represent information useful for correcting context-based predictions. Data from the context model NN (916) and the super neural network can be combined by an entropy parameter NN (917). The entropy parameter NN (917) can generate parameters such as mean and scale parameters for an entropy model such as a conditional Gaussian entropy model (e.g., GMM).

[0150] Reference Fig. 9 , at the encoder side, the quantized potential from the quantizer (912) is fed into the context model NN (916). On the decoder side, the quantized potential from the entropy decoder (914) is fed into the context model NN (916). The context model NN (916) can be implemented using a neural network such as a CNN. The context model NN (916) can be based on the context Generate output o cm,i , the context is the quantized potential available to the context model NN (916). Context The output of the context model NN (916) may include a previously quantized potential at the encoder side or a previously entropy decoded quantized potential at the decoder side. cm,i With input (for example, ) can be described using Formula 4.

[0151]

[0152] Here, parameter θ3 represents parameters such as the weights used in the convolution kernel in the context model NN (916) and the bias (if a bias is used in the context model NN (916)).

[0153] Output o from context model NN (916) cm,i and the output o from the super decoder (925) hc is fed to the entropy parameter NN (917) to generate the output o ep The entropy parameter NN (917) may be implemented using a neural network such as a CNN. The output of the entropy parameter NN (917) is ep With input (for example, o cm,i and hc ) can be described using Formula 5.

[0154] o ep =f4(o cm,i , o hc ;θ4) Formula 5

[0155] Wherein the parameter θ4 represents the following parameters, such as the weights used in the convolution kernel in the entropy parameter NN (917) and the bias (if a bias is used in the entropy parameter NN (917)). The output o of the entropy parameter NN (917) ep can be used to determine (eg, adjust) an entropy model, and thus, the adjusted entropy model can be used, for example, via an output from the super decoder (925). hc It depends on the input image x. In this example, the output o ep Includes parameters for tuning entropy models (e.g., GMM), such as mean and scale parameters. Fig. 9, the entropy encoder (913) and the entropy decoder (914) can use an entropy model (e.g., a conditional entropy model) in entropy encoding and entropy decoding, respectively.

[0156] The second sub-NN (952) can be described below. Potential y can be fed into a super encoder (921) to generate a super potential z. In an example, the super encoder (921) is implemented using a neural network such as a CNN. The relationship between the super potential z and the potential y can be described using Equation 6.

[0157] z=f5(y;θ5) Formula 6

[0158] Wherein, parameter θ5 represents parameters such as the weights used in the convolution kernel in the super encoder (921) and the bias (if a bias is used in the super encoder (921)).

[0159] The super potential z is quantized by a quantizer (922) to generate a quantized potential For example, the quantized latent Compression is performed to generate side information such as coded bits (932) from a hyperneural network. The entropy encoder (923) can use entropy coding techniques such as Huffman coding, arithmetic coding, etc. In an example, the entropy encoder (923) uses arithmetic coding and is an arithmetic encoder. In an example, the side information such as the coded bits (932) can be transmitted in a coded bit stream, for example, with the coded image (931).

[0160] The side information, such as the coded bits (932), may be decompressed (e.g., entropy decoded) by an entropy decoder (924) to generate an output. The entropy decoder (924) may use an entropy coding technique, such as Huffman coding, arithmetic coding, etc. In an example, the entropy decoder (924) uses arithmetic decoding and is an arithmetic decoder. In an example, lossless compression is used in the entropy encoder (923), lossless decompression is used in the entropy decoder (924), and noise, such as due to the transmission of the side information, may be ignored. The output from the entropy decoder (924) may be a quantized potential. The super decoder (925) can process the quantized latent Decode to generate output o hc Output hc With quantified potential The relationship between can be described using Equation 7.

[0161]

[0162] Wherein, parameter θ6 represents parameters such as the weights used in the convolution kernel in the super decoder (925) and the bias (if a bias is used in the super decoder (925)).

[0163] As described above, compressed bits or coded bits (932) can be added to the coded bitstream as side information, which enables the entropy decoder (914) to use a conditional entropy model. Thus, the entropy model can be image-dependent and spatially adaptive, and can therefore be more accurate than a fixed entropy model.

[0164] The NIC framework (900) may be appropriately adjusted to, for example, omit Fig. 9 One or more components shown in Fig. 9 One or more components shown in and / or including Fig. 9 In the example, the NIC framework using the fixed entropy model includes the first sub-NN (951) but does not include the second sub-NN (952). In the example, the NIC framework includes the components of the NIC framework (900) except the entropy encoder (923) and the entropy decoder (924).

[0165] In an embodiment, a neural network such as a CNN is used to implement Fig. 9 One or more components in the NIC framework (900) shown. Each NN-based component (e.g., main encoder network (911), main decoder network (915), context model NN (916), entropy parameter NN (917), super encoder (921) or super decoder (925)) in a NIC framework (e.g., NIC framework (900)) can include any suitable architecture (e.g., having any suitable combination of layers), include any suitable type of parameters (e.g., weights, biases, a combination of weights and biases, etc.) and include any suitable number of parameters.

[0166] In an embodiment, the main encoder network (911), the main decoder network (915), the context model NN (916), the entropy parameter NN (917), the super encoder (921) and the super decoder (925) are implemented using corresponding CNNs.

[0167] Fig.10 An exemplary CNN of a main encoder network (911) according to an embodiment of the present disclosure is shown. For example, the main encoder network (911) includes four groups of layers, where each group of layers includes a convolutional layer 5×5 c192 s2 followed by a GDN layer. Fig.10 One or more of the layers shown in may be modified and / or omitted. Additional layers may be added to the main encoder network (911).

[0168] Fig.11 An exemplary CNN of the main decoder network (915) according to an embodiment of the present disclosure is shown. For example, the main decoder network (915) includes three groups of layers, where each group of layers includes a deconvolution layer 5x5 c192s2 followed by an IGDN layer. In addition, the three groups of layers are followed by a deconvolution layer 5x5 c3 s2 followed by an IGDN layer. Fig.11 One or more of the layers shown in may be modified and / or omitted. Additional layers may be added to the main decoder network (915).

[0169] Fig.12 An exemplary CNN of a super encoder (921) according to an embodiment of the present disclosure is shown. For example, the super encoder (921) includes a convolutional layer 3x3 c192 s1 followed by a leaky ReLU, a convolutional layer 5x5 c192 s2 followed by a leaky ReLU, and a convolutional layer 5x5 c192 s2. Fig.12 One or more of the layers shown in may be modified and / or omitted. Additional layers may be added to the super encoder (921).

[0170] Fig.13 An exemplary CNN of a super decoder (925) according to an embodiment of the present disclosure is shown. For example, the super decoder (925) includes a deconvolution layer 5x5 c192 s2 followed by a leaky ReLU, a deconvolution layer 5x5 c288 s2 followed by a leaky ReLU, and a deconvolution layer 3x3 c384 s1. Fig.13 One or more of the layers shown in may be modified and / or omitted. Additional layers may be added to the super encoder (925).

[0171] Fig.14 An exemplary CNN of a context model NN (916) according to an embodiment of the present disclosure is shown. For example, the context model NN (916) includes a masked convolution 5x5 c384 s1 for context prediction, and thus the context in Equation 4 Include limited context (e.g., 5×5 convolution kernel). Can be modified Fig.14 Additional layers may be added to the context model NN (916).

[0172] Fig.15 An exemplary CNN of an entropy parameter NN (917) according to an embodiment of the present disclosure is shown. For example, the entropy parameter NN (917) includes a convolution layer 1x1 c640 s1 followed by a leaky ReLU, a convolution layer 1x1 c512 s1 followed by a leaky ReLU, and a convolution layer 1x1 c384 s1. The following may be modified and / or omitted. Fig.15Additional layers may be added to the entropy parameter NN (917).

[0173] As reference Figures 10 to 15 As described, the NIC framework (900) can be implemented using a CNN. The NIC framework (900) can be appropriately adjusted so that one or more components (e.g., (911), (915), (916), (917), (921), and / or (925)) in the NIC framework (900) are implemented using any appropriate type of neural network (e.g., a CNN-based or non-CNN-based neural network). One or more other components of the NIC framework (900) can be implemented using a neural network.

[0174] The NIC framework (900) including a neural network (e.g., CNN) can be trained to learn parameters used in the neural network. For example, when using CNN, parameters represented by θ1 to θ6, such as weights and biases used in convolution kernels in the main encoder network (911) (if biases are used in the main encoder network (911)), weights and biases used in convolution kernels in the main decoder network (915) (if biases are used in the main decoder network (915)), weights and biases used in convolution kernels in the super encoder (921) (if biases are used in the super encoder (921)), weights and biases used in convolution kernels in the super decoder (925) (if biases are used in the super decoder (925)), weights and biases used in convolution kernels in the context model NN (916) (if biases are used in the context model NN (916)), and weights and biases used in convolution kernels in the entropy parameter NN (917) (if biases are used in the entropy parameter NN (917)), can be learned separately during the training process.

[0175] In the example, refer to Fig.10 , the main encoder network (911) includes four convolutional layers, each of which has a 5×5 convolution kernel and 192 channels. Therefore, the number of weights used in the convolution kernels in the main encoder network (911) is 19200 (i.e., 4×5×5×192). The parameters used in the main encoder network (911) include 19200 weights and an optional bias. When a bias and / or additional NN is used in the main encoder network (911), additional parameters may be included.

[0176] Reference Fig. 9, the NIC framework (900) includes at least one component or module built on a neural network. At least one component may include one or more of a main encoder network (911), a main decoder network (915), a super encoder (921), a super decoder (925), a context model NN (916), and an entropy parameter NN (917). At least one component can be trained separately. In an example, a training process is used to learn the parameters of each component separately. At least one component can be jointly trained as a group. In an example, a training process is used to jointly learn the parameters of a subset of at least one component. In an example, a training process is used to learn the parameters of all at least one component, and is therefore referred to as E2E optimization.

[0177] During the training process of one or more components in the NIC framework (900), the weights (or weight coefficients) of the one or more components may be initialized. In an example, the weights are initialized based on a pre-trained corresponding neural network model (e.g., a DNN model, a CNN model). In an example, the weights are initialized by setting the weights to random numbers.

[0178] For example, after initializing the weights, a set of training images can be used to train one or more components. The set of training images can include any suitable image of any suitable size. In some examples, the set of training images includes original images, natural images, computer-generated images, etc. in the spatial domain. In some examples, the set of training images includes residual images with residual data in the spatial domain. The residual data can be calculated by a residual calculator (e.g., residual calculator (723)). In some examples, the training images in the set of training images (e.g., original images and / or residual images including residual data) can be divided into blocks of suitable size, and these blocks and / or images can be used to train the neural network in the NIC framework. Therefore, the original image, the residual image, the block from the original image and / or the block from the residual image can be used to train the neural network in the NIC framework.

[0179] For the sake of brevity, the following training process is described using training images as an example. The description can be appropriately adapted to the training images. A training image t in the set of training images can be obtained by Fig. 9 The encoding process in the to generate a compressed representation (e.g., such as encoded information to a bitstream). The encoded information can be obtained by Fig. 9 The decoding process described in is used to compute and reconstruct the reconstructed image .

[0180] For the NIC framework (900), two competing objectives such as reconstruction quality and bit consumption are balanced. Quality loss function (e.g., distortion or distortion loss) can be used to indicate reconstruction quality, such as reconstruction (e.g., reconstructed image ) and the original image (e.g., training image t). The rate (or rate loss) R can be used to indicate the bit consumption of the compressed representation. In an example, the rate loss R also includes side information used, for example, in determining the context model.

[0181] For neural image compression, a differentiable approximation of quantization can be used in E2E optimization. In various examples, during the training process of neural network-based image compression, noise injection is used to simulate quantization, and therefore, quantization is simulated by noise injection rather than performed by a quantizer (e.g., quantizer (912)). Therefore, training using noise injection can approximate the quantization error in a varying manner. A bits per pixel (BPP) estimator can be used to simulate an entropy encoder, and therefore, entropy encoding is simulated by a BPP estimator rather than performed by an entropy encoder (e.g., (913)) and an entropy decoder (e.g., (914)). Therefore, the rate loss R in the loss function L shown in Equation 1 during training can be estimated, for example, based on noise injection and a BPP estimator. Typically, a higher rate R can allow lower distortion D, while a lower rate R can result in higher distortion D. Therefore, the trade-off hyperparameter λ in Equation 1 can be used to optimize a joint RD loss L, where L, which is the sum of λD and R, can be optimized. The training process can be used to adjust the parameters of one or more components (e.g., (911)(915)) in the NIC framework (900) so that the joint RD loss L is minimized or optimized.

[0182] Various models can be used to determine the distortion loss D and the rate loss R, and thus determine the joint RD loss L in Equation 1. In the example, the distortion loss It is expressed as Peak Signal-to-Noise Ratio (PSNR), which is a metric based on mean square error, multi-scale structural similarity (MS-SSIM) quality index, a weighted combination of PSNR and MS-SSIM, etc.

[0183] In an example, the goal of the training process is to train an encoding neural network (e.g., encoding DNN) such as a video encoder to be used on the encoder side and a decoding neural network (e.g., decoding DNN) such as a video decoder to be used on the decoder side. Fig. 9 , the encoding neural network may include a main encoder network (911), a super encoder (921), a super decoder (925), a context model NN (916), and an entropy parameter NN (917). The decoding neural network may include a main decoder network (915), a super decoder (925), a context model NN (916), and an entropy parameter NN (917). The video encoder and / or video decoder may include other components based on NN and / or not based on NN.

[0184] The NIC framework (e.g., NIC framework (900)) can be trained in an E2E manner. In an example, the encoding neural network and the decoding neural network are jointly updated in an E2E manner based on back-propagation gradients during training.

[0185] After the parameters of the neural network in the NIC framework (900) are trained, one or more components in the NIC framework (900) can be used to encode and / or decode an image. In an embodiment, on the encoder side, the video encoder is configured to encode the input image x into an encoded image (931) to be transmitted in a bitstream. The video encoder may include multiple components in the NIC framework (900). In an embodiment, on the decoder side, the corresponding video decoder is configured to decode the encoded image (931) in the bitstream into a reconstructed image The video decoder may include multiple components in the NIC framework (900).

[0186] In an example, for example, when content-adaptive online training is employed, the video encoder includes all components in the NIC framework (900).

[0187] Fig.16A An exemplary video encoder (1600A) according to an embodiment of the present disclosure is shown. The video encoder (1600A) includes a reference Fig. 9 The main encoder network (911), quantizer (912), entropy encoder (913) and second sub-NN (952) are described, and the detailed description is omitted for the purpose of simplicity. Fig. 16B An exemplary video decoder (1600B) according to an embodiment of the present disclosure is shown. The video decoder (1600B) may correspond to the video encoder (1600A). The video decoder (1600B) may include a main decoder network (915), an entropy decoder (914), a context model NN (916), an entropy parameter NN (917), an entropy decoder (924), and a super decoder (925). FIG. 16A to FIG. 16B On the encoder side, the video encoder (1600A) can generate a coded image (931) and coded bits (932) to be transmitted in a bitstream. On the decoder side, the video decoder (1600B) can receive the coded image (931) and coded bits (932) and decode the coded image (931) and coded bits (932).

[0188] Figure 17 to Figure 18 An exemplary video encoder (1700) and a corresponding video decoder (1800) according to an embodiment of the present disclosure are shown respectively. Fig.17, the encoder (1700) includes a main encoder network (911), a quantizer (912) and an entropy encoder (913). Fig. 9 Describes an example of a main encoder network (911), a quantizer (912), and an entropy encoder (913). Fig.18 , the video decoder (1800) includes a main decoder network (915) and an entropy decoder (914). Fig. 9 Describes an example of a main decoder network (915) and an entropy decoder (914). Fig.17 and Fig.18 The video encoder (1700) can generate a coded image (931) to be transmitted in a bitstream. The video decoder (1800) can receive the coded image (931) and decode the coded image (931).

[0189] As described above, a NIC framework (900) including a video encoder and a video decoder can be trained based on images and / or blocks in a set of training images. In some examples, one or more images to be compressed (e.g., encoded) and / or transmitted have characteristics that are significantly different from the set of training images. Therefore, encoding and decoding one or more images using a video encoder and a video decoder trained based on the set of training images, respectively, may result in a relatively poor RD loss L (e.g., relatively large distortion and / or relatively large bit rate). Therefore, aspects of the present disclosure describe a content-adaptive online training method for a NIC.

[0190] In order to distinguish between a training process based on a set of training images and a content-adaptive online training process based on one or more images to be compressed (e.g., encoded) and / or transmitted, the NIC framework (900), the video encoder, and the video decoder trained by the set of training images are respectively referred to as the pre-trained NIC framework (900), the pre-trained video encoder, and the pre-trained video decoder. The parameters in the pre-trained NIC framework (900), the parameters in the pre-trained video encoder, or the parameters in the pre-trained video decoder are respectively referred to as NIC pre-trained parameters, encoder pre-trained parameters, and decoder pre-trained parameters. In the example, the NIC pre-trained parameters include encoder pre-trained parameters and decoder pre-trained parameters. In the example, the encoder pre-trained parameters and the decoder pre-trained parameters do not overlap, wherein the encoder pre-trained parameters are not included in the decoder pre-trained parameters. For example, the encoder pre-trained parameters in (1700) (e.g., the pre-trained parameters in the main encoder network (911)) and the decoder pre-trained parameters in (1800) (e.g., the pre-trained parameters in the main decoder network (915)) do not overlap. In an example, the encoder pre-trained parameters and the decoder pre-trained parameters overlap, wherein at least one of the encoder pre-trained parameters is included in the decoder pre-trained parameters. For example, the encoder pre-trained parameters in (1600A) (e.g., pre-trained parameters in the context model NN (916)) and the decoder pre-trained parameters in (1600B) (e.g., pre-trained parameters in the context model NN (916)) overlap. The NIC pre-trained parameters can be obtained based on blocks and / or images in a set of training images.

[0191] The content adaptive online training process may be referred to as a fine-tuning process and is described below. One or more of the NIC pre-trained parameters in the pre-trained NIC framework (900) may be further trained (e.g., fine-tuned) based on one or more images to be encoded and / or transmitted, wherein one or more images may be different from the set of training images. One or more pre-trained parameters used in the NIC pre-trained parameters may be fine-tuned by optimizing the joint RD loss L based on one or more images. One or more pre-trained parameters that have been fine-tuned by one or more images are referred to as one or more replacement parameters or one or more fine-tuning parameters. In an embodiment, after one or more pre-trained parameters in the NIC pre-trained parameters have been fine-tuned (e.g., replaced) by one or more replacement parameters, the neural network update information is encoded into the bitstream to indicate one or more replacement parameters or a subset of one or more replacement parameters. In an example, the NIC framework (900) is updated (or fine-tuned), wherein one or more pre-trained parameters are replaced by one or more replacement parameters, respectively.

[0192] In a first case, the one or more pre-trained parameters include a first subset of the one or more pre-trained parameters and a second subset of the one or more pre-trained parameters. The one or more replacement parameters include a first subset of the one or more replacement parameters and a second subset of the one or more replacement parameters.

[0193] A first subset of one or more pre-trained parameters is used in a pre-trained video encoder and replaced by a first subset of one or more replacement parameters, for example, during a training process. Thus, the pre-trained video encoder is updated to an updated video encoder through the training process. The neural network update information may indicate a second subset of one or more replacement parameters to replace the second subset of one or more replacement parameters. One or more images may be encoded using the updated video encoder and transmitted in a bitstream along with the neural network update information.

[0194] At the decoder side, the second subset of the one or more pre-trained parameters is used in the pre-trained video decoder. In an embodiment, the pre-trained video decoder receives the neural network update information and decodes the neural network update information to determine the second subset of the one or more replacement parameters. When the second subset of the one or more pre-trained parameters in the pre-trained video decoder is replaced by the second subset of the one or more replacement parameters, the pre-trained video decoder is updated to an updated video decoder. The updated video decoder may be used to decode the one or more encoded images.

[0195] FIG. 16A to FIG. 16B An example of the first case is shown. For example, one or more pre-trained parameters include N1 pre-trained parameters in the pre-trained context model NN (916) and N2 pre-trained parameters in the pre-trained main decoder network (915). Therefore, a first subset of one or more pre-trained parameters includes N1 pre-trained parameters, and a second subset of one or more pre-trained parameters is the same as the one or more pre-trained parameters. Therefore, the N1 pre-trained parameters in the pre-trained context model NN (916) can be replaced by N1 corresponding replacement parameters, so that the pre-trained video encoder (1600A) can be updated to the updated video encoder (1600A). The pre-trained context model NN (916) is also updated to the updated context model NN (916). On the decoder side, N1 pre-trained parameters may be replaced by N1 corresponding replacement parameters, and N2 pre-trained parameters may be replaced by N2 corresponding replacement parameters, the pre-trained context model NN (916) is updated to the updated context model NN (916), and the pre-trained main decoder network (915) is updated to the updated main decoder network (915). Therefore, the pre-trained video decoder (1600B) can be updated to the updated video decoder (1600B).

[0196] In a second case, one or more pre-trained parameters are not used in the pre-trained video encoder on the encoder side. Instead, one or more pre-trained parameters are used in the pre-trained video decoder on the decoder side. Therefore, the pre-trained video encoder is not updated and continues to be a pre-trained video encoder after the training process. In an embodiment, the neural network update information indicates one or more replacement parameters. One or more images may be encoded using the pre-trained video encoder and transmitted in the bitstream together with the neural network update information.

[0197] On the decoder side, the pre-trained video decoder may receive the neural network update information and decode the neural network update information to determine one or more replacement parameters. When one or more pre-trained parameters in the pre-trained video decoder are replaced by one or more replacement parameters, the pre-trained video decoder is updated to an updated video decoder. The updated video decoder may be used to decode one or more encoded images.

[0198] FIG. 16A to FIG. 16B An example of the second case is shown. For example, one or more pre-trained parameters include N2 pre-trained parameters in the pre-trained main decoder network (915). Therefore, one or more pre-trained parameters are not used in the pre-trained video encoder on the encoder side (e.g., the pre-trained video encoder (1600A)). Therefore, the pre-trained video encoder (1600A) continues to be a pre-trained video encoder after the training process. On the decoder side, the N2 pre-trained parameters can be replaced by N2 corresponding replacement parameters, which updates the pre-trained main decoder network (915) to the updated main decoder network (915). Therefore, the pre-trained video decoder (1600B) can be updated to the updated video decoder (1600B).

[0199] In a third case, one or more pre-trained parameters are used in a pre-trained video encoder and are replaced by one or more replacement parameters, for example, during a training process. Thus, the pre-trained video encoder is updated to an updated video encoder through a training process. One or more images can be encoded using the updated video encoder and transmitted in a bitstream. No neural network update information is encoded in the bitstream. On the decoder side, the pre-trained video decoder is not updated and remains a pre-trained video decoder. One or more encoded images can be decoded using the pre-trained video decoder.

[0200] FIG. 16A to FIG. 16BAn example of the third case is shown. For example, one or more pre-trained parameters are in the pre-trained main encoder network (911). Therefore, one or more pre-trained parameters in the pre-trained main encoder network (911) can be replaced by one or more replacement parameters, so that the pre-trained video encoder (1600A) can be updated to the updated video encoder (1600A). The pre-trained main encoder network (911) is also updated to the updated main encoder network (911). On the decoder side, the pre-trained video decoder (1600B) is not updated.

[0201] In various examples such as those described in the first case, the second case, and the third case, video decoding may be performed by pre-trained decoders having different capabilities, including decoders with and without the ability to update pre-trained parameters.

[0202] In an example, compression performance can be improved by encoding one or more images using an updated video encoder and / or an updated video decoder compared to encoding one or more images using a pre-trained video encoder and a pre-trained video decoder. Thus, the content-adaptive online training method can be used to adapt a pre-trained NIC framework (e.g., pre-trained NIC framework (900)) to target image content (e.g., one or more images to be transmitted) and thus fine-tune the pre-trained NIC framework. Thus, the video encoder on the encoder side and / or the video decoder on the decoder side can be updated.

[0203] The content-adaptive online training method can be used as a pre-processing step (eg, a pre-encoding step) for improving the compression performance of a pre-trained E2E NIC compression method.

[0204] In an embodiment, the one or more images include a single input image, and the fine-tuning process is performed on the single input image. The NIC framework (900) is trained and updated (e.g., fine-tuned) based on the single input image. The updated video encoder on the encoder side and / or the updated video decoder on the decoder side can be used to encode the single input image and optionally other input images. The neural network update information can be encoded into the bitstream together with the encoded single input image.

[0205] In an embodiment, the one or more images include multiple input images, and the fine-tuning process is performed on the multiple input images. The NIC framework (900) is trained and updated (e.g., fine-tuned) based on the multiple input images. The updated video encoder on the encoder side and / or the updated decoder on the decoder side can be used to encode the multiple input images and optionally other input images. The neural network update information can be encoded into the bitstream together with the encoded multiple input images.

[0206] The rate loss R may increase as the neural network update information is signaled in the bitstream. When the one or more images include a single input image, the neural network update information is signaled for each encoded image, and a first increase in the rate loss R is used to indicate an increase in the rate loss R due to the signaling of the neural network update information for each image. When the one or more images include multiple input images, the neural network update information is signaled for the multiple input images and shared by the multiple input images, and a second increase in the rate loss R is used to indicate an increase in the rate loss R due to the signaling of the neural network update information for each image. Because the neural network update information is shared by multiple input images, the second increase in the rate loss R may be less than the first increase in the rate loss R. Therefore, in some examples, it may be advantageous to fine-tune the NIC framework using multiple input images.

[0207] In an embodiment, the one or more pre-trained parameters to be updated are in one component of the pre-trained NIC framework (900). Thus, one component of the pre-trained NIC framework (900) is updated based on the one or more replacement parameters, and other components of the pre-trained NIC framework (900) are not updated.

[0208] A component may be a pre-trained context model NN (916), a pre-trained entropy parameter NN (917), a pre-trained main encoder network (911), a pre-trained main decoder network (915), a pre-trained super encoder (921), or a pre-trained super decoder (925). The pre-trained video encoder and / or the pre-trained video decoder may be updated based on which of the components in the pre-trained NIC framework (900) are updated.

[0209] In the example, one or more pre-trained parameters to be updated are in the pre-trained context model NN (916), and thus the pre-trained context model NN (916) is updated while the remaining components (911), (915), (921), (917) and (925) are not updated. In the example, the pre-trained video encoder on the encoder side and the pre-trained video decoder on the decoder side include the pre-trained context model NN (916), and thus both the pre-trained video encoder and the pre-trained video decoder are updated.

[0210] In the example, one or more pre-trained parameters to be updated are in the pre-trained super decoder (925), and thus the pre-trained super decoder (925) is updated while the remaining components (911), (915), (916), (917) and (921) are not updated. Therefore, the pre-trained video encoder is not updated, while the pre-trained video decoder is updated.

[0211] In an embodiment, one or more pre-trained parameters to be updated are in multiple components of the pre-trained NIC framework (900). Therefore, multiple components of the pre-trained NIC framework (900) are updated based on one or more replacement parameters. In an example, multiple components of the pre-trained NIC framework (900) include all components configured with a neural network (e.g., DNN, CNN). In an example, multiple components of the pre-trained NIC framework (900) include CNN-based components: a pre-trained main encoder network (911), a pre-trained main decoder network (915), a pre-trained context model NN (916), a pre-trained entropy parameter NN (917), a pre-trained super encoder (921), and a pre-trained super decoder (925).

[0212] As described above, in an example, one or more pre-trained parameters to be updated are in a pre-trained video encoder of a pre-trained NIC framework (900). In an example, one or more pre-trained parameters to be updated are in a pre-trained video decoder of a NIC framework (900). In an example, one or more pre-trained parameters to be updated are in a pre-trained video encoder and a pre-trained video decoder of a pre-trained NIC framework (900).

[0213] The NIC framework (900) may be based on a neural network, for example, one or more components in the NIC framework (900) may include a neural network, such as a CNN, a DNN, etc. As described above, a neural network may be specified by different types of parameters, such as weights, biases, etc. Each neural network-based component in the NIC framework (900) (e.g., a context model NN (916), an entropy parameter NN (917), a main encoder network (911), a main decoder network (915), a super encoder (921), or a super decoder (925)) may be configured with appropriate parameters, such as corresponding weights, biases, or a combination of weights and biases. When (one or more) CNNs are used, the weights may include elements in a convolution kernel. One or more types of parameters may be used to specify a neural network. In an embodiment, one or more pre-trained parameters to be updated are (one or more) bias terms, and only (one or more) bias terms are replaced by one or more replacement parameters. In an embodiment, one or more pre-trained parameters to be updated are weights, and only weights are replaced by one or more replacement parameters. In an embodiment, the one or more pre-trained parameters to be updated include weights and (one or more) bias terms, and all pre-trained parameters including weights and (one or more) bias terms are replaced by one or more replacement parameters. In an embodiment, other parameters may be used to specify the neural network, and other parameters may be fine-tuned.

[0214] The fine-tuning process may include multiple stages (e.g., iterations), wherein one or more pre-trained parameters are updated during the iterative fine-tuning process. The process may stop when the training loss has been or will be unchanged. In an example, the fine-tuning process stops when the training loss (e.g., RD ​​loss L) is below a first threshold. In an example, the fine-tuning process stops when the difference between two consecutive training losses is below a second threshold.

[0215] Along with the loss function (e.g., RD ​​loss L), two hyperparameters (e.g., step size and maximum number of steps) can be used in the fine-tuning process. The maximum number of iterations can be used as a threshold for the maximum number of iterations to terminate the fine-tuning process. In the example, the fine-tuning process stops when the number of iterations reaches the maximum number of iterations.

[0216] The step size may indicate a learning rate for an online training process (e.g., an online fine-tuning process). The step size may be used in a back propagation calculation or a gradient descent algorithm performed during the fine-tuning process. Any suitable method may be used to determine the step size. In an embodiment, different step sizes are used for images with different types of content to achieve optimal results. Different types may refer to different variances. In an example, the step size is determined based on the variance of the image used to update the NIC framework. For example, the step size of an image with a high variance is greater than the step size of an image with a low variance, where the high variance is greater than the low variance.

[0217] In an embodiment, a first step size may be used to run a certain number of iterations (e.g., 100). Then, a second step size (e.g., the first step size plus or minus the size increment) may be used to run a certain number of iterations. The results from the first step size and the second step size may be compared to determine the step size to use. More than two step sizes may be tested to determine the optimal step size.

[0218] The step size may vary during the fine-tuning process. The step size may have an initial value at the beginning of the fine-tuning process, and in the later stage of the fine-tuning process, for example, after a certain number of iterations, the initial value may be reduced (e.g., halved) to achieve finer adjustments. During iterative online training, the step size or learning rate may be changed by a scheduler. The scheduler may include a parameter adjustment method for adjusting the step size. The scheduler may determine the value of the step size so that the step size may increase, decrease, or remain constant within a number of intervals. In the example, the scheduler changes the learning rate in each step. A single scheduler or multiple different schedulers may be used for different images. Therefore, multiple sets of replacement parameters may be generated based on multiple schedulers, and a set of replacement parameters from multiple sets of replacement parameters with better compression performance (e.g., less RD loss) may be selected.

[0219] At the end of the fine-tuning process, one or more update parameters may be calculated for the corresponding one or more replacement parameters. In an embodiment, the one or more update parameters are calculated as the difference between the one or more replacement parameters and the corresponding one or more pre-trained parameters. In an embodiment, the one or more update parameters are respectively one or more replacement parameters.

[0220] In an embodiment, one or more update parameters may be generated from one or more replacement parameters, for example using a linear or nonlinear transformation, and the one or more update parameters are representative parameters generated based on the one or more replacement parameters. The one or more replacement parameters are converted into one or more update parameters for better compression.

[0221] The first subset of the one or more update parameters corresponds to the first subset of the one or more replacement parameters, and the second subset of the one or more update parameters corresponds to the second subset of the one or more replacement parameters.

[0222] In an example, the one or more update parameters may be compressed, for example, using LZMA2, which is a variant of a Lempel-Ziv-Markov chain algorithm (LZMA), a bzip2 algorithm, or the like. In an example, compression is omitted for the one or more update parameters. In some embodiments, the one or more update parameters or a second subset of the one or more update parameters may be encoded into a bitstream as neural network update information, wherein the neural network update information indicates the one or more replacement parameters or a second subset of the one or more replacement parameters.

[0223] After the fine-tuning process, in some examples, a pre-trained video encoder on the encoder side can be updated or fine-tuned based on (i) a first subset of one or more replacement parameters or (ii) one or more replacement parameters. An input image (e.g., one of the one or more images used for the fine-tuning process) can be encoded into a bitstream using the updated video encoder. Thus, the bitstream includes both the encoded image and the neural network update information.

[0224] If applicable, in an example, the neural network update information is decoded (e.g., decompressed) by a pre-trained video decoder to obtain one or more update parameters or a second subset of one or more update parameters. In an example, the one or more replacement parameters or a second subset of one or more replacement parameters may be obtained based on a relationship between the one or more update parameters and the one or more replacement parameters. As described above, the pre-trained video decoder may be fine-tuned, and the decoded update video may be used to decode the encoded image.

[0225] The NIC framework can include any type of neural network and use any neural network-based image compression method, such as a contextual hyper-prior encoder-decoder framework (e.g., the NIC framework shown in FIG. (9)), a scaled hyper-prior encoder-decoder framework, a Gaussian mixture likelihood framework and variants of the Gaussian mixture likelihood framework, an RNN-based recursive compression method and variants of the RNN-based recursive compression method, etc.

[0226] Compared with the related E2E image compression methods, the content adaptive online training method and apparatus in the present disclosure may have the following advantages. An adaptive online training mechanism is used to improve NIC coding efficiency. The use of a flexible and general framework may adapt to various types of pre-trained frameworks and quality metrics. For example, by using online training using images to be encoded and transmitted, certain pre-trained parameters in various types of pre-trained frameworks may be replaced.

[0227] Fig.19 A flowchart outlining a process (S1900) according to an embodiment of the present disclosure is shown. The process (S1900) may be used to encode an image such as an original image or a residual image. In various embodiments, the process (S1900) is performed by a processing circuit system, the processing circuit system including, for example, a processing circuit system in a terminal device (310), (320), (330), and (340), a processing circuit system that performs the functions of a video encoder (1600A), and a processing circuit system that performs the functions of a video encoder (1700). In an example, the processing circuit system performs a combination of the functions of (i) one of the video encoders (403), (603), and (703) and (ii) one of the video encoder (1600A) and the video encoder (1700). In some embodiments, the process (S1900) is implemented in software instructions, so that when the processing circuit system executes the software instructions, the processing circuit system performs the process (S1900). The process starts at (S1901). In an example, the NIC framework is based on a neural network. In the example, the NIC frame is referenced Fig. 9 The NIC framework (900) described herein can be based on, for example, Figures 10 to 15 As described above, the video encoder (e.g., (1600A) or (1700)) and the corresponding video decoder (e.g., (1600B) or (1800)) may include multiple components in the NIC framework. The neural network-based NIC framework is pre-trained to pre-train the video encoder and the video decoder. Processing (S1900) proceeds to (S1910).

[0228] At (S1910), a fine-tuning process is performed on the NIC framework based on one or more images (or input images). The input image can be any suitable image with any suitable size. In some examples, the input image includes an original image in a spatial domain, a natural image, a computer-generated image, etc.

[0229] In some examples, the input image includes residual data in the spatial domain, for example, calculated by a residual calculator (e.g., residual calculator (723)). Components in various devices may be appropriately combined to implement (S1910), for example, referring to Figure 7 and Fig. 9 , the residual data from the residual calculator are combined into an image and fed to the main encoder network (911) in the NIC framework.

[0230] As described above, one or more parameters (e.g., one or more pretrained parameters) in one or more pretrained neural networks in a NIC framework (e.g., a pretrained NIC framework) can be updated to one or more replacement parameters, respectively. In an embodiment, in (S1910), one or more parameters in one or more neural networks are updated, for example, during the training process described in each step.

[0231] In an embodiment, at least one neural network in a video encoder (e.g., a pre-trained video encoder) is configured with a first subset of one or more pre-trained parameters, and thus the at least one neural network in the video encoder can be updated based on the corresponding first subset of one or more replacement parameters. In an example, the first subset of one or more replacement parameters includes all of the one or more replacement parameters. In an example, when the first subset of one or more pre-trained parameters is replaced by the first subset of one or more replacement parameters, respectively, the at least one neural network in the video encoder is updated. In an example, the at least one neural network in the video encoder is iteratively updated in a fine-tuning process. In an example, none of the one or more pre-trained parameters are included in the video encoder, so the video encoder is not updated and the pre-trained video encoder remains.

[0232] At (S1920), one of the one or more images may be encoded using a video encoder having at least one updated neural network. In an example, after updating at least one neural network in the video encoder, one of the one or more images is encoded.

[0233] Step (S1920) may be modified appropriately. For example, when none of the one or more replacement parameters is included in at least one neural network in the video encoder, the video encoder is not updated, and thus a pre-trained video encoder (e.g., a video encoder including at least one pre-trained neural network) may be used to encode one of the one or more images.

[0234] At (S1930), neural network update information indicating a second subset of the one or more replacement parameters may be encoded into the bitstream. In an example, the second subset of the one or more replacement parameters will be used to update at least one neural network in a video decoder on the decoder side. For example, if the second subset of the one or more replacement parameters does not include parameters and the neural network update information is not signaled in the bitstream, step (S1930) may be omitted and no neural network in the video decoder is updated.

[0235] At (S1940), a bitstream including the encoded images in one or more images and the neural network update information may be transmitted. Step (S1940) may be modified appropriately. For example, if step (S1930) is omitted, the bitstream does not include the neural network update information. Process (S1900) proceeds to (S1999) and terminates.

[0236] The process (S1900) can be appropriately adapted to various scenarios, and the steps in the process (S1900) can be adjusted accordingly. One or more of the steps in the process (S1900) can be modified, omitted, repeated, and / or combined. The process (S1900) can be implemented in any suitable order. Additional steps can be added. For example, in addition to encoding one of the one or more images, other images in the one or more images, such as the remaining images, are encoded in (S1920) and transmitted in (S1940).

[0237] In some examples of the process (S1900), one of the one or more images is encoded by an updated video encoder and transmitted in a bitstream.Since the fine-tuning process is based on the one or more images, the fine-tuning process is based on a context to be encoded and is therefore context-based.

[0238] In some examples, the neural network update information also indicates what parameters the second subset of one or more pre-trained parameters (or the corresponding second subset of one or more replacement parameters) are, so that the corresponding pre-trained parameters in the video decoder can be updated. The neural network update information may indicate component information (e.g., (915)), layer information (e.g., fourth layer DeConv: 5×5 c3 s2), channel information (e.g., second channel), etc. of the second subset of one or more pre-trained parameters. Therefore, referring to Fig.11 , the second subset of one or more replacement parameters includes the convolution kernel of the second channel of DeConv: 5×5 c3 s2 in the main decoder network (915). Therefore, the convolution kernel of the second channel of DeConv: 5×5 c3 s2 in the pre-trained main decoder network (915) is updated. In some examples, component information (e.g., (915)), layer information (e.g., fourth layer DeConv: 5×5 c3 s2), channel information (e.g., second channel), etc. of the second subset of one or more pre-trained parameters are predetermined and stored in the pre-trained video decoder, so no signaling is required.

[0239] Fig. 20 A flowchart is shown outlining a process (2000) according to an embodiment of the present disclosure. The process (2000) can be used in the reconstruction of an encoded image. In various embodiments, the process (2000) is performed by a processing circuit system, the processing circuit system including, for example, a processing circuit system in a terminal device (310), (320), (330) and (340), a processing circuit system that performs the functions of a video decoder (1600B), and a processing circuit system that performs the functions of a video decoder (1800). In an example, the processing circuit system performs a combination of the functions of (i) one of a video decoder (410), a video decoder (510) and a video decoder (810) and (ii) one of a video decoder (1600B) or a video decoder (1800). In some embodiments, the process (2000) is implemented as software instructions, so that when the processing circuit system executes the software instructions, the processing circuit system performs the process (2000). The process starts at (S2001). In an example, the NIC framework is based on a neural network. In an example, the NIC framework is a reference to Fig. 9 The NIC framework (900) described herein can be based on, for example, Figures 10 to 15 As described above, the video decoder (e.g., (1600B) or (1800)) may include multiple components in the NIC framework. The neural network-based NIC framework may be pre-trained. The video decoder may be pre-trained using pre-trained parameters. Process (2000) proceeds to (S2010).

[0240] At (S2010), neural network update information in the encoded bitstream may be decoded. The neural network update information may be used for a neural network in a video decoder. The neural network may be configured with pre-trained parameters. The neural network update information may correspond to an encoded image to be reconstructed and indicate a replacement parameter corresponding to a pre-trained parameter in the pre-trained parameters.

[0241] In the example, the pre-trained parameter is a pre-trained bias term.

[0242] In the example, the pre-trained parameters are pre-trained weight coefficients.

[0243] In an embodiment, the video decoder includes a plurality of neural networks. The plurality of neural networks include a neural network. The neural network update information may indicate update information for one or more remaining neural networks in the plurality of neural networks. For example, the neural network update information further indicates one or more replacement parameters for one or more remaining neural networks in the plurality of neural networks. The one or more replacement parameters correspond to one or more corresponding pre-trained parameters for the one or more remaining neural networks. In an example, each of the pre-trained parameter and the one or more pre-trained parameters is a corresponding pre-trained bias term. In an example, each of the pre-trained parameter and the one or more pre-trained parameters is a corresponding pre-trained weight coefficient. In an example, the pre-trained parameter and the one or more pre-trained parameters include one or more pre-trained bias terms and one or more pre-trained weight coefficients in the plurality of neural networks.

[0244] In an example, the neural network update information indicates update information for a subset of the plurality of neural networks, while a remaining subset of the plurality of neural networks is not updated.

[0245] In the example, the video decoder is Fig.18 The video decoder (1800) shown. The neural network is the main decoder network (915).

[0246] In the example, the video decoder is Fig. 16B The video decoder (1600B) shown. The multiple neural networks in the video decoder include a main decoder network (915), a context model NN (916), an entropy parameter NN (917), and a super decoder (925). The neural network is one of the main decoder network (915), the context model NN (916), the entropy parameter NN (917), and the super decoder (925). For example, the neural network is the context model NN (916). The neural network update information also indicates one or more replacement parameters for one or more remaining neural networks (e.g., the main decoder network (915), the entropy parameter NN (917), and / or the super decoder (925)) in the video decoder (1600B).

[0247] In an example, the neural network update information indicates a plurality of replacement parameters corresponding to a plurality of pre-trained parameters in the pre-trained parameters for the neural network. The plurality of pre-trained parameters include pre-trained parameters. The plurality of pre-trained parameters include one or more pre-trained bias terms and one or more pre-trained weight coefficients. The neural network in the video decoder may be updated based on the plurality of replacement parameters including the replacement parameters.

[0248] At (S2020), replacement parameters may be determined based on the neural network update information. In an embodiment, the updated parameters are obtained from the neural network update information. In an example, the updated parameters may be obtained from the neural network update information by decompression. In an example, the neural network update information indicates that the updated parameters are the difference between the replacement parameters and the pre-trained parameters, and the replacement parameters may be calculated based on the sum of the updated parameters and the pre-trained parameters. In an embodiment, the replacement parameters are determined as updated parameters. In an embodiment, the updated parameters are representative parameters generated based on the replacement parameters on the encoder side (e.g., using a linear transformation or a nonlinear transformation), and the replacement parameters are obtained based on the representative parameters.

[0249] At (S2030), the neural network in the video decoder may be updated (or fine-tuned) based on the replacement parameters, for example, by replacing the pre-trained parameters with the replacement parameters in the neural network. If the video decoder includes multiple neural networks, and the neural network update information indicates update information (e.g., additional replacement parameters) for the multiple neural networks, the multiple neural networks may be updated. For example, the neural network update information also includes one or more replacement parameters for one or more remaining neural networks in the video decoder, and the one or more remaining neural networks may be updated based on the one or more replacement parameters.

[0250] At (S2040), the encoded image in the bitstream may be decoded by an updated video decoder, for example based on an updated neural network. The output image generated at (S2040) may be any suitable image having any suitable size. In some examples, the output image includes a reconstructed original image in the spatial domain, a natural image, a computer-generated image, etc.

[0251] In some examples, the output image of the video decoder includes residual data in the spatial domain, so further processing can be used to generate a reconstructed image based on the output image. For example, the reconstruction module (874) is configured to combine the residual data and the prediction result (output by the inter-frame or intra-frame prediction module) in the spatial domain to form a reconstructed block, which can be part of the reconstructed image. Additional appropriate operations such as deblocking operations can be performed to improve visual quality. The components in various devices can be appropriately combined to implement (S2040), for example, referring to Figure 8 and Fig. 9 , the residual data and the corresponding prediction results from the main decoder network (915) in the video decoder are fed to the reconstruction module (874) to generate a reconstructed image.

[0252] In an example, the bitstream also includes one or more coded bits used to determine a context model for decoding a coded image. The video decoder may include a main decoder network (e.g., (911)), a context model network (e.g., (916)), an entropy parameter network (e.g., (917)), and a super decoder network (e.g., (925)). The neural network is one of the main decoder network, the context model network, the entropy parameter NN, and the super decoder network. The super decoder network may be used to decode one or more coded bits. The entropy model (e.g., context model) may be determined based on quantized potential and decoded bits of the coded image available to the context model network using the context model network and the entropy parameter network. The coded image may be decoded using the main decoder network and the entropy model.

[0253] Processing (2000) proceeds to (S2099) and terminates.

[0254] The process (2000) may be appropriately adapted to various scenarios, and the steps in the process (2000) may be adjusted accordingly. One or more of the steps in the process (2000) may be modified, omitted, repeated, and / or combined. The process (2000) may be implemented in any suitable order. Additional steps may be added.

[0255] For example, at (S2040), one or more additional encoded images in the encoded bitstream are decoded based on the updated neural network. Therefore, the encoded image and the one or more additional encoded images can share the same neural network update information.

[0256] In an embodiment, a device for video decoding includes: a neural network update information decoding module, an update module and an image decoding module. The neural network update information decoding module is used to decode neural network update information for a neural network in a video decoder in a coded bitstream, the neural network is configured with pre-trained parameters, and the neural network update information corresponds to a coded image to be reconstructed and indicates replacement parameters corresponding to pre-trained parameters in the pre-trained parameters. The update module is used to update the neural network in the video decoder based on the replacement parameters. The image decoding module is used to decode the coded image based on the updated neural network for the coded image.

[0257] In some examples, the neural network update information also includes one or more replacement parameters for one or more remaining neural networks in the video decoder, and the update module is used to update the one or more remaining neural networks based on the one or more replacement parameters.

[0258] In some examples, the coded bitstream also indicates one or more coded bits, and the one or more coded bits are used to determine a context model for decoding the coded image. The video decoder includes a main decoder network, a context model network, an entropy parameter network, and a super decoder network. The neural network is one of the main decoder network, the context model network, the entropy parameter network, and the super decoder network. The image decoding module is used to: decode one or more coded bits using the super decoder network, determine the context model based on the quantized potential of the coded image available to the context model network and one or more decoded bits using the context model network and the entropy parameter network, and decode the coded image using the main decoder network and the context model.

[0259] In some examples, the pre-trained parameters are pre-trained bias terms. In some examples, the pre-trained parameters are pre-trained weight coefficients.

[0260] In some examples, the neural network update information indicates a plurality of replacement parameters corresponding to a plurality of pretrained parameters in pretrained parameters for the neural network, the plurality of pretrained parameters include a pretrained parameter, and the plurality of pretrained parameters include one or more pretrained bias terms and one or more pretrained weight coefficients, and the update module is used to update the neural network in the video decoder based on the plurality of replacement parameters including the replacement parameters.

[0261] In some examples, the neural network update information indicates a difference between the replacement parameter and the pre-trained parameter. The device also includes a determination module for determining the replacement parameter based on the difference and the sum of the pre-trained parameter.

[0262] In some examples, the image decoding module is used to decode another encoded image in the encoded bitstream based on the updated neural network.

[0263] In an embodiment, a computer device includes a processor and a memory. The memory is used to store program codes and transmit the program codes to the processor. The processor is used to execute the above-mentioned method for video decoding according to instructions in the program codes.

[0264] In an embodiment, a non-transitory computer-readable storage medium storing a program stores the program. The program can be executed by at least one processor to perform the above-mentioned method for video decoding.

[0265] The embodiments of the present disclosure may be used alone or in combination in any order. In addition, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.

[0266] The present disclosure does not impose any restrictions on the methods used for encoders (such as neural network-based encoders), decoders (such as neural network-based decoders). The neural networks used in encoders, decoders, etc. can be any suitable type of neural network, such as DNN, CNN, etc.

[0267] Therefore, the content-adaptive online training method of the present disclosure can adapt to different types of NIC frameworks, such as different types of encoding DNNs, decoding DNNs, encoding CNNs, decoding CNNs, etc.

[0268] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. Fig.21 A computer system (2100) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0269] Computer software may be encoded using any suitable machine code or computer language, which may be subjected to mechanisms such as assembly, compilation, linking, etc. to create code comprising instructions that may be executed directly by one or more computer central processing units (CPU), graphics processing units (GPU), etc., or through interpretation, microcode execution, etc.

[0270] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0271] Fig.21 The components for the computer system (2100) shown in the example are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing embodiments of the present disclosure. Nor should the configuration of components be interpreted as having any dependency or requirement related to any one or combination of components shown in the exemplary embodiment of the computer system (2100).

[0272] The computer system (2100) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to inputs implemented by one or more human users through, for example, tactile inputs (e.g., keystrokes, swipes, data glove movements), audio inputs (e.g., voice, tapping), visual inputs (e.g., gestures), olfactory inputs (not depicted). The human-machine interface devices may also be used to capture certain media that are not necessarily directly related to intentional human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and videos (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0273] The input human-machine interface devices may include one or more of the following (only one of each is depicted): keyboard (2101), mouse (2102), trackpad (2103), touch screen (2110), data gloves (not shown), joystick (2105), microphone (2106), scanner (2107) and camera (2108).

[0274] The computer system (2100) may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate one or more senses of a human user through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include: tactile output devices (e.g., tactile feedback through a touch screen (2110), a data glove (not shown), or a joystick (2105), but there may also be tactile feedback devices that are not used as input devices); audio output devices (e.g., speakers (2109), headphones (not depicted)); visual output devices (e.g., screens (2110), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which may be able to output two-dimensional visual output or more than three-dimensional output through methods such as stereoscopic image output; virtual reality glasses (not depicted); holographic displays and cigarette cans (not depicted)); and printers (not depicted).

[0275] The computer system (2100) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (2120) with CD / DVD etc. media (2121), thumb drives (2122), removable hard drives or solid-state drives (2123), traditional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD based devices such as security dongles (not depicted), etc.

[0276] Those skilled in the art will also appreciate that the term "computer-readable media" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0277] The computer system (2100) may also include an interface (2154) to one or more communication networks (2155). The network may be, for example, wireless, wired, optical. The network may also be a local area network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter attached to some common data port or peripheral bus (2149) (e.g., a USB port of the computer system (2100)); other networks are typically integrated into the core of the computer system (2100) by attaching to a system bus as described below (e.g., integrated into a PC computer system via an Ethernet interface, or integrated into a smart phone computer system via a cellular network interface). The computer system (2100) can communicate with other entities by using any of these networks. Such communications may be one-way receive-only (e.g., broadcast television), one-way send-only (e.g., CANbus to certain CANbus devices), or two-way, such as to other computer systems using local or wide area digital networks. Specific protocols and protocol stacks may be used on each of these networks and network interfaces as described above.

[0278] The above-mentioned human-machine interface device, human-accessible storage device, and network interface may be attached to the core ( 2140 ) of the computer system ( 2100 ).

[0279] The core (2140) may include one or more central processing units (CPUs) (2141), graphics processing units (GPUs) (2142), dedicated programmable processing units in the form of field programmable gate areas (FPGAs) (2143), hardware accelerators (2144) for certain tasks, graphics adapters (2150), etc. These devices as well as read-only memory (ROM) (2145), random access memory (2146), internal large-capacity storage devices (2147) such as internal non-user accessible hard drives, SSDs, etc. can be connected through a system bus (2148). In some computer systems, the system bus (2148) can be accessed in the form of one or more physical plugs to achieve expansion through additional CPUs, GPUs, etc. Peripheral devices can be attached to the system bus (2148) of the core directly or through a peripheral bus (2149). In an example, a screen (2110) can be connected to a graphics adapter (2150). The architecture of the peripheral bus includes PCI, USB, etc.

[0280] The CPU (2141), GPU (2142), FPGA (2143) and accelerator (2144) can execute certain instructions, which can be combined to form the computer code mentioned above. The computer code can be stored in ROM (2145) or RAM (2146). Transition data can also be stored in RAM (2146), while permanent data can be stored in, for example, an internal mass storage device (2147). Fast storage and retrieval of any storage device in the storage device can be achieved by using a cache memory, which can be closely associated with one or more CPUs (2141), GPUs (2142), mass storage devices (2147), ROMs (2145), RAMs (2146), etc.

[0281] The computer readable medium may have computer code thereon for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or the medium and computer code may be of a type well known and available to those skilled in the art of computer software.

[0282] As an example and not limitation, a computer system (2100) having an architecture, in particular a core (2140), can provide functions provided as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such a computer-readable medium can be a medium associated with a user-accessible mass storage device as described above, as well as certain storage devices of the core (2140) having a non-transitory nature, such as a mass storage device (2147) or ROM (2145) inside the core. Software implementing various embodiments of the present disclosure can be stored in such a device and executed by the core (2140). Depending on specific needs, the computer-readable medium may include one or more memory devices or chips. The software can enable the core (2140) - in particular the processor therein (including a CPU, GPU, FPGA, etc.) - to perform specific processing or specific parts of specific processing described herein, including defining data structures stored in RAM (2146) and modifying such data structures according to processing defined by the software. Additionally or alternatively, the computer system may provide functionality due to logic hardwired or otherwise embodied in circuits (e.g., accelerator (2144)) that may operate in place of or in conjunction with software to perform specific processing or specific portions of specific processing described herein. Where appropriate, references to software may include logic, and conversely references to logic may include software. Where appropriate, references to computer-readable media may include circuits (e.g., integrated circuits (ICs)) storing software for execution, circuits implementing logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.

[0283] Appendix A: Acronyms

[0284] JEM: Joint Development Model

[0285] VVC: Versatile Video Coding

[0286] BMS: Benchmark Set

[0287] MV: Motion Vector

[0288] HEVC: High Efficiency Video Codec

[0289] SEI: Supplemental Enhancement Information

[0290] VUI: Video Availability Information

[0291] GOPs: Group of Pictures

[0292] TUs: Transformation Units

[0293] PUs: Prediction Units

[0294] CTUs: Coding Tree Units

[0295] CTBs: Coding Tree Blocks

[0296] PBs: prediction blocks

[0297] HRD: Hypothetical Reference Decoder

[0298] SNR: Signal to Noise Ratio

[0299] CPUs: Central Processing Units

[0300] GPUs: Graphics Processing Units

[0301] CRT: cathode ray tube

[0302] LCD: Liquid Crystal Display

[0303] OLED: Organic Light Emitting Diode

[0304] CD: Compact Disc

[0305] DVD: Digital Video Disc

[0306] ROM: Read Only Memory

[0307] RAM: Random Access Memory

[0308] ASIC: Application-Specific Integrated Circuit

[0309] PLD: Programmable Logic Device

[0310] LAN: Local Area Network

[0311] GSM: Global System for Mobile Communications

[0312] LTE: Long Term Evolution

[0313] CANBus: Controller Area Network Bus

[0314] USB: Universal Serial Bus

[0315] PCI: Peripheral Component Interconnect

[0316] FPGA: Field Programmable Gate Array

[0317] SSD: Solid State Drive

[0318] IC: Integrated Circuit

[0319] CU: Coding Unit

[0320] NIC: Neural Image Compression

[0321] RD: Rate Distortion

[0322] E2E: End to End

[0323] ANN: Artificial Neural Network

[0324] DNN: Deep Neural Network

[0325] CNN: Convolutional Neural Network

[0326] Although the present disclosure has described several exemplary embodiments, there are changes, permutations, and various substitute equivalents that fall within the scope of the present disclosure. It will therefore be appreciated that, although not explicitly shown or described herein, those skilled in the art will be able to conceive of many systems and methods that implement the principles of the present disclosure and are therefore within its spirit and scope.

Claims

1. A method for video decoding, characterized in that: The method comprises: decoding neural network update information for a neural network in a video decoder in a coded bitstream, the neural network being configured with pre-trained parameters, the neural network update information corresponding to a coded image to be reconstructed and indicating replacement parameters corresponding to pre-trained parameters in the pre-trained parameters; Updating a neural network in the video decoder based on the replacement parameters; and decoding the encoded image based on the updated neural network for the encoded image; The coded bitstream further indicates one or more coded bits, the one or more coded bits being used to determine a context model for decoding the coded image, The video decoder comprises a main decoder network, an entropy decoder, a context model network, an entropy parameter network and a super decoder network, the neural network is one of the main decoder network, the context model network, the entropy parameter network and the super decoder network, The method further comprises: Decode the one or more coded bits using the super decoder network to obtain an output o hc ;as well as generating, by the entropy decoder, a quantized latent representation of the encoded image based on the encoded image; Using the context model network based on the quantized latent representation of the encoded image, an output is obtained. cm,i ; Use the entropy parameters of the network based on the output o hc and the output o cm,i , and get the output o ep , and through the output o ep Determine the context model; Decoding the encoded image includes decoding the encoded image using the primary decoder network and the context model.

2. The method according to claim 1, characterized in that The pre-training parameter is a pre-training bias item.

3. The method according to claim 1, characterized in that The pre-training parameter is a pre-training weight coefficient.

4. The method according to any one of claims 1 to 3, characterized in that The neural network update information indicates a plurality of replacement parameters corresponding to a plurality of pre-trained parameters of the pre-trained parameters for the neural network, the plurality of pre-trained parameters including one or more pre-trained bias terms and one or more pre-trained weight coefficients, and The updating includes updating the neural network in the video decoder based on the plurality of replacement parameters including the replacement parameter.

5. The method according to any one of claims 1 to 3, characterized in that The neural network update information indicates a difference between the replacement parameters and the pre-trained parameters, and The method further comprises determining the replacement parameter based on a sum between the pre-trained parameter and the difference.

6. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: Another encoded image in the encoded bitstream is decoded based on the updated neural network.

7. A device for video decoding, characterized in that The device comprises: a neural network update information decoding module for decoding neural network update information for a neural network in a video decoder in a coded bitstream, the neural network being configured with pre-trained parameters, the neural network update information corresponding to a coded image to be reconstructed and indicating replacement parameters corresponding to pre-trained parameters in the pre-trained parameters; an updating module, configured to update a neural network in the video decoder based on the replacement parameters; and an image decoding module for decoding the encoded image based on the updated neural network for the encoded image; The coded bitstream further indicates one or more coded bits, the one or more coded bits being used to determine a context model for decoding the coded image, The video decoder comprises a main decoder network, an entropy decoder, a context model network, an entropy parameter network and a super decoder network, the neural network is one of the main decoder network, the context model network, the entropy parameter network and the super decoder network, The device also includes: Decode the one or more coded bits using the super decoder network to obtain an output o hc ;as well as generating, by the entropy decoder, a quantized latent representation of the encoded image based on the encoded image; Using the context model network based on the quantized latent representation of the encoded image, an output is obtained. cm,i ; Use the entropy parameters of the network based on the output o hc and the output o cm,i , and get the output o ep , and through the output o ep Determine the context model; Decoding the encoded image includes decoding the encoded image using the primary decoder network and the context model.

8. The device according to claim 7, characterized in that The pre-training parameter is a pre-training bias item.

9. The device according to claim 7, characterized in that The pre-training parameter is a pre-training weight coefficient.

10. The device according to any one of claims 7 to 9, characterized in that The neural network update information indicates a plurality of replacement parameters corresponding to a plurality of pre-trained parameters of the pre-trained parameters for the neural network, the plurality of pre-trained parameters including one or more pre-trained bias terms and one or more pre-trained weight coefficients, and The update module is used to update the neural network in the video decoder based on the multiple replacement parameters including the replacement parameters.

11. The device according to any one of claims 7 to 9, characterized in that The neural network update information indicates a difference between the replacement parameters and the pre-trained parameters, and The device further comprises a determination module for determining the replacement parameter according to a sum between the pre-trained parameter and the difference.

12. The device according to any one of claims 7 to 9, characterized in that: The image decoding module is used for: Another encoded image in the encoded bitstream is decoded based on the updated neural network.

13. A computer device, characterized in that: The computer device comprises a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the method according to any one of claims 1 to 6 according to the instructions in the program code.

14. A non-transitory computer-readable storage medium storing a program, wherein the program can be executed by at least one processor to perform the method of any one of claims 1 to 6.

15. A method for storing a video bitstream, characterized in that: The video bit stream is decoded according to the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Methods And Apparatuses For Learned Image Compression

    US20200160565A1

  • An apparatus, a method and a computer program for video coding and decoding

    WO2020165493A1