Efficient neural network decoder for image compression
By introducing core decoder and hyperdecoder into image compression, the potential representation of images is achieved using multi-layer neural networks, the complexity and efficiency balance problems in the prior art are solved, and the encoding performance of image compression is significantly improved.
Patent Information
- Application Number
- CN202480004267.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-18
- Filing Date
- 2024-04-19
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively balance the complexity and coding efficiency of neural networks in image compression, resulting in poor coding performance.
A method and system for neural image compression, including core decoders and hyperdecoders, is proposed to achieve a balance between network complexity and coding efficiency through these components. The core decoder includes a neural network with three upsampling stages, while the hyperdecoder generates potential representation samples of the image through multiple upsampling stages and convolutional neural network layers.
Through this method and system, the encoding performance of image compression can be significantly improved, rate distortion performance can be improved, and a better balance between training efficiency and coding gain can be found.
Smart Images

Figure CN119998826A_ABST
Abstract
Description
[0001] Cross-references
[0002] This application is based on the U.S. provisional patent application with application number 63 / 460,889, application date April 20, 2023, and name “Efficient Neural Network Decoder for Image Compression” and the U.S. patent application with application number 18 / 639,613, application date April 18, 2024, and name “Efficient Neural Network Decoder for Image Compression”, and claims the priority of the above patent applications. The entire contents of the above patent applications are hereby introduced into this application as a reference. Technical Field
[0003] The present disclosure relates generally to encoding of images, and in particular to methods and systems for neural image compression. Background Art
[0004] Artificial Intelligence (AI) technology (including but not limited to Deep Neural Network (DNN)) can be applied to various aspects of image or video compression. For example, various encoding tools can be assisted by pre-trained AI models. For another example, an end-to-end AI-based video or image encoder and / or video or image decoder can be based on DNN. This end-to-end encoding and / or decoding can be called Neural Image Compression (NIC). Compared with traditional video or image codecs, end-to-end NIC systems can be globally optimized through an automatic training process, without relying on tuning a large number of independent encoding tools one by one, nor being limited by the inability to utilize the optimization correlation between these encoding tools. Therefore, the end-to-end NIC system can help improve encoding performance (e.g., rate-distortion performance) through a single-module optimization process. In order to achieve this optimization, it is necessary to consider balancing the complexity, training efficiency, and coding gain of the NIC model. Summary of the invention
[0005] The present disclosure relates generally to image coding, and in particular to methods and systems for NIC. The disclosed NIC decoder / encoder may include various neural network components configured to achieve a balance between network complexity and coding efficiency. Implementations of such NIC decoder / encoder specifically include a core decoder and a hyper decoder, each of which includes a neural network architecture suitable for implementing a lightweight decoder / encoder.
[0006] In some example implementations, a method for decoding an encoded bitstream of an image is disclosed. The decoder may include a super decoder, a core decoder, and a context model. The method may include: generating a super data item set according to the encoded bitstream by the super decoder; generating a context parameter set by the context model; generating a potential representation sample of the image by processing the encoded bitstream and an entropy parameter set generated according to the super data item set and the context parameter set; and reconstructing an image sample according to the potential representation sample of the image by the core decoder. The core decoder includes a neural network having three upsampling stages or fewer upsampling stages.
[0007] In the above example implementation, the core decoder includes three upsampling stages, and each pair of consecutive upsampling stages in the three upsampling stages are connected via at least a pruning neural network layer and an activation neural network layer.
[0008] In any of the above example implementations, the single-dimensional upsampling ratios of the three upsampling stages are 2, 2, and 4, or 2, 4, and 2, or 4, 2, and 2, respectively.
[0009] In any of the above example implementations, no convolutional neural network layers are arranged between the three upsampling stages.
[0010] In any of the above example embodiments, the core decoder includes two upsampling stages connected by a pruning neural network layer, an activation neural network layer, and a convolutional neural network layer.
[0011] In any of the above example implementations, the single-dimensional upsampling ratios of the two upsampling stages are 4 and 4 respectively.
[0012] In any of the above example embodiments, the super decoder comprises at least two upsampling stages, and a first upsampling stage in sequence is directly connected to the entropy decoded data items from the encoded bitstream.
[0013] In any of the above example implementations, each pair of adjacent upsampling stages is connected by at least a pruning neural network layer, an activation neural network layer, and a convolutional neural network layer.
[0014] In any of the above example implementations, at least one pair of adjacent upsampling stages does not include a pruned neural network layer between them.
[0015] In any of the above example embodiments, at least one of the at least two upsampling stages comprises a convolutional neural network layer.
[0016] In any of the above example embodiments, at least one of the at least two upsampling stages includes a pixel shuffle neural network having a convolutional neural network layer.
[0017] In any of the above example embodiments, at least one of the at least two upsampling stages includes a transposed convolutional neural network layer.
[0018] In any of the above example embodiments, a coded stream is generated for the Y component, U component, and V component of the image; and the Y component and the U and V components of the image are decoded using a separate model of the super decoder and a separate model of the decoder. Alternatively, a coded stream is generated for the R component, G component, and B component of the image; and the R component, G component, and B component of the image are decoded using a separate model of the super decoder and a separate model of the core decoder.
[0019] Aspects of the present disclosure also provide an electronic decoding device or apparatus or an electronic encoding device or apparatus, the device or apparatus comprising a circuit or processor configured to execute any of the above method embodiments.
[0020] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions, which, when executed by an electronic device, causes the electronic device to perform any of the above-mentioned method embodiments.
[0021] Aspects of the present disclosure also provide a non-transitory computer-readable recording medium for storing the above-mentioned code stream. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Further features, properties and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0023] Figure 1 A schematic diagram showing a simplified block diagram of a communication system (100) according to an example embodiment.
[0024] Figure 2 A schematic diagram showing a simplified block diagram of a communication system (200) according to an example embodiment is shown.
[0025] Figure 3 A schematic diagram showing a simplified block diagram of a video decoder according to an example embodiment.
[0026] Figure 4 A schematic diagram showing a simplified block diagram of a video encoder according to an example embodiment.
[0027] Figure 5 A block diagram of a video encoder according to another example embodiment is shown.
[0028] Figure 6 A block diagram of a video decoder according to another example embodiment is shown.
[0029] Figure 7 An example Neural Image Compression (NIC) model is shown.
[0030] Figure 8 Shows that it is possible to Figure 7 A box of an example superdecoder used in the NIC model.
[0031] Fig. 9 Shows that it is possible to Figure 7 A box of an example decoder used in the NIC model.
[0032] Fig.10 An example logic flow of a method for a NIC is shown.
[0033] Fig.11 is a schematic diagram of a computer system according to an example embodiment of the present disclosure. DETAILED DESCRIPTION
[0034] Throughout the specification and claims, in addition to the meanings explicitly given, terms may have subtle meanings that are suggested or implied from the context. The phrases "in one embodiment / implementation" or "in some embodiments / implementations" used herein do not necessarily refer to the same embodiment / implementation, and the phrases "in another embodiment / implementation" or "in other embodiments" used herein do not necessarily refer to different embodiments. For example, the claimed subject matter is intended to include, in whole or in part, a combination of exemplary embodiments / implementations.
[0035] Typically, terms can be understood at least in part from usage in context. For example, as used herein, terms such as "and", "or" or "and / or" can include various context-dependent meanings. Typically, if "or" is used to associate a list such as A, B, or C, it is intended to represent A, B, and C (herein for inclusive meanings) as well as A, B, or C (herein for exclusive meanings). In addition, as used herein, the terms "one or more", "at least one", "one", "an" or "the" depend at least in part on the context and can be used as a singular meaning or a plural meaning. In addition, the term "based on" or "determined by..." can be understood as not necessarily intended to convey a set of exclusive factors, but may allow for the presence of other factors that are not necessarily explicitly described, which also depends at least in part on the context.
[0036] Figure 1 A simplified block diagram of a communication system (100) according to an embodiment of the present disclosure is shown. The communication system (100) includes a plurality of terminal devices, for example, terminal device 110, terminal device 120, terminal device 130, and terminal device 140, which can communicate with each other via, for example, a network (150). Figure 1 In the example of , a first pair of terminal devices (110) and a terminal device (120) can perform unidirectional data transmission. For example, the terminal device (110) can encode video / image data (e.g., a stream of video / image pictures collected by the terminal device (110)) in the form of one or more encoded code streams for transmission over the network (150). The encoded video / image data is transmitted in the form of one or more encoded video / image code streams. The terminal device (120) can receive encoded video / image data or image data from the network (150), decode the encoded video / image data or image data to restore the video or image picture, and display the video or image picture based on the restored video or image data. Unidirectional data transmission can be implemented in applications such as media services.
[0037] In another example, the second pair of terminal devices (130) and terminal devices (140) can perform bidirectional transmission of encoded video / image data, which can be performed, for example, during a video conferencing application. For bidirectional data transmission, in one example, each of the terminal devices (130) and the terminal devices (140) can encode video / image data (e.g., a stream of video / image picture streams collected by the terminal devices) to transmit the encoded video / image data to the other terminal device (130) and the terminal device (140) through the network (150), and can also receive the encoded video / image data from the other terminal device (130) and the terminal device (140) through the network (150). Each of the terminal devices (130) and (140) can also receive encoded video / image data transmitted by another terminal device among the terminal devices (130) and (140), and can decode the encoded video / image data to restore the video / image picture, and can display the video / image picture on an accessible display device based on the restored video / image data.
[0038] exist Figure 1 In the example of, the terminal device can be implemented as a server, a personal computer and a smart phone, but the applicability of the basic principles disclosed in the present application may not be limited to this. Embodiments of the present disclosure can be implemented in desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated video conferencing equipment and / or similar devices. Network (150) represents any number or type of network that transmits encoded video / image data between terminal devices (110), terminal devices (120), terminal devices (130) and terminal devices (140), including, for example, wired (wired) and / or wireless communication networks. Communication network (150) can exchange data in circuit switching, packet switching and / or other types of channels. Representative networks may include telecommunication networks, local area networks, wide area networks and / or the Internet. For the purposes of this application, unless explicitly explained herein, the architecture and topology of network (150) may be irrelevant to the operations disclosed in this application.
[0039] As an example of application of the disclosed subject matter, Figure 2 The placement of the video / image encoder and the video / image decoder in a video / image streaming environment is shown. The disclosed subject matter can be equally applicable to other video / image applications, including, for example, video conferencing, digital TV broadcasting, gaming, virtual reality, storing compressed video / images on digital media including CDs, DVDs, memory sticks, etc., and the like.
[0040] like Figure 2As shown, the video / image streaming system may include a video / image acquisition subsystem (213), which may include a video / image source (201), such as a digital camera, which creates an uncompressed video / image picture or image stream (202). In the example, the video / image picture stream (202) includes samples recorded by the digital camera of the video / image source 201. Compared to the encoded video / image data (204) (or the encoded video / image code stream), the video / image picture stream (202) is depicted as a thick line to emphasize the high data volume and can be processed by an electronic device (220), which includes a video / image encoder (203) coupled to the video / image source (201). The video / image encoder (203) may include hardware, software, or a combination of hardware and software to implement or implement various aspects of the disclosed subject matter as described in more detail below. Compared to the uncompressed video / image picture stream (202), the encoded video / image data (204) (or the encoded video / image code stream (204)) is depicted as a thin line to emphasize the lower amount of data, and it can be stored on the streaming server (205) for future use or directly stored to a downstream video / image device (not shown). One or more streaming client subsystems, such as Figure 2 The client subsystem (206) and the client subsystem (208) in the video transmission server (205) can access the streaming server (205) to retrieve the copy (207) and the copy (209) of the encoded video / image data (204). The client subsystem (206) may include, for example, a video / image decoder (210) in the electronic device (230). The video / image decoder (210) decodes the incoming copy (207) of the encoded video / image data and creates an output stream (211) of uncompressed video / image pictures that can be presented on a display (212) (e.g., a display screen) or other presentation device (not depicted). The video / image decoder 210 can be configured to perform some or all of the various functions described in the present disclosure. In some streaming transmission systems, the encoded video / image data (204), the video / image data (207), and the video / image data (209) (e.g., video / image code streams) can be encoded according to certain video / image encoding / compression standards.
[0041] It should be noted that the electronic device (220) and the electronic device (230) may include other components (not shown). For example, the electronic device (220) may include a video / image decoder (not shown), and the electronic device (230) may also include a video / image encoder (not shown).
[0042] Figure 3A block diagram of an example video / image decoder (310) of an electronic device (330) according to any of the embodiments disclosed below in the present application is shown. The electronic device (330) may include a receiver (331) (e.g., a receiving circuit). The video / image decoder (310) may be used to replace Figure 2 A video / image decoder (210) in the example of FIG.
[0043] Below for Figures 3 to 6 The disclosure describes an example video encoding / decoding system, wherein the video being encoded or decoded comprises a sequence of images. Aspects of such a video encoder / decoder applicable to a single image (particularly in intra-frame coding mode) may be applied to encoding / decoding of still images.
[0044] like Figure 3 As shown, a receiver (331) may receive one or more encoded video sequences from a channel (301). To prevent network jitter and / or process playback timing, a buffer memory (315) may be arranged between the receiver (331) and an entropy decoder / parser (320) (hereinafter referred to as "parser (320)"). The parser (320) may reconstruct symbols (321) based on the encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (310) and potential information for controlling a presentation device such as a display (312) (e.g., a display screen). The parser (320) may parse / entropy decode the encoded video sequence. The parser (320) may extract a subgroup parameter set for at least one of the subgroups of pixels in the video decoder from the encoded video sequence. Subgroups may include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CU), blocks, Transform Units (TU), Prediction Units (PU), etc. The parser (320) may also extract information from the coded video sequence, such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, motion vectors, etc. The reconstruction of the symbols (321) may involve a plurality of different processing or functional units. The units involved and the manner in which these units are involved may be controlled by subgroup control information parsed by the parser (320) from the coded video sequence.
[0045] The first unit may include a scaler / inverse transform unit (351). The scaler / inverse transform unit (351) may receive quantized transform coefficients as symbols (321) and control information from the parser (320), including information indicating which type of inverse transform to use, block size, quantization factors / parameters, quantization scaling matrices, etc. The scaler / inverse transform unit (351) may output a block including sample values, which may be input into an aggregator (355).
[0046] In some cases, the output samples of the scaler / inverse transform unit (351) may belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed image, but may use predictive information from a previously reconstructed portion of a current image. Such predictive information may be provided by an intra-image prediction unit (352). In some cases, the intra-image prediction unit (352) may generate a block of the same size and shape as the block being reconstructed using reconstructed surrounding block information stored in a current image buffer (358). For example, the current image buffer (358) buffers a partially reconstructed current image and / or a fully reconstructed current image. In some embodiments, the aggregator (355) may add the prediction information generated by the intra-frame prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351) on a per-sample basis.
[0047] In other cases, the output samples of the scaler / inverse transform unit (351) may belong to an inter-frame coded and potentially motion compensated block. In this case, the motion compensated prediction unit (353) may access the reference image memory (357) based on the motion vector to extract samples for inter-frame image prediction. After the extracted reference samples are motion compensated according to the symbols (321) belonging to the block, these samples can be added to the output of the scaler / inverse transform unit (351) (the output of unit 351 can be called residual samples or residual signal) through an aggregator (355) to generate output sample information.
[0048] The output samples of the aggregator (355) may be subjected to various loop filtering techniques in the loop filter unit (356), which may include multiple types of loop filters. The output of the loop filter unit (356) may be a sample stream that may be output to the rendering device (312) and stored in the reference image memory (357) for subsequent inter-frame image prediction.
[0049] Figure 4 A block diagram of an example video encoder (403) according to an example embodiment disclosed in the present application is shown. The video encoder (403) may be included in an electronic device (420). The electronic device (420) may also include a transmitter (440) (e.g., a transmission circuit). The video encoder (403) may be used to replace Figure 4 A video encoder (403) in an example.
[0050] The video encoder (403) may receive video samples from a video source (401). According to some exemplary embodiments, the video encoder (403) may encode and compress images of a source video sequence into an encoded video sequence (443) in real time or under any other time constraints required by an application. Implementing an appropriate encoding speed constitutes a function of a controller (450). In some embodiments, the controller (450) is functionally coupled to and controls other functional units, as described below. The parameters set by the controller (450) may include rate control related parameters (picture skipping, quantizer, lambda value for rate distortion optimization techniques, etc.), picture size, picture group GOP layout, maximum motion vector search range, etc.
[0051] In some example embodiments, the video encoder (403) may be configured to operate in an encoding loop. The encoding loop may include a source encoder (430) and a (local) decoder (433) embedded in the video encoder (403). Because in the video compression techniques considered in the presently disclosed subject matter, any compression between symbols and the encoded video code stream in entropy coding can be lossless, even if the embedded decoder 433 processes the video stream encoded by the source encoder 430 without entropy coding, the decoder (433) reconstructs the symbols to create sample data in a manner similar to that which the (remote) decoder would create. At this point, it can be observed that any decoder technology other than parsing / entropy decoding, which may only exist in a decoder, may also need to exist in a corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter may sometimes focus on the operation of the decoder associated with the decoding portion of the encoder. Because the encoder technology is reciprocal to the decoder technology described in full, the description of the encoder technology can be simplified. The encoder is described in more detail only in certain areas or aspects below.
[0052] During operation, in some example embodiments, the source encoder (430) may perform motion compensated predictive coding, which predictively encodes an input picture with reference to one or more previously encoded pictures from a video sequence designated as "reference pictures."
[0053] The local video decoder (433) can decode the encoded video data of the picture that can be designated as the reference picture. The local video decoder (433) replicates the decoding process that can be performed by the video decoder on the reference picture and can cause the reconstructed reference picture to be stored in the reference picture cache (434). In this way, the video encoder (403) can store a copy of the reconstructed reference picture locally that has common content (absent transmission errors) with the reconstructed reference picture obtained by the far-end (remote) video decoder.
[0054] The predictor (435) may perform prediction search for the encoding engine (432). That is, for a new image to be encoded, the predictor (435) may search the reference image memory (434) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that may serve as appropriate prediction references for the new image.
[0055] The controller (450) may manage encoding operations of the source encoder (430), including, for example, setting parameters and subgroup parameters for encoding video data.
[0056] The outputs of all of the above functional units may be entropy encoded in an entropy encoder (445). The transmitter (440) may buffer the encoded video sequence created by the entropy encoder (445) in preparation for transmission over a communication channel (460), which may be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (440) may combine the encoded video data from the video encoder (403) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0057] The controller (450) may manage the operation of the video encoder (403). During encoding, the controller (450) may assign a certain coded picture type to each coded picture, but this may affect the coding techniques that may be applied to the corresponding picture. For example, pictures may generally be classified as one of the following picture types: intra pictures (I pictures), predicted pictures (P pictures), bidirectionally predicted pictures (B pictures), and multi-predicted pictures. As described in further detail below, a source picture may generally be spatially subdivided into a plurality of sample coding blocks.
[0058] Figure 5 A diagram of an example video encoder (503) according to another example embodiment disclosed herein is shown. The video encoder (503) is configured to receive a processed block (e.g., a predicted block) of sample values within a current video image in a video image sequence and encode the processed block into an encoded image that is part of an encoded video sequence. The example video encoder (503) may be used in place of Figure 4A video encoder (403) in an example.
[0059] For example, the video encoder (503) receives a matrix of sample values of a processing block. The video encoder (503) then uses, for example, rate-distortion optimization (RDO) to determine whether the best encoding mode for the processing block is intra mode, inter mode, or bi-prediction mode.
[0060] exist Figure 5 In the example of , the video encoder (503) includes Figure 5 An inter-frame encoder (530), an intra-frame encoder (522), a residual calculator (523), a switch (526), a residual encoder (524), a general controller (521), and an entropy encoder (525) coupled together are shown in the example arrangement of.
[0061] The inter-frame encoder (530) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference image (e.g., blocks in a previous image and a subsequent image in display order), generate inter-frame prediction information (e.g., redundant information description, motion vector, merge mode information according to an inter-frame coding technique), and calculate an inter-frame prediction result (e.g., a prediction block) based on the inter-frame prediction information using any suitable technique.
[0062] The intra encoder (522) is configured to receive samples of a current block (e.g., a processing block), compare the block with encoded blocks in the same image, and generate quantized coefficients after transformation, and in some cases also generate intra prediction information (e.g., intra prediction direction information based on one or more intra coding techniques).
[0063] The general controller (521) can be configured to determine general control data and control other components of the video encoder (503) based on the general control data, for example, to determine a prediction mode for a block and provide a control signal to the switch (526) based on the prediction mode.
[0064] The residual calculator (523) may be configured to calculate the difference (residual data) between the received block and the prediction result of the block selected from the intra encoder (522) or the inter encoder (530). The residual encoder (524) may be configured to encode the residual data to generate a transform coefficient. The transform coefficient is then subjected to a quantization process to obtain a quantized transform coefficient. In various example embodiments, the video encoder (503) further includes a residual decoder (528). The residual decoder (528) is configured to perform an inverse transform and generate decoded residual data. The entropy encoder (525) may be configured to organize the code stream into a structure including encoded blocks and perform entropy encoding.
[0065] Figure 6 A diagram of an example video decoder (610) according to another embodiment of the present disclosure is shown. The video decoder (610) is configured to receive an encoded image as part of an encoded video sequence and decode the encoded image to generate a reconstructed image. In an example, the video decoder (610) may be used instead of Figure 4 A video decoder (410) in an example of FIG.
[0066] exist Figure 6 In the example of FIG. 6 , the video decoder ( 610 ) includes Figure 6 An entropy decoder (671), an inter-frame decoder (680), a residual decoder (673), a reconstruction module (674), and an intra-frame decoder (672) coupled together are shown in the example arrangement of.
[0067] The entropy decoder (671) may be configured to reconstruct certain symbols from the encoded image, which symbols represent syntax elements constituting the encoded image. The inter-frame decoder (680) may be configured to receive inter-frame prediction information and generate an inter-frame prediction result based on the inter-frame prediction information. The intra-frame decoder (672) may be configured to receive intra-frame prediction information and generate a prediction result based on the intra-frame prediction information. The residual decoder (673) may be configured to perform inverse quantization to extract dequantized transform coefficients, and process the dequantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The reconstruction module (674) may be configured to combine the residual output by the residual decoder (673) with the prediction result (which may be output by the inter-frame prediction module or the intra-frame prediction module) in the spatial domain to form a reconstructed block, which forms a part of the reconstructed image, and the reconstructed image forms a part of the reconstructed video.
[0068] It should be noted that the video encoder (203), video encoder (403) and video encoder (503) and video decoder (210), video decoder (310) and video decoder (610) may be implemented using any suitable technology. In some example embodiments, the video encoder (203), video encoder (403) and video encoder (503) and video decoder (210), video decoder (310) and video decoder (610) may be implemented using one or more integrated circuits. In another embodiment, the video encoder (203), video encoder (403) and video encoder (503) and video decoder (210), video decoder (310) and video decoder (610) may be implemented using one or more processors that execute software instructions.
[0069] Therefore, the above-mentioned example video encoder and decoder implementation can include multiple independent video data processing components. Each of these components can be supported by one or more coding tools. Therefore, the improvement of these encoders and / or decoders can involve the optimization of one or more independent coding tools or data processing components. However, it may be difficult and inefficient to perform system-level optimization. Therefore, due to the association between these components, the coding gain that may be obtained will be difficult to achieve.
[0070] In some other example embodiments, instead of using Figures 3 to 6 Instead of the example video encoding and decoding architecture shown, a video encoding / decoding system can be constructed based on the inference of an AI model. Such a video encoding / decoding system may include one or more AI models that can be trained jointly or jointly in an end-to-end (E2E) manner. In particular, such an AI model can be implemented based on a neural network such as a deep learning neural network (DNN), and accordingly, such an E2E AI-based video encoding / decoding system can be referred to as a neural image compression (NIC).
[0071] The general example framework of NIC is described below. For an input image x, the goal of NIC is to use the image as input to a DNN encoder to compute a compressed representation This compressed representation is compact for storage and transmission. Then, use As input to the DNN decoder to reconstruct the image In some example embodiments, the NIC method may employ a variational autoencoder (VAE) structure, in which a DNN encoder directly uses the entire image x as its input, which is passed through a set of network layers to compute an output compressed representation These network layers contain connected neuron units and work like black boxes. Accordingly, the example DNN decoder can represent the entire compressed representation As its input, this compressed representation is passed through another set of network layers to compute the reconstructed These network layers contain another set of connected neuron units and work like another black box. During the iterative training process, the rate-distortion (RD) loss can be optimized to achieve the distortion loss of the reconstructed image using a trade-off hyperparameter λ. With compact representation Therefore, the loss function of the example can be expressed as:
[0072]
[0073] Figure 7 An example architecture of a generalized NIC model is shown. The NIC model includes a first subnetwork containing a core autoencoder for learning a quantized latent representation of an image and a subnetwork for learning a probabilistic model on the quantized latent representation for entropy coding. Figure 7 In, the first subnetwork of the core autoencoder includes an "encoder" box and a "decoder" box. The second subnetwork includes a "context model", a supernetwork (including a "superencoder" box and a "superdecoder" box). The context model can be used as an autoregressive model to process the quantized latent representation. The supernetwork can be configured to learn a representation of information that corrects context-based predictions. The output of the context model and the output of the supernetwork can be combined by an "entropy parameter" network to generate, for example, mean and scale parameters of a conditional Gaussian entropy model.
[0074] exist Figure 7 AE stands for the Arithmetic Encoding block, which generates a compressed representation of the symbols from the quantizer. Figure 7 denoted by the “Q” box in the figure. For decoding, once any information that depends on the quantized latent representation is decoded, it can be used by the decoder box. The context model should only access the decoded quantized latent representation (arrows between Arithmetic Decoding (AD) boxes).
[0075] exist Figure 7 In further detail of an example implementation of , the input of the encoder is represented by the input image x 702. The output of the encoder 704 is a latent representation represented by y, which becomes a quantized latent representation of the input image after passing through a quantizer Q. The latent representation y is input into the hyperencoder to generate a hyper latent z, which is quantized by the quantizer to generate a quantized hyper latent The quantized latent representation may be processed by the arithmetic encoder AE 708 To generate coded bits 730. The quantized super latent representation may also be encoded by AE 716 To generate the super potential coded bits 720. The arithmetic decoder 724 may be configured to decode the coded bits 720 to generate the reconstructed quantized super potential The arithmetic encoder 710 and arithmetic decoder 724 for the super latent representation may be assisted by the factorized entropy model 718. The reconstructed super latent representation may be processed and decoded by the super decoder 726. The output of the super decoder 726 and the output of the context model 710 may form an entropy parameter 728 to assist the arithmetic decoder AD 732 in processing the coded bits 730 to generate a decoded quantized latent representation. The decoded quantized latent representation will be processed by the decoder 734 to generate a reconstructed image The encoder 704 and the decoder 734 may alternatively be referred to as a core encoder and a core decoder, respectively (to distinguish from the super encoder 712 and the super decoder 726).
[0076] Thus, to encode the input image 702 to generate bits 730 and 720, it will involve Figure 7 All blocks except decoder 734 in FIG. 7 include a decoding branch including factorized entropy model network 718, AD 724, super decoder 726, entropy parameter network 728, and AD 732 to reconstruct the quantized latent representation Feedback to context model 710. The decoding path for bit 730 and bit 720 involves the above decoding branch and additionally includes context model 710 (which only needs to access the reconstructed quantized latent representation ).
[0077] exist Figure 7 In some example embodiments, the encoder 704 includes a set of visual modules (such as convolutional networks or visual transformers) having downsampling capabilities, and the decoder 734 may include a set of visual modules (such as convolutional networks or visual transformers) having upsampling capabilities.
[0078] The further disclosure below relates to various example implementations of the above-described superdecoder 726 and decoder 734. These examples provide, among other things, various lightweight superdecoder and / or decoder structures for decoding the encoded bits 720 and the encoded bits 730. Moreover, because, among other things, the superdecoder 726 is part of the encoding process, these example implementations additionally help balance the complexity of the network involved in the decoding branch with the encoding efficiency (in terms of the size of the bits 720 and the bits 730).
[0079] Figure 8 800 is an example implementation of the super decoder 726 as a neural network 800. The data processing path in the example super decoder 800 is shown as flowing from bottom to top. The input 801 of the super decoder 800 can be a quantized super latent The super decoder 800 may include a first convolutional layer 802, followed by an upsampling layer 804, a cropping layer 806, and an activation function layer 808. These convolutions, upsampling, cropping, and activations may be repeated one or more times to recover the Figure 7The resolution is reduced by downsampling in the corresponding super encoder 712, for example, Figure 8 The exemplary single repetition shown in 812 to 818 in . The final layer of the exemplary super decoder 800 may include a convolution layer 822 and an activation function layer 814. The term "layer" may alternatively be referred to as a "network", and thus, a convolution layer, an upsampling layer, a cropping layer, an activation function layer, etc. may alternatively be referred to as a convolution network, an upsampling network, a cropping network, an activation function network, etc.
[0080] In some example embodiments, the Figure 8 The first convolutional layer or network 802 (as shown by the dashed outline of 802). Thus, the resulting example network includes two upsampling stages and two convolutional layers, rather than three convolutional layers. In some other example embodiments, there may be more than two upsampling stages (each upsampling stage, for example, includes an upsampling layer, a cropping layer, and an activation function layer), and a convolutional layer may be included between each adjacent upsampling stage. These upsampling stages are followed by a final convolutional layer 822 and an activation function layer and 824 before outputting the learned parameters. Upsampling networks such as 804 and 814 can be implemented using any type of upsampling method. For example, upsampling can be performed using one or more convolutional networks. For another example, upsampling can be performed based on pixel shuffling. In addition, activation function layers such as 808 and 818 can be implemented based on any suitable activation form. By removing Figure 8 The first convolutional layer 802 in the CNN helps reduce the complexity of the super decoder. However, the information embedded in the neural network parameters associated with the first convolutional layer 802 can be captured by the network parameters in the upsampling stage (by using a convolutional network or a pixel shuffle network). Therefore, removing the first convolutional layer can reduce the complexity of the model and reduce the inference time without sacrificing performance too much.
[0081] In some other example embodiments, one or more of the cropped layers or cropped networks 806 and 816 may be removed. Specifically, as a result of cropping (reducing the number of data points), the amount of computational gain for the super decoder 800 may be negligible (e.g., edge points will be zero-filled, so removing these edge points may not significantly save the computation of the network s in the super decoder 800), and removing such cropped layers will make the model smaller and have fewer model parameters, thereby improving the training process and inference time.
[0082] Fig. 9 7 shows an example implementation of a decoder 734 as a neural network 900. The data processing path in the example decoder 900 is shown as flowing from bottom to top. The input 901 of the decoder neural network 900 can be a quantized potential The super decoder neural network 900 may include multiple upsampling stages, for example, including a first upsampling network 902, a second upsampling network 912, and a third upsampling network 924. Each upsampling network may be followed by a cropping network and then an activation function network. In other words, each pair of consecutive upsampling stages may be connected at least by a cropping network and an activation network. For example, the last upsampling stage 934 may be followed by a cropping network 934. For example, the upsampling network 902 may be followed by a cropping network 904 and then an activation function network 906. Similarly, the upsampling network 912 may be followed by a cropping network 914 and then an activation function network 916. One or more convolutional networks with activation networks may also be included and arranged between the upsampling stages. Fig. 9 , it is shown that the convolution network 922 and the activation network 924 are located between the second upsampling stage and the third upsampling stage.
[0083] exist Fig. 9 In the example decoder neural network 900 of , each upsampling stage is configured to increase the resolution of the output image by inserting new pixels between existing pixels in each previous stage. Each upsampling stage can be characterized by an upsampling factor, which represents the ratio between the number of pixels after the upsampling process and the number of pixels before the upsampling process in a single dimension. A 2D image is upsampled by a factor of 2 when the size of the image is doubled in height and width. For example, when the upsampling factor used is 3, an image that was originally 200×200 pixels will be upsampled to 400×400 pixels. Similarly, if a 2D image is upsampled by a factor of 4, it means that the size of the image is quadrupled in both height and width. For example, when the upsampling factor used is 4, an image that was originally 200×200 pixels will be upsampled to 800×800 pixels.
[0084] The upsampling process in the decoder can be performed in stages. Fig. 9 In the example of , three upsampling stages are involved. The three upsampling stages can be configured to achieve a combined upsampling of a factor of 16. In an example, the distribution of upsampling factors in the three upsampling stages can be 2, 2, and 4. Alternatively, the distribution of upsampling factors can be 2, 4, and 2, or 4, 2, and 2. Upsampling stage.
[0085] In some other example embodiments, 4 upsampling stages are involved to achieve an upsampling factor of 16, wherein each upsampling stage is configured to achieve an upsampling factor of 2. Compared to the four-stage implementation, the three-stage implementation helps to reduce the overall complexity of the decoder network, thereby facilitating the training of the model and reducing the inference time. Similar to the above three-stage implementation, each pair of consecutive upsampling stages can be connected at least by a cropping network and an activation network. The last upsampling stage can be followed by a cropping network. In some example embodiments, a convolutional network and a cropping layer can also be included between consecutive pairs of upsampling stages in a four-stage decoder.
[0086] In some example embodiments, the decoder may include only two upsampling stages. The two stages may be connected via at least a pruning network and an activation function network. The second upsampling stage may be followed by a pruning network.
[0087] Each of the above upsampling stages can be based on any upsampling process. For example, each upsampling stage (or Fig. 9 The upsampling method layer) can be implemented based on pixel shuffling with a regular convolutional network (for changing channels), or a transposed convolutional network, or other suitable upsampling networks.
[0088] exist Fig. 9 In the example of , a conventional convolutional layer or convolutional network 922 is added between some adjacent upsampling stages, for example, between the second upsampling stage and the third upsampling stage. In some other example embodiments, such a conventional convolutional layer or convolutional network may be added between the first upsampling stage and the second upsampling stage. In some other example embodiments, such a conventional convolutional layer or convolutional network may be removed (e.g. Fig. 9 ), so that the information carried in the model parameters of these layers can be embedded into the parameters of other network layers during the training process, thereby reducing the overall complexity of the decoder neural network 900. Similarly, the activation function layer or activation function network 924 between the corresponding upsampling stages can also be removed together with the convolution layer 922 to further reduce the complexity of the model (as shown by the dotted outline of element 924). After removing the convolution layer 922 and the activation function layer 924, Fig. 9 The example three-stage decoder neural network 900 will include three upsampling stages with pruning networks and activation function networks (including 902, 904, 906, 912, 914, 924, 932 and 934) in between.
[0089] In some example embodiments, the encoder, super encoder, super decoder, and decoder (e.g. Figure 7As shown) can be implemented using different designs for different image components (such as Y and UV components or R, G, and B components). Each of the encoder, super encoder, super decoder, and decoder can be designed as a separate network for a color component. In some embodiments, the UV (or chrominance component) of the image can be encoded, super encoded, decoded, and super decoded by the network. For example, for a YUV image, two groups of networks can be designed and trained, where one group of networks is used for the Y component and one group of networks is used for the U and V components. For another example, for an RGB image, three groups of networks can be designed and trained, where one group of networks is used for each of the R, G, and B components. In Figure 7 In this paper, these independent networks for color components can refer to each other in terms of processing context models and entropy parameters during encoding, decoding, super-encoding and super-decoding.
[0090] Figure 7 The components of can be trained jointly using training images, or can be trained in stages, where in each training stage, the model parameters of some components are fixed while the model parameters of other components can be optimized. The optimized models are shuffled between training stages, which can be performed iteratively.
[0091] The various embodiments above, particularly those related to the neural network architecture for the super decoder and decoder, help reduce model complexity and achieve a balance between model complexity and a reasonable compression rate.
[0092] Fig.10 An example logic flow 1000 according to the above-described embodiment is shown. The logic flow is performed by a decoder including a super decoder, a core decoder, and a context model to decode an encoded bitstream of an image. The logic flow 1300 starts from S1001. In S1010, a super data item set is generated according to the encoded bitstream by the super decoder. In S1020, a context parameter set is generated by the context model. In S1030, a potential representation sample of the image is generated by processing the encoded bitstream and an entropy parameter set generated according to the super data item set and the context parameter set. In S1040, an image sample is reconstructed according to the potential representation sample of the image by a core decoder, wherein the core decoder includes a neural network having three upsampling stages or less. The logic flow 1300 stops at S1099.
[0093] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. Fig.11 A computer system (1100) suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0094] Computer software may be encoded using any suitable machine code or computer language, which may be assembled, compiled, linked or similarly constructed to create code comprising instructions that may be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or through interpretation, microcode, etc.
[0095] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, IoT devices, etc.
[0096] Fig.11 The components of the computer system (1100) shown are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing embodiments of the present disclosure. Neither should the configuration of the components be interpreted as having any dependency or requirement relating to any one or combination of components shown in the exemplary embodiment of the computer system (1100).
[0097] The computer system (1100) may include certain human-machine interface input devices. The input human-machine interface devices may include one or more of the following (only one of each is shown): keyboard (1101), mouse (1102), touch pad (1103), touch screen (1110), data gloves (not shown), joystick (1105), microphone (1106), scanner (1107), camera (1108).
[0098] The computer system (1100) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate one or more senses of a human user, for example, through tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., touch screen (1110), tactile feedback of data gloves (not shown) or joystick (1105), but may also be tactile feedback devices that are not input devices), audio output devices (e.g., speakers (1109), headphones (not shown)), visual output devices (e.g., screens (1110) including cathode ray tube (CRT) screens, liquid crystal display (LCD) screens, plasma screens, organic light-emitting diode (OLED) screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which are capable of outputting two-dimensional visual outputs or more than three-dimensional outputs through devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays and smoke boxes (not depicted), and printers (not depicted).
[0099] The computer system (1100) may also include human-accessible storage devices and their associated media: for example, optical media including CD / DVD read-only memory (ROM) / RW (1120) having media (1121) such as compact disc (CD) / digital video disc (DVD), thumb drive (1122), removable hard disk drive or solid state drive (1123), traditional magnetic media such as magnetic tapes and floppy disks (not shown), dedicated ROM / application-specific integrated circuit (ASIC) / programmable logic device (PLD) based devices such as security software dogs (not shown), etc.
[0100] Those skilled in the art should also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transient signals.
[0101] The computer system (1100) may also include an interface (1154) to one or more communication networks (1155). The network may be, for example, a wireless network, a wired network, an optical network. The network may further be a local network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless local area networks (LAN), cellular networks including global system for mobile communications (GSM), third generation (3G), fourth generation (4G), fifth generation (5G), long term evolution (LTE), etc., television wired or wireless wide area digital networks including cable television, satellite television and terrestrial broadcast television, vehicle and industrial television including controller area network bus (CANbus), etc.
[0102] The above-mentioned human-machine interface device, human-machine accessible storage device, and network interface may be attached to the kernel ( 1140 ) of the computer system ( 1100 ).
[0103] The core (1140) may include one or more CPUs (1141), GPUs (1142), dedicated programmable processing units in the form of Field Programmable Gate Areas (FPGA) (1143), hardware accelerators (1144) for certain tasks, graphics adapters (1150), etc. These devices as well as ROM (1145), random-access memory (RAM) (1146), internal mass storage (1147) such as internal non-user accessible hard disk drives, solid state drives (SSD), etc. can be connected via a system bus (1148). In some computer systems, the system bus (1148) can be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be connected directly to the core's system bus (1148) or connected to the core's system bus via a peripheral bus (1149). In one example, a screen (1110) can be connected to a graphics adapter (1150). The structure of the peripheral bus includes Peripheral Component Interconnection (PCI), USB, etc.
[0104] The computer readable medium may have thereon computer codes for performing various computer-implemented operations. The medium and computer codes may be those specially designed and constructed for the purposes of the present disclosure, or the medium and computer codes may be of a type well known and available to those skilled in the art of computer software.
[0105] Although the present disclosure has described a number of exemplary embodiments, there are modifications, permutations, and various replacement equivalents that fall within the scope of the present disclosure. Therefore, it should be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and therefore fall within the spirit and scope of the present disclosure.
Claims
1. A method for a decoder to decode an encoded bitstream of an image, the decoder comprising a super decoder, a core decoder and a context model, the method comprising: generating a set of super data items according to the encoded bitstream by the super decoder; Generate a context parameter set using the context model; Generate a potential representation sample of the image by processing the encoded code stream and an entropy parameter set generated according to the super data item set and the context parameter set; as well as reconstructing, by the core decoder, image samples from the latent representation samples of the image, Wherein, the core decoder comprises a neural network having three upsampling stages or less.
2. The method according to claim 1, wherein: The core decoder comprises three upsampling stages, wherein each pair of consecutive upsampling stages in the three upsampling stages are connected via at least a pruning neural network layer and an activation neural network layer.
3. The method according to claim 2, wherein: The single-dimensional upsampling ratios of the three upsampling stages are 2, 2 and 4 respectively.
4. The method according to claim 2, wherein: The single-dimensional upsampling ratios of the three upsampling stages are 2, 4 and 2 respectively.
5. The method according to claim 2, wherein: The single-dimensional upsampling ratios of the three upsampling stages are 4, 2 and 2 respectively.
6. The method according to claim 2, wherein: There are no convolutional neural network layers arranged between the three upsampling stages.
7. The method according to claim 1, wherein: The core decoder includes two upsampling stages connected by a pruning neural network layer, an activation neural network layer, and a convolutional neural network layer.
8. The method according to claim 7, wherein: The single-dimensional upsampling ratios of the two upsampling stages are 4 and 4 respectively.
9. The method according to any one of claims 1 to 8, wherein: The super decoder comprises at least two upsampling stages, and a first upsampling stage in sequence is directly connected to entropy decoded data items from the encoded bitstream.
10. The method according to claim 9, wherein: Each pair of adjacent upsampling stages are connected by at least a pruning neural network layer, an activation neural network layer, and a convolutional neural network layer.
11. The method according to claim 9, wherein: At least one pair of adjacent upsampling stages does not include pruned neural network layers.
12. The method according to claim 9, wherein: At least one of the at least two upsampling stages comprises a convolutional neural network layer.
13. The method according to claim 9, wherein: At least one of the at least two upsampling stages comprises a pixel shuffle neural network having a convolutional neural network layer.
14. The method according to claim 9, wherein: At least one of the at least two upsampling stages comprises a transposed convolutional neural network layer.
15. The method according to any one of claims 1 to 8, wherein: Generate the encoded bitstream for the Y component, U component and V component of the image; as well as The Y component and the U and V components of the image are decoded using a separate model of the super decoder and a separate model of the core decoder.
16. The method according to any one of claims 1 to 8, wherein: Generate the encoded code stream for the R component, the G component and the B component of the image; as well as The R component, the G component, and the B component of the image are decoded using a separate model of the super decoder and a separate model of the core decoder.
17. A decoder for decoding an encoded code stream of an image, the decoder comprising a memory and at least one processor, the memory being used to store computer instructions, the at least one processor being configured to execute the computer instructions to: generating a set of super data items according to the encoded bitstream by a super decoder of the decoder; generating a set of context parameters by a context model of the decoder; Generate a potential representation sample of the image by processing the encoded code stream and an entropy parameter set generated according to the super data item set and the context parameter set; and reconstructing, by a core decoder of the decoder, image samples from the latent representation samples of the image, in, The core decoder includes a neural network having three upsampling stages or less.
18. The decoder of claim 17, comprising three upsampling stages, wherein each pair of consecutive upsampling stages in the three upsampling stages are connected at least by a pruning neural network layer and an activation neural network layer.
19. The decoder of claim 17, comprising at least two upsampling stages, and a first upsampling stage in sequence is directly connected to the entropy decoded data items from the encoded bitstream.
20. A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to: Generate a set of super data items according to the encoded bitstream of the image by a super decoder of the decoder; generating a set of context parameters by a context model of the decoder; Generate a potential representation sample of the image by processing the encoded code stream and an entropy parameter set generated according to the super data item set and the context parameter set; and reconstructing, by a core decoder of the decoder, image samples from the latent representation samples of the image, in, The core decoder includes a neural network having three upsampling stages or less.