Learning image compression based on fast residual channel attention network

By introducing FRCAN and DSC networks into the post-processing network of the video encoding system, combining residual learning and channel attention, the problem of limited video compression performance in the prior art is solved, and a more efficient compression and decompression process is achieved, and image quality is improved.

CN120035840APending Publication Date: 2025-05-23GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280100889.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-13
Filing Date
2022-12-01
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing CNN-based video compression methods have performance limitations in extracting low-level local detail recovery, resulting in poor image quality after compression.

Method used

A video encoding system with mixed network structure is designed, combining residual learning and channel attention to be applied to the post-processing network of video encoding network. The system includes FRCAN and DSC networks for generating information image features and compensating for feature loss during compression.

Benefits of technology

The compression speed and decompression speed of the video encoding system are improved, which are about 1.5 times and 1.23 times respectively, and the performance of PSNR and MS-SSIM is improved at high bit rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120035840A_ABST
    Figure CN120035840A_ABST
Patent Text Reader

Abstract

According to an aspect of the present disclosure, a video post-processing method may include a processor receiving a plurality of input feature maps associated with an image. A plurality of input feature maps may be generated by a video preprocessing network. A video post-processing method may include a processor inputting a plurality of input feature maps into a first depth separable convolution (DSC) network of a fast residual channel attention network (FRCAN) component. A video post-processing method may include a processor outputting a first set of output feature maps from a first DSC network of an FRCAN component.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] The disclosed embodiments relate to video coding.

[0002] Digital video has become mainstream and is widely used in various applications such as digital television, video phones, and teleconferencing. These digital video applications are feasible due to advances in computing and communication technology and efficient video coding techniques. Various video coding techniques can be used to compress video data, so that the video data can be encoded using one or more video coding standards. Exemplary video coding standards may include, but are not limited to, versatile video coding (H.266 / VVC), high-efficiency video coding (H.265 / HEVC), advanced video coding (H.264 / AVC), moving picture expert group (MPEG) coding, etc. Summary of the invention

[0003] According to one aspect of the present disclosure, a video post-processing method is provided. The method may include a processor receiving a plurality of input feature maps associated with an image. The plurality of input feature maps may be generated by a video pre-processing network. The method may include the processor inputting the plurality of input feature maps into a first depth-wise separable convolutional (DSC) network of a fast residual channel attention network (FRCAN) component. The method may include the processor outputting a first set of output feature maps of a first DSC network from the FRCAN component.

[0004] According to another aspect of the present disclosure, a video post-processing system is provided. The system may include a memory for storing instructions. The system may include a processor coupled to the memory, which is used to receive multiple input feature maps associated with an image when executing the instructions. The multiple input feature maps can be generated by a video pre-processing network. The system may include a processor coupled to the memory, which is used to input the multiple input feature maps into a first DSC network of a FRCAN component when executing the instructions. The system may include a processor coupled to the memory, which is used to output a first group of output feature maps from the first DSC network of the FRCAN component when executing the instructions.

[0005] According to another aspect of the present disclosure, a video compression method is provided. The method may include a processor using a preprocessing network to preprocess an input image to generate an encoded image. The method may include a processor using a postprocessing network to postprocess the encoded image to generate a decoded compressed image. The preprocessing network and the postprocessing network may be asymmetric.

[0006] According to another aspect of the present disclosure, a video compression system is provided. The system may include a memory for storing instructions. The system may include a processor coupled to the memory, which, when executing the instructions, is used to pre-process an input image using a pre-processing network to generate a coded image. The system may include a processor coupled to the memory, which, when executing the instructions, is used to post-process the coded image using a post-processing network to generate a decoded compressed image. The pre-processing network and the post-processing network may be asymmetric.

[0007] These illustrative embodiments are mentioned not to limit or define the present disclosure, but to provide examples to aid understanding of the present disclosure. Other embodiments are described in the detailed description and further description is provided. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The accompanying drawings, which are incorporated herein and constitute a part of the specification, illustrate embodiments of the present disclosure and, together with the description, further explain the principles of the present disclosure and enable one skilled in the relevant art to make and use the present disclosure.

[0009] Figure 1 A block diagram of an exemplary encoding system according to some embodiments of the present disclosure is shown.

[0010] Figure 2 A block diagram of an exemplary decoding system according to some embodiments of the present disclosure is shown.

[0011] Figure 3 A detailed block diagram of an exemplary video encoding network according to some embodiments of the present disclosure is shown.

[0012] Figure 4 The deployment of some embodiments of the present disclosure is shown in Figure 3 Detailed block diagram of an exemplary channel autoregressive entropy model in an exemplary video coding network of FIG.

[0013] Figure 5 It shows that according to some embodiments of the present disclosure, Figure 2 An exemplary decoder of performs depthwise separable convolution (DSC).

[0014] Figure 6 The present disclosure shows that the DSC network is deployed in the Figure 2Residual channel attention block (RCAB) of an exemplary FRCAN component in an exemplary decoder of FIG.

[0015] Figure 7 The deployment of some embodiments of the present disclosure is shown in Figure 2 An exemplary residual upsampling component in an exemplary decoder of .

[0016] Figure 8 The deployment of some embodiments of the present disclosure is shown in Figure 2 An exemplary residual dense block (RRDB) component in an exemplary decoder of .

[0017] Fig. 9 Showing the use of some aspects of the present disclosure Figure 3 Graphical representation of the peak signal-to-noise ratio (PSNR) rate-distortion (RD) performance of video coding for an exemplary video coding network implementation of FIG.

[0018] Fig.10 According to some aspects of the present disclosure, Figure 3 Graphical representation of the Multi-Scale Structural Similarity (MS-SSIM) RD performance of video coding achieved by an exemplary video coding network.

[0019] Fig.11 A graphical representation of PSNR versus bit rate for video compression using and without an exemplary FRCAN is shown in accordance with aspects of the present disclosure.

[0020] Fig.12 A graphical representation of MS-SSIM versus bit rate for video compression using and without an exemplary FRCAN is shown in accordance with aspects of the present disclosure.

[0021] Fig.13 A flowchart illustrating a first exemplary video encoding method according to some aspects of the present disclosure is shown.

[0022] Fig.14 A flowchart of a second exemplary video encoding method according to some aspects of the present disclosure is shown.

[0023] Embodiments of the present disclosure will be described with reference to the accompanying drawings. DETAILED DESCRIPTION

[0024] Although some configurations and arrangements are discussed, it should be understood that this is for illustrative purposes only. Those skilled in the art will recognize that other configurations and arrangements can be used without departing from the spirit and scope of the present disclosure. Those skilled in the art will appreciate that the present disclosure can also be used in a variety of other applications.

[0025] Note that the phrases "one embodiment", "embodiment", "example embodiment", "some embodiments", "certain embodiments", etc. mentioned in the specification indicate that the described embodiments may include certain features, structures, or characteristics, but not every embodiment necessarily includes the certain features, structures, or characteristics. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a certain feature, structure, or characteristic is described in conjunction with an embodiment, whether or not it is explicitly described, the relevant technicians should know that such feature, structure, or characteristic can be implemented in conjunction with other embodiments.

[0026] In general, terms can be understood at least in part based on usage in context. For example, depending at least in part on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in a singular sense, or can be used to describe a combination of features, structures, or characteristics in a plural sense. Similarly, depending at least in part on the context, terms such as "one," "an," or "the" can also be understood to express singular or plural usage. In addition, also depending at least in part on the context, the term "based on" can be understood to not necessarily express a set of exclusive factors, but can allow for the presence of other factors that are not necessarily explicitly described.

[0027] Various aspects of the video encoding system will now be described with reference to various devices and methods. These devices and methods will be described in the following detailed description and illustrated in the accompanying drawings by various modules, components, circuits, steps, operations, processes, algorithms, etc. (collectively referred to as "elements"). These elements can be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether these elements are implemented as hardware, firmware, or software depends on the specific application and design constraints imposed on the overall system.

[0028] The techniques described herein can be used in various video coding applications. As described herein, video coding includes encoding and decoding a video. The encoding and decoding of a video can be performed in units of blocks. For example, a coding block, a transform block, or a prediction block can be subjected to encoding / decoding processing, such as transformation, quantization, prediction, loop filtering, reconstruction, etc. As described herein, the block to be encoded / decoded will be referred to as a "current block". For example, a current block may represent a coding block, a transform block, or a prediction block according to a current encoding / decoding process. In addition, it should be understood that the term "unit" used in the present disclosure represents a basic unit for performing a specific encoding / decoding process, and the term "block" represents a sample array of a predetermined size. Unless otherwise specified, "block", "unit", and "component" can be used interchangeably.

[0029] Current image compression methods can be divided into two categories: traditional image compression (such as JPEG, JPEG2000, BPG) and recent deep learning-based image compression.

[0030] Traditional image compression uses module-based encoder / decoder (codec) blocks to eliminate spatial redundancy and improve image coding efficiency. To this end, these methods may use fixed transform matrices, intra-frame prediction units, quantization units, adaptive arithmetic encoders, and various deblocking filters or loop filters. With the rapid development of new image formats and the popularity of high-resolution mobile devices, new video coding technologies need to be developed to replace existing image compression standards.

[0031] In recent years, variational auto-encoder (VAE)-based learned image compression (also known as “convolutional neural network (CNN)-based image compression”) has achieved better rate-distortion performance than traditional image compression methods in terms of peak-signal-to-noise ratio (PSNR) and multi-scale-structural similarity (MS-SSIM), showing great potential for practical compression uses.

[0032] For encoding, VAE-based image compression methods use linear and nonlinear parametric transformations to map images to latent spaces. After quantization, the entropy estimation model predicts the distribution of the latent data, and then the lossless context-based adaptive binary arithmetic coding (CABAC) or range encoder compresses the latent data into the bitstream. At the same time, hyper-prior, autoregressive prior, and Gaussian mixture model (GMM) allow the entropy estimation component to accurately predict the distribution of the latent data and achieve better rate-distortion (RD) performance. For decoding, the lossless CABAC or range encoder decompresses the bitstream; the decompressed latent data is then mapped to the reconstructed image through linear and nonlinear parametric synthesis transformations. Combined with the above-mentioned sequential units, these models can be trained end-to-end.

[0033] A core problem of existing CNN-based compression methods is that the original convolutional layers are designed for high-level global feature extraction rather than low-level local detail restoration, which inevitably limits further performance improvements.

[0034] In order to overcome these challenges and other challenges of CNN-based compression, the present disclosure provides an exemplary CNN-based compression network designed with a hybrid network structure. The hybrid network structure can be based on residual learning and channel attention, which is applied to a post-processing network (also referred to as a "decoder" in an end-to-end video coding network based on deep learning (also referred to as a "video coding system" herein). To this end, the present disclosure provides FRCAN and DSC networks with channel attention (CA) layers to increase the processing speed of the post-processing network while generating information features in the image. By deploying a DSC network in a post-processing network with residual learning for upsampling, the video coding system of the present disclosure captures image features lost during decoding, thereby reducing the compression ratio. Compared with existing end-to-end image coding systems, the proposed video coding system increases the compression speed and decompression speed by approximately 1.5 times and 1.23 times, respectively, while achieving gains in PSNR and MS-SSIM at high bit rates.

[0035] In addition, the exemplary video coding system of the present disclosure uses an asymmetric encoding and decoding framework to improve encoding efficiency and decoding quality. The asymmetric encoding and decoding framework can refer to different types of convolutions performed in the encoder and decoder. For example, the encoder can use standard convolutions, while the decoder uses depth-separable convolutions. This framework has two advantages: 1) simplifies encoding, thereby increasing encoding speed and reducing compressed bitstream; 2) compensates for information lost during compression and improves the quality of decoded images by using a complex DSC network.

[0036] In addition, the DSC network deployed in the post-processing network described in this paper improves the network speed while reducing the network parameters. The residual learning performed by the post-processing network can generate additional features that have been lost, thereby reducing the bitstream after image compression. For example, the CA network of the FRCAN component can capture information features that may be omitted by the feature map of the pre-processing network, thereby improving runtime latency and visual quality. Figures 1 to 14 Further details of an exemplary video encoding system of the present disclosure are provided.

[0037] Figure 1 A block diagram of an exemplary encoding system 100 is shown according to some embodiments of the present disclosure. Figure 2A block diagram of an exemplary decoding system 200 according to some embodiments of the present disclosure is shown. Each system 100 or 200 can be applied to or integrated into various systems and devices capable of data processing, such as computers and wireless communication devices. For example, the system 100 or 200 can be all or part of a mobile phone, a desktop computer, a laptop computer, a tablet computer, a car computer, a game console, a printer, a positioning device, a wearable electronic device, a smart sensor, a virtual reality (VR) device, an augmented reality (AR) device, or any other suitable electronic device with data processing capabilities. Figure 1 and Figure 2 As shown, system 100 or 200 may include processor 102, memory 104, and interface 106. These components are shown as being interconnected via a bus, but other connection types are also allowed. It is understood that system 100 or 200 may include any other suitable components for performing the functions described herein.

[0038] Processor 102 may include a microprocessor, such as a graphics processing unit (GPU), an image signal processor (ISP), a central processing unit (CPU), a digital signal processor (DSP), a tensor processing unit (TPU), a vision processing unit (VPU), a neural processing unit (NPU), a synergistic processing unit (SPU) or a physical processing unit (PPU), a microcontroller unit (MCU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic device (PLD), a state machine, gated logic, discrete hardware circuits, and other suitable hardware for performing the various functions described in the present disclosure. Although Figure 1 and Figure 2Only one processor is shown, but it is understood that multiple processors may be included. Processor 102 may be a hardware device having one or more processing cores. Processor 102 may execute software. Whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise, software should be broadly understood as instructions, instruction sets, codes, code segments, program codes, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, processes, functions, etc. Software may include computer instructions written in an interpreted language, a compiled language, or machine code. Other techniques for directing hardware are also permitted under the broad category of software.

[0039] Memory 104 may broadly include memory (also referred to as main / system memory) and storage (also referred to as secondary memory). For example, memory 104 may include random-access memory (RAM), read-only memory (ROM), static RAM (SRAM), dynamic RAM (DRAM), ferroelectric RAM (FRAM), electrically erasable programmable ROM (EEPROM), compact disc read-only memory (CD-ROM) or other optical storage, hard disk drive (HDD) (such as magnetic disk storage or other magnetic storage device), flash drive, solid-state drive (SSD), or any other medium that can be used to carry or store the desired program code in the form of instructions that can be accessed and executed by processor 102. In a broad sense, memory 104 can be implemented by any computer-readable medium, such as a non-transitory computer-readable medium. Although Figure 1 and Figure 2 Only one memory is shown, but it will be appreciated that multiple memories may be included.

[0040] The interface 106 may broadly include a data interface and a communication interface for receiving and sending signals in the process of receiving and sending information with other external network elements. For example, the interface 106 may include an input / output (I / O) device and a wired or wireless transceiver. Figure 1 and Figure 2 Only one memory is shown, but it will be appreciated that multiple interfaces may be included.

[0041] The processor 102, memory 104, and interface 106 can be implemented in various forms in the system 100 or 200 to perform video encoding functions. In some embodiments, the processor 102, memory 104, and interface 106 of the system 100 or 200 are implemented (e.g., integrated) on one or more system-on-chips (SoCs). In one example, the processor 102, memory 104, and interface 106 can be integrated on an application processor (AP) SoC, which is responsible for application processing in an operating system (OS) environment, including running video encoding and decoding applications. In another example, the processor 102, memory 104, and interface 106 can be integrated on a dedicated processor chip such as a GPU or ISP chip dedicated to image and video processing in a real-time operating system (RTOS) for video encoding.

[0042] like Figure 1 As shown, in the encoding system 100, the processor 102 may include one or more modules, such as an encoder 101 (also referred to herein as a "pre-processing network"). Figure 1 The encoder 101 is shown to be located within a processor 102, but it is understood that the encoder 101 may include one or more submodules that may be implemented on different processors that are close to or far from each other. The encoder 101 (and any corresponding submodules or subunits) may be a hardware unit (e.g., part of an integrated circuit) of the processor 102 designed for use with other components, or a software unit implemented by the processor 102 by executing at least a portion of a program (i.e., instructions). The instructions of the program may be stored on a computer-readable medium such as a memory 104, and when executed by the processor 102, a process having one or more functions related to video encoding may be performed, such as image partitioning, inter-frame prediction, intra-frame prediction, transform, quantization, filtering, entropy coding, etc., as described in detail below.

[0043] Similarly, if Figure 2 As shown, in the decoding system 200, the processor 102 may include one or more modules, such as a decoder 201 (also referred to herein as a "post-processing network"). Figure 2The decoder 201 is shown as being located within one processor 102, but it is understood that the decoder 201 may include one or more submodules that may be implemented on different processors that are close to or far away from each other. The decoder 201 (and any corresponding submodules or subunits) may be a hardware unit (e.g., part of an integrated circuit) of the processor 102 designed for use with other components, or a software unit implemented by the processor 102 by executing at least a portion of a program (i.e., instructions). The instructions of the program may be stored on a computer-readable medium such as the memory 104, and when executed by the processor 102 may perform a process having one or more functions related to video decoding, such as entropy decoding, inverse quantization, inverse transformation, inter-frame prediction, intra-frame prediction, filtering, as described in detail below. Figure 3 As shown, the encoder 101 and the decoder 201 can be designed with an asymmetric encoding / decoding framework, wherein the encoder 101 implements a standard CNN and the decoder 201 adopts a DSC network.

[0044] Figure 3 A detailed block diagram of an exemplary video encoding network 300 (hereinafter referred to as “video encoding network 300 ”) according to some embodiments of the present disclosure is shown. Figure 4 The deployment of some embodiments of the present disclosure is shown in Figure 3 Detailed block diagram 400 of an exemplary channel autoregressive entropy model 301 (hereinafter referred to as “entropy model 301”) in an exemplary video coding network of FIG. Figure 3 and Figure 4 Will be described together.

[0045] like Figure 3 As shown, the video coding network 300 may include, for example, an encoder 101, a decoder 201, and a channel autoregressive entropy model 301. The encoder 101 may receive an image from a video. The encoder 101 may include a convolution downsampling component 302, a generalized divisive normalization (GDN) component 304, and a window attention mechanism (WAM) component 306. The convolution downsampling component 302 may be responsible for extracting input image features. The GDN component 304 may be used to normalize intermediate features and improve nonlinearity. The WAM component 306 may focus on high contrast areas and use more bits in these complex areas. In addition, WAM reconstructed images may improve image clarity in terms of texture details. The encoder 101 may output multiple feature maps, which may be input into the entropy model 301. The structure of the entropy model 301 is as shown in FIG. Figure 4 shown.

[0046] refer to Figure 4, by using the residual prediction module of the latent representation and the channel-based conditional context, the network architecture of the entropy model 301 can be optimized. Compared with the existing context entropy model, the RD performance can be improved by optimizing the entropy model 301 while minimizing the serial processing.

[0047] Back to Figure 3 , after entropy modeling, the feature map can be input to the decoder 201. The decoder 201 may include, for example, a WAM component 306, an upsampled residual component 308, a FRCAN component 310, and a RRDB component 312. The residual upsample component 308, the FRCAN component 310, and the RRDB component 312 may be responsible for generating more features to compensate for feature loss during encoding. The WAM component 306 in the encoder 101 and the decoder 201 may have the same range.

[0048] By training the video coding network 300 using the asymmetric encoder 101 and decoder 201, different properties can be balanced to minimize a loss function, which is a weighted sum of terms that measure image reconstruction quality and compression rate. Loss function of the image compression model generated by the video coding network 300 It can be expressed as the following expression (1). where λ controls the trade-off between compression rate and distortion, and R is the latent data and The bit rate, is the original image x and the reconstructed image Distortion between.

[0049] In addition, the video coding network 300 of the present disclosure uses an asymmetric coding (e.g., preprocessing) and decoding (e.g., post-processing) framework to enhance coding efficiency and decoding quality. The asymmetric coding and decoding framework may refer to the type of convolution performed in the encoder and decoder. The framework has two advantages: 1) simplifies coding, thereby increasing coding speed and reducing compressed bitstreams; 2) compensates for information lost during compression and improves the quality of decoded images by using a complex DSC network. The DSC network deployed in the FRCAN component 310 and the RRDB component 312 increases network speed while reducing network parameters. The residual learning performed by the decoder 201 can generate additional features that have been lost, thereby reducing the code stream after image compression. For example, the channel attention block of the RCAB (CA_Block) 608 in the FRCAN component 310 can improve runtime latency and visual quality by weighted capture of information features that may be omitted from the feature map of the preprocessing network.

[0050] Figure 5 It shows that according to some embodiments of the present disclosure, Figure 2Detailed description of the exemplary decoder 201 of FIG. 5 shows a depthwise separable convolution 500 performed by the decoder 201. In order to improve the performance of the FRCAN component 310 and the RRDB component 312 in the decoder 201, a depthwise separable convolution may be performed.

[0051] Still refer to Figure 5 , depthwise separable convolution can be divided into two processes: 1) channel-by-channel convolution and 2) point-wise convolution. One convolution kernel / layer of depthwise separable convolution is responsible for one channel, and one channel is convolved by only one convolution kernel / layer. The number of feature map channels generated in this process can be the same as the number of input channels. After the depthwise convolution is completed, the number of feature maps is the same as the number of channels in the input layer; therefore, the feature map may not be enlarged. In addition, this operation convolves each channel of the input layer independently, and does not effectively utilize the feature information of different channels at the same spatial position.

[0052] Therefore, point convolution can be performed to merge these feature maps to generate a new feature map. Point convolution is similar to the standard convolution operation, and its convolution kernel size is 1×1×M, where M is the number of channels in the previous layer. The point convolution operation merges the maps in the depth direction to generate a new feature map. There are multiple output feature maps with multiple convolution kernels. The shape of the convolution kernel is 1×1×the number of input channels×the number of output channels. With the same input, 4 feature maps can be obtained through point convolution. The number of parameters of depthwise separable convolution is about one-third of the number of parameters of traditional convolution. Therefore, under the premise of the same parameters, the number of layers of the neural network based on depthwise separable convolution can be deeper.

[0053] Figure 6 Showing some embodiments of the present disclosure Figure 3 Block diagram 600 of a residual channel attention block (RCAB) of an exemplary FRCAN component 310 of FIG. Figure 6 , the convolutional layer in front of the CA layer 608 is replaced by a simplified residual-in-residual dense block (RRDB). By using a dense residual structure, the FRCAN component 310 can generate informative image features to compensate for feature loss during compression, thereby improving the quality of the image generated after decompression. The dense residual structure may include, for example, multiple DSC networks 602, Relu layers 604, and multiple LeakyRelU layers 606. Four residual channel attention blocks (RCABs) (compared to 12 RCABs in other systems) can be merged to form one of the FRCAN components 310, thereby achieving runtime reduction and quality enhancement.

[0054] Figure 7 According to some embodiments of the present disclosure Figure 3 Detailed block diagram 700 of the residual upsampling component 308 of FIG. Figure 7 , using the sub-pixel convolution layer 702, the residual upsampling component 308 can perform a mapping from a small rectangle to a large rectangle, thereby improving the resolution. The LeakyRelu layer 704 fixes the neuron death in the Relu. It has a small positive slope in the negative area, so it can be back-propagated even for negative input values. The residual structure restores more features. The convolution layer 706 performs pixel-based convolution. The GDN layer 708 can normalize the intermediate features and improve nonlinearity.

[0055] Figure 8 According to some embodiments of the present disclosure Figure 3 A block diagram 800 of an exemplary RRDB component 312 is shown. Figure 8 As shown, the DSC network 802 replaces the standard convolutional network to reduce processing latency. The RRDB component 312 also includes a LeakyRelu layer 804.

[0056] Fig. 9 According to some aspects of the present disclosure Figure 3 A graphical representation 900 of the PSNRRD performance of video encoding achieved by the video encoding network 300 is shown. Fig.10 According to some aspects of the present disclosure Figure 3 A graphical representation 1000 of the MS-SSIM RD performance of video coding achieved by the video coding network 300 . Fig.11 A graphical representation 1100 showing PSNR versus bit rate for video compression using the FRCAN component 310 and without the FRCAN component 310 in accordance with aspects of the present disclosure. Fig.12 A graphical representation 1200 of MS-SSIM versus bit rate for video compression using the exemplary FRCAN component 310 and without the FRCAN component 310 is shown in accordance with aspects of the present disclosure.

[0057] Fig.13 1300 according to some embodiments of the present disclosure. The method 1300 may be performed by an apparatus (e.g., decoder 201, video encoding network 300, or any other suitable video decoding and / or compression system). The method 1300 may include operations 1302 to 1312 as described below. It is understood that some of the operations may be optional, and some of the operations may be performed simultaneously or in different order. Fig.13 Execute in the order shown.

[0058] refer to Fig.13At 1302, the device may receive a plurality of input feature maps associated with an image from a video preprocessing network. Figure 3 , after entropy modeling, the feature map can be input into the decoder 201 .

[0059] At 1304, the device may apply a WAM network to the plurality of input feature maps. Figure 3 , the WAM component 306 of the decoder 201 can focus on high contrast areas and use more bits in these complex areas. In addition, the WAM reconstructed image can improve the image clarity in terms of texture details.

[0060] At 1306, the device may apply a residual upsampling network to the plurality of input feature maps after the WAM network. Figure 3 , the residual upsampling component 308 may be responsible for generating more features to compensate for the feature loss during encoding. Figure 7 , using the sub-pixel convolution layer 702, the residual upsampling component 308 can perform a mapping from a small rectangle to a large rectangle, thereby improving the resolution. The LeakyRelu layer 704 fixes the neuron death in the Relu. It has a small positive slope in the negative area, so it can be back-propagated even for negative input values. The residual structure restores more features. The convolution layer 706 performs pixel-based convolution. The GDN layer 708 can normalize the intermediate features and improve nonlinearity.

[0061] At 1308, the apparatus may apply the first DSC network of FRCAN after the residual upsampling network. Figure 6 , the convolutional layer before the CA layer 608 is replaced by a simplified RRDB. By using a dense residual structure, the FRCAN component 310 can generate informative image features to compensate for feature loss during compression, thereby improving the quality of the image generated after decompression. The dense residual structure can include, for example, multiple DSC networks 602, Relu layers 604, and multiple LeakyRelU layers 606. Four RCAB components (compared to 12 RCAB components in other systems) can be merged to form one of the FRCAN components 310, thereby achieving runtime reduction and quality enhancement.

[0062] At 1310, the device may apply a second DSC network of RRDB components after FRCAN. Figure 8 , the DSC network 802 replaces the standard convolutional network to reduce processing delay.

[0063] At 1312, the device may generate a compressed image. Figure 3 , the decoder 201 outputs a compressed image after post-processing is completed.

[0064] Fig.14 1400 according to some embodiments of the present disclosure. The method 1400 may be performed by an apparatus (e.g., the video encoding network 300 or any other suitable video decoding and / or compression system). The method 1400 may include operations 1402 to 1408 as described below. It is understood that some of the operations may be optional, and some of the operations may be performed simultaneously or in different modes. Fig.14 Execute in the order shown.

[0065] refer to Fig.14 At 1402, the device may use a preprocessing network to preprocess the input image to generate an encoded image. Figure 3 , encoder 101 (e.g., preprocessing network) can receive images from video. Encoder 101 can include convolution downsampling component 302, GDN component 304 and WAM component 306. Convolution downsampling component 302 can be responsible for extracting input image features. GDN component 304 can be used to normalize intermediate features and improve nonlinearity. WAM component 306 can focus on high contrast areas and use more bits in these complex areas. In addition, WAM reconstructed images can improve image clarity in terms of texture details. Encoder 101 can output multiple feature maps, which can be input to entropy model 301.

[0066] At 1404, the device may post-process the encoded image using a post-processing network to generate a decoded compressed image. Figure 3 , after entropy modeling, the feature map can be input to the decoder 201. The decoder 201 may include, for example, a WAM component 306, an upsampled residual component 308, a FRCAN component 310, and a RRDB component 312. The residual upsample component 308, the FRCAN component 310, and the RRDB component 312 may be responsible for generating more features to compensate for feature loss during encoding. The WAM component 306 in the encoder 101 and the decoder 201 may have the same range.

[0067] At 1406, the apparatus may identify, based on post-processing, a set of features to be omitted from a feature map generated by a pre-processing network during pre-processing of the input image. Figure 3, the video coding network 300 of the present disclosure uses an asymmetric coding (e.g., preprocessing) and decoding (e.g., post-processing) framework to enhance coding efficiency and decoding quality. The asymmetric coding and decoding framework may refer to the type of convolution performed in the encoder and decoder. The framework has two advantages: 1) simplifies coding, thereby increasing coding speed and reducing compressed bitstream; 2) compensates for information lost during compression and improves the quality of decoded images by using a complex DSC network. The DSC network deployed in the FRCAN component 310 and the RRDB component 312 increases network speed while reducing network parameters. The residual learning performed by the decoder 201 can generate additional features that have been lost, thereby reducing the code stream after image compression. For example, the channel attention block (CA_Block) 608 of the RCAB in the FRCAN component 310 can capture information features that may be omitted in the preprocessing network by weighting.

[0068] At 1408, the device may indicate a set of features to be omitted from the feature map generated by the preprocessing network. Figure 3 , the decoder 201 can indicate to the encoder 101 which features to omit from preprocessing. This is because the decoder 201 can identify during training which features the DSC network of the decoder 201 can capture and which would be redundant if the encoder 101 also captured in the feature map of the encoder 101.

[0069] In various aspects of the present disclosure, the functions described herein may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as instructions on a non-transitory computer-readable medium. Computer-readable media include computer storage media. Storage media may be any computer program that can be processed by a processor (e.g., Figure 1 and Figure 2 As an example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, HDD (e.g., disk storage or other magnetic storage device), flash drive, SSD, or any other medium that can be used to carry or store the required program code in the form of instructions or data structures and can be accessed by a processing system (e.g., a mobile device or a computer). Disks and optical disks used in this article include CDs, laser optical disks, optical disks, digital video discs (DVDs), and floppy disks, where disks typically reproduce data magnetically, while optical disks use lasers to reproduce data optically. Combinations of the above should also be included within the scope of computer-readable media.

[0070] According to one aspect of the present disclosure, a video post-processing method is provided. The method may include a processor receiving a plurality of input feature maps associated with an image. The plurality of input feature maps may be generated by a video pre-processing network. The method may include the processor inputting the plurality of input feature maps into a first DSC network of a FRCAN component. The method may include the processor outputting a first set of output feature maps from the first DSC network of the FRCAN component.

[0071] In some embodiments, the method may include a processor applying a depth convolution and a point convolution to a plurality of input feature maps using a first DSC network. In some embodiments, the method may include a processor generating a first set of output feature maps based on the depth convolution and the point convolution after the depth convolution.

[0072] In some embodiments, the method may include a processor inputting the first set of output feature maps into a residual upsampling component. In some embodiments, the method may include a processor upsampling the first set of output feature maps based on a residual upsampling network of the residual upsampling component to generate a set of upsampled feature maps.

[0073] In some embodiments, the method may include the processor inputting a set of upsampled feature maps into a second DSC network of the RRDB component. In some embodiments, the method may include the processor outputting a second set of output feature maps from the second DSC network of the RRDB component.

[0074] In some embodiments, the method may include the processor applying a depth convolution and a point convolution to a set of upsampled feature maps using a second DSC network. In some embodiments, the method may include the processor generating a second set of output feature maps based on the depth convolution and the point convolution after the depth convolution.

[0075] In some embodiments, the method may include the processor inputting the second set of output feature maps to the WAM component. In some embodiments, the method may include the processor outputting a set of enhanced feature maps from the WAM component.

[0076] In some embodiments, the method may include a processor generating a compressed image based on a set of enhanced feature maps.

[0077] According to another aspect of the present disclosure, a video post-processing system is provided. The system may include a memory for storing instructions. The system may include a processor coupled to the memory, which is used to receive multiple input feature maps associated with an image when executing the instructions. The multiple input feature maps can be generated by a video pre-processing network. The system may include a processor coupled to the memory, which is used to input the multiple input feature maps into a first DSC network of a FRCAN component when executing the instructions. The system may include a processor coupled to the memory, which is used to output a first group of output feature maps from the first DSC network of the FRCAN component when executing the instructions.

[0078] In some embodiments, the processor coupled to the memory can also be used to use the first DSC network to apply depth convolution and point convolution to multiple input feature maps in sequence when executing the instructions. In some embodiments, the processor coupled to the memory can also be used to generate a first set of output feature maps based on depth convolution and point convolution after depth convolution when executing the instructions.

[0079] In some embodiments, the processor coupled to the memory may also be used to input the first set of output feature maps into the residual upsampling component when executing the instructions. In some embodiments, the processor coupled to the memory may also be used to upsample the first set of output feature maps based on the residual upsampling network of the residual upsampling component when executing the instructions to generate a set of upsampled feature maps.

[0080] In some embodiments, the processor coupled to the memory can also be used to input a set of upsampled feature maps into the second DSC network of the RRDB component when executing the instructions. In some embodiments, the processor coupled to the memory can also be used to output a second set of output feature maps from the second DSC network of the RRDB component when executing the instructions.

[0081] In some embodiments, the processor coupled to the memory can also be used to apply depth convolution and point convolution to a set of upsampled feature maps using a second DSC network when executing the instructions. In some embodiments, the processor coupled to the memory can also be used to generate a second set of output feature maps based on the depth convolution and the point convolution after the depth convolution when executing the instructions.

[0082] In some embodiments, the processor coupled to the memory may also be used to input the second set of output feature maps to the WAM component when executing the instructions. In some embodiments, the processor coupled to the memory may also be used to output a set of enhanced feature maps from the WAM component when executing the instructions.

[0083] In some embodiments, the processor coupled to the memory, when executing the instructions, may also be used to generate a compressed image based on a set of enhanced feature maps.

[0084] According to another aspect of the present disclosure, a video compression method is provided. The method may include a processor using a preprocessing network to preprocess an input image to generate an encoded image. The method may include a processor using a postprocessing network to postprocess the encoded image to generate a decoded compressed image. The preprocessing network and the postprocessing network may be asymmetric.

[0085] In some embodiments, the method may include a processor identifying a set of features to be omitted from a feature map generated by a preprocessing network during preprocessing of an input image. In some embodiments, a post-processing network may be used to identify a set of features to be omitted from the feature map. In some embodiments, the method may include a processor indicating a set of features to be omitted from a feature map generated by a preprocessing network. In some embodiments, a post-processing network may be used to capture a set of features to be omitted from a feature map generated by a preprocessing network.

[0086] In some embodiments, the processor preprocesses the input image using the preprocessing network to generate the encoded image, which may include applying a standard convolution to the input image using a standard convolution component. In some embodiments, the processor preprocesses the input image using the preprocessing network to generate the encoded image, which may include applying GDN to the input image using a GDN component after applying the standard convolution using the standard convolution component. In some embodiments, the processor preprocesses the input image using the preprocessing network to generate the encoded image, which may include applying a first WAM to the input image using a first WAM component after applying GDN using the GDN component. In some embodiments, the processor postprocesses the encoded image using the postprocessing network to generate the decoded compressed image, which may include applying a second WAM to a set of feature maps generated by the preprocessing network. In some embodiments, the processor postprocesses the encoded image using the postprocessing network to generate the decoded compressed image, which may include applying a first depth separable convolution (DSC) network of the FRCAN component to a set of feature maps after applying the second WAM. In some embodiments, the processor postprocesses the encoded image using the postprocessing network to generate the decoded compressed image, which may include applying a second DSC network of the RRDB to a set of feature maps after applying the first DSC network.

[0087] According to another aspect of the present disclosure, a video compression system is provided. The system may include a memory for storing instructions. The system may include a processor coupled to the memory, which, when executing the instructions, is used to pre-process an input image using a pre-processing network to generate a coded image. The system may include a processor coupled to the memory, which, when executing the instructions, is used to post-process the coded image using a post-processing network to generate a decoded compressed image. The pre-processing network and the post-processing network may be asymmetric.

[0088] In some embodiments, the processor coupled to the memory, when executing the instructions, may also be used to identify a set of features to be omitted from a feature map generated by a preprocessing network during preprocessing of an input image. In some embodiments, a post-processing network may be used to identify a set of features to be omitted from a feature map. In some embodiments, the processor coupled to the memory, when executing the instructions, may also be used to indicate a set of features to be omitted from a feature map generated by a preprocessing network. In some embodiments, a post-processing network may be used to capture a set of features to be omitted from a feature map generated by a preprocessing network.

[0089] In some embodiments, a processor coupled to a memory may be used to preprocess an input image using a preprocessing network to generate a coded image by applying a standard convolution to the input image using a standard convolution component. In some embodiments, a processor coupled to a memory may be used to preprocess an input image using a preprocessing network to generate a coded image by applying a GDN to the input image using a GDN component after applying a standard convolution using a standard convolution component. In some embodiments, a processor coupled to a memory may be used to preprocess an input image using a preprocessing network to generate a coded image by applying a first WAM to the input image using a first WAM component after applying a GDN using a GDN component. In some embodiments, a processor coupled to a memory may be used to postprocess an input image using a postprocessing network to generate a decoded compressed image by applying a second WAM to a set of feature maps generated by the preprocessing network. In some embodiments, a processor coupled to a memory may be used to postprocess an encoded image using a postprocessing network to generate a decoded compressed image by applying a first DSC network of a FRCAN component to a set of feature maps after applying the second WAM. In some embodiments, a processor coupled to a memory may be used to post-process an encoded image using a post-processing network to generate a decoded compressed image by applying a second DSC network of RRDB to a set of feature maps after applying a first DSC network.

[0090] The above description of the embodiments will reveal the general nature of the present disclosure, and others can easily modify and / or adjust these embodiments for various applications without departing from the general concept of the present disclosure by applying knowledge within the technical scope of the art, without excessive experimentation. Therefore, based on the teachings and guidance provided herein, such adjustments and modifications are intended to belong to the meaning and scope of the equivalents of the disclosed embodiments. It should be understood that the wording or terminology herein is for description rather than for limitation, and therefore, those skilled in the art should interpret the terms or wording of this specification in accordance with the teachings and guidance.

[0091] The disclosed embodiments have been described above with the aid of functional building blocks, which illustrate the implementation of their specific functions and their relationships. For ease of description, the boundaries of these functional building blocks are arbitrarily defined herein. As long as their specific functions and their relationships are properly performed, other boundaries may also be defined.

[0092] The Summary and Abstract sections may set forth one or more but not all exemplary embodiments of the present disclosure as contemplated by the inventor(s), and thus, are not intended to limit the present disclosure and the appended claims in any way.

[0093] Various functional blocks, modules, and steps are disclosed above. The arrangements provided are illustrative and not restrictive. Therefore, the functional blocks, modules, and steps may be reordered or combined in a manner different from the examples provided above. Likewise, some embodiments include only a subset of the functional blocks, modules, and steps, and any such subset is permitted.

[0094] The breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents.

Claims

1. A video post-processing method, include: The processor receives a plurality of input feature maps associated with an image, the plurality of input feature maps generated by a video preprocessing network; The processor inputs the plurality of input feature maps into a first depthwise separable convolutional (DSC) network of a fast residual channel attention (FRCAN) component; and The processor outputs a first set of output characteristic maps of the first DSC network from the FRCAN component.

2. The method according to claim 1, further comprising: include: The processor applies depth convolution and point convolution to the plurality of input feature maps in sequence using the first DSC network; as well as The processor generates the first set of output feature maps based on the depthwise convolution and the pointwise convolution following the depthwise convolution.

3. The method according to claim 1, further comprising: include: The processor inputs the first set of output feature maps into a residual upsampling component; as well as The processor upsamples the first set of output feature maps based on a residual upsampling network of the residual upsampling component to generate a set of upsampled feature maps.

4. The method according to claim 3, further comprising: include: The processor inputs the set of upsampled feature maps into a second DSC network of a residual dense block (RRDB) component; as well as The processor outputs a second set of output feature maps of the second DSC network from the RRDB component.

5. The method according to claim 4, further comprising: include: The processor applies a depth convolution and a point convolution to the set of upsampled feature maps using the second DSC network; as well as The processor generates the second set of output feature maps based on the depthwise convolution and the pointwise convolution following the depthwise convolution.

6. The method according to claim 4, further comprising: include: The processor inputs the second set of output feature maps into a window attention mechanism (WAM) component; as well as The processor outputs a set of enhanced feature maps from the WAM component.

7. The method according to claim 6, further comprising: include: The processor generates a compressed image based on the set of enhanced feature maps.

8. A video post-processing system, include: A memory for storing instructions; as well as A processor, coupled to the memory and configured, when executing the instructions, to: Receiving a plurality of input feature maps associated with an image, the plurality of input feature maps generated by a video preprocessing network; Inputting the plurality of input feature maps into a first depthwise separable convolutional (DSC) network of a fast residual channel attention (FRCAN) component; as well as A first set of output feature maps of the first DSC network from the FRCAN component is output.

9. The system according to claim 8, in, The processor coupled to the memory, when executing the instructions, is further configured to: Applying depthwise convolution and pointwise convolution to the plurality of input feature maps using the first DSC network; and The first set of output feature maps is generated based on the depthwise convolution and the pointwise convolution after the depthwise convolution.

10. The system according to claim 8, in, The processor coupled to the memory, when executing the instructions, is further configured to: Inputting the first set of output feature maps into a residual upsampling component; as well as The first set of output feature maps is upsampled based on a residual upsampling network of the residual upsampling component to generate a set of upsampled feature maps.

11. The system according to claim 10, in, The processor coupled to the memory, when executing the instructions, is further configured to: Inputting the set of upsampled feature maps into a second DSC network of a residual dense block (RRDB) component; and A second set of output feature maps of the second DSC network from the RRDB component is output.

12. The system according to claim 11, in, The processor coupled to the memory, when executing the instructions, is further configured to: Applying depthwise convolution and pointwise convolution to the set of upsampled feature maps using the second DSC network; and The second set of output feature maps is generated based on the depthwise convolution and the pointwise convolution after the depthwise convolution.

13. The system according to claim 11, in, The processor coupled to the memory, when executing the instructions, is further configured to: Inputting the second set of output feature maps into a window attention mechanism (WAM) component; and Output is a set of enhanced feature maps from the WAM component.

14. The system according to claim 13, in, The processor coupled to the memory, when executing the instructions, is further configured to: A compressed image is generated based on the set of enhanced feature maps.

15. A video compression method, include: The processor preprocesses the input image using a preprocessing network to generate an encoded image; as well as The processor post-processes the encoded image using a post-processing network to generate a decoded compressed image, Wherein, the pre-processing network and the post-processing network are asymmetric.

16. The method according to claim 15, further comprising: include: the processor identifying a set of features to be omitted in a feature map generated by the pre-processing network during the pre-processing of the input image, the set of features to be omitted in the feature map being identified using the post-processing network; as well as The processor indicates the set of features to be omitted from the feature map generated by the preprocessing network, The post-processing network is used to capture the set of features to be omitted in the feature map generated by the pre-processing network.

17. The method according to claim 15, in: The processor preprocesses the input image using the preprocessing network to generate an encoded image, comprising: Applying a standard convolution to the input image using a standard convolution component; applying a generalized divisive normalization (GDN) to the input image using a GDN component after applying the standard convolution using the standard convolution component; and After applying the GDN using the GDN component, applying a first window attention module (WAM) to the input image using a first WAM component, and The processor uses the post-processing network to post-process the encoded image to generate the decoded compressed image, comprising: Applying a second WAM to a set of feature maps generated by the preprocessing network; applying a first depthwise separable convolutional (DSC) network of a fast residual channel attention (FRCAN) component to the set of feature maps after applying the second WAM; and After applying the first DSC network, a second DSC network of residual dense blocks (RRDB) is applied to the set of feature maps.

18. A video compression system, include: A memory for storing instructions; as well as A processor, coupled to the memory and configured, when executing the instructions, to: Preprocessing the input image using a preprocessing network to generate an encoded image; as well as post-processing the encoded image using a post-processing network to generate a decoded compressed image, Wherein, the pre-processing network and the post-processing network are asymmetric.

19. The system according to claim 18, in, The processor coupled to the memory, when executing the instructions, is further configured to: identifying a set of features to be omitted from a feature map generated by the pre-processing network during the pre-processing of the input image, the set of features to be omitted from the feature map being identified using the post-processing network; as well as indicating the set of features to be omitted in the feature map generated by the preprocessing network, The post-processing network is used to capture the set of features to be omitted in the feature map generated by the pre-processing network.

20. The system according to claim 18, in: The processor coupled to the memory is configured to preprocess the input image using the preprocessing network to generate an encoded image as follows: Applying a standard convolution to the input image using a standard convolution component; After applying the standard convolution using the standard convolution component, applying a generalized divisive normalization (GDN) to the input image using a GDN component; as well as After applying the GDN using the GDN component, applying a first window attention module (WAM) to the input image using a first WAM component, and The processor coupled to the memory is configured to post-process the encoded image using the post-processing network to generate the decoded compressed image as follows: Applying a second WAM to a set of feature maps generated by the preprocessing network; applying a first depthwise separable convolutional (DSC) network of a fast residual channel attention (FRCAN) component to the set of feature maps after applying the second WAM; and After applying the first DSC network, a second DSC network of residual dense blocks (RRDB) is applied to the set of feature maps.