Coding method and device
By introducing prediction refinement networks, two-stage residual processing and loop filtering networks into video encoding technology, the problem of low inter prediction and residual coding efficiency in the prior art is solved, and more efficient video compression and better image quality are achieved.
Patent Information
- Application Number
- CN202110169583.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-07
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-02-07
AI Technical Summary
Existing video encoding technologies have challenges in improving encoding efficiency, especially in inter-frame prediction and residual coding.
The prediction refinement network (PR-Net) is used to refine the inter-predicted images to improve prediction performance, and further improve the quality of the reconstructed images through two-stage residual processing and loop filtering network.
Improves encoding efficiency, improves inter-frame prediction performance and reconstructs images, and achieves more efficient video compression.
Smart Images

Figure CN114915783B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of video or image compression based on artificial intelligence (AI), and in particular to a coding method and device. Background Art
[0002] Video coding (video encoding and decoding) is widely used in digital video applications such as broadcast digital television, video transmission over the Internet and mobile networks, real-time conversation applications such as video chat and video conferencing, DVD and Blu-ray Discs, video content acquisition and editing systems, and security applications in camcorders.
[0003] In recent years, the application of deep learning in the field of image or video coding has gradually become a trend. Common end-to-end video coding technology designs an overall trainable network for modules such as intra-frame prediction, inter-frame prediction, residual coding, and bit rate control. Among them, how to improve coding efficiency is the key technology. Summary of the invention
[0004] The present application provides a coding method and device, which can improve coding efficiency.
[0005] In a first aspect, the present application provides an encoding method, comprising: obtaining an inter-frame prediction image of a current image, wherein the current image is not the first frame image of a video sequence; inputting the inter-frame prediction image into a prediction refinement network to obtain an enhanced inter-frame prediction image; and obtaining a reconstructed image of the current image based on the enhanced inter-frame prediction image.
[0006] The input of the prediction refinement network (PR-Net) is the inter-frame prediction image of the current image, and the output of PR-Net is the enhanced inter-frame prediction image of the current image. It is precisely because PR-Net can achieve better prediction performance that the enhanced inter-frame prediction image obtained after the inter-frame prediction image of the current image passes through PR-Net has better performance than the original inter-frame prediction image, such as the edge part of the image and the part with more texture details.
[0007] In a possible implementation, obtaining the inter-frame predicted image of the current image includes: obtaining a reconstructed image of a previous frame image, the encoding order of the previous frame image is just earlier than the current image; obtaining a quantized compressed motion vector MV of the current image; and inputting the reconstructed image of the previous frame image and the quantized compressed MV of the current image into a motion compensation network to obtain the inter-frame predicted image.
[0008] In a possible implementation, the obtaining of the reconstructed image of the current image based on the enhanced inter-frame prediction image includes: obtaining an intermediate reconstructed image based on the current image and the enhanced inter-frame prediction image; obtaining a reconstructed residual image based on the current image and the intermediate reconstructed image; inputting the reconstructed residual image into a refined residual encoding network and a refined residual decoding network to obtain a quantized compressed reconstructed residual image; and obtaining the reconstructed image based on the quantized compressed reconstructed residual image and the intermediate reconstructed image.
[0009] The present application forms a two-level residual processing. On the one hand, parameter sharing is not implemented for the two-level residual network during training to ensure that the two learned residual networks are different. On the other hand, the two-level residual processing constitutes a scalable residual, which can be used to fine-tune modeling of the reconstructed image, and use a new residual coding module to encode the residual of the reconstructed image to further improve performance.
[0010] In a possible implementation, the obtaining of the reconstructed image based on the quantized compressed reconstructed residual image and the intermediate reconstructed image includes: obtaining a pre-filtered reconstructed image based on the quantized compressed reconstructed residual image and the intermediate reconstructed image; and inputting the pre-filtered reconstructed image into a loop filtering network to obtain the reconstructed image.
[0011] The loop filter network can further improve the quality of the reconstructed image.
[0012] In a possible implementation manner, acquiring the reconstructed image according to the quantized compressed reconstructed residual image and the intermediate reconstructed image includes: acquiring the reconstructed image according to the quantized compressed reconstructed residual image and the intermediate reconstructed image.
[0013] In a possible implementation, the obtaining of the intermediate reconstructed image based on the current image and the enhanced inter-frame prediction image includes: taking the difference between the pixel values at corresponding positions in the current image and the enhanced inter-frame prediction image to obtain a residual image; inputting the residual image into a residual encoding network and a residual decoding network to obtain a quantized compressed residual image; and summing the pixel values at corresponding positions in the quantized compressed residual image and the enhanced inter-frame prediction image to obtain the intermediate reconstructed image.
[0014] In a possible implementation manner, acquiring the reconstructed residual image according to the current image and the intermediate reconstructed image includes: obtaining the reconstructed residual image by subtracting pixel values at corresponding positions in the current image and the intermediate reconstructed image.
[0015] In a possible implementation, obtaining the pre-filtering reconstructed image based on the quantized compressed reconstructed residual image and the intermediate reconstructed image includes: summing the pixel values at corresponding positions in the quantized compressed reconstructed residual image and the intermediate reconstructed image to obtain the pre-filtering reconstructed image.
[0016] In a possible implementation, acquiring the reconstructed image based on the quantized compressed reconstructed residual image and the intermediate reconstructed image includes: summing pixel values at corresponding positions in the quantized compressed reconstructed residual image and the intermediate reconstructed image to obtain the reconstructed image.
[0017] In a possible implementation, obtaining the quantized compressed motion vector MV of the current image includes: inputting the reconstructed image of the previous frame image and the current image into a motion estimation network to obtain the MV of the current image; and inputting the MV of the current image into an MV encoding network and an MV decoding network to obtain the quantized compressed MV of the current image.
[0018] This application inputs the predicted image of the current image into the prediction refinement network to obtain an enhanced inter-frame prediction image, which can have better prediction performance than the original inter-frame prediction image. In addition, through two-level residual processing, on the one hand, parameter sharing is not implemented for the two-level residual network during training to ensure that the two residual networks learned are different. On the other hand, the two-level residual processing constitutes a scalable residual, which can be used to finely model the reconstructed image, and use a new residual coding module to encode the residual of the reconstructed image to further improve performance. Coupled with the loop filtering network, the quality of the reconstructed image can be further improved.
[0019] In a second aspect, the present application provides an encoding device, comprising: an acquisition module, used to acquire an inter-frame prediction image of a current image, wherein the current image is not the first frame image of a video sequence; a prediction module, used to input the inter-frame prediction image into a prediction refinement network to obtain an enhanced inter-frame prediction image; and a reconstruction module, used to acquire a reconstructed image of the current image based on the enhanced inter-frame prediction image.
[0020] In a possible implementation, the acquisition module is specifically used to obtain a reconstructed image of a previous frame image, where the encoding order of the previous frame image is just earlier than the current image; obtain a quantized compressed motion vector MV of the current image; and input the reconstructed image of the previous frame image and the quantized compressed MV of the current image into a motion compensation network to obtain the inter-frame predicted image.
[0021] In a possible implementation, the reconstruction module is specifically used to obtain an intermediate reconstructed image based on the current image and the enhanced inter-frame prediction image; obtain a reconstructed residual image based on the current image and the intermediate reconstructed image; input the reconstructed residual image into a refined residual encoding network and a refined residual decoding network to obtain a quantized compressed reconstructed residual image; and obtain the reconstructed image based on the quantized compressed reconstructed residual image and the intermediate reconstructed image.
[0022] In a possible implementation, the reconstruction module is specifically configured to obtain a pre-filtering reconstructed image based on the quantized compressed reconstructed residual image and the intermediate reconstructed image; and input the pre-filtering reconstructed image into a loop filtering network to obtain the reconstructed image.
[0023] In a possible implementation manner, the reconstruction module is specifically configured to obtain the reconstructed image according to the quantized compressed reconstructed residual image and the intermediate reconstructed image.
[0024] In a possible implementation, the reconstruction module is specifically used to obtain a residual image by taking the difference between the pixel values at corresponding positions in the current image and the enhanced inter-frame prediction image; inputting the residual image into a residual encoding network and a residual decoding network to obtain a quantized compressed residual image; and summing the pixel values at corresponding positions in the quantized compressed residual image and the enhanced inter-frame prediction image to obtain the intermediate reconstructed image.
[0025] In a possible implementation manner, the reconstruction module is specifically configured to obtain the reconstructed residual image by subtracting pixel values at corresponding positions in the current image and the intermediate reconstructed image.
[0026] In a possible implementation, the reconstruction module is specifically configured to sum pixel values at corresponding positions in the quantized compressed reconstructed residual image and the intermediate reconstructed image to obtain the pre-filtering reconstructed image.
[0027] In a possible implementation manner, the reconstruction module is specifically configured to sum pixel values at corresponding positions in the quantized compressed reconstructed residual image and the intermediate reconstructed image to obtain the reconstructed image.
[0028] In a possible implementation, the acquisition module is specifically used to input the reconstructed image of the previous frame image and the current image into a motion estimation network to obtain the MV of the current image; and input the MV of the current image into an MV encoding network and an MV decoding network to obtain a quantized compressed MV of the current image.
[0029] In a third aspect, the present application provides an encoder, comprising: one or more processors; a non-transitory computer-readable storage medium, coupled to the processor and storing a program executed by the processor, wherein when the program is executed by the processor, the encoder performs a method according to any one of the above-mentioned first aspects.
[0030] In a fourth aspect, the present application provides a computer program product, characterized in that it includes program code, which, when executed on a computer or a processor, is used to execute the method according to any one of the above-mentioned first aspects.
[0031] In a fifth aspect, the present application provides a non-transitory computer-readable storage medium, characterized in that it includes program code, which, when executed by a computer device, is used to execute the method according to any one of the above-mentioned first aspects.
[0032] In a sixth aspect, the present application provides a non-transitory storage medium, characterized in that it includes a bit stream encoded according to any method described in the first aspect above. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is an exemplary block diagram of a decoding system 10 according to an embodiment of the present application;
[0034] Figure 2 is an exemplary block diagram of a video encoder 20 according to an embodiment of the present application;
[0035] Figure 3 is an exemplary block diagram of a device 300 according to an embodiment of the present application;
[0036] Figure 4a-4e An exemplary architecture of a neural network according to an embodiment of the present application;
[0037] Figure 5 is an exemplary structural diagram of an end-to-end video coding architecture according to an embodiment of the present application;
[0038] Figure 6 A flowchart of a process 600 of an encoding method according to an embodiment of the present application;
[0039] Figure 7 This is an exemplary architecture of PR-Net according to an embodiment of the present application;
[0040] Figure 8 An exemplary architecture of a loop filter network according to an embodiment of the present application;
[0041] Fig. 9 This is an exemplary structural diagram of the encoding device 900 according to an embodiment of the present application. DETAILED DESCRIPTION
[0042] The embodiments of the present application provide an AI-based video image compression technology, and in particular, provide an end-to-end video encoding technology based on a neural network to improve a traditional hybrid video encoding and decoding system.
[0043] Video coding generally refers to processing a sequence of images that form a video or video sequence. In the field of video coding, the terms "picture", "frame" or "image" can be used as synonyms. Video coding (or commonly referred to as coding) includes two parts: video encoding and video decoding. Video coding is performed on the source side and generally includes processing (e.g., compressing) the original video image to reduce the amount of data required to represent the video image (thereby more efficiently storing and / or transmitting). Video decoding is performed on the destination side and generally includes inverse processing relative to the encoder to reconstruct the video image. The "encoding" of the video image (or commonly referred to as the image frame) involved in the embodiment should be understood as the "encoding" or "decoding" of the video image or video sequence. The encoding part and the decoding part are also collectively referred to as codec (encoding and decoding, CODEC).
[0044] In the case of lossy video coding, further compression is performed through quantization, etc. to reduce the amount of data required to represent the video image, but the encoder side and the decoder side cannot fully reconstruct the video image, that is, the quality of the reconstructed image is worse than that of the original video image.
[0045] End-to-end video coding techniques usually encode at the frame level. In other words, the encoder usually encodes the video at the frame level, for example, generating a predicted image through spatial (intra-frame) prediction and temporal (inter-frame) prediction; subtracting the predicted image from the current image to obtain a residual image; transforming the residual image in the transform domain and quantizing the residual image to reduce the amount of data to be transmitted (compressed), while the decoder side applies the inverse processing relative to the encoder to the encoded or compressed image frame to reconstruct the current image frame for representation. In addition, the encoder needs to repeat the processing steps of the decoder so that the encoder and the decoder generate the same prediction (e.g., intra-frame prediction and inter-frame prediction) and / or reconstructed image for encoding subsequent image frames.
[0046] In the following embodiment of the decoding system 10, the encoder 20 and the decoder 30 are based on Figures 1 to 3 Give a description.
[0047] Figure 1 1 is an exemplary block diagram of a decoding system 10 of an embodiment of the present application, for example, a video decoding system 10 (or simply a decoding system 10) that can utilize the techniques of the present application. The video encoder 20 (or simply an encoder 20) and the video decoder 30 (or simply a decoder 30) in the video decoding system 10 represent devices that can be used to perform various techniques according to various examples described in the present application.
[0048] like Figure 1 As shown, the decoding system 10 includes a source device 12, which is used to provide encoded image data 21 such as encoded images to a destination device 14 for decoding the encoded image data 21.
[0049] The source device 12 includes an encoder 20 , and may additionally or optionally include an image source 16 , a preprocessor (or a preprocessing unit) 18 such as an image preprocessor, and a communication interface (or a communication unit) 22 .
[0050] The image source 16 may include or may be any type of image capture device for capturing real-world images, etc., and / or any type of image generation device, such as a computer graphics processor for generating computer-animated images or any type of device for acquiring and / or providing real-world images, computer-generated images (e.g., screen content, virtual reality (VR) images, and / or any combination thereof (e.g., augmented reality (AR) images). The image source may be any type of memory or storage that stores any of the above images.
[0051] In order to distinguish the processing performed by the pre-processor (or pre-processing unit) 18 , the image (or image data) 17 may also be referred to as a raw image (or raw image data) 17 .
[0052] The preprocessor 18 is used to receive (raw) image data 17 and preprocess the image data 17 to obtain a preprocessed image (or preprocessed image data) 19. For example, the preprocessing performed by the preprocessor 18 may include trimming, color format conversion (e.g., from RGB to YCbCr), color adjustment, or denoising. It is understood that the preprocessing unit 18 may be an optional component.
[0053] The video encoder (or encoder) 20 is used to receive the pre-processed image data 19 and provide the encoded image data 21 (hereinafter referred to as Figure 2 etc. for further description).
[0054] The communication interface 22 in the source device 12 may be used to receive the encoded image data 21 and send the encoded image data 21 (or any other processed version) to the destination device 14 via the communication channel 13 for storage or direct reconstruction.
[0055] The destination device 14 includes a decoder 30 and, in addition or alternatively, may include a communication interface (or communication unit) 28 , a post-processor (or post-processing unit) 32 , and a display device 34 .
[0056] The communication interface 28 in the destination device 14 is used to receive the encoded image data 21 (or any other processed version) directly from the source device 12 or from any other source device such as a storage device, for example, the storage device is a encoded image data storage device, and provide the encoded image data 21 to the decoder 30.
[0057] The communication interface 22 and the communication interface 28 can be used to send or receive encoded image data (or encoded data) 21 through a direct communication link between the source device 12 and the destination device 14, such as a direct wired or wireless connection, or through any type of network, such as a wired network, a wireless network or any combination thereof, any type of private network and public network or any type of combination thereof.
[0058] For example, the communication interface 22 may be used to encapsulate the encoded image data 21 into a suitable format such as a message, and / or process the encoded image data using any type of transmission coding or processing for transmission over a communication link or network.
[0059] The communication interface 28 corresponds to the communication interface 22 , for example, and can be used to receive transmission data and process the transmission data using any type of corresponding transmission decoding or processing and / or decapsulation to obtain the encoded image data 21 .
[0060] Both the communication interface 22 and the communication interface 28 can be configured as follows Figure 1 A unidirectional communication interface, or a bidirectional communication interface, indicated by an arrow pointing from the source device 12 to the corresponding communication channel 13 of the destination device 14, and can be used to send and receive messages, etc. to establish a connection, confirm and exchange any other information related to the communication link and / or data transmission such as encoded image data transmission, etc.
[0061] The video decoder (or decoder) 30 is used to receive the encoded image data 21 and provide decoded image data (or decoded image data) 31 .
[0062] The post-processor 32 is used to post-process the decoded image data 31 (also called reconstructed image data) such as the decoded image to obtain the post-processed image data 33 such as the post-processed image. The post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color adjustment, cropping or resampling, or any other processing for generating the decoded image data 31 for display by the display device 34 or the like.
[0063] The display device 34 is used to receive the post-processed image data 33 to display the image to a user or viewer, etc. The display device 34 can be or include any type of display for displaying the reconstructed image, such as an integrated or external display screen or display. For example, the display screen can include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display screen.
[0064] The decoding system 10 also includes a training engine 25, which is used to train one or more neural networks, which can be used to implement the functions of one or more functional modules in the encoder 20 or the decoder 30, for example, a neural network for implementing intra-frame prediction, a neural network for implementing inter-frame prediction, a neural network for implementing filtering processing, and the like.
[0065] The training engine 25 can be trained in the cloud to obtain a target model, and then the decoding system 10 downloads and uses the target model from the cloud; or, the training engine 25 can be trained in the cloud to obtain a target model and use the target model, and the decoding system 10 directly obtains the processing result from the cloud. For example, the training engine 25 is trained to obtain a target model with a filtering function, and the decoding system 10 downloads the target model from the cloud, and then the loop filter 220 in the encoder 20 or the loop filter 320 in the decoder 30 can filter the input reconstructed image according to the target model to obtain a filtered image. For another example, the training engine 25 is trained to obtain a target model with a filtering function, and the decoding system 10 does not need to download the target model from the cloud. The encoder 20 or the decoder 30 transmits the reconstructed image to the cloud, and the cloud performs filtering on the reconstructed image through the target model to obtain a filtered image and transmit it to the encoder 20 or the decoder 30.
[0066] although Figure 1 The source device 12 and the destination device 14 are shown as independent devices, but the device embodiment may also include the source device 12 and the destination device 14 or the functions of the source device 12 and the destination device 14 at the same time, that is, the source device 12 or the corresponding function and the destination device 14 or the corresponding function at the same time. In these embodiments, the source device 12 or the corresponding function and the destination device 14 or the corresponding function can be implemented using the same hardware and / or software or by separate hardware and / or software or any combination thereof.
[0067] According to the description, Figure 1 The presence and (exact) division of different units or functions in the source device 12 and / or the destination device 14 shown may vary depending on the actual device and application, which will be obvious to the skilled person.
[0068] Source device 12 and destination device 14 may include any of a variety of devices, including any type of handheld or fixed device, such as a notebook or laptop computer, a mobile phone, a smart phone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (e.g., a content service server or a content distribution server), a broadcast receiving device, a broadcast transmitting device, etc., and may not use or use any type of operating system. In some cases, source device 12 and destination device 14 may be equipped with components for wireless communication. Therefore, source device 12 and destination device 14 may be wireless communication devices.
[0069] In some cases, Figure 1 The video decoding system 10 shown is merely exemplary, and the techniques provided herein may be applicable to video encoding settings (e.g., video encoding or video decoding) that do not necessarily include any data communication between an encoding device and a decoding device. In other examples, data is retrieved from a local memory, sent over a network, and so on. A video encoding device may encode data and store the data in a memory, and / or a video decoding device may retrieve data from a memory and decode the data. In some examples, encoding and decoding are performed by devices that do not communicate with each other but only encode data to a memory and / or retrieve and decode data from a memory.
[0070] It should be understood that for the examples described with reference to the video encoder 20 in the embodiments of the present application, the video decoder 30 can be used to perform the reverse process. With respect to the signaling syntax elements, the video decoder 30 can be used to receive and parse such syntax elements and decode the related video data accordingly. In some examples, the video encoder 20 can entropy encode the syntax elements into an encoded video bitstream. In such examples, the video decoder 30 can parse such syntax elements and decode the related video data accordingly.
[0071] Encoders and encoding methods
[0072] Figure 2 FIG. 2 is an exemplary block diagram of a video encoder 20 according to an embodiment of the present application. Figure 2In the example of , the video encoder 20 includes an input terminal (or input interface) 201, a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a loop filter 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy coding unit 270, and an output terminal (or output interface) 272. The mode selection unit 260 may include an inter-frame prediction unit 244, an intra-frame prediction unit 254, and a segmentation unit 262. The inter-frame prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). Figure 2 The illustrated video encoder 20 may also be referred to as a hybrid video encoder or a hybrid video codec-based video encoder.
[0073] The residual calculation unit 204, the transform processing unit 206, the quantization unit 208, and the mode selection unit 260 constitute the forward signal path of the encoder 20, while the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, the inter-frame prediction unit 244, and the intra-frame prediction unit 254 constitute the backward signal path of the encoder, wherein the backward signal path of the encoder 20 corresponds to the signal path of the decoder. The inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer 230, the inter-frame prediction unit 244, and the intra-frame prediction unit 254 also constitute the "built-in decoder" of the video encoder 20.
[0074] image
[0075] The encoder 20 may be configured to receive an image (or image data) 17, such as an image in a sequence of images forming a video or a video sequence, via an input 201 or the like. The received image or image data may also be a pre-processed image (or pre-processed image data) 19. For simplicity, the following description uses image 17. Image 17 may also be referred to as a current image or an image to be encoded.
[0076] A (digital) image can be considered as a two-dimensional array or matrix of pixels with intensity values. The pixels in the array can also be called pixels (pixels or pels) (short for picture elements). The number of pixels in the array or image in the horizontal and vertical directions (or axes) determines the size and / or resolution of the image. In order to represent color, three color components are usually used, that is, the image can be represented as or include three pixel arrays. In the RBG format or color space, the image includes corresponding red, green and blue pixel arrays. However, in video coding, each pixel is usually represented in a luminance / chrominance format or color space, such as YCbCr, including a luminance component indicated by Y (sometimes also indicated by L) and two chrominance components indicated by Cb and Cr. The luminance (luma) component Y represents the brightness or grayscale level intensity (for example, the two are the same in grayscale images), while the two chrominance (abbreviated as chroma) components Cb and Cr represent the chrominance or color information components. Accordingly, an image in YCbCr format includes a luminance pixel array of luminance pixel values (Y) and two chrominance pixel arrays of chrominance values (Cb and Cr). An image in RGB format can be converted or transformed into YCbCr format and vice versa, a process also referred to as color conversion or transformation. If the image is black and white, the image may include only a luminance pixel array. Accordingly, an image may be, for example, a luminance pixel array in a monochrome format or a luminance pixel array and two corresponding chrominance pixel arrays in 4:2:0, 4:2:2 and 4:4:4 color formats.
[0077] In one embodiment, Figure 2 The video encoder 20 shown is used to encode and predict the image 17.
[0078] Residual calculation
[0079] The residual calculation unit 204 is used to calculate the residual image 205 (the predicted image 265 is introduced in detail later) based on the image 17 and the predicted image 265 in the following manner: for example, the pixel value of the predicted image 265 is subtracted from the pixel value of the image 17 pixel by pixel (pixel by pixel) to obtain the residual image 205 in the pixel domain.
[0080] Transform
[0081] The transform processing unit 206 is used to perform discrete cosine transform (DCT) or discrete sine transform (DST) on the pixel values of the residual image 205 to obtain transform coefficients 207 in the transform domain. The transform coefficients 207 may also be called transform residual coefficients, representing the residual image 205 in the transform domain.
[0082] The transform processing unit 206 may be used to apply an integerized approximation of the DCT / DST. This integerized approximation is typically scaled by a factor compared to an orthogonal DCT transform. In order to maintain the norm of the residual block processed by the forward transform and the inverse transform, other scaling factors are used as part of the transform process. The scaling factor is typically selected based on certain constraints, such as whether the scaling factor is a power of 2 for the shift operation, the bit depth of the transform coefficients, a trade-off between accuracy and implementation cost, etc. For example, a specific scaling factor is specified for the inverse transform on the encoder 20 side by the inverse transform processing unit 212 (and for the corresponding inverse transform on the decoder 30 side by, for example, the inverse transform processing unit 312), and correspondingly, a corresponding scaling factor may be specified for the forward transform on the encoder 20 side by the transform processing unit 206.
[0083] In one embodiment, the video encoder 20 (correspondingly, the transform processing unit 206) may be used to output transform parameters such as one or more transform types, for example, directly output or output after encoding or compression by the entropy coding unit 270, for example, so that the video decoder 30 may receive and use the transform parameters for decoding.
[0084] Quantification
[0085] The quantization unit 208 is used to quantize the transform coefficient 207 by, for example, scalar quantization or vector quantization to obtain a quantized transform coefficient 209 . The quantized transform coefficient 209 may also be referred to as a quantized residual coefficient 209 .
[0086] The quantization process may reduce the bit depth associated with some or all of the transform coefficients 207. For example, an n-bit transform coefficient may be rounded down to an m-bit transform coefficient during quantization, where n is greater than m. The degree of quantization may be modified by adjusting a quantization parameter (QP). For example, for scalar quantization, different degrees of scaling may be applied to achieve finer or coarser quantization. A smaller quantization step size corresponds to finer quantization, while a larger quantization step size corresponds to coarser quantization. A suitable quantization step size may be indicated by a quantization parameter (QP). For example, a quantization parameter may be an index into a predefined set of suitable quantization step sizes. For example, a smaller quantization parameter may correspond to fine quantization (a smaller quantization step size), a larger quantization parameter may correspond to coarse quantization (a larger quantization step size), or vice versa. Quantization may include dividing by a quantization step size, while a corresponding or inverse dequantization performed by the inverse quantization unit 210 or the like may include multiplying by a quantization step size. In general, the quantization step size may be calculated based on the quantization parameter using a fixed-point approximation of an equation involving division. Other scaling factors may be introduced for quantization and dequantization to recover the norm of the residual block that may be modified due to the scaling used in the fixed-point approximation of the equations for the quantization step size and the quantization parameter. In one exemplary implementation, the scaling of the inverse transform and dequantization may be combined. Alternatively, a custom quantization table may be used and indicated from the encoder to the decoder in the bitstream, etc. Quantization is a lossy operation, where the larger the quantization step size, the greater the loss.
[0087] In one embodiment, the video encoder 20 (correspondingly, the quantization unit 208) may be configured to output a quantization parameter (QP), for example, directly output or output after being encoded or compressed by the entropy coding unit 270, for example, so that the video decoder 30 may receive and use the quantization parameter for decoding.
[0088] Dequantization
[0089] The inverse quantization unit 210 is used to perform inverse quantization of the quantization unit 208 on the quantized coefficients to obtain dequantized coefficients 211, for example, performing an inverse quantization scheme of the quantization scheme performed by the quantization unit 208 according to or using the same quantization step size as the quantization unit 208. The dequantized coefficients 211 may also be referred to as dequantized residual coefficients 211, corresponding to the transform coefficients 207, but due to the loss caused by quantization, the dequantized coefficients 211 are usually not completely the same as the transform coefficients.
[0090] Inverse Transform
[0091] The inverse transform processing unit 212 is used to perform an inverse transform of the transform performed by the transform processing unit 206, such as an inverse discrete cosine transform (DCT) or an inverse discrete sine transform (DST), to obtain a reconstructed residual image 213 (or a corresponding dequantized coefficient 213) in the pixel domain. The reconstructed residual image 213 may also be referred to as a transformed image 213.
[0092] reconstruction
[0093] The reconstruction unit 214 (e.g., the summer 214) is used to add the transformed image 213 (i.e., the reconstructed residual image 213) to the predicted image 265 to obtain the reconstructed image 215 in the pixel domain, for example, by adding the pixel point values of the reconstructed residual image 213 and the pixel point values of the predicted image 265.
[0094] Filtering
[0095] The loop filter unit 220 (or simply "loop filter" 220) is used to filter the reconstructed image 215 to obtain a filtered image 221, or is generally used to filter the reconstructed pixel points to obtain filtered pixel point values. For example, the loop filter unit is used to smoothly perform pixel conversion or improve video quality. The loop filter unit 220 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination. For example, the loop filter unit 220 may include a deblocking filter, an SAO filter, and an ALF filter. The order of the filtering process may be a deblocking filter, an SAO filter, and an ALF filter. For another example, a process called luma mapping with chromascaling (LMCS) (i.e., an adaptive in-loop shaper) is added. This process is performed before deblocking. For another example, the deblocking filtering process may also be applied to internal sub-block edges, such as affine sub-block edges, ATMVP sub-block edges, sub-block transform (SBT) edges, and intra sub-partition (ISP) edges. Figure 2 2. Although shown as a loop filter in FIG. 2, in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filtered image 221 may also be referred to as a filtered reconstructed image 221.
[0096] In one embodiment, the video encoder 20 (correspondingly, the loop filter unit 220) may be used to output loop filter parameters (e.g., SAO filter parameters, ALF filter parameters, or LMCS parameters), for example, directly or after being entropy encoded by the entropy encoding unit 270, such that the decoder 30 may receive and use the same or different loop filter parameters for decoding.
[0097] Decoded Image Buffer
[0098] The decoded picture buffer (DPB) 230 may be a reference picture memory that stores reference picture data for use by the video encoder 20 when encoding video data. The DPB 230 may be formed by any of a variety of memory devices, such as a dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of storage devices. The decoded picture buffer 230 may be used to store one or more filtered images 221. The decoded picture buffer 230 may also be used to store the same current image or different images such as previously reconstructed images, such as previously reconstructed and filtered images 221, and may provide a complete previously reconstructed image, such as for inter-frame prediction. The decoded picture buffer 230 may also be used to store one or more unfiltered reconstructed images 215, or generally store unfiltered reconstructed pixels, such as reconstructed images 215 that have not been filtered by the loop filter unit 220, or reconstructed images that have not undergone any other processing.
[0099] Mode selection (segmentation and prediction)
[0100] The mode selection unit 260 includes an inter-frame prediction unit 244 and an intra-frame prediction unit 254, which are used to receive or obtain image 17 and reconstructed image data from the decoded image buffer 230 or other buffers (e.g., column buffers, not shown in the figure), such as filtered and / or unfiltered reconstructed images of the current image and / or one or more previously decoded images. The reconstructed image data is used as reference image data required for prediction such as inter-frame prediction or intra-frame prediction to obtain a predicted image 265 or a predicted value 265.
[0101] The mode selection unit 260 may be used to determine a current image prediction mode (including non-segmentation) and a prediction mode (eg, intra-frame or inter-frame prediction mode) and generate a corresponding prediction image 265 to calculate the residual image 205 and reconstruct the reconstructed image 215 .
[0102] In one embodiment, the mode selection unit 260 may be used to select a prediction mode that provides the best match or minimum residual (minimum residual means better compression in transmission or storage), or provides minimum signaling overhead (minimum signaling overhead means better compression in transmission or storage), or considers or balances both of the above. The mode selection unit 260 may be used to determine the prediction mode based on rate distortion optimization (RDO), that is, to select the prediction mode that provides minimum rate distortion optimization. The terms "best", "lowest", "optimal", etc. herein do not necessarily mean "best", "lowest", "optimal" in general, but may also refer to situations where termination or selection criteria are met, for example, values exceeding or below a threshold or other restrictions may result in a "suboptimal selection" but reduce complexity and processing time.
[0103] The prediction process performed by video encoder 20 (eg, performed by inter-prediction unit 244 and intra-prediction unit 254) will be described in detail below.
[0104] As described above, the video encoder 20 is operable to determine or select the best or optimal prediction mode from a (predetermined) set of prediction modes. The set of prediction modes may include, for example, an intra prediction mode and / or an inter prediction mode.
[0105] Intra-frame prediction
[0106] The intra prediction mode set may include 35 different intra prediction modes, for example, non-directional modes like DC (or mean) mode and plane mode, or directional modes as defined in HEVC, or may include 67 different intra prediction modes, for example, non-directional modes like DC (or mean) mode and plane mode, or directional modes as defined in VVC. For example, several traditional angle intra prediction modes are adaptively replaced with wide-angle intra prediction modes for non-square blocks defined in VVC. For another example, in order to avoid the division operation of DC prediction, only the longer side is used to calculate the average value of non-square blocks. In addition, the intra prediction result of the plane mode can also be modified using the position-dependent intra prediction combination (PDPC) method.
[0107] The intra prediction unit 254 is configured to generate an intra prediction image 265 using reconstructed pixels of neighboring blocks of a current image according to an intra prediction mode in an intra prediction mode set.
[0108] The intra-frame prediction unit 254 (or generally the mode selection unit 260) is also used to output intra-frame prediction parameters (or generally information indicating the selected intra-frame prediction mode of the block) in the form of syntax elements 266 to the entropy coding unit 270 for inclusion in the encoded image data 21, so that the video decoder 30 can perform operations, such as receiving and using the prediction parameters for decoding.
[0109] Inter prediction
[0110] In a possible implementation, the set of inter-prediction modes depends on the available reference images (i.e., at least part of the previously decoded images stored in the DBP 230 as mentioned above) and other inter-prediction parameters, such as a search window area around the area of the current image to search for the best matching reference image, and / or on pixel interpolation, such as whether to perform half-pixel, quarter-pixel and / or 1 / 16 interpolation.
[0111] The inter-frame prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (both in Figure 2 ). The motion estimation unit may be configured to receive or acquire the image 17 and the decoded image 231, or at least one or more previously reconstructed images, e.g., reconstructed images of one or more previously decoded images 231, for motion estimation. For example, a video sequence may include the current image and the previously decoded image 231, or in other words, the current image and the previously decoded image 231 may be part of or form a sequence of images forming the video sequence.
[0112] For example, the encoder 20 may be used to select a reference image from the same or different image among a plurality of other images, and provide the offset (spatial offset) between the position (x, y coordinates) of the reference image (or reference image index) and the position of the current image as an inter-frame prediction parameter to the motion estimation unit. The offset is also called a motion vector (MV).
[0113] The motion compensation unit is used to obtain, such as receiving, inter-frame prediction parameters, and perform inter-frame prediction based on or using the inter-frame prediction parameters to obtain an inter-frame prediction image. The motion compensation performed by the motion compensation unit may include extracting or generating a prediction image based on a motion / block vector determined by motion estimation, and may also include performing interpolation with sub-pixel accuracy. Interpolation filtering can generate pixel points of other pixels from pixel points of known pixels, thereby potentially increasing the number of candidate prediction images that can be used to encode the image. Once a motion vector corresponding to a current image is received, the motion compensation unit can locate the prediction image pointed to by the motion vector in one of the reference image lists.
[0114] The motion compensation unit may also generate syntax elements associated with the picture for use by video decoder 30 when decoding the pictures of the video slice.
[0115] Entropy Coding
[0116] The entropy coding unit 270 is used to apply an entropy coding algorithm or scheme (e.g., a variable length coding (VLC) scheme, a context adaptive VLC scheme (CALVC), an arithmetic coding scheme, a binarization algorithm, a context adaptive binary arithmetic coding (CABAC), a syntax-based context-adaptive binary arithmetic coding (SBAC), a probability interval partitioning entropy (PIPE) coding, or other entropy coding methods or techniques) to the quantized residual coefficients 209, inter-frame prediction parameters, intra-frame prediction parameters, loop filter parameters, and / or other syntax elements to obtain coded image data 21 that can be output in the form of a coded bit stream 21 through an output terminal 272, so that the video decoder 30, etc. can receive and use the parameters for decoding. The coded bit stream 21 can be transmitted to the video decoder 30, or it can be stored in a memory for later transmission or retrieval by the video decoder 30.
[0117] Other structural variations of the video encoder 20 may be used to encode the video stream. For example, a non-transform based encoder 20 may directly quantize the residual signal without the transform processing unit 206. In another implementation, the encoder 20 may have the quantization unit 208 and the inverse quantization unit 210 combined into a single unit.
[0118] It should be noted that the functional modules included in the above-mentioned video encoder 20, such as the mode selection unit 260, the intra-frame prediction unit 254, the inter-frame prediction unit 244, the entropy coding unit 270, the loop filter 220, etc., one or more of which can respectively adopt the neural network trained by the training engine 25 to implement the function, and the processing objects of these neural networks are all at the frame level.
[0119] It should be understood that the processing result of the current step can be further processed in the encoder 20 and the decoder 30 and then output to the next step. For example, after interpolation filtering, motion vector derivation or loop filtering, the processing result of interpolation filtering, motion vector derivation or loop filtering can be further operated, such as clipping or shifting operation.
[0120] Although the above embodiments mainly describe video coding, it should be noted that the embodiments of the decoding system 10, the encoder 20 and the decoder 30 and other embodiments described herein can also be used for still image processing or coding, that is, the processing or coding of a single image in video coding that is independent of any previous or consecutive images. In general, if the image processing is limited to a single image 17, the inter-frame prediction unit 244 (encoder) may not be available. All other functions (also called tools or techniques) of the video encoder 20 and the video decoder 30 can also be used for still image processing, such as residual calculation 204, transformation 206, quantization 208, inverse quantization 210, (inverse) transformation 212, intra-frame prediction 254 and / or loop filtering 220, entropy coding 270.
[0121] Figure 3 FIG. 3 is an exemplary block diagram of an apparatus 300 according to an embodiment of the present application. The apparatus 300 may be used as Figure 1 Either or both of the source device 12 and the destination device 14 in .
[0122] The processor 302 in the device 300 may be a central processing unit. Alternatively, the processor 302 may be any other type of device or devices that are currently available or will be developed in the future and are capable of manipulating or processing information. Although a single processor such as the processor 302 shown in the figure may be used to implement the disclosed implementation, using more than one processor is faster and more efficient.
[0123] In one implementation, the memory 304 in the apparatus 300 may be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may be used as the memory 304. The memory 304 may include code and data 306 accessed by the processor 302 via a bus 312. The memory 304 may also include an operating system 308 and an application 310, which includes at least one program that allows the processor 302 to perform the methods described herein. For example, the application 310 may include applications 1 to N, and also include a video decoding application that performs the methods described herein.
[0124] The apparatus 300 may also include one or more output devices, such as a display 318. In one example, the display 318 may be a touch-sensitive display that combines a display with a touch-sensitive element that may be used to sense touch input. The display 318 may be coupled to the processor 302 via the bus 312.
[0125] Although bus 312 in device 300 is described herein as a single bus, bus 312 may include multiple buses. In addition, auxiliary storage may be directly coupled to other components of device 300 or accessed through a network, and may include a single integrated unit such as a memory card or multiple units such as multiple memory cards. Therefore, device 300 may have a variety of configurations.
[0126] Since the embodiments of the present application involve the application of neural networks, in order to facilitate understanding, some nouns or terms used in the embodiments of the present application are explained below, and these nouns or terms are also considered as part of the content of the invention.
[0127] (1) Neural Network
[0128] Neural Network (NN) is a machine learning model. A neural network can be composed of neural units. A neural unit can refer to an operation unit with xs and intercept 1 as input. The output of the operation unit can be:
[0129]
[0130] Where s=1, 2, ...n, n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal in the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolution layer. The activation function can be a sigmoid function. A neural network is a network formed by connecting many of the above-mentioned single neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.
[0131] (2) Deep Neural Networks
[0132] A deep neural network (DNN), also known as a multi-layer neural network, can be understood as a neural network with many hidden layers. There is no special metric for "many" here. From the position of different layers of DNN, the neural network inside DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the layers in between are all hidden layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the i+1-th layer. Although DNN looks complicated, the work of each layer is actually not complicated. Simply put, it is the following linear relationship expression: in, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also called coefficient), and α() is the activation function. Each layer is just an input vector After such a simple operation, the output vector Since DNN has many layers, the coefficient W and the offset vector The definition of these parameters in DNN is as follows: Take coefficient W as an example: Assume that in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer number of the coefficient W, while the subscripts correspond to the output third layer index 2 and the input second layer index 4. In summary, the coefficients from the kth neuron in the L-1th layer to the jth neuron in the Lth layer are defined as It should be noted that the input layer does not have a W parameter. In a deep neural network, more hidden layers allow the network to better describe complex situations in the real world. Theoretically, the more parameters a model has, the higher its complexity and the greater its "capacity", which means it can complete more complex learning tasks. Training a deep neural network is the process of learning the weight matrix, and its ultimate goal is to obtain the weight matrix of all layers of the trained deep neural network (the weight matrix formed by many layers of vector W).
[0133] (3) Convolutional Neural Network
[0134] A convolutional neural network (CNN) is a deep neural network with a convolutional structure. It is a deep learning architecture. A deep learning architecture refers to multiple levels of learning at different levels of abstraction through machine learning algorithms. As a deep learning architecture, CNN is a feed-forward artificial neural network in which each neuron can respond to the image input into it. A convolutional neural network contains a feature extractor consisting of a convolutional layer and a pooling layer. The feature extractor can be regarded as a filter, and the convolution process can be regarded as using a trainable filter to convolve an input image or convolution feature plane (feature map).
[0135] A convolutional layer is a layer of neurons in a convolutional neural network that performs convolution on the input signal. A convolutional layer can include many convolution operators, also known as kernels. In image processing, the convolution operator is essentially a weight matrix, which is usually predefined. During the convolution operation on the image, the weight matrix is usually processed horizontally on the input image one pixel after another (or two pixels after two pixels... depending on the value of the stride) to complete the work of extracting specific features from the image. The size of the weight matrix should be related to the size of the image. It should be noted that the depth dimension of the weight matrix is the same as the depth dimension of the input image. During the convolution operation, the weight matrix will extend to the entire depth of the input image. Therefore, convolution with a single weight matrix will produce a convolution output with a single depth dimension, but in most cases, instead of using a single weight matrix, multiple weight matrices of the same size (rows × columns) are applied, that is, multiple isotype matrices. The output of each weight matrix is stacked to form the depth dimension of the convolution image, where the dimension can be understood as being determined by the "multiple" mentioned above. Different weight matrices can be used to extract different features in the image, for example, one weight matrix is used to extract image edge information, another weight matrix is used to extract specific colors of the image, and another weight matrix is used to blur unnecessary noise points in the image. The multiple weight matrices have the same size (rows × columns), and the size of the feature maps extracted by the multiple weight matrices of the same size is also the same. The extracted multiple feature maps of the same size are then merged to form the output of the convolution operation. The weight values in these weight matrices need to be obtained through a lot of training in practical applications. The weight matrices formed by the weight values obtained through training can be used to extract information from the input image, so that the convolutional neural network can make correct predictions. When the convolutional neural network has multiple convolutional layers, the initial convolutional layer often extracts more general features, which can also be called low-level features; as the depth of the convolutional neural network deepens, the features extracted by the later convolutional layers become more and more complex, such as high-level semantic features. Features with higher semantics are more suitable for the problem to be solved.
[0136] Since it is often necessary to reduce the number of training parameters, it is often necessary to periodically introduce a pooling layer after the convolution layer. It can be a convolution layer followed by a pooling layer, or multiple convolution layers followed by one or more pooling layers. In the image processing process, the only purpose of the pooling layer is to reduce the spatial size of the image. The pooling layer can include an average pooling operator and / or a maximum pooling operator to sample the input image to obtain an image of smaller size. The average pooling operator can calculate the pixel values in the image within a specific range to produce an average value as the result of average pooling. The maximum pooling operator can take the pixel with the largest value in the range as the result of maximum pooling within a specific range. In addition, just as the size of the weight matrix used in the convolution layer should be related to the image size, the operator in the pooling layer should also be related to the image size. The size of the image output after processing by the pooling layer can be smaller than the size of the image input to the pooling layer. Each pixel in the image output by the pooling layer represents the average value or maximum value of the corresponding sub-region of the image input to the pooling layer.
[0137] After being processed by the convolution layer / pooling layer, the convolution neural network is not sufficient to output the required output information. Because as mentioned above, the convolution layer / pooling layer will only extract features and reduce the parameters brought by the input image. However, in order to generate the final output information (the required class information or other related information), the convolution neural network needs to use the neural network layer to generate one or a group of outputs of the required number of classes. Therefore, the neural network layer may include multiple hidden layers, and the parameters contained in the multiple hidden layers can be pre-trained according to the relevant training data of the specific task type. For example, the task type may include image recognition, image classification, image super-resolution reconstruction, etc.
[0138] Optionally, after the multiple hidden layers in the neural network layer, an output layer of the entire convolutional neural network is also included. The output layer has a loss function similar to the classification cross entropy, which is specifically used to calculate the prediction error. Once the forward propagation of the entire convolutional neural network is completed, the back propagation will begin to update the weight values and biases of the aforementioned layers to reduce the loss of the convolutional neural network and the error between the result output by the convolutional neural network through the output layer and the ideal result.
[0139] (4) Recurrent Neural Network
[0140] Recurrent neural networks (RNN) are used to process sequence data. In the traditional neural network model, from the input layer to the hidden layer and then to the output layer, the layers are fully connected, and the nodes within each layer are disconnected. Although this ordinary neural network has solved many difficult problems, it is still powerless for many problems. For example, if you want to predict the next word in a sentence, you generally need to use the previous word, because the previous and next words in a sentence are not independent. The reason why RNN is called a recurrent neural network is that the current output of a sequence is also related to the previous output. The specific manifestation is that the network will remember the previous information and apply it to the calculation of the current output, that is, the nodes between the hidden layers are no longer disconnected but connected, and the input of the hidden layer includes not only the output of the input layer but also the output of the hidden layer at the previous moment. In theory, RNN can process sequence data of any length. The training of RNN is the same as the training of traditional CNN or DNN. The same error back propagation algorithm is used, but there is a difference: that is, if the RNN is expanded, the parameters, such as W, are shared; this is not the case with the traditional neural network mentioned above. And when using the gradient descent algorithm, the output of each step depends not only on the network at the current step, but also on the state of the network at the previous steps. This learning algorithm is called the Back Propagation Through Time (BPTT).
[0141] Since we already have convolutional neural networks, why do we need recurrent neural networks? The reason is simple. In convolutional neural networks, there is a premise that the elements are independent of each other, and the input and output are also independent, such as cats and dogs. But in the real world, many elements are interconnected, such as the change of stocks over time, or someone said: I like traveling, and my favorite place is Yunnan. I must go there if I have the chance in the future. Fill in the blank here, humans should all know to fill in "Yunnan". Because humans will make inferences based on the content of the context, but how can machines do this? RNN came into being. RNN is designed to give machines the ability to remember like humans. Therefore, the output of RNN needs to rely on the current input information and historical memory information.
[0142] (5) Loss Function
[0143] In the process of training a deep neural network, because we hope that the output of the deep neural network is as close as possible to the value we really want to predict, we can compare the predicted value of the current network with the target value we really want, and then update the weight vector of each layer of the neural network according to the difference between the two (of course, there is usually an initialization process before the first update, that is, pre-configuring parameters for each layer in the deep neural network). For example, if the predicted value of the network is high, adjust the weight vector to make it predict a lower value, and keep adjusting until the deep neural network can predict the target value we really want or a value very close to the target value we really want. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or objective function, which are important equations used to measure the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference, so the training of the deep neural network becomes a process of minimizing this loss as much as possible.
[0144] (6) Back propagation algorithm
[0145] Convolutional neural networks can use the error back propagation (BP) algorithm to correct the size of the parameters in the initial super-resolution model during the training process, so that the reconstruction error loss of the super-resolution model becomes smaller and smaller. Specifically, the forward transmission of the input signal to the output will generate error loss, and the error loss information is back-propagated to update the parameters in the initial super-resolution model, so that the error loss converges. The back propagation algorithm is a back propagation movement dominated by error loss, aiming to obtain the optimal parameters of the super-resolution model, such as the weight matrix.
[0146] (7) Generative Adversarial Networks
[0147] Generative adversarial networks (GAN) are a type of deep learning model. The model includes at least two modules: one is the generative model and the other is the discriminative model. These two modules learn from each other through competition to produce better output. Both the generative model and the discriminative model can be neural networks, specifically deep neural networks or convolutional neural networks. The basic principle of GAN is as follows: Taking the GAN that generates images as an example, suppose there are two networks, G (Generator) and D (Discriminator), where G is a network that generates images. It receives a random noise z and generates images through this noise, denoted as G(z); D is a discriminative network that is used to determine whether a picture is "real". Its input parameter is x, which represents a picture. The output D(x) represents the probability that x is a real picture. If it is 1, it means that it is 100% a real picture, and if it is 0, it means that it cannot be a real picture. In the process of training the generative adversarial network, the goal of the generative network G is to generate realistic images as much as possible to deceive the discriminative network D, while the goal of the discriminative network D is to distinguish the images generated by G from the real images as much as possible. In this way, G and D constitute a dynamic "game" process, which is also the "adversary" in the "generative adversarial network". The final result of the game is that, under ideal conditions, G can generate images G(z) that are "real" enough, while D has difficulty determining whether the images generated by G are real, that is, D(G(z)) = 0.5. In this way, an excellent generative model G is obtained, which can be used to generate images.
[0148] The following will be combined Figure 4a-4e Describes the structure of the target model (also called a neural network). Figure 4a-4e This is an exemplary architecture of a neural network according to an embodiment of the present application.
[0149] like Figure 4a As shown in the figure, the neural network includes, in the order of processing, a 3×3 convolutional layer (3×3Conv), an activation layer (Relu), a block processing layer (Res-Block), ..., a block processing layer, a 3×3 convolutional layer, an activation layer, and a 3×3 convolutional layer. The original matrix of the input neural network is processed by the above layers and then added to the original matrix to obtain the final output matrix.
[0150] like Figure 4bAs shown, the neural network includes, in the order of processing, two 3×3 convolutional layers and activation layers, one block processing layer, ..., block processing layer, 3×3 convolutional layer, activation layer and 3×3 convolutional layer. The first matrix passes through one 3×3 convolutional layer and activation layer, and the second matrix passes through another 3×3 convolutional layer and activation layer. The two processed matrices are combined (contact) and then processed through the block processing layer, ..., block processing layer, 3×3 convolutional layer, activation layer and 3×3 convolutional layer to obtain a matrix, which is then added to the first matrix to obtain the final output matrix.
[0151] like Figure 4c As shown, the neural network includes, in the order of processing, two 3×3 convolutional layers and activation layers, one block processing layer, ..., block processing layer, 3×3 convolutional layer, activation layer and 3×3 convolutional layer. Before inputting into the neural network, the first matrix and the second matrix are multiplied, and then the first matrix passes through one 3×3 convolutional layer and activation layer, and the multiplied matrix passes through another 3×3 convolutional layer and activation layer. The two processed matrices are added and then passed through the block processing layer, ..., block processing layer, 3×3 convolutional layer, activation layer and 3×3 convolutional layer to obtain a matrix, which is then added to the first matrix to obtain the final output matrix.
[0152] like Figure 4d As shown in the figure, the block processing layers include, in order of processing, a 3×3 convolution layer, an activation layer, and a 3×3 convolution layer. After the input matrix is processed by these three layers, the processed matrix is added to the initial input matrix to obtain the output matrix. Figure 4e As shown in FIG. 1 , the block processing layers include, in order of processing, a 3×3 convolution layer, an activation layer, a 3×3 convolution layer, and an activation layer. After the input matrix is processed by the 3×3 convolution layer, the activation layer, and the 3×3 convolution layer, the processed matrix is added to the initial input matrix and then passed through an activation layer to obtain the output matrix.
[0153] It should be noted that if Figure 4a-4e Only several exemplary architectures of the neural network are shown, which does not constitute a limitation on the neural network architecture. The number of layers, layer structure, addition, multiplication or merging and other processing included in the neural network, as well as the number and size of input and / or output matrices can be determined according to actual conditions, and this application does not make specific limitations on this. In addition, the image can be represented in the form of a matrix, and the element (i, j) in the matrix corresponds to the pixel point in the i-th row and j-th column of the image, and the value of the element (i, j) is the pixel value of the corresponding pixel point, such as the chromaticity value, brightness value, etc. of the corresponding pixel point.
[0154] Figure 5 is an exemplary structural diagram of an end-to-end video coding architecture according to an embodiment of the present application, such as Figure 5 As shown,
[0155] A. Current image x t When it is the first frame, it is input to the intranet coding module, and the reconstructed image of the image is obtained by the intranet coding module and stored in the decoded picture buffer (DPB). The motion estimation network (ME-Net) can obtain the reconstructed image of the current image from the DPB for subsequent image coding.
[0156] B. Current image x t When it is not the first frame image, it is input to ME-Net. In addition, the reconstructed image of the previous frame image is extracted from DPB and also input to ME-Net. ME-Net outputs the motion vector (MV) v of the current image t . MVv of the current image t Sequentially pass through the MV Encoder network (to achieve MV dimension transformation, from one-dimensional vector to multi-dimensional vector in three-dimensional space) (to obtain m t ), entropy coding module (implements quantization, transformation and entropy coding) (obtains ) and MV Decoder (to achieve MV dimension transformation, from multi-dimensional vector in three-dimensional space to one-dimensional vector), to obtain the quantized compressed MV of the current image Optional, the quantized compressed MV of the current image It can be stored in a decoded MV buffer (Decoded MV Buffer) for subsequent processing.
[0157] Reconstructed image of the previous frame and the quantized compressed MV of the current image Input to the motion compensation network (MC-Net), MC-Net outputs the inter-frame prediction image of the current image Inter-frame prediction image of the current image Input to the prediction refinement network (PR-Net), PR-Net outputs an enhanced inter-frame prediction image of the current image
[0158] Current image x t Input to residual 1, in addition, the enhanced inter-frame prediction image of the current image It is also input to residual device 1, which outputs the residual image r tThe residual image r t After passing through the residual encoding network (Residual Encoder) (to achieve the dimensional transformation of the residual image) (to obtain y t ), entropy coding module (implements quantization, transformation and entropy coding) (obtains ) and residual decoding network (Residual Decoder) (to achieve the dimensional transformation of the residual image), and obtain the quantized compressed residual image of the current image The quantized compressed residual image of the current image Input to the adder 1, in addition, the enhanced inter-frame prediction image of the current image It is also input to stacker 1, which outputs an intermediate reconstructed image.
[0159] Current image x t The intermediate reconstructed image is also input to residual device 2, and residual device 2 outputs the reconstructed residual image r t ′. Reconstruct the residual image r t ′ passes through the refined residual encoding network (C2F ResidualEncoder) in turn (to achieve the dimensional transformation of the reconstructed residual image) (to obtain y t '), entropy coding module (implementing quantization, transformation and entropy coding) (obtaining ) and the refined residual decoding network (C2F Residual Decoder) (to achieve the dimensional transformation of the residual image) to obtain the quantized compressed residual image of the current image The quantized compressed residual image of the current image The image is input to the stacker 2. In addition, the intermediate reconstructed image is also input to the stacker 2. The stacker 2 outputs the reconstructed image before filtering.
[0160] The reconstructed image before filtering is input to the loop filter network (Loop Filter), and the loop filter network outputs the reconstructed image of the current image. Optionally, the reconstructed image of the current image can be stored in the DPB for subsequent image encoding.
[0161] In the above architecture, the prediction refinement network (PR-Net), the refined residual encoding network (C2F ResidualEncoder), the refined residual decoding network (C2F Residual Decoder) and the loop filter network (Loop Filter) are processing modules added to the original end-to-end encoding technology of this application, which are explained in the following embodiments.
[0162] based on Figure 5 The architecture shown in this application provides an encoding method.
[0163] Figure 6 600 is a flowchart of a process 600 of an encoding method according to an embodiment of the present application. The process 600 may be performed by the video encoder 20. The process 600 is described as a series of steps or operations. It should be understood that the process 600 may be performed in various orders and / or may occur simultaneously, not limited to Figure 6 The execution order shown. Process 600 may include:
[0164] Step 601: Obtain an inter-frame prediction image of the current image.
[0165] The current image of the present application is not the first frame of the video sequence, that is, before encoding the current image, at least one frame has been encoded and a reconstructed image of the image has been obtained. The first frame of the video sequence can be encoded and reconstructed by intra-frame prediction and encoding. The process can refer to the relevant technology and will not be repeated here.
[0166] In a possible implementation, the encoder may obtain the inter-frame prediction image of the current image by the following steps:
[0167] (1) The encoder obtains the reconstructed image of the previous frame from the decoded picture buffer (DPB). The encoding order of the previous frame is just earlier than the current image. That is, in the video sequence, according to the encoding order, the previous frame is encoded earlier than the current image, and there is no other image between the previous frame and the current image.
[0168] (2) The encoder inputs the reconstructed image of the previous frame and the current image into the motion estimation network to obtain the motion vector (MV) of the current image, and then inputs the MV of the current image into the MV encoding network and the MV decoding network to obtain the quantized compressed MV of the current image.
[0169] The motion estimation network (ME-Net) can be pre-trained. For example, ME-Net is a multi-scale convolutional neural network, which includes five spatial resolution scales, for example, H×W, H / 2×W / 2, H / 4×W / 4, H / 8×W / 8, H / 16×W / 16, H represents the height of the image, W represents the width of the image, and each scale is composed of five convolutional layers, and the output of the small-scale convolutional layer is the input of the large-scale convolutional layer. It should be understood that the ME-Net can also adopt other architectures, and this application does not specifically limit this.
[0170] The input of ME-Net is the reconstructed image of the previous frame and the current image, and the output of ME-Net is the MV of the current image, which may refer to the motion vector from the reconstructed image of the previous frame to the current image.
[0171] The MV encoding network and the MV decoding network are paired, and entropy coding processing is included between the MV encoding network and the MV decoding network. The MV encoding network is used to perform dimension conversion on the input MV to obtain the MV in the three-dimensional space corresponding to the MV. At this time, the input of the MV encoding network is the MV of the current image, and the output of the MV encoding network is the MV in the three-dimensional space of the MV of the current image. The encoder performs entropy coding on the MV in the three-dimensional space to obtain a quantized and compressed MV in the three-dimensional space. The MV decoding network is used to perform dimension conversion on the input MV in the three-dimensional space to obtain the MV corresponding to the MV in the three-dimensional space. At this time, the input of the MV decoding network is the quantized and compressed MV in the three-dimensional space, and the output of the MV decoding network is the quantized and compressed MV of the current image.
[0172] (3) The reconstructed image of the previous frame and the quantized compressed MV of the current image are input into the motion compensation network to obtain the inter-frame prediction image of the current image.
[0173] The motion compensation network (MC-Net) can be pre-trained. For example, based on the principle of bilinear interpolation, MC-Net can use bilinear difference to obtain sub-pixels of different pixel precisions of the reference image (for example, 1 / 2 pixel, 1 / 4 pixel, etc.), and then obtain the pixel value of the corresponding precision position after interpolation according to the sub-pixel precision of MV. This process can be called warping. It should be understood that the MC-Net can also be implemented based on other technologies, and this application does not make specific limitations on this.
[0174] Step 602: Input the inter-frame prediction image into the prediction refinement network to obtain an enhanced inter-frame prediction image.
[0175] The prediction refinement network (PR-Net) can be pre-trained. For example, PR-Net can be a neural network based on a residual attention network (RA-Net). PR-Net can be composed of N (the value of N can be customized, for example, N = 5) residual attention blocks (RAB) connected in series. In computer vision research, the residual attention block can generate a binary attention channel for each pixel of the input image, and use the attention channel to dynamically adjust the quality of the predicted image during the deep network learning process. Figure 7 The exemplary architecture of PR-Net of the embodiment of the present application is as follows: Figure 7 As shown, PR-Net includes 5 RABs, each of which can be composed of two convolutional layers (k3c64s1), an attention unit (Channel Attention) and a convolutional layer (k3c64s1) executed sequentially, and RAB also includes a residual connection from input to output. The attention unit (Channel Attention) can include a pooling layer (Ave Pool), a convolutional layer (k1c4s1), a convolutional layer (k1c64s1), a neuron (Sigmoid) and a convolutional layer (k3c64s1). The number of channels in the network is, for example, 64. This method can achieve better performance prediction for the edge parts and parts with more texture details in the image. It should be understood that the PRN can also adopt other architectures, and this application does not specifically limit this.
[0176] The input of PR-Net is the inter-frame prediction image of the current image, and the output of PR-Net is the enhanced inter-frame prediction image of the current image. It is precisely because PR-Net can achieve better performance prediction, so the enhanced inter-frame prediction image obtained after the inter-frame prediction image of the current image passes through PR-Net has better performance than the original inter-frame prediction image, such as the edge part of the image and the part with more texture details.
[0177] Step 603: Obtain a reconstructed image of the current image according to the enhanced inter-frame prediction image.
[0178] Based on the above Figure 5 The structure of the video encoder in the illustrated embodiment may include the following steps to obtain a reconstructed image of a current image:
[0179] (1) The encoder obtains an intermediate reconstructed image based on the current image and the enhanced inter-frame prediction image.
[0180] The encoder can obtain a residual image by taking the difference between the pixel values at corresponding positions in the current image and the enhanced inter-frame prediction image. The residual image is then input into the residual encoding network and the residual decoding network to obtain a quantized compressed residual image. Similar to the principles of the above-mentioned MV encoding network and MV decoding network, the residual encoding network and the residual decoding network are also paired, and entropy coding processing is included between the residual encoding network and the residual decoding network. It will not be repeated here. The encoder sums the pixel values at corresponding positions in the quantized compressed residual image and the enhanced inter-frame prediction image to obtain an intermediate reconstructed image.
[0181] (2) Obtain a reconstructed residual image based on the current image and the intermediate reconstructed image.
[0182] The encoder can obtain a reconstructed residual image by subtracting the pixel values at corresponding positions in the current image and the intermediate reconstructed image.
[0183] (3) The reconstructed residual image is input into the refined residual encoding network and the refined residual decoding network to obtain a quantized and compressed reconstructed residual image.
[0184] Similar to the principle of the above-mentioned MV encoding network and MV decoding network, the refined residual encoding network and the refined residual decoding network are also paired, and entropy coding processing is included between the refined residual encoding network and the refined residual decoding network. The present application forms a two-level residual processing through step (3) and step (1). On the one hand, parameter sharing is not implemented for the two-level residual network during training to ensure that the two residual networks learned are different. On the other hand, the two-level residual processing constitutes a scalable residual, which can be used to refine the modeling of the reconstructed image, and use a new residual coding module to encode the residual of the reconstructed image to further improve performance.
[0185] (4) Obtain a reconstructed image based on the quantized compressed reconstructed residual image and the intermediate reconstructed image.
[0186] The encoder can sum the pixel values of the corresponding positions in the quantized compressed reconstructed residual image and the intermediate reconstructed image to obtain the pre-filtered reconstructed image, and then input the pre-filtered reconstructed image into the loop filter network to obtain the reconstructed image.
[0187] The loop filter network (Loop Filter) can be pre-trained. For example, Figure 8 The exemplary architecture of the loop filter network of the embodiment of the present application is as follows: Figure 8As shown, the loop filter network is a fully convolutional network with 9 layers of convolution, with three convolution layers in a group, each group includes a convolution layer (k3c64s1), a convolution layer (k3c64s1 and k5c64s1), and a merging layer (Concat). The convolution kernel size of each layer is 3×3, the number of channels is 64, and there is also a residual connection directly connected from the input to the output, and the number of channels of the input layer and the output layer is 3. In this application, the loop filter network can directly improve the image quality.
[0188] Optionally, the encoder may also sum the pixel values at corresponding positions in the quantized compressed reconstructed residual image and the intermediate reconstructed image to obtain a pre-filtered reconstructed image, and the pre-filtered reconstructed image is the reconstructed image.
[0189] This application inputs the predicted image of the current image into the prediction refinement network to obtain an enhanced inter-frame prediction image, which can have better prediction performance than the original inter-frame prediction image. In addition, through two-level residual processing, on the one hand, parameter sharing is not implemented for the two-level residual network during training to ensure that the two residual networks learned are different. On the other hand, the two-level residual processing constitutes a scalable residual, which can be used to finely model the reconstructed image, and use a new residual coding module to encode the residual of the reconstructed image to further improve performance. Coupled with the loop filtering network, the quality of the reconstructed image can be further improved.
[0190] Fig. 9 900 is an exemplary structural diagram of an encoding device 900 according to an embodiment of the present application. The encoding device 900 comprises: an acquisition module 901, a prediction module 902 and a reconstruction module 903, wherein:
[0191] The acquisition module 901 is used to obtain the inter-frame prediction image of the current image, where the current image is not the first frame image of the video sequence; the prediction module 902 is used to input the inter-frame prediction image into the prediction refinement network to obtain an enhanced inter-frame prediction image; and the reconstruction module 903 is used to obtain the reconstructed image of the current image based on the enhanced inter-frame prediction image.
[0192] In one possible implementation, the acquisition module 901 is specifically used to obtain a reconstructed image of a previous frame image, where the encoding order of the previous frame image is just earlier than the current image; obtain a quantized compressed motion vector MV of the current image; and input the reconstructed image of the previous frame image and the quantized compressed MV of the current image into a motion compensation network to obtain the inter-frame predicted image.
[0193] In a possible implementation, the reconstruction module 903 is specifically used to obtain an intermediate reconstructed image based on the current image and the enhanced inter-frame prediction image; obtain a reconstructed residual image based on the current image and the intermediate reconstructed image; input the reconstructed residual image into a refined residual encoding network and a refined residual decoding network to obtain a quantized compressed reconstructed residual image; and obtain the reconstructed image based on the quantized compressed reconstructed residual image and the intermediate reconstructed image.
[0194] In a possible implementation, the reconstruction module 903 is specifically configured to obtain a pre-filtering reconstructed image according to the quantized compressed reconstructed residual image and the intermediate reconstructed image; and input the pre-filtering reconstructed image into a loop filtering network to obtain the reconstructed image.
[0195] In a possible implementation manner, the reconstruction module 903 is specifically configured to obtain the reconstructed image according to the quantized compressed reconstructed residual image and the intermediate reconstructed image.
[0196] In a possible implementation, the reconstruction module 903 is specifically used to obtain a residual image by taking the difference between the pixel values at corresponding positions in the current image and the enhanced inter-frame prediction image; inputting the residual image into the residual encoding network and the residual decoding network to obtain a quantized compressed residual image; and summing the pixel values at corresponding positions in the quantized compressed residual image and the enhanced inter-frame prediction image to obtain the intermediate reconstructed image.
[0197] In a possible implementation, the reconstruction module 903 is specifically configured to obtain the reconstructed residual image by subtracting pixel values at corresponding positions in the current image and the intermediate reconstructed image.
[0198] In a possible implementation, the reconstruction module 903 is specifically configured to sum the pixel values at corresponding positions in the quantized compressed reconstructed residual image and the intermediate reconstructed image to obtain the pre-filtering reconstructed image.
[0199] In a possible implementation, the reconstruction module 903 is specifically configured to sum the pixel values at corresponding positions in the quantized compressed reconstructed residual image and the intermediate reconstructed image to obtain the reconstructed image.
[0200] In a possible implementation, the acquisition module 901 is specifically used to input the reconstructed image of the previous frame image and the current image into a motion estimation network to obtain the MV of the current image; and input the MV of the current image into an MV encoding network and an MV decoding network to obtain a quantized compressed MV of the current image.
[0201] The device of this embodiment can be used to perform Figure 6The technical solution of the method embodiment shown has similar implementation principles and technical effects, which will not be repeated here.
[0202] In the implementation process, each step of the above method embodiment can be completed by an integrated logic circuit of hardware in a processor or an instruction in the form of software. The processor can be a general-purpose processor, a digital signal processor (digital signal processor, DSP), an application-specific integrated circuit (application-specific integrated circuit, ASIC), a field programmable gate array (field programmable gate array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware coding processor to be executed, or the hardware and software modules in the coding processor are combined to be executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0203] The memory mentioned in the above embodiments may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0204] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0205] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0206] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0207] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0208] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0209] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks or optical disks.
[0210] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A coding method, characterized in that: include: Obtaining an inter-frame prediction image of a current image, where the current image is not a first frame image of a video sequence; Inputting the inter-frame prediction image into a prediction refinement network to obtain an enhanced inter-frame prediction image; Acquire a reconstructed image of the current image according to the enhanced inter-frame prediction image; The step of obtaining a reconstructed image of the current image according to the enhanced inter-frame prediction image comprises: Acquire an intermediate reconstructed image according to the current image and the enhanced inter-frame prediction image; Acquire a reconstructed residual image according to the current image and the intermediate reconstructed image; Inputting the reconstructed residual image into a refined residual encoding network and a refined residual decoding network to obtain a quantized compressed reconstructed residual image; The reconstructed image is acquired according to the quantized compressed reconstructed residual image and the intermediate reconstructed image.
2. The method according to claim 1, characterized in that The obtaining of the inter-frame prediction image of the current image comprises: Acquire a reconstructed image of a previous frame of image, where the encoding order of the previous frame of image is just earlier than that of the current image; Obtaining a quantized compressed motion vector MV of the current image; The reconstructed image of the previous frame image and the quantized compressed MV of the current image are input into a motion compensation network to obtain the inter-frame prediction image.
3. The method according to claim 1, characterized in that The step of obtaining the reconstructed image according to the quantized compressed reconstructed residual image and the intermediate reconstructed image comprises: Acquire a pre-filtering reconstructed image according to the quantized compressed reconstructed residual image and the intermediate reconstructed image; The pre-filtered reconstructed image is input into a loop filter network to obtain the reconstructed image.
4. The method according to claim 1, characterized in that: The step of obtaining the reconstructed image according to the quantized compressed reconstructed residual image and the intermediate reconstructed image comprises: Acquire a pre-filtering reconstructed image according to the quantized compressed reconstructed residual image and the intermediate reconstructed image; The pre-filtering reconstructed image is determined as the reconstructed image.
5. The method according to claim 1, 3 or 4, characterized in that: The step of acquiring an intermediate reconstructed image according to the current image and the enhanced inter-frame prediction image comprises: Difference is calculated between pixel values at corresponding positions in the current image and the enhanced inter-frame prediction image to obtain a residual image; Inputting the residual image into a residual encoding network and a residual decoding network to obtain a quantized compressed residual image; The intermediate reconstructed image is obtained by summing pixel values at corresponding positions in the quantized compressed residual image and the enhanced inter-frame prediction image.
6. The method according to claim 1, 3 or 4, characterized in that: The step of obtaining a reconstructed residual image according to the current image and the intermediate reconstructed image comprises: The reconstructed residual image is obtained by subtracting pixel values at corresponding positions in the current image and the intermediate reconstructed image.
7. The method according to claim 3, characterized in that The step of obtaining the reconstructed image before filtering according to the quantized compressed reconstructed residual image and the intermediate reconstructed image comprises: The pre-filtering reconstructed image is obtained by summing the pixel values at corresponding positions in the quantized compressed reconstructed residual image and the intermediate reconstructed image.
8. The method according to claim 2, characterized in that: The step of obtaining the quantized compressed motion vector MV of the current image comprises: Inputting the reconstructed image of the previous frame image and the current image into a motion estimation network to obtain the MV of the current image; The MV of the current image is input into the MV encoding network and the MV decoding network to obtain the quantized compressed MV of the current image.
9. An encoding device, characterized in that: include: An acquisition module, used for acquiring an inter-frame prediction image of a current image, where the current image is not the first frame image of a video sequence; A prediction module, configured to input the inter-frame prediction image into a prediction refinement network to obtain an enhanced inter-frame prediction image; A reconstruction module, used for obtaining a reconstructed image of the current image according to the enhanced inter-frame prediction image; The reconstruction module is specifically used to obtain an intermediate reconstructed image based on the current image and the enhanced inter-frame prediction image; obtain a reconstructed residual image based on the current image and the intermediate reconstructed image; input the reconstructed residual image into a refined residual encoding network and a refined residual decoding network to obtain a quantized compressed reconstructed residual image; and obtain the reconstructed image based on the quantized compressed reconstructed residual image and the intermediate reconstructed image.
10. The device according to claim 9, characterized in that The acquisition module is specifically used to acquire a reconstructed image of a previous frame image, the encoding order of the previous frame image being just earlier than the current image; and to acquire a quantized compressed motion vector MV of the current image; The reconstructed image of the previous frame image and the quantized compressed MV of the current image are input into a motion compensation network to obtain the inter-frame prediction image.
11. The device according to claim 9, characterized in that The reconstruction module is specifically used to obtain a pre-filtering reconstructed image based on the quantized compressed reconstructed residual image and the intermediate reconstructed image; and input the pre-filtering reconstructed image into a loop filtering network to obtain the reconstructed image.
12. The device according to claim 9, characterized in that The reconstruction module is specifically used to obtain a pre-filtering reconstructed image according to the quantized compressed reconstructed residual image and the intermediate reconstructed image; and determine the pre-filtering reconstructed image as the reconstructed image.
13. The device according to claim 9, 11 or 12, characterized in that The reconstruction module is specifically used to obtain a residual image by taking the difference between the pixel values at corresponding positions in the current image and the enhanced inter-frame prediction image; inputting the residual image into the residual encoding network and the residual decoding network to obtain a quantized compressed residual image; and summing the pixel values at corresponding positions in the quantized compressed residual image and the enhanced inter-frame prediction image to obtain the intermediate reconstructed image.
14. The device according to claim 9, 11 or 12, characterized in that The reconstruction module is specifically used to obtain the reconstructed residual image by subtracting the pixel values at corresponding positions in the current image and the intermediate reconstructed image.
15. The device according to claim 11, characterized in that The reconstruction module is specifically used to sum the pixel values at corresponding positions in the quantized compressed reconstructed residual image and the intermediate reconstructed image to obtain the pre-filtering reconstructed image.
16. The device according to claim 10, characterized in that The acquisition module is specifically used to input the reconstructed image of the previous frame image and the current image into a motion estimation network to obtain the MV of the current image; and input the MV of the current image into an MV encoding network and an MV decoding network to obtain a quantized compressed MV of the current image.
17. An encoder, characterized in that: include: one or more processors; A non-transitory computer-readable storage medium, coupled to the processor and storing a program executed by the processor, wherein the program, when executed by the processor, causes the encoder to perform the method according to any one of claims 1-8.
18. A computer program product, characterized in that The method comprises a program code which, when executed on a computer or a processor, is used to perform the method according to any one of claims 1 to 8.
19. A non-transitory computer-readable storage medium, characterized in that: The method comprises a program code which, when executed by a computer device, is used to perform the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Video denoising method based on cascaded deep residual network
CN110930327A
Method and apparatus for video encoding and video decoding based on neural network
US20190246102A1