Encoding and decoding methods, apparatus, and related equipment

The decoding method improves decoding performance and reduces complexity by generating modified components from initial components with different resolutions using Haar wavelet transforms and feature enhancement, effectively addressing high computational complexity in neural network-based image decoding.

JP7893986B2Active Publication Date: 2026-07-22HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
Filing Date
2024-03-08
Publication Date
2026-07-22

AI Technical Summary

Technical Problem

Existing encoding and decoding methods based on neural networks face challenges with high computational complexity and decreased decoding performance, particularly in restoring the resolution of chromaticity signals in YUV color space during image compression.

Method used

A decoding method that involves generating a modified component from initial components with different resolutions using Haar wavelet transforms and feature enhancement processes, such as concatenation and residual block networks, to improve decoding performance and reduce complexity.

Benefits of technology

The method enhances decoding performance, reduces computational complexity, and recovers compression defects, making it suitable for cost and latency-sensitive equipment under various bitrates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007893986000031
    Figure 0007893986000031
  • Figure 0007893986000032
    Figure 0007893986000032
  • Figure 0007893986000033
    Figure 0007893986000033
Patent Text Reader

Abstract

The present invention provides a decoding method, apparatus, and device thereof, which includes the steps of: decoding a bitstream corresponding to a current image block to obtain a reconstructed image block, the reconstructed image block including a first initial component and a second initial component, where the resolution of the first initial component is equal to or greater than the resolution of the second initial component; generating an adjusted component corresponding to the second initial component based on the first initial component and the second initial component; and performing a feature enhancement process on the adjusted component to obtain a restored target component corresponding to the second initial component. The technical solution of the present invention can improve decoding performance and reduce decoding complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of encoding and decoding, and particularly relates to an encoding and decoding method, apparatus, and device thereof.

Background Art

[0002] In order to achieve the purpose of space saving, all moving pictures are encoded and then transmitted. The complete video encoding process may include prediction, transformation, quantization, entropy encoding, filtering, etc. For the prediction process, it may include intra prediction and inter prediction. Inter prediction utilizes the temporal correlation of the video, predicts the current pixel using the pixels of adjacent encoded images, and achieves the purpose of efficiently removing the redundancy in the temporal domain of the video. Intra prediction utilizes the spatial correlation of the video, predicts the current pixel using the pixels of the encoded blocks in the current frame image, and achieves the purpose of removing the redundancy in the spatial domain of the video.

[0003] With the rapid development of deep learning, deep learning has achieved success in many high-level computer vision tasks such as image classification and object detection, and its application is also progressing in the field of encoding and decoding. That is, an image can be encoded and decoded using a neural network. However, although the encoding and decoding methods based on neural networks show high performance potential, there are still problems such as a decrease in decoding performance and high computational complexity.

Summary of the Invention

[0004] In view of such a situation, the present invention provides an encoding and decoding method, apparatus, and device thereof that can improve decoding performance.

[0005] According to a first aspect, the present invention provides a decoding method, and the method includes: A step of decoding a bitstream corresponding to the current image block to obtain a reconstructed image block, wherein the reconstructed image block includes a first initial component and a second initial component, and the resolution of the first initial component is greater than or equal to the resolution of the second initial component. A step of generating a modified component corresponding to the second initial component based on the first initial component and the second initial component, The process includes the step of performing a feature enhancement process on the adjusted component to obtain a restored target component corresponding to the second initial component.

[0006] In some embodiments, the reconstructed image block is a reconstructed image block in YUV format, the first initial component is a luminance component, and the second initial components are a chromaticity U component and a chromaticity V component.

[0007] In some embodiments, the step of generating a modified component corresponding to the second initial component based on the first initial component and the second initial component is: The steps include obtaining a first image feature corresponding to the first initial component, The steps include obtaining a second image feature corresponding to the second initial component, The step of generating the adjusted component based on the first image feature and the second image feature is included, The aforementioned first image feature includes an image feature that has not been processed by a neural network, The second set of image features includes image features that have not been processed by a neural network.

[0008] In some embodiments, the first image feature includes a Haar wavelet frequency domain feature, and the second image feature includes a Haar wavelet frequency domain feature.

[0009] In some embodiments, the step of generating a modified component corresponding to the second initial component based on the first initial component and the second initial component is: The steps include performing at least one Haar wavelet transform on the first initial component and obtaining multiple frequency bands after the transformation as a first feature map, The steps include performing at least one Haar wavelet transform on the second initial component and obtaining multiple frequency bands after the transformation as a second feature map, The process includes the step of performing a concatenation operation on the first feature map and the second feature map to obtain the adjusted component.

[0010] In some embodiments, performing a concatenation operation on the first and second feature maps includes performing a concatenation operation on the first and second feature maps along the channel dimension.

[0011] In some embodiments, performing at least one Haar wavelet transform on the first initial component and obtaining a plurality of frequency bands after the transform includes performing two Haar wavelet transforms on the first initial component and obtaining a plurality of frequency bands obtained by the two Haar wavelet transforms. Performing at least one Haar wavelet transform on the second initial component and obtaining multiple frequency bands after the transform includes performing two Haar wavelet transforms on the second initial component and obtaining multiple frequency bands obtained from the two Haar wavelet transforms.

[0012] In some embodiments, performing multiple Haar wavelet transforms on the first initial component and obtaining multiple frequency bands after the transformation includes performing two Haar wavelet transforms on the first initial component and obtaining multiple frequency bands obtained from the two Haar wavelet transforms. Performing multiple Haar wavelet transforms on the second initial component and obtaining multiple frequency bands after the transformation includes performing two Haar wavelet transforms on the second initial component and obtaining multiple frequency bands obtained from the two Haar wavelet transforms.

[0013] In some embodiments, performing two Haar wavelet transforms on the first initial component and obtaining a plurality of frequency bands obtained by the two Haar wavelet transforms includes: performing a first Haar wavelet transform on the first initial component and obtaining a plurality of frequency bands obtained by the first Haar wavelet transform; and performing a second Haar wavelet transform on the plurality of frequency bands obtained by the first Haar wavelet transform and obtaining a plurality of frequency bands obtained by the two Haar wavelet transforms. Performing two Haar wavelet transforms on the second initial component and obtaining multiple frequency bands obtained from the two Haar wavelet transforms includes performing a first Haar wavelet transform on the second initial component and obtaining multiple frequency bands obtained from the first Haar wavelet transform, and performing a second Haar wavelet transform on the multiple frequency bands obtained from the first Haar wavelet transform and obtaining multiple frequency bands obtained from the two Haar wavelet transforms.

[0014] In some embodiments, the step of performing feature enhancement processing on the adjusted component to obtain the restored target component corresponding to the second initial component includes the step of performing feature enhancement processing on the adjusted component using a plurality of residual block networks to obtain the target component.

[0015] According to a second aspect, the present invention provides an encoding method, which is The step includes encoding a first initial component and a second initial component of the current image block to obtain a bitstream corresponding to the current image block, The resolution of the first initial component is greater than or equal to the resolution of the second initial component.

[0016] In some embodiments, the current image block is in YUV format, the first initial component is the luminance component, and the second initial components are the chromaticity U component and the chromaticity V component.

[0017] According to a third aspect, the present invention provides a decoding device, said device A decoding module configured to decode a bitstream corresponding to the current image block to obtain a reconstructed image block, wherein the reconstructed image block includes a first initial component and a second initial component, and the resolution of the first initial component is greater than or equal to the resolution of the second initial component. A decision module configured to generate an adjusted component corresponding to the second initial component based on the first initial component and the second initial component, The system includes a processing module configured to perform feature enhancement processing on the adjusted component to obtain a restored target component corresponding to the second initial component.

[0018] According to a fourth aspect, the present invention provides an encoding device, which is The system includes an encoding module configured to encode a first initial component and a second initial component of the current image block to obtain a bitstream corresponding to the current image block, The resolution of the first initial component is greater than or equal to the resolution of the second initial component.

[0019] According to a fifth aspect, the present invention provides a decoding-side device including a processor and a machine-readable storage medium storing machine-executable instructions executable by the processor. The processor is configured to implement the method described in any embodiment of the first aspect by executing the machine-executable instructions.

[0020] According to a sixth aspect, the present invention provides an encoding-side device including a processor and a machine-readable storage medium storing machine-executable instructions executable by the processor. The processor is configured to implement the method described in any embodiment of the second aspect by executing the machine-executable instructions.

[0021] According to a seventh aspect, the present invention provides a machine-readable storage medium storing some computer instructions. When the computer instructions are executed by a processor, the method described in any embodiment of the first aspect is implemented, or the method described in any embodiment of the second aspect is implemented.

[0022] According to an eighth aspect, the present invention provides a computer program product including a computer program. When the computer program is executed by a processor, the method described in any embodiment of the first aspect is implemented, or the method described in any embodiment of the second aspect is implemented.

Advantages of the Invention

[0023] As can be seen from the above technical proposals, in the embodiments of the present invention, after decoding and obtaining a reconstructed image block, the reconstructed image block includes a first initial component and a second initial component, the resolution of the first initial component is greater than or equal to the resolution of the second initial component, and an adjusted component corresponding to the second initial component is generated based on the first and second initial components; that is, the second initial component is auxiliaryly adjusted based on the first initial component to obtain the adjusted component, and the target component corresponding to the second initial component is determined based on the adjusted component. This improves decoding performance, reduces decoding complexity and computational complexity, provides a low-complexity and highly efficient post-processing network model framework that has relatively low resource requirements and can be applied to cost / latency sensitive equipment, and recovers compression defects under multiple bitrates. [Brief explanation of the drawing]

[0024] [Figure 1] This is a diagram of the 3D feature matrix in one embodiment of the present invention. [Figure 2] This is a flowchart of the decoding method in one embodiment of the present invention. [Figure 3] This is a schematic diagram of the encoding-side processing process in one embodiment of the present invention. [Figure 4] This is a schematic diagram of the decoding process in one embodiment of the present invention. [Figure 5A] This is a structural diagram of end-to-end image compression in one embodiment of the present invention. [Figure 5B] This is a structural diagram of deep learning-based post-processing in one embodiment of the present invention. [Figure 6A] This is a diagram illustrating the structure of the post-processing network in one embodiment of the present invention. [Figure 6B] This is a diagram illustrating the structure of the post-processing network in one embodiment of the present invention. [Figure 6C] This is a diagram illustrating the structure of the post-processing network in one embodiment of the present invention. [Figure 6D] This is a diagram illustrating the structure of the post-processing network in one embodiment of the present invention. [Figure 6E]This is a diagram illustrating the structure of the post-processing network in one embodiment of the present invention. [Figure 6F] This is a diagram illustrating the structure of the post-processing network in one embodiment of the present invention. [Figure 6G] This is a diagram illustrating the structure of the post-processing network in one embodiment of the present invention. [Figure 6H] This is a diagram illustrating the structure of the post-processing network in one embodiment of the present invention. [Figure 6I] This is a diagram illustrating the structure of the post-processing network in one embodiment of the present invention. [Figure 7A] This is a diagram illustrating the upsampling structure in one embodiment of the present invention. [Figure 7B] This is a diagram illustrating the upsampling structure in one embodiment of the present invention. [Figure 7C] This is a diagram illustrating the upsampling structure in one embodiment of the present invention. [Figure 7D] This is a diagram illustrating the structure of the post-processing network in one embodiment of the present invention. [Figure 7E] This is a structural diagram of the feature enhancement process in one embodiment of the present invention. [Figure 7F] This is a structural diagram of the feature enhancement process in one embodiment of the present invention. [Figure 8] This is a hardware structure diagram of the decoding device in one embodiment of the present invention. [Modes for carrying out the invention]

[0025] The terms used in the embodiments of the present invention are for illustrative purposes only and do not limit the invention. The singular terms “one,” “the said,” and “the said” used in the embodiments and claims of the present invention are also intended to include the plural form unless the context clearly indicates otherwise. It should be understood that “and / or” in the terms used herein means any or all possible combination including one or more items listed in association. In the embodiments herein, terms such as “first,” “second,” and “third” may be used to describe various types of information, but it should be understood that this information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing the scope of the embodiments of the present invention, depending on the context, first information may be called second information, and similarly, second information may be called first information. Furthermore, the word “if” used may be interpreted as “when,” “in the case,” or “in response to having decided.”

[0026] Embodiments of the present invention provide a decoding method, which relates to the following concepts.

[0027] JPEG (Joint PhotographicExperts Group): JPEG is a standard for compressing continuous-tone still images, and its file extension can be .jpg or .jpeg. JPEG is a common image file format. JPEG employs an integrated coding scheme combining Predictive Coding (DPCM), Discrete Cosine Transform (DCT), and Entropy Coding to remove redundant image and color data. It belongs to the category of lossy compression formats, allowing images to be compressed into a small memory space, but this results in degradation of image data quality. In particular, applying excessively high compression ratios leads to a degradation of the final decompressed image quality. Therefore, the use of high compression ratios is not recommended when pursuing high-quality images.

[0028] JPEG-AI (Joint PhotographicExperts Group Artificial Intelligence): The scope of JPEG-AI is to develop a learning-based image coding standard that provides a single-stream, compact, compressed region representation, significantly improving compression efficiency compared to existing image coding standards at comparable subjective image quality, and effectively enhancing performance in image processing and computer vision tasks. JPEG-AI is aimed at a wide range of applications, including cloud storage, visual management, autonomous vehicles and equipment, image acquisition, storage and management, real-time management of visual data, and media distribution. The goal of JPEG-AI is to design coding / decoding solutions that significantly improve compression efficiency while maintaining the same subjective image quality, providing effective compressed region processing for machine learning-based image processing and computer vision tasks. JPEG-AI needs to provide coding and decoding suitable for implementation in hardware and software, supporting 8-bit and 10-bit depths, and performing highly efficient coding and progressive decoding of images using text and graphics.

[0029] Post-processing filter: The encoding side encodes and compresses the image to form a series of bitstreams, and the decoding side decodes from the bitstream to reconstruct the image. However, due to the characteristics of the encoding algorithm, after decoding from the bitstream and reconstructing the image, defects such as block artifacts, image artifacts, and chromaticity shifts occur in the image. The purpose of the post-processing filter is to improve the defects caused by the encoding and compression of the image, to restore the image data to the greatest extent possible, and to improve the subjective image quality after decoding and reconstructing the image.

[0030] Image Super-Resolution: Super-resolution refers to a technique that restores low-resolution images to high-resolution images, enhances image detail, and improves subjective image quality. Image super-resolution is a technique used in computer vision image processing to improve image resolution.

[0031] YUV Color Space: Each pixel in a color image can usually be described by several independent physical quantities, which constitute the spatial coordinate system, and this is the color space of the image. The YUV color space is a color space in which the pixels of a color image are described by three attributes: Y (luminance), U (chromaticity), and V (chromaticity).

[0032] Entropy coding: Entropy coding is an encoding method that follows the entropy principle in the encoding process, resulting in no information loss. Information entropy represents the average amount of information (a measure of uncertainty) in the information source. Entropy coding methods include, but are not limited to, Shannon coding, Huffman coding, and arithmetic coding.

[0033] Neural Network (NN): A neural network refers to an artificial neural network, which is a computational model composed of numerous nodes (called neurons) connected to one another. In a neural network, neuron processing units can represent different objects, such as features, alphabets, concepts, or several meaningful abstract modes. Processing units in a neural network can be divided into three types: input units, output units, and hidden units. Input units receive external signals and data, output units realize the output of the processing results, and hidden units are located between the input and output units and are not observable from outside the system. The connection weights between neurons reflect the strength of the connections between units, and the representation and processing of information are reflected in the connection relationships of the processing units. A neural network is a non-programmable, brain-like information processing method, and its essence is to acquire parallel and distributed information processing capabilities through the transformation and dynamic behavior of the neural network, mimicking the information processing capabilities of the human brain's nervous system to different degrees and hierarchies. In the field of video processing, commonly used neural networks include, but are not limited to, convolutional neural networks (CNNs), recurrent neural networks (RNNs), and fully connected networks.

[0034] Convolutional Neural Networks (CNNs): Convolutional neural networks are feedforward neural networks and are one of the most representative network structures in deep learning techniques. The artificial neurons in a convolutional neural network can respond to surrounding units within a certain coverage area and have excellent representation for large-scale image processing. The basic structure of a convolutional neural network consists of two layers. One is the feature extraction layer (also called the convolutional layer), where the input of each neuron is connected to the local receptive field of the previous layer, and the local features are extracted. Once the local features are extracted, their positional relationship to other features is also determined accordingly. The second is the feature mapping layer (also called the activation layer), where each computational layer of the neural network consists of multiple feature mappings, each feature mapping is a plane, and the weights of all neurons in the plane are equal. The feature mapping structure can utilize functions such as the Sigmoid function, ReLU (Rectified Linear Unit) function, Leaky-ReLU function, PReLU (Parametric ReLU) function, and GDN (Generalized Divisive Normalization) function as activation functions for the convolutional network. Furthermore, because neurons on a single mapping surface share weights, the number of free parameters in the network can be reduced.

[0035] For example, one advantage of convolutional neural networks compared to image processing algorithms is that they can avoid complex pre-processing steps for images (such as extracting artificial features), directly inputting the original image and performing end-to-end learning. Another advantage of convolutional neural networks compared to general neural networks is that general neural networks all employ a fully connected architecture, meaning all neurons from the input layer to the hidden layer are interconnected. This results in a huge number of parameters, making network training time-consuming and ultimately difficult. Convolutional neural networks avoid this difficulty through methods such as local connections and weight sharing.

[0036] Deconvolution: Also known as transposed convolution, deconvolution layers operate similarly to convolutional layers. The main difference is that deconvolution layers, through padding, can make the output size larger than the input size (although it may be maintained at the same size). A stride of 1 indicates that the output size is equal to the input size, while a stride of N indicates that the width of the output features is N times the width of the input features, and the height of the output features is N times the height of the input features.

[0037] Generalization Ability: Generalization ability refers to a machine learning algorithm's ability to adapt to untrained samples. The goal of learning is to learn the underlying rules of data pairs, and the trained network can produce appropriate outputs for data other than the training set that shares the same rules. This ability may be called generalization ability.

[0038] Feature: The feature of the present invention is a C × W × H three-dimensional feature matrix. As shown in Figure 1, this is a diagram of the three-dimensional feature matrix, where C represents the number of channels, H represents the feature height, and W represents the feature width. The three-dimensional feature matrix may be the input to a neural network or the output of a neural network.

[0039] Rate-Distortion Optimized: There are two metrics for evaluating encoding efficiency: bitrate and PSNR (Peak Signal to Noise Ratio). A smaller bitstream results in greater compression, and a higher PSNR results in better reconstructed image quality. When selecting a mode, the cost function is essentially a combined evaluation of both. For example, the cost corresponding to a mode is J(mode) = D + λ × R, where D represents distortion, which can usually be evaluated using the SSE metric, which is the mean square sum of the differences between the reconstructed image block and the source image. To consider the cost, the SAD metric may also be used, where SAD is the sum of the absolute differences between the reconstructed image block and the source image, λ is the Lagrange multiplier, and R is the actual number of bits required to encode the image block in that mode, including the total number of bits required for encoding such as mode information, motion information, and residuals. When selecting a mode, comparing and evaluating encoding modes using rate-distortion optimization usually guarantees optimal encoding performance.

[0040] For each module on the encoding side, a variety of encoding tools exist, and each tool has multiple modes. Often, the encoding tool that yields optimal encoding performance differs depending on the video sequence. Therefore, during the encoding process, Rate-Distortion Optimize (RDO) is usually used to compare the encoding performance of different tools or modes and select the optimal mode. After determining the optimal tool or mode, the decision information is transmitted by encoding marker information into the bitstream. Although this method results in high encoding complexity, it can adaptively select the optimal mode combination for different content and obtain optimal encoding performance. On the decoding side, the relevant mode information can be obtained by directly analyzing the marker information, and the impact of complexity is small.

[0041] In deep learning-based end-to-end image compression technology, the image is converted to the YUV color space, the UV color signal is downsampled by half, and then the Y and UV signals are compressed and encoded. The decoding side performs decoding and reconstruction based on the bitstream data to obtain the reconstructed Y and UV signals. Since the UV color signal is compressed and encoded after being downsampled by half, meaning the resolution of the UV color signal is halved, it is possible to restore the resolution of the UV color signal in the post-processing stage. However, the computational complexity and resource requirements of this resolution restoration process are high, and it is not possible to restore the resolution of the UV color signal using cost- and latency-sensitive equipment.

[0042] Based on the above findings, this embodiment proposes a neural network-based compression defect correction method. It constructs a low-complexity and efficient post-processing network model framework in the field of end-to-end image compression. This framework can enhance the UV signal based on the Y signal, is low-complexity, high-performance, and can recover compression defects under multiple bitrates.

[0043] The decoding method in the embodiment of the present invention will be described in detail below by combining several specific examples.

[0044] Example 1: An embodiment of the present invention provides a decoding method, which is a flowchart of the decoding method as shown in Figure 2, and the method may be applied to the decoding side (also called a video decoder), and the method may include steps 201 to 203.

[0045] In step 201, the bitstream corresponding to the current image block is decoded to obtain the reconstructed image block.

[0046] For example, a reconstructed image block may include a first initial component and a second initial component, where the resolution of the second initial component is the resolution after downsampling, and the resolution of the first initial component may be greater than or equal to the resolution of the second initial component. For example, the resolution of the first initial component may be the resolution before downsampling (i.e., the original resolution) or the resolution after downsampling, as long as it is greater than or equal to the resolution of the second initial component, and is not limited to these.

[0047] For example, if the reconstructed image block is in YUV format, the first initial component may be the luminance component, and the second initial component may be the chromaticity U component and / or the chromaticity V component. Alternatively, if the reconstructed image block is in RGB format, the first initial component may be the G component, and the second initial component may be the R component and / or the B component.

[0048] In step 202, a modified component corresponding to the second initial component is generated based on the first and second initial components. The reconstruction quality of the modified component may be better than that of the second initial component, the reconstruction quality of the modified component may be equivalent to that of the second initial component, and the reconstruction quality of the second initial component may be better than that of the modified component.

[0049] For example, a first image feature corresponding to a first initial component may be obtained, and this first image feature may include, but is not limited to, image features that have not been processed by a neural network and / or image features obtained by a neural network. A second image feature corresponding to a second initial component may be obtained, and this second image feature may include, but is not limited to, image features that have not been processed by a neural network and / or image features obtained by a neural network. A modified component may be generated based on the first and second image features.

[0050] For example, if the first image feature includes an image feature that has not been processed by a neural network, the first image feature may include, but is not limited to, at least one of the following: a texture feature, a subjective feature, a frequency domain feature, or a histogram feature. If the second image feature includes an image feature that has not been processed by a neural network, the second image feature may include, but is not limited to, at least one of the following: a texture feature, a subjective feature, a frequency domain feature, or a histogram feature. The above are just some examples and are not limited to them.

[0051] If a first image feature includes a first feature map obtained by a neural network, and a second image feature includes a second feature map obtained by a neural network, the step of generating an adjusted component corresponding to the second initial component based on the first and second initial components includes, but is not limited to, the steps of inputting the first initial component into a first neural network to obtain a first feature map, inputting the second initial component into a second neural network to obtain a second feature map, performing an addition operation on the first and second feature maps to obtain an adjusted component, or performing a concatenation operation on the first and second feature maps to obtain an adjusted component.

[0052] For example, if a first image feature includes a weight coefficient map obtained by a neural network and a second image feature includes a second feature map obtained by a neural network, the step of generating an adjusted component corresponding to the second initial component based on the first and second initial components includes, but is not limited to, inputting the first initial component into a first neural network, obtaining a first feature map corresponding to the first initial component, performing guided filtering on the first feature map, obtaining a weight coefficient map, inputting the second initial component into a second neural network, obtaining a second feature map, and generating an adjusted component based on the weight coefficient map and the second feature map.

[0053] For example, performing guided filtering on a first feature map to obtain a weight coefficient map may include, but is not limited to, performing a convolution operation on the first feature map to obtain a convolved feature map, performing weight mapping on the convolved feature map to obtain a weight coefficient map.

[0054] Exemplary examples include, but are not limited to, performing weight mapping on a convolutional feature map to obtain a weight coefficient map, performing a pooling operation on the convolutional feature map to obtain the feature map after the pooling operation, performing a fully connected operation and a ReLU (Rectified Linear Unit) activation operation on the feature map after the pooling operation to obtain the feature map after the ReLU activation operation, performing a fully connected operation and a Sigmoid activation operation on the feature map after the ReLU activation operation to obtain the feature map after the Sigmoid activation operation, and generating a weight coefficient map based on the convolutional feature map and the feature map after the Sigmoid activation operation, and obtaining a weight coefficient map by multiplying the convolutional feature map and the feature map after the Sigmoid activation operation.

[0055] For example, performing weight mapping on a convolutional feature map to obtain a weight coefficient map may include, but is not limited to, performing a convolution operation and a sigmoid activation operation on a convolutional feature map to obtain a weight coefficient map.

[0056] For example, generating an adjusted component based on the weight coefficient map and the second feature map may include, but is not limited to, performing a multiplication operation on the weight coefficient map and the second feature map to obtain the multiplied feature map, performing a convolution operation on the first feature map to obtain the convolved feature map, performing an addition operation on the multiplied feature map and the convolved feature map to obtain the added feature map, and performing a concatenation operation on the added feature map and the convolved feature map to obtain the adjusted component.

[0057] For example, if a first image feature includes K first feature maps obtained by a neural network, and a second image feature includes K+1 second feature maps obtained by a neural network, the step of generating an adjusted component corresponding to the second initial component based on the first and second initial components includes, but is not limited to, inputting the first initial components of the first feature map and the first second feature map into the neural network to obtain the first first feature map, and inputting the second initial components of the first second feature map into the neural network to obtain the first second feature map. For the i-th first feature map and the i-th second feature map, i is any integer between 2 and K, and K may be a positive integer greater than 1. In this case, the (i-1)th first feature map may be input into the neural network to obtain the i-th first feature map, feature fusion may be performed on the (i-1)th first feature map and the (i-1)th second feature map to obtain the fused features, the fused features may be input into the neural network to obtain the i-th second feature map, and after obtaining the last second feature map (i.e., the last second feature map in K+1 second feature maps), adjusted components may be generated based on the last second feature map.

[0058] Exemplary examples of feature fusion of a first feature map (e.g., the (i-1)th first feature map) and a second feature map (e.g., the (i-1)th second feature map) to obtain a fused feature include, but are not limited to, performing an additive operation on the first and second feature maps to obtain a fused feature, or performing a concatenation operation on the first and second feature maps to obtain a fused feature, or performing guided filtering on the first feature map to obtain a weight coefficient map corresponding to the first feature map, and then generating a fused feature based on the weight coefficient map and the second feature map. The above are just some examples and are not limited thereto.

[0059] For example, inputting a first initial component into a first neural network to obtain a first feature map includes, but is not limited to, extracting M-dimensional (where M is a positive integer) image features from the first initial component, concatenating the M-dimensional image features to obtain a concatenated multidimensional feature, and inputting the multidimensional feature into the first neural network to obtain a first feature map.

[0060] Exemplary, an M-dimensional image feature includes, but is not limited to, at least one of the following: an image feature obtained by convolving the first initial component using one convolution kernel; an image feature obtained by cascading the first initial component using two convolution kernels; a first-order spatial domain feature obtained by processing the first initial component using the Sobel operator; a second-order spatial domain feature obtained by processing the first initial component using the Laplacian operator; and a frequency domain feature obtained by performing a Fourier transform on the first initial component. The above are just some examples of image features and are not limited to them.

[0061] For example, inputting a first initial component into a first neural network to obtain a first feature map includes, but is not limited to, performing a wavelet transform on the first initial component to obtain multiple frequency bands after the wavelet transform, and inputting multiple frequency bands or a portion of multiple frequency bands into the first neural network to obtain a first feature map. Inputting a second initial component into a second neural network to obtain a second feature map includes, but is not limited to, performing a wavelet transform on the second initial component to obtain multiple frequency bands after the wavelet transform, and inputting multiple frequency bands or a portion of multiple frequency bands into the second neural network to obtain a second feature map.

[0062] When performing a wavelet transform on the first initial component to obtain multiple frequency bands after the wavelet transform, one wavelet transform may be performed on the first initial component once to obtain multiple frequency bands after the wavelet transform, or multiple wavelet transforms may be performed on the first initial component to obtain multiple frequency bands after the wavelet transform. When performing multiple wavelet transforms on the first initial component, first a wavelet transform is performed on the first initial component to obtain multiple frequency bands after the wavelet transform, then a target frequency band (all or some frequency bands in the multiple frequency bands) is selected from the multiple frequency bands, a wavelet transform is performed on the target frequency band to obtain multiple frequency bands after the wavelet transform, and this process is repeated until multiple frequency bands after the wavelet transform are obtained. When performing a wavelet transform on the second initial component to obtain multiple frequency bands after the wavelet transform, one wavelet transform may be performed on the second initial component once to obtain multiple frequency bands after the wavelet transform, or multiple wavelet transforms may be performed on the second initial component to obtain multiple frequency bands after the wavelet transform.

[0063] For example, before inputting the first initial component into the first neural network to obtain the first feature map, the first initial component may be preprocessed, and the preprocessed first initial component may be obtained. Here, the preprocessed first initial component is used to input into the first neural network to obtain the first feature map. Here, preprocessing the first initial component and obtaining the preprocessed first initial component includes, but is not limited to, performing edge enhancement on the first initial component to obtain the edge-enhanced image features, performing multiscale feature extraction on the edge-enhanced image features to obtain multiscale features, and determining the preprocessed first initial component based on the multiscale features, for example, by setting the multiscale features as the preprocessed first initial component.

[0064] For example, performing multiscale feature extraction on edge-enhanced image features to obtain multiscale features includes, but is not limited to, performing a convolution operation on the edge-enhanced image features to obtain the convolutional features, performing a downsampling operation on the convolutional features to obtain the downsampled features, performing a channel conversion on the downsampled features to obtain the channel conversion features, performing an upsampling operation on the channel conversion features to obtain the upsampled features, generating multiscale features based on the upsampled features and the channel conversion features, and then inputting the multiscale features as the first initial component after preprocessing into a first neural network.

[0065] In step 203, feature enhancement processing is performed on the adjusted component to obtain the restored target component corresponding to the second initial component.

[0066] In one possible embodiment, before generating a modified component corresponding to the second initial component based on the first and second initial components, the second initial component may be upsampled to obtain the upsampled second initial component. Here, the resolution of the upsampled second initial component may be equal to the resolution of the first initial component.

[0067] In another embodiment, feature enhancement processing may be performed on the adjusted component, and before obtaining the restored target component corresponding to the second initial component (i.e., before step 203), the adjusted component may be upsampled to obtain the upsampled adjusted component. The resolution of the upsampled adjusted component may be equal to the resolution of the first initial component.

[0068] In another embodiment, after performing feature enhancement on the adjusted component to obtain the restored target component corresponding to the second initial component (i.e., after step 203), the target component may be upsampled to obtain the upsampled target component. Here, the resolution of the upsampled target component may be equal to the resolution of the first initial component.

[0069] For example, performing feature enhancement on the adjusted component to obtain the reconstructed target component corresponding to the second initial component includes, but is not limited to, performing feature enhancement on the adjusted component using at least one residual block network to obtain the target component corresponding to the second initial component, or performing feature enhancement on the adjusted component using a U-Net (an encoder-decoder structure network in which the first half is feature extraction and the second half is upsampling) network to obtain the target component corresponding to the second initial component.

[0070] For illustrative purposes, the execution order described above is merely an example provided for illustrative purposes, and in actual applications, the order of execution between steps may be changed and is not limiting. Furthermore, in other embodiments, the steps of the corresponding method may not necessarily be performed in the order shown and described herein, and the number of steps included in such methods may be more or fewer than those described herein. Also, a single step described herein may be broken down into multiple steps and described in other embodiments, and multiple steps described herein may be combined into a single step and described in other embodiments.

[0071] As can be seen from the above technical proposals, in the embodiments of the present invention, after decoding and obtaining a reconstructed image block, the reconstructed image block includes a first initial component and a second initial component, the resolution of the second initial component is the resolution after downsampling, and the resolution of the first initial component is greater than or equal to the resolution of the second initial component. Based on the first and second initial components, an adjusted component corresponding to the second initial component is generated; that is, the second initial component is auxiliaryly adjusted based on the first initial component to obtain the adjusted component, and the target component corresponding to the second initial component is determined based on the adjusted component. This improves decoding performance, reduces decoding complexity and computational complexity, provides a low-complexity and highly efficient post-processing network model framework that has relatively low resource requirements and can be applied to cost / latency sensitive equipment, and recovers compression defects under multiple bitrates.

[0072] Example 2: For the processing process on the encoding side (also called the video encoder), please refer to Figure 3. Figure 3 is merely one example of the processing process on the encoding side, and the processing process is not limited to this example.

[0073] The encoding side may, after obtaining the current image block x (which may be the original image block x, i.e., the input image block), perform a feature transformation on the current image block x using an analytical transformation network (i.e., a neural network) to obtain the image features y corresponding to the current image block x. Here, performing a feature transformation on the current image block x using an analytical transformation network means transforming the current image block x into image features y in the latent domain, thereby making all subsequent processes operable in the latent domain.

[0074] The image may be divided into one image block or into multiple image blocks. When the image is divided into one image block, the current image block x can be considered as the image itself, that is, the encoding process for the image block can be applied directly to the image.

[0075] The encoding side obtains image features y, then performs a coefficient hyperparameter feature transformation on image features y to obtain coefficient hyperparameter feature z. For example, image features y may be input to a hyperparameter coding network (i.e., a neural network), and the hyperparameter coding network may perform a coefficient hyperparameter feature transformation on image features y to obtain coefficient hyperparameter feature z. Here, the hyperparameter coding network may be a pre-trained neural network, and its training process is not limited; it just needs to be able to perform a coefficient hyperparameter feature transformation on image features y. Here, image features y from the latent domain are processed by the hyperparameter coding network to obtain hyperprior latent information z.

[0076] The encoding side may, after obtaining the coefficient hyperparameter feature z, quantize the coefficient hyperparameter feature z to obtain the hyperparameter quantization feature corresponding to the coefficient hyperparameter feature z; that is, the Q operation in Figure 3 represents the quantization process. After obtaining the hyperparameter quantization feature corresponding to the coefficient hyperparameter feature z, the hyperparameter quantization feature is encoded to obtain Bitstream#1 (i.e., the first bitstream) corresponding to the current image block. That is, the AE operation in Figure 3 represents an encoding process such as the entropy encoding process. Alternatively, the encoding side may directly encode the coefficient hyperparameter feature z to obtain Bitstream#1 corresponding to the current image block. Here, the hyperparameter quantization feature or coefficient hyperparameter feature z included in Bitstream#1 is mainly used to obtain the mean and the parameters of the probability distribution model.

[0077] The encoding side may obtain Bitstream#1 corresponding to the current image block and then send Bitstream#1 corresponding to the current image block to the decoding side. For the decoding side's processing process for Bitstream#1 corresponding to the current image block, please refer to the subsequent embodiments.

[0078] The encoding side may obtain Bitstream#1 corresponding to the current image block, then decode Bitstream#1 to obtain the hyperparameter quantization feature; that is, AD in Figure 3 represents the decoding process. Next, the encoding side may perform inverse quantization on the hyperparameter quantization feature to obtain the coefficient hyperparameter feature z_hat. The coefficient hyperparameter feature z_hat may be the same as or different from the coefficient hyperparameter feature z. The IQ operation in Figure 3 is the inverse quantization process. Alternatively, the encoding side may obtain Bitstream#1 corresponding to the current image block, then decode Bitstream#1 to obtain the coefficient hyperparameter feature z_hat; in this case, the inverse quantization process of the coefficient hyperparameter feature z_hat is not performed.

[0079] A fixed probability density model encoding method may be used for the encoding process of Bitstream#1, and a fixed probability density model decoding method may be used for the decoding process of Bitstream#1; however, the encoding and decoding processes are not limited to this.

[0080] The encoding side may, after obtaining the coefficient hyperparameter feature z_hat, perform context-based prediction based on the coefficient hyperparameter feature z_hat of the current image block and the residual feature y_hat of the previous image block (see subsequent examples for the process of determining the residual feature y_hat), to obtain a predicted value mu (i.e., mean mu) corresponding to the current image block. For example, the coefficient hyperparameter feature z_hat and the residual feature y_hat may be input to a mean prediction network, and the mean prediction network may determine the predicted value mu based on the coefficient hyperparameter feature z_hat and the residual feature y_hat; however, this prediction process is not limited to this. Here, for the context-based prediction process, the input includes the coefficient hyperparameter feature z_hat and the decoded residual feature y_hat, and a more accurate predicted value mu is obtained by integrating both inputs. The predicted value mu is used to obtain the residual by subtracting it from the original feature and to obtain the reconstructed y by adding it to the decoded residual.

[0081] Note that the mean prediction network is a selectable neural network; in other words, it is not necessary to have a mean prediction network. That is, it is not necessary to determine the predicted value mu using the mean prediction network. The dashed box in Figure 3 indicates that the mean prediction network is selectable.

[0082] The encoding side may, after obtaining the image feature y, determine the residual feature r based on the image feature y and the predicted value mu. For example, the residual feature r is defined as the difference between the image feature y and the predicted value mu. Subsequently, feature processing is performed on the residual feature r to obtain the image feature s. This feature processing process is not limited and can be any feature processing method. In this case, it is necessary to deploy an mean prediction network, which provides the predicted value mu. Alternatively, the encoding side may, after obtaining the image feature y, perform feature processing on the image feature y to obtain the image feature s. This feature processing process is not limited and can be any feature processing method. In this case, it is not necessary to deploy an mean prediction network. The dashed box indicates that the residual process is a selectable process.

[0083] The encoding side may, after obtaining image features s, quantize the image features s to obtain image quantized features corresponding to the image features s; that is, the Q operation in Figure 3 represents the quantization process. After obtaining the image quantized features corresponding to the image features s, the encoding side may encode the image quantized features to obtain Bitstream#2 (i.e., the second bitstream) corresponding to the current image block; that is, the AE operation in Figure 3 represents an encoding process such as an entropy encoding process. Alternatively, the encoding side may directly encode the image features s to obtain Bitstream#2 corresponding to the current image block; in this case, the quantization process of image features s is not performed.

[0084] The encoding side may obtain Bitstream#2 corresponding to the current image block and then send Bitstream#2 corresponding to the current image block to the decoding side. For the encoding side's processing process for Bitstream#2 corresponding to the current image block, please refer to the subsequent embodiments.

[0085] The encoding side may obtain Bitstream#2 corresponding to the current image block, then decode Bitstream#2 to obtain the image quantization features; that is, AD in Figure 3 represents the decoding process. Next, the encoding side may dequantize the image quantization features to obtain image features s'. Image features s' may be the same as or different from image features s. The IQ operation in Figure 3 is the dequantization process. Alternatively, the encoding side may obtain Bitstream#2 corresponding to the current image block, then decode Bitstream#2 to obtain image features s'; in this case, the dequantization process of the image quantization features is not performed.

[0086] The encoding side may, after obtaining image features s', perform feature reconstruction (i.e., the reverse process of feature processing) on ​​image features s' to obtain residual features r_hat. This feature reconstruction process is not limited and can be any feature reconstruction method, and residual features r_hat may be the same as or different from residual features r. After obtaining residual features r_hat, the encoding side may determine image features y_hat based on residual features r_hat and predicted values ​​mu, and image features y_hat may be the same as or different from image features y. For example, image features y_hat may be the sum of residual features r_hat and predicted values ​​mu. In this case, it is necessary to set up an mean prediction network, which provides the predicted values ​​mu. Alternatively, the encoding side may, after obtaining image features s', perform feature reconstruction (i.e., the reverse process of feature processing) on ​​image features s' to obtain image features y_hat, and image features y_hat may be the same as or different from image features y. In this case, it is not necessary to set up an mean prediction network. The dashed box indicates that the residual process is a selectable process.

[0087] The encoding side may obtain the image feature y_hat, then perform a composite transformation on the image feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the image feature y_hat can be input to a composite transformation network, the composite transformation network can perform a composite transformation on the image feature y_hat to obtain the reconstructed image block x_hat, and thus the image reconstruction process is complete.

[0088] In one embodiment, when the encoding side encodes image quantization features or image features s to obtain Bitstream#2 corresponding to the current image block, it is necessary to first determine a probability distribution model and then encode the image quantization features or image features s based on that probability distribution model. Furthermore, when decoding Bitstream#2, the encoding side is also required to first determine a probability distribution model and then decode Bitstream#2 based on that probability distribution model.

[0089] To obtain a probability distribution model, as shown in Figure 3, the encoding side may obtain the coefficient hyperparameter feature z_hat, then perform an inverse coefficient hyperparameter feature transformation on the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter p. For example, the coefficient hyperparameter feature z_hat may be input to a probabilistic hyperparameter decoding network, the probabilistic hyperparameter decoding network may perform an inverse coefficient hyperparameter feature transformation on the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter p, and then a probability distribution model may be generated based on the probability distribution parameter p. Here, the probabilistic hyperparameter decoding network may be a pre-trained neural network, and the learning process of this probabilistic hyperparameter decoding network is not limited as long as it can perform an inverse coefficient hyperparameter feature transformation on the coefficient hyperparameter feature z_hat.

[0090] In one possible embodiment, the encoding process may be performed by a deep learning model or a neural network model. This enables an end-to-end image compression and encoding process. This encoding process is not limited to this method.

[0091] Example 3: For the processing process of the decoding side (also called the video decoder), please refer to Figure 4. Figure 4 is merely one example of the processing process of the decoding side, and the processing process is not limited to this.

[0092] The decoding side may obtain Bitstream#1 corresponding to the current image block, then decode Bitstream#1 to obtain the hyperparameter quantization feature, i.e., AD in Figure 4 represents the decoding process. Next, the decoding side may dequantize the hyperparameter quantization feature to obtain the coefficient hyperparameter feature z_hat. The coefficient hyperparameter feature z_hat may be the same as or different from the coefficient hyperparameter feature z. The IQ operation in Figure 4 is the dequantization process. Alternatively, the decoding side may obtain Bitstream#1 corresponding to the current image block, then decode Bitstream#1 to obtain the coefficient hyperparameter feature z_hat, in which case the dequantization process of the coefficient hyperparameter feature z_hat is not performed.

[0093] The decoding process for Bitstream#1 may, but is not limited to, a decoding method based on a fixed probability density model.

[0094] The image may be divided into one image block or into multiple image blocks. If the image is divided into one image block, the current image block x can be considered as the image itself, that is, the decoding process for the image block can be applied directly to the image.

[0095] The decoding side may, after obtaining the coefficient hyperparameter feature z_hat, perform context-based prediction based on the coefficient hyperparameter feature z_hat of the current image block and the residual feature y_hat of the previous image block (see subsequent examples for the process of determining the residual feature y_hat) to obtain a predicted value mu (i.e., mean mu) corresponding to the current image block. For example, the coefficient hyperparameter feature z_hat and the residual feature y_hat may be input to a mean prediction network, and the mean prediction network may determine the predicted value mu based on the coefficient hyperparameter feature z_hat and the residual feature y_hat; however, this prediction process is not limited to this. Here, for the context-based prediction process, the input includes the coefficient hyperparameter feature z_hat and the decoded residual feature y_hat, and a more accurate predicted value mu is obtained by integrating both inputs.

[0096] Note that the mean prediction network is a selectable neural network; that is, it is not necessary to have a mean prediction network. In other words, it is not necessary to determine the predicted value mu using the mean prediction network. The dashed box in Figure 4 indicates that the mean prediction network is selectable.

[0097] The decoding side may obtain Bitstream#2 corresponding to the current image block, then decode Bitstream#2 to obtain the image quantization features; that is, AD in Figure 4 represents the decoding process. Next, the decoding side may dequantize the image quantization features to obtain image features s'. Image features s' may be the same as or different from image features s. The IQ operation in Figure 4 is the dequantization process. Alternatively, the decoding side may obtain Bitstream#2 corresponding to the current image block, then decode Bitstream#2 to obtain image features s'; in this case, the dequantization process of the image quantization features is not performed.

[0098] The decoding side may, after obtaining image features s', perform feature reconstruction (i.e., the reverse process of feature processing) on ​​image features s' to obtain residual features r_hat. Residual features r_hat may be the same as or different from residual features r. After obtaining residual features r_hat, the decoding side may determine image features y_hat based on residual features r_hat and predicted values ​​mu. Image features y_hat may be the same as or different from image features y. For example, image features y_hat may be the sum of residual features r_hat and predicted values ​​mu. In this case, an mean prediction network needs to be set up, and the mean prediction network provides the predicted value mu. Alternatively, the decoding side may, after obtaining image features s', perform feature reconstruction on image features s' to obtain image features y_hat. Image features y_hat may be the same as or different from image features y. In this case, an mean prediction network does not need to be set up. The dashed box indicates that the residual process is a selectable process.

[0099] The decoding side may, after obtaining the image feature y_hat, perform a composite transformation on the image feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the image feature y_hat can be input to a composite transformation network, the composite transformation network can perform a composite transformation on the image feature y_hat to obtain the reconstructed image block x_hat, and thus the image reconstruction process is completed.

[0100] In one embodiment, when decoding Bitstream#2, the decoding side must first determine a probability distribution model and then decode Bitstream#2 based on that probability distribution model. To obtain the probability distribution model, as shown in Figure 4, the decoding side may obtain the coefficient hyperparameter feature z_hat, then perform an inverse transformation of the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter p. For example, the coefficient hyperparameter feature z_hat may be input to a probabilistic hyperparameter decoding network, the probabilistic hyperparameter decoding network may perform an inverse transformation of the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter p, and after obtaining the probability distribution parameter p, a probability distribution model may be generated based on the probability distribution parameter p. Here, the probabilistic hyperparameter decoding network may be a pre-trained neural network, and the learning process of this probabilistic hyperparameter decoding network is not limited; it is sufficient that an inverse transformation of the coefficient hyperparameter feature z_hat can be performed to obtain the probability distribution parameter p.

[0101] In one possible embodiment, the decoding process may be performed by a deep learning model or a neural network model. This enables an end-to-end image compression and encoding process. This decoding process is not limited to this method.

[0102] Example 4: As shown in Figure 5A, this is a structural diagram of end-to-end image compression based on deep learning. On the encoding side, the image is converted to the YUV color space, the UV signal is downsampled by half, and then the Y signal and UV signal are compressed and encoded respectively, and the Y signal bitstream and UV signal bitstream are sent to the decoding side. On the decoding side, the Y signal bitstream is decoded and reconstructed to obtain the reconstructed Y signal, and the UV signal bitstream is decoded and reconstructed to obtain the reconstructed UV signal. The UV signal is compressed and encoded after being downsampled by half, meaning the resolution of the UV signal is halved. Therefore, the resolution of the reconstructed UV signal is halved, meaning the resolution of the reconstructed Y signal is H × W, and the resolution of the reconstructed UV signal is (H / 2) × (W / 2), meaning the resolution of the reconstructed UV signal is half the resolution of the reconstructed Y signal.

[0103] Since the resolution of the reconstructed UV signal is half that of the reconstructed Y signal, the resolution of the UV signal may be restored in the post-processing stage, i.e., the resolution of the UV signal may be restored to H × W. However, when restoring the resolution of the UV signal, the computational complexity of the resolution restoration process is high, and the resource requirements are high, making it impossible for cost / delay-sensitive equipment to restore the resolution of the UV color signal.

[0104] Based on the above findings, embodiments of the present invention provide a low-complexity and efficient post-processing method in the field of end-to-end image compression. In the post-processing process, the UV signal can be enhanced based on the Y signal, and this method is low in complexity, high in performance, improves decoding performance, reduces decoding complexity, reduces computational complexity, has low resource requirements, can be used by cost / latency-sensitive equipment, and can restore compression defects under multiple bitrates.

[0105] In embodiments of the present invention, the decoding side can receive a bitstream corresponding to the current image block, and then decode the bitstream corresponding to the current image block to obtain a reconstructed image block. For example, the decoding side may use the processing flow of Embodiment 3 to decode the bitstream corresponding to the current image block to obtain a reconstructed image block, and the decoding process for this reconstructed image block is not limited.

[0106] After obtaining a reconstructed image block, if the reconstructed image block is in YUV format, the chromaticity U component (i.e., the U component) may be enhanced based on the luminance component (i.e., the Y component), the chromaticity V component (i.e., the V component) may be enhanced based on the Y component, or both the U and V components may be enhanced simultaneously based on the Y component. If the reconstructed image block is in RGB format, the R component may be enhanced based on the G component, the B component may be enhanced based on the G component, or both the R and B components may be enhanced simultaneously based on the G component. For the sake of explanation, this embodiment uses a reconstructed image block in YUV format as an example, so the reconstructed image block may include Y, U, and V components, and the resolution of the Y component may be greater than the resolution of the U component. For example, the resolution of the U component may be the resolution before downsampling, and the resolution of the Y component may be the original resolution or the resolution after downsampling, but in any case the resolution of the Y component is greater than the resolution of the U component. The resolution of the Y component may be greater than the resolution of the V component. For example, the resolution of the V component may be the resolution before downsampling, and the resolution of the Y component may be the original resolution or the resolution after downsampling. In any case, the resolution of the Y component is greater than the resolution of the V component. For example, let the resolution of the Y component be H × W, the resolution of the U component be (H / 2) × (W / 2), and the resolution of the V component be (H / 2) × (W / 2).

[0107] Figure 5B is a diagram of the structure of a deep learning-based post-processing. The decoding side may decode the bitstream corresponding to the current image block to obtain a reconstructed image block. If the reconstructed image block is a reconstructed image block in YUV format, the reconstructed image block may include Y components, U components, and V components, where the Y component is represented as the Y initial component (i.e., the first initial component in the embodiment described above), the U component is represented as the U initial component (i.e., the second initial component in the embodiment described above), and the V component is represented as the V initial component (i.e., the second initial component in the embodiment described above). Based on this, the U initial component may be enhanced based on the Y initial component, the V initial component may be enhanced based on the Y initial component, or the U initial component and V initial component may be enhanced based on the Y initial component. Here, the resolution of the Y initial component may be greater than the resolution of the U initial component, and the resolution of the Y initial component may be greater than the resolution of the U initial component. For example, as shown in Figure 5B, the resolution of the initial Y component is H × W, the resolution of the initial U component is (H / 2) × (W / 2), and the resolution of the initial V component is (H / 2) × (W / 2).

[0108] In this embodiment, as shown in Figure 5B, with respect to the preprocessing module, the preprocessing module is an optional module. If the preprocessing module is deployed after decoding to obtain the initial Y component, the preprocessing module may preprocess the initial Y component to obtain the preprocessed initial Y component, and this preprocessed initial Y component may be input to the Y auxiliary UV module and the YUV signal enhancement module. If the preprocessing module is not deployed, the initial Y component is directly input to the Y auxiliary UV module and the YUV signal enhancement module. In subsequent embodiments, the initial Y component obtained by the Y auxiliary UV module and the YUV signal enhancement module may be the preprocessed initial Y component or the original initial Y component obtained by decoding, and is not limited to these.

[0109] Because there is a certain correlation between the three components of YUV, the compression loss of the Y component is smaller than the compression loss of the UV component during encoding. However, a certain enhancement may be applied to the Y component to better enhance the UV component. Therefore, a preprocessing module may be used to preprocess the initial Y component and then perform feature enhancement on the initial Y component.

[0110] In this embodiment, as shown in Figure 5B, with respect to the Y auxiliary UV module, after decoding and obtaining the initial Y component, initial U component, and initial V component, the loss of the initial U component and initial V component is large, while the initial Y component retains a relatively large amount of image detail information. Furthermore, there is a certain correlation between the initial U component, initial V component, and initial Y component. Therefore, by utilizing the information of the initial Y component, it is possible to improve the reconstruction quality of the initial U component and initial V component.

[0111] Based on this, the Y auxiliary UV module may utilize information from the initial Y component to improve the reconstruction quality of the initial U component, utilize information from the initial Y component to improve the reconstruction quality of the initial V component, or utilize information from the initial Y component to improve the reconstruction quality of both the initial U component and the initial V component.

[0112] In this embodiment, as shown in Figure 5B, with respect to the resolution conversion module, if the resolution of the initial U component is smaller than the resolution of the initial Y component, for example, if the resolution of the initial U component is the resolution obtained by applying 2x downsampling, and the resolution of the initial U component is half the resolution of the original signal, for example, (H / 2) × (W / 2), then the resolution conversion module upsamples the resolution of the initial U component to restore the U component having the resolution of the original signal, that is, upsamples the resolution of the initial U component to obtain a U component with a resolution of H × W. If the resolution of the initial V component is smaller than the resolution of the initial Y component, for example, if the resolution of the initial V component is the resolution obtained by applying 2x downsampling, and the resolution of the initial V component is half the resolution of the original signal, for example, (H / 2) × (W / 2), then the resolution conversion module upsamples the resolution of the initial V component to restore the V component having the resolution of the original signal, that is, upsamples the resolution of the initial V component to obtain a V component with a resolution of H × W.

[0113] The resolution conversion module may be located at position 1, that is, the resolution conversion module upsamples the resolution of the initial U component to obtain a U component with a resolution of H×W, inputs the U component with a resolution of H×W to the Y auxiliary UV module, and the resolution conversion module upsamples the resolution of the initial V component to obtain a V component with a resolution of H×W, inputs the V component with a resolution of H×W to the Y auxiliary UV module. Alternatively, the resolution conversion module may be located at position 2, that is, the resolution conversion module upsamples the resolution of the U component output from the Y auxiliary UV module to obtain a U component with a resolution of H×W, inputs the U component with a resolution of H×W to the YUV signal enhancement module, and the resolution conversion module upsamples the resolution of the V component output from the Y auxiliary UV module to obtain a V component with a resolution of H×W, inputs the V component with a resolution of H×W to the YUV signal enhancement module. Alternatively, the resolution conversion module may be located at position 3, that is, the resolution conversion module upsamples the resolution of the U component output from the YUV signal enhancement module to obtain a U component with a resolution of H × W, and finally outputs this H × W U component. The resolution conversion module also upsamples the resolution of the V component output from the YUV signal enhancement module to obtain a V component with a resolution of H × W, and finally outputs this H × W V component.

[0114] Regarding the resolution conversion module, it can appropriately reduce the resolution of the feature map and effectively reduce the computational complexity of the network. On the other hand, in the JPEG-AI framework, the resolution conversion module upsamples the resolution of the reconstructed UV component to restore a reconstructed U / V signal that has the same resolution as the original signal.

[0115] In this embodiment, as shown in Figure 5B, with respect to the YUV signal enhancement module, in the image coding compression process, in addition to the reduction in resolution of the UV signal, both the compression algorithm and the convolution process cause additional information loss to the YUV signal. Therefore, the YUV signal enhancement module can enhance the Y component, thereby approximating the reconstructed Y component to the original signal before compression and improving subjective quality; the YUV signal enhancement module can enhance the U component, thereby approximating the reconstructed U component to the original signal before compression and improving subjective quality; and the YUV signal enhancement module can enhance the V component, thereby approximating the reconstructed V component to the original signal before compression and improving subjective quality.

[0116] As an example, as shown in Figure 5B, three network branches (each processing the Y, U, and V components respectively) can be integrated into a single network branch. That is, the computational complexity of the network is reduced by processing the Y, U, and V components simultaneously using a single network and outputting the reconstructed Y, U, and V components simultaneously.

[0117] Example 5: In Example 4, the Y auxiliary UV module can utilize information from the Y initial component to supplementarily improve the reconstruction quality of the U initial component and / or the V initial component. For example, it can generate an adjusted component corresponding to the U initial component based on the Y initial component and the U initial component, and / or generate an adjusted component corresponding to the V initial component based on the Y initial component and the V initial component. Here, the reconstruction quality of the adjusted component corresponding to the U initial component is superior to the reconstruction quality of the U initial component and can improve the reconstruction quality of the U initial component. The reconstruction quality of the adjusted component corresponding to the U initial component may be equal to or lower than the reconstruction quality of the U initial component, but is not limited thereto. The reconstruction quality of the adjusted component corresponding to the V initial component is superior to the reconstruction quality of the V initial component and can improve the reconstruction quality of the V initial component. The reconstruction quality of the adjusted component corresponding to the V initial component may be equal to or lower than the reconstruction quality of the V initial component, but is not limited thereto.

[0118] In one possible embodiment, the Y auxiliary UV module may acquire a first image feature corresponding to the initial Y component, and the first image feature may be an image feature that has not been processed by a neural network. For example, the first image feature may include, but is not limited to, at least one of texture features, subjective features, frequency domain features, or histogram features. The Y auxiliary UV module may acquire a second image feature corresponding to the initial U component, and the second image feature may be an image feature that has not been processed by a neural network. For example, the second image feature may include, but is not limited to, at least one of texture features, subjective features, frequency domain features, or histogram features. The Y auxiliary UV module may generate an adjusted component based on the first and second image features. For example, an adjusted component corresponding to the initial U component may be obtained by generating a texture fusion feature based on the texture feature corresponding to the initial Y component and the texture feature corresponding to the initial U component, and then replacing the texture feature corresponding to the initial U component with the texture fusion feature. Alternatively, for example, a subjective fusion feature may be generated based on the subjective feature corresponding to the initial Y component and the subjective feature corresponding to the initial U component, and the adjusted component corresponding to the initial U component may be obtained by substituting the subjective feature corresponding to the initial U component with the subjective fusion feature. Alternatively, for example, a frequency domain fusion feature may be generated based on the frequency domain feature corresponding to the initial Y component and the frequency domain feature corresponding to the initial U component, and the adjusted component corresponding to the initial U component may be obtained by substituting the frequency domain feature corresponding to the initial U component with the frequency domain fusion feature. Alternatively, for example, a histogram fusion feature may be generated based on the histogram feature corresponding to the initial Y component and the histogram feature corresponding to the initial U component, and the adjusted component corresponding to the initial U component may be obtained by substituting the histogram feature corresponding to the initial U component with the histogram fusion feature.Alternatively, for example, a texture fusion feature may be generated based on the texture feature corresponding to the initial Y component and the texture feature corresponding to the initial U component; a frequency domain fusion feature may be generated based on the frequency domain feature corresponding to the initial Y component and the frequency domain feature corresponding to the initial U component; the texture feature corresponding to the initial U component may be replaced with the texture fusion feature; and the frequency domain feature corresponding to the initial U component may be replaced with the frequency domain fusion feature to obtain the adjusted component corresponding to the initial U component. The above are just a few examples and are not limiting.

[0119] The Y Auxiliary UV module may acquire a first image feature corresponding to the Y initial component, and this first image feature may be an image feature that has not been processed by a neural network. For example, the first image feature may include, but is not limited to, at least one of texture features, subjective features, frequency domain features, or histogram features. The Y Auxiliary UV module may acquire a second image feature corresponding to the V initial component, and this second image feature may be an image feature that has not been processed by a neural network. For example, the second image feature may include, but is not limited to, at least one of texture features, subjective features, frequency domain features, or histogram features. The Y Auxiliary UV module may generate an adjusted component corresponding to the V initial component based on the first and second image features. For the generation method, please refer to the processing process for the U initial component.

[0120] In another embodiment, the Y auxiliary UV module may acquire a first image feature corresponding to the initial Y component, and the first image feature may be an image feature acquired by a neural network, i.e., an image feature output by a neural network. For example, the first image feature may include, but is not limited to, a first feature map, a weight coefficient map, etc. The Y auxiliary UV module may acquire a second image feature corresponding to the initial U component, and the second image feature may be an image feature acquired by a neural network, i.e., an image feature output by a neural network. For example, the second image feature may include, but is not limited to, a second feature map, etc. The Y auxiliary UV module may generate an adjusted component corresponding to the initial U component based on the first and second image features. For example, the adjusted component corresponding to the initial U component may be generated based on a first feature map corresponding to the initial Y component and a second feature map corresponding to the initial U component. Alternatively, for example, the adjusted component corresponding to the initial U component may be generated based on a weight coefficient map corresponding to the initial Y component and a second feature map corresponding to the initial U component. The above are just some examples and are not limiting.

[0121] The Y auxiliary UV module may acquire a first image feature corresponding to the initial Y component, and this first image feature may be an image feature acquired by a neural network, i.e., an image feature output by a neural network. For example, the first image feature may include, but is not limited to, a first feature map, a weight coefficient map, etc. The Y auxiliary UV module may acquire a second image feature corresponding to the initial V component, and this second image feature may be an image feature acquired by a neural network, i.e., an image feature output by a neural network. For example, the second image feature may include, but is not limited to, a second feature map, etc. The Y auxiliary UV module may generate an adjusted component corresponding to the initial V component based on the first and second image features. For the generation method, please refer to the processing process for the initial U component.

[0122] Example 6: In order to enable the UV component to make full use of the information from the Y component, a feature map with multiple channels is obtained by performing convolution on the Y component at the input stage of the post-processing network. These feature maps contain various types of features from the Y component. By adding (Add) or concatenating (Concat) these various types of features with the feature map obtained by convolving the UV component, the subsequent neural network is enabled to learn the various features from the Y component, thereby improving the reconstruction quality of the UV component.

[0123] For example, refer to the structural diagram of the post-processing network shown in Figure 6A. The Y auxiliary UV module may, after obtaining the initial Y component, initial U component, and initial V component, input the initial Y component into a first neural network (e.g., CNN 1) to obtain a first feature map, which may be called the Y Feature Map. Alternatively, it may input the initial U component into a second neural network (e.g., CNN 2) to obtain a second feature map corresponding to the initial U component, which may be called the U Feature Map. Similarly, it may input the initial V component into a second neural network (e.g., CNN 2) to obtain a second feature map corresponding to the initial V component, which may be called the V Feature Map.

[0124] After obtaining the first feature map and the second feature map corresponding to the initial U component, an adjusted component corresponding to the initial U component may be generated based on the first and second feature maps. For example, an addition operation may be performed on the first and second feature maps to obtain the adjusted component shown in Figure 6A.

number

number

[0125] Performing a concatenation operation on the first and second feature maps refers to performing a concatenation operation on the first and second feature maps along the channel dimension.

[0126] For example, in order to correctly implement the above operation, the number of convolutional kernels in the second neural network may be the same as the number of convolutional kernels in the first neural network. For example, CNN 1 and CNN 2 have the same number of convolutional kernels. The processing process shown in Figure 6A is:

number

number

number

number

number

number

number

number

[0127] Example 7: Guide filtering can be introduced to utilize the information from each feature map of the Y component. That is, the features generated by the Y component are used to guide the UV component to generate a feature map with higher expressive power. By using an attention mechanism to weight the pixel values ​​in the UV feature map with the Y component's feature map as a guide variable, important features in the UV component are enhanced, general features in the UV component are suppressed, and efficient conversion and fusion of features between the Y and UV components are achieved.

[0128] For example, see the structural diagram of the post-processing network shown in Figure 6B. The Y auxiliary UV module may, after obtaining the initial Y component, initial U component, and initial V component, input the initial Y component into a first neural network (e.g., CNN 1) to obtain a first feature map, which may be called the Y Feature Map. Alternatively, it may input the initial U component into a second neural network (e.g., CNN 2) to obtain a second feature map corresponding to the initial U component, which may be called the U Feature Map. Alternatively, it may input the initial V component into a second neural network (e.g., CNN 2) to obtain a second feature map corresponding to the initial V component, which may be called the V Feature Map.

[0129] After obtaining a first feature map and a second feature map corresponding to the initial U component, guide filtering may be performed on the first feature map to obtain a weight coefficient map, and adjusted components corresponding to the initial U component may be generated based on the weight coefficient map and the second feature map. For example, an output feature map may be obtained by performing at least one operation of multiplication, addition, or concatenation based on the weight coefficient map and the second feature map. For example, an output feature map may be obtained by performing a multiplication operation based on the weight coefficient map and the second feature map, or by performing both multiplication and addition operations based on the weight coefficient map and the second feature map, or by performing all of multiplication, addition, and concatenation operations based on the weight coefficient map and the second feature map. After obtaining the output feature map, the output feature map may be input to a third neural network (e.g., CNN 3) to obtain adjusted components corresponding to the initial U component.

[0130] After obtaining a first feature map and a second feature map corresponding to the initial V components, guide filtering may be performed on the first feature map to obtain a weight coefficient map, and adjusted components corresponding to the initial V components may be generated based on the weight coefficient map and the second feature map. For example, at least one operation of multiplication, addition, or concatenation may be performed based on the weight coefficient map and the second feature map to obtain an output feature map, and after obtaining the output feature map, the output feature map may be input to a third neural network (e.g., CNN 3) to obtain adjusted components corresponding to the initial V components.

[0131] For example, the processing process shown in Figure 6B is:

number

number

number

number

number

number

number

number

[0132] In one embodiment, the guide filtering embodiment is shown in Figure 6C. Specifically, a convolution operation is performed on the first feature map (Y Feature Map) to obtain a convolutional feature map (hereinafter referred to as convolutional feature map A). ​​For example, the first feature map may be input to a CNN network, the CNN network may perform a convolution operation on the first feature map to obtain convolutional feature map A, and then weight mapping may be performed on convolutional feature map A to obtain a weight coefficient map after weight mapping.

[0133] When generating an adjusted component corresponding to the initial U component based on a weight coefficient map and a second feature map corresponding to the initial U component, a multiplication operation may be performed on the weight coefficient map and the second feature map corresponding to the initial U component (U Feature Map) to obtain the multiplied feature map. Then, an addition operation is performed on the multiplied feature map and the convolutional feature map A to obtain the added feature map. Subsequently, a concatenation operation is performed on the added feature map and the convolutional feature map A to obtain the output feature map, and this output feature map is input to a third neural network (e.g., CNN 3) to obtain the adjusted component corresponding to the initial U component.

[0134] When generating adjusted components corresponding to the initial components of V based on a weight coefficient map and a second feature map corresponding to the initial components of V, a multiplication operation may be performed on the weight coefficient map and the second feature map corresponding to the initial components of V (V Feature Map) to obtain the multiplied feature map. Then, an addition operation is performed on the multiplied feature map and the convolutional feature map A to obtain the added feature map. Subsequently, a concatenation operation is performed on the added feature map and the convolutional feature map A to obtain the output feature map, and this output feature map is input to a third neural network (e.g., CNN 3) to obtain adjusted components corresponding to the initial components of V.

[0135] As shown in Figure 6C, after processing the first feature map of the Y component with several convolutional layers, the convolutional feature map A is obtained. Weight mapping is used to map the convolutional feature map A to specific weight coefficients, i.e., a weight coefficient map. The weight coefficient map is multiplied with the second feature map of the UV component to obtain a UV component feature guide map. Then, the convolutional feature map A and the UV component feature guide map are added together, and finally, the convolutional feature map A is concatenated with the guide map to output a fused feature map (Output Feature Map). For example, the guide filtering process is:

number

number

number

number

number

[0136] In one possible embodiment, during the weight mapping process, weight mapping is performed on the convolutional feature map A to obtain a weight coefficient map after weight mapping. An embodiment of weight mapping is shown in Figure 6D, which shows two types of weight mapping methods. In the first type of weight mapping method, a pooling operation is performed on the convolutional feature map A to obtain a feature map after the pooling operation. For example, a pooling operation is performed on a convolutional feature map A with size C×H×W to obtain a feature map after the pooling operation with size C×1×1. That is, the resolution of the feature map is reduced to 1×1 by the "Pooling" method.

[0137] Next, a fully connected operation and a ReLU activation operation are performed on the feature map after the pooling operation, and the feature map after the ReLU activation operation is obtained. For example, a fully connected operation (i.e., FC operation) and a ReLU activation operation (i.e., ReLU operation) are performed on a feature map after the pooling operation with a size of C×1×1, and the feature map after the ReLU activation operation with a size of C×1×1 is obtained. A fully connected operation and a sigmoid activation operation are performed on the feature map after the ReLU activation operation, and the feature map after the sigmoid activation operation is obtained. For example, a fully connected operation (i.e., FC operation) and a sigmoid activation operation (i.e., sigmoid operation) are performed on a feature map after the ReLU activation operation with a size of C×1×1, and the feature map after the sigmoid activation operation with a size of C×1×1 is obtained.

[0138] Then, a weight coefficient map is generated based on the convolutional feature map A and the feature map after the sigmoid activation operation. For example, the weight coefficient map is obtained by multiplying the convolutional feature map A and the feature map after the sigmoid activation operation. For example, a multiplication operation is performed on the convolutional feature map A, which has a size of C×H×W, and the feature map after the sigmoid activation operation, which has a size of C×1×1, to obtain a weight coefficient map of size C×H×W. Clearly, after reducing the resolution of the feature map to 1×1 using the "Pooling" method, it passes through several fully connected layers, and finally, the weight coefficients of each channel in the original input feature map are obtained using the sigmoid activation function. These weight coefficients are then used to assign different weights to the channels of the original input feature map, and finally, a weight coefficient map (also called a weight feature map) is obtained.

[0139] In the second type of weight mapping scheme, a convolution operation may be performed on the convolutional feature map A to obtain the convolutional feature map B. For example, a convolutional feature map A with size C×H×W is input to a CNN network, and the CNN network performs a convolution operation on the convolutional feature map A through several convolutional layers to obtain a convolutional feature map B with size C×H×W. Then, a sigmoid activation operation is performed on the convolutional feature map B to obtain a weight coefficient map. For example, an activation operation is performed on the convolutional feature map B with size C×H×W using a sigmoid activation function to obtain a weight coefficient map with size C×H×W. Clearly, the original input feature map is processed through several convolutional layers, and then the weight coefficient map is finally obtained using a sigmoid activation function.

[0140] The difference between the first type of weight mapping method and the second type of weight mapping method is that the first type assigns different weights to the channels, while the second type assigns different weights to each pixel of each channel feature map.

[0141] Example 8: In Examples 6 and 7, Y component feature extraction and information fusion were performed at the input stage of the post-processing network. However, in Example 8, unlike these, Y component feature extraction and information fusion are distributed to each stage of the post-processing network, and the UV component features and Y component features from different stages are fused, thereby improving the UV component enhancement capability of the post-processing network. For example, please refer to the structural diagram of the post-processing network shown in Figure 6E. The Y auxiliary UV module acquires the initial Y component, initial U component, and initial V component, and then distributes Y component feature extraction and information fusion to each stage of the post-processing network. Here,

number

[0142] In one possible embodiment, multiple first feature maps and multiple second feature maps are involved in fusing the Y component feature map with the UV component feature maps at each stage of the network. For example, K first feature maps and K+1 second feature maps are used, where K is a positive integer greater than 1. In this case, after obtaining the initial Y and U components, the initial Y component is input into the neural network to obtain the first first feature map, and the initial U component is input into the neural network to obtain the first second feature map. For the i-th first feature map and the i-th second feature map, the value of i ranges from 2 to K. The (i-1)th first feature map is input into the neural network to obtain the i-th first feature map, feature fusion is performed on the (i-1)th first feature map and the (i-1)th second feature map to obtain the fused features, and the fused features are input into the neural network to obtain the i-th second feature map. After obtaining the last second feature map (i.e., the K+1th second feature map), the adjusted component corresponding to the initial U component may be generated based on the last second feature map.

[0143] For example, let's take the case where K is 3. The initial Y component is input to neural network a1 to obtain the first feature map b1. That is, neural network a1 performs several convolution operations on the initial Y component to obtain the first feature map b1. The initial U component is input to neural network c1 to obtain the second feature map d1. That is, neural network c1 performs several convolution operations on the initial U component to obtain the second feature map d1.

[0144] The first feature map b1 is input to the neural network a2 to obtain the first feature map b2; that is, the neural network a2 performs several convolution operations on the first feature map b1 to obtain the first feature map b2. Feature fusion is performed on the first feature map b1 and the second feature map d1 to obtain the fused features, and these fused features are input to the neural network c2 to obtain the second feature map d2.

[0145] The first feature map b2 is input to the neural network a3 to obtain the first feature map b3; that is, the neural network a3 performs several convolution operations on the first feature map b2 to obtain the first feature map b3. Feature fusion is performed on the first feature map b2 and the second feature map d2 to obtain the fused features, and these fused features are input to the neural network c3 to obtain the second feature map d3.

[0146] Then, feature fusion is performed on the first feature map b3 and the second feature map d3 to obtain the fused features, and these fused features are input into the neural network c4 to obtain the second feature map d4. After obtaining the second feature map d4, an adjusted component corresponding to the initial U component may be generated based on the second feature map d4. For example, the second feature map d4 may be the adjusted component corresponding to the initial U component, or an adjusted component corresponding to the initial U component may be obtained after performing operations on the second feature map d4.

[0147] In the process described above, obtaining the fused features by performing feature fusion on the first feature map and the second feature map includes, but is not limited to, performing an addition operation on the first feature map and the second feature map to obtain the fused features, or performing a concatenation operation on the first feature map and the second feature map to obtain the fused features, or performing guide filtering on the first feature map to obtain a weight coefficient map corresponding to the first feature map, and generating the fused features based on the weight coefficient map and the second feature map. For details on the implementation method of feature fusion, please refer to Examples 6 and 7, which will not be explained again here.

[0148] After obtaining the initial Y and V components, for the first first feature map and the first second feature map, the initial Y component is input into the neural network to obtain the first first feature map, and the initial V component is input into the neural network to obtain the first second feature map. For the i-th first feature map and the i-th second feature map, the range of i is from 2 to K. The (i-1)th first feature map is input into the neural network to obtain the i-th first feature map, feature fusion is performed on the (i-1)th first feature map and the (i-1)th second feature map to obtain the fused features, and the fused features are input into the neural network to obtain the i-th second feature map. After obtaining the last second feature map, an adjusted component corresponding to the initial V component may be generated based on the last second feature map. For the method of obtaining the adjusted component corresponding to the initial V component, please refer to the processing process for the initial U component.

[0149] Example 9: In Examples 6, 7, and 8, the fusion method for the Y component guide map was designed. However, in Example 9, unlike these, detailed feature extraction of the Y component is performed from multiple dimensions, and then fused into the post-processing network, thereby improving the network's ability to enhance the UV component.

[0150] Please refer to the structural diagram of the post-processing network shown in Figure 6F. The Y Auxiliary UV module obtains the Y initial component, U initial component, and V initial component, and then first performs multidimensional feature extraction on the Y initial component. For example, it extracts M-dimensional image features (where M is a positive integer) from the Y initial component, and then concatenates these M-dimensional image features to obtain the concatenated multidimensional features. After obtaining the multidimensional features, the multidimensional features may be input into the first neural network to obtain the first feature map (Y Feature Map). Alternatively, the Y Auxiliary UV module may input the U initial component into the second neural network (e.g., CNN 2) to obtain a second feature map corresponding to the U initial component, and this second feature map may be called the U Feature Map. The V initial component may also be input into the second neural network (e.g., CNN 2) to obtain a second feature map corresponding to the V initial component, and this second feature map may be called the V Feature Map.

[0151] After obtaining the first feature map and the second feature map corresponding to the initial U component, feature fusion may be performed on the first and second feature maps to obtain the adjusted component corresponding to the initial U component. Alternatively, after obtaining the first feature map and the second feature map corresponding to the initial V component, feature fusion may be performed on the first and second feature maps to obtain the adjusted component corresponding to the initial V component. For details on this feature fusion process, please refer to Examples 6, 7, and 8.

[0152] For example, the processing process shown in Figure 6F is:

number

number

number

[0153] For example, please refer to the structural diagram of the multidimensional feature extraction process shown in Figure 6G. The initial Y component is taken as input, and each branch feature extraction structure is used to extract features of the initial Y component. For example, the initial Y component is input to the branch 1 feature extraction structure, and the first-dimensional features of the initial Y component are extracted from the branch 1 feature extraction structure. The initial Y component is then input to the branch 2 feature extraction structure, and the second-dimensional features of the initial Y component are extracted from the branch 2 feature extraction structure. In this way, the initial Y component is input to the branch M feature extraction structure, and the M-th dimension features of the initial Y component are extracted from the branch M feature extraction structure. Through this process, M-dimensional image features can be extracted from the initial Y component, and these M-dimensional image features can be concatenated to obtain a concatenated multidimensional feature. After obtaining the concatenated multidimensional feature, the multidimensional feature may be used as the first feature map, or the multidimensional feature may be input to the first neural network to obtain the first feature map.

[0154] Figure 6G shows the structure of the multidimensional feature extraction process. Specifically, first, the Y component is passed through different feature extraction channels to obtain feature information for each dimension, and finally, this feature information is concatenated and output as a Y component feature map. The multidimensional feature extraction process is as follows:

number

number

number

[0155] In one embodiment, the M-dimensional image features include, but are not limited to, at least one of the following: an image feature obtained by convolving the initial Y component using one convolution kernel; an image feature obtained by cascading the initial Y component using two convolution kernels; a first-order spatial domain feature obtained by processing the initial Y component using the Sobel operator; a second-order spatial domain feature obtained by processing the initial Y component using the Laplacian operator; and a frequency domain feature obtained by performing a Fourier transform on the initial Y component. The above image features are merely examples, and the type of image feature is not limited in this embodiment.

[0156] Figure 6H shows the structure diagram of the multidimensional feature extraction process. Branch 1 feature extraction structure (i.e., the first branch) acquires feature information using a single 3x3 convolution. For example, the initial Y component is convolved using a single 3x3 convolution kernel, and the image features after convolution are obtained. Branch 2 feature extraction structure (i.e., the second branch) acquires feature information after channel expansion and compression by employing a 1x1 and 3x3 cascaded convolution method. For example, a cascaded convolution is performed on the initial Y component using a 1x1 and 3x3 convolution kernel, and the image features after cascaded convolution are obtained. Branch 3 feature extraction structure (i.e., the third branch) acquires first-order spatial domain features using the Sobel operator. For example, the initial Y component is processed using the Sobel operator, and first-order spatial domain features are obtained. Branch 4 feature extraction structure (i.e., the fourth branch) acquires second-order spatial domain features using the Laplacian operator. For example, the initial Y component is processed using the Laplacian operator, and second-order spatial domain features are obtained. The branch 5 feature extraction structure (i.e., the 5th branch) uses the Fourier transform to obtain frequency-domain features. For example, the Fourier transform is performed on the initial Y component to obtain frequency-domain features. After obtaining the features of the five dimensions described above, these features can be concatenated to obtain multi-dimensional features.

[0157] Example 10: While Examples 6 through 9 mainly focused on "which Y component features to select and how to fuse the Y component features with the UV component features," Example 10 differs in that its core focus is on "how to enhance the features of the Y component." For example, since a certain amount of loss occurs in the image details of the Y component after the YUV components of an image have been encoded and compressed, a certain amount of enhancement may be performed on the Y component and its features first. For example, since there is a certain correlation between the three components of YUV, the compression loss of the Y component is smaller than the compression loss of the UV component on the encoding side, but a certain amount of enhancement may be performed on the Y component in order to better enhance the UV component. Therefore, preprocessing may be performed on the initial Y component using a preprocessing module, and then feature enhancement processing may be performed on the initial Y component.

[0158] For example, the preprocessing module may first obtain the initial Y component, then perform edge enhancement on the initial Y component to obtain the image features after edge enhancement, and then perform multiscale feature extraction on the image features after edge enhancement to obtain multiscale features. After obtaining the multiscale features, the multiscale features may be used as the preprocessed initial Y component, and the preprocessed initial Y component may be input to the Y auxiliary UV module, which in turn inputs the preprocessed initial Y component to the first neural network to obtain the first feature map (Y Feature Map). Alternatively, the Y auxiliary UV module may input the initial U component to a second neural network (e.g., CNN 2) to obtain a second feature map corresponding to the initial U component, and this second feature map may be called the U Feature Map. The initial V component may also be input to a second neural network (e.g., CNN 2) to obtain a second feature map corresponding to the initial V component, and this second feature map may be called the V Feature Map.

[0159] After obtaining the first feature map and the second feature map corresponding to the initial U component, feature fusion may be performed on the first and second feature maps to obtain the adjusted component corresponding to the initial U component. Alternatively, after obtaining the first feature map and the second feature map corresponding to the initial V component, feature fusion may be performed on the first and second feature maps to obtain the adjusted component corresponding to the initial V component. For details on this feature fusion process, please refer to Examples 6, 7, and 8.

[0160] The preprocessing module performs multiscale feature extraction on edge-enhanced image features to obtain multiscale features. This involves performing a convolution operation on the edge-enhanced image features to obtain the convolutional features, performing a downsampling operation on the convolutional features to obtain the downsampled features, performing a channel conversion on the downsampled features to obtain the channel-converted features, performing an upsampling operation on the channel-converted features to obtain the upsampled features, and then generating multiscale features based on the upsampled features and the channel-converted features.

[0161] Please refer to the structural diagram of the post-processing network shown in Figure 6I. Figure 6I shows an example of a multiscale enhancement method for the Y component. In the downsampling portion, as you proceed from D1 to D3, the resolution of the feature map gradually decreases and the number of channels gradually increases. In the upsampling portion, as you proceed from U3 to U1, the resolution of the feature map gradually increases and the number of channels remains constant. After channel conversion, D3, D2, and D1 are added to the upsampled U3, U2, and U1 feature maps, respectively, to obtain the multiscale features after multiscale enhancement.

[0162] First, edge enhancement is performed on the initial Y component, and the image features after edge enhancement are obtained. For example, image sharpening enhancement is performed on the initial Y component, and the image features after edge enhancement are obtained. The edge enhancement method is not limited to this.

[0163] We will use two downsampling operations and two upsampling operations as an example, but the number of downsampling and upsampling operations can be increased and is not limited to this. After obtaining the edge-enhanced image features, the edge-enhanced image features are output to a CNN to obtain feature D1 with a scale of C × H × W, a downsampling operation is performed on feature D1 to obtain feature D2 with a scale of 4C × (H / 2) × (W / 2), a downsampling operation is performed on feature D2 to obtain feature D3 with a scale of 16C × H / 4 × W / 4.

[0164] Feature D3 is subjected to channel transformation to obtain feature U3 with a scale of T×H / 4×W / 4. Feature U3 is upsampled, feature D2 is subjected to channel transformation, and the upsampled feature U3 and the channel-transformed feature D2 are added together to obtain feature U2 with a scale of T×(H / 2)×(W / 2). Feature U2 is upsampled, feature D1 is subjected to channel transformation, and the upsampled feature U2 and the channel-transformed feature D1 are added together to obtain feature U1 with a scale of T×H×W. Clearly, feature U1 is a multiscale feature after multiscale enhancement, i.e., a preprocessed initial Y component, and the preprocessing module outputs this preprocessed initial Y component. By performing multiscale feature extraction on the Y component, the final output first feature map integrates feature information from different scales and contributes to the enhancement of the UV component.

[0165] Example 11: In Example 4, with respect to the resolution conversion module, the resolution conversion module upsamples the resolution of the U component to restore the U component having the resolution of the original signal, for example, obtaining a U component with a resolution of H × W. The resolution conversion module upsamples the resolution of the initial V component to restore the V component having the resolution of the original signal, for example, obtaining a V component with a resolution of H × W. For example, in the JPEG-AI framework, the original UV component is downsampled by 2x and then compressed and encoded on the encoding side, so the resolution of the UV component decoded and reconstructed on the decoding side is half the resolution of the original signal. In the post-processing network, there is a step to upsample the UV component so that a reconstructed UV component with the same resolution as the original signal is restored. Exemplary methods by which the resolution conversion module upsamples the resolution of the U component and / or V component include, but are not limited to, interpolation sampling, pixel-shuffle upsampling, and deconvolution (DeConv). The above are just some examples, and this embodiment does not limit the method of upsampling.

[0166] For example, upsampling the resolution of the UV component affects the computational complexity of the network model; that is, increasing the resolution increases the computational complexity. Since the position of the resolution conversion module also affects the computational complexity of the network model, its position can be designed to reduce the computational complexity of the network model.

[0167] Case 1: As shown in Figure 5B, when the resolution conversion module is positioned at position 1, the resolution conversion module upsamples the resolution of the initial U component, obtains a U component with a resolution of H × W, and inputs this U component with a resolution of H × W to the Y auxiliary UV module. The resolution conversion module then upsamples the resolution of the initial V component, obtains a V component with a resolution of H × W, and inputs this V component with a resolution of H × W to the Y auxiliary UV module. In this case, for Examples 5 to 10, the initial U component is the upsampled U component, the initial V component is the upsampled V component, the resolution of the upsampled U component is equal to the resolution of the initial Y component, and the resolution of the upsampled V component is equal to the resolution of the initial Y component. Thus, Examples 5 to 10 are executed based on the initial Y component and the upsampled U component, and Examples 5 to 10 are executed based on the initial Y component and the upsampled V component.

[0168] As can be seen from the above, in Case 1, upsampling can be performed at the input terminal of the network, and as shown in Figure 7A, this is a structural diagram of upsampling at the input terminal of the network. In other words, at the input terminal of the network, the initial U component can be upsampled to obtain the upsampled U component, for example, the U component with a resolution of H × W, and at the input terminal of the network, the initial V component can be upsampled to obtain the upsampled V component, for example, the V component with a resolution of H × W.

[0169] Case 2: As shown in Figure 5B, when the resolution conversion module is positioned at position 2, the resolution conversion module upsamples the resolution of the U component output from the Y auxiliary UV module, obtains a U component with a resolution of H × W, and inputs this U component with a resolution of H × W to the YUV signal enhancement module. The resolution conversion module also upsamples the resolution of the V component output from the Y auxiliary UV module, obtains a V component with a resolution of H × W, and inputs this V component with a resolution of H × W to the YUV signal enhancement module. In this case, for Examples 5 to 10, the initial U component is the U component before upsampling, and the initial V component is the V component before upsampling. The resolution of the U component before upsampling is smaller than the resolution of the initial Y component, and the resolution of the V component before upsampling is smaller than the resolution of the initial Y component. Thus, Examples 5 to 10 are executed based on the initial Y component and the U component before upsampling, and Examples 5 to 10 are executed based on the initial Y component and the V component before upsampling. After obtaining the adjusted components corresponding to the U / V components based on Examples 5 to 10, the adjusted components may be upsampled to obtain the adjusted components after upsampling, and the adjusted components after upsampling may be input to the YUV signal enhancement module, and the resolution of the adjusted components after upsampling will be equal to the resolution of the initial Y component.

[0170] As can be seen from the above, in Case 2, stepwise upsampling is possible between network layers, and as shown in Figure 7B, it is a structural diagram of stepwise upsampling between network layers. In other words, in multiple network layers, the initial U component can be upsampled to obtain the U component after upsampling, and the resolution of the U component after upsampling is equal to the resolution of the initial Y component, for example, the U component with a resolution of H × W. In multiple network layers, the initial V component can be upsampled to obtain the V component after upsampling, and the resolution of the V component after upsampling is equal to the resolution of the initial Y component, for example, the V component with a resolution of H × W.

[0171] Case 3: As shown in Figure 5B, when the resolution conversion module is positioned at position 3, the resolution conversion module upsamples the resolution of the U component output from the YUV signal enhancement module to obtain a U component with a resolution of H × W, and finally outputs this H × W U component. The resolution conversion module also upsamples the resolution of the V component output from the YUV signal enhancement module to obtain a V component with a resolution of H × W, and finally outputs this H × W V component. In this case, for Examples 5 to 10, the initial U component is the U component before upsampling, and the initial V component is the V component before upsampling. The resolution of the U component before upsampling is smaller than the resolution of the initial Y component, and the resolution of the V component before upsampling is smaller than the resolution of the initial Y component. Thus, Examples 5 to 10 are executed based on the initial Y component and the U component before upsampling, and Examples 5 to 10 are executed based on the initial Y component and the V component before upsampling. Based on Examples 5 to 10, after obtaining the adjusted components corresponding to the U / V components, the adjusted components are input to the YUV signal enhancement module, and the adjusted components are the U / V components before upsampling, and the resolution of the U / V components before upsampling is smaller than that of the initial Y component. In this way, the YUV signal enhancement module enhances the signal based on the U / V components before upsampling and obtains the U / V components after signal enhancement. After obtaining the U / V components after signal enhancement, the U / V components are upsampled to obtain the U / V components after upsampling, and the resolution of the U / V components after upsampling is equal to the resolution of the initial Y component.

[0172] As can be seen from the above, in Case 3, upsampling can be performed at the network output terminal, and as shown in Figure 7C, this is a structural diagram of upsampling at the network output terminal. In other words, the U component can be upsampled at the network output terminal to obtain the upsampled U component, and the resolution of the upsampled U component is equal to the resolution of the initial Y component, for example, the U component with a resolution of H × W. The V component can be upsampled at the network output terminal to obtain the upsampled V component, and the resolution of the upsampled V component is equal to the resolution of the initial Y component, for example, the V component with a resolution of H × W.

[0173] Example 12: In Example 4, with respect to the resolution conversion module, in the first embodiment, the resolution conversion module can appropriately reduce the resolution of the feature map by using lossless resolution conversion, thereby effectively reducing the computational complexity of the network. In the second embodiment, the resolution conversion module can restore the U and V components to have the same resolution as the original signal by upsampling the resolution of the reconstructed U and V components. For the functionality of the second embodiment, please refer to Example 11.

[0174] The functionality of the first embodiment may be implemented by a resolution conversion module or by a Y auxiliary UV module. For example, to restore the original resolution, a wavelet transform (e.g., a Haar wavelet transform) can be used to reduce the resolution of the original image by half, thus reducing the computational complexity of the network by approximately one-quarter. If the hardware requirements for computational complexity are high, the resolution can be further reduced by using multiple successive wavelet transforms. For some frequency bands after the wavelet transform, all frequency bands may be used, or some frequency bands may be selected.

[0175] For example, refer to the structural diagram of the post-processing network shown in Figure 7D. A wavelet transform (DWT transform in Figure 7D) may be performed on the initial Y component to obtain multiple frequency bands after the wavelet transform, and these multiple frequency bands or a portion of the frequency bands (i.e., the Y subband) may be input into the first neural network to obtain the first feature map (Y Feature Map). A wavelet transform (DWT transform in Figure 7D) may be performed on the initial U component to obtain multiple frequency bands after the wavelet transform, and these multiple frequency bands or a portion of the frequency bands (i.e., the U subband) may be input into the second neural network to obtain the second feature map (U Feature Map). A wavelet transform may be performed on the initial V component to obtain multiple frequency bands after the wavelet transform, and these multiple frequency bands or a portion of the frequency bands (i.e., the V subband) may be input into the second neural network to obtain the second feature map (V Feature Map). After obtaining the first feature map and the second feature map corresponding to the initial U component, feature fusion may be performed on the first and second feature maps to obtain the adjusted component corresponding to the initial U component. Alternatively, after obtaining the first feature map and the second feature map corresponding to the initial V component, feature fusion may be performed on the first and second feature maps to obtain the adjusted component corresponding to the initial V component. For details on this feature fusion process, please refer to Examples 6, 7, and 8.

[0176] When performing feature fusion on the first and second feature maps to obtain the adjusted component corresponding to the initial U component, the fused features may first be obtained using Examples 6, 7, and 8, then the fused features may be input into a neural network (CNN 3) to obtain the output features (U subband) of the neural network, and then the inverse wavelet transform (i.e., the inverse operation of the wavelet transform, IDWT transform in Figure 7D) may be performed on these output features to obtain the adjusted component of the U component. When performing feature fusion on the first and second feature maps to obtain the adjusted component corresponding to the initial V component, the fused features may first be obtained using Examples 6, 7, and 8, then the fused features may be input into a neural network (CNN 3) to obtain the output features (V subband) of the neural network, and then the inverse wavelet transform (i.e., the inverse operation of the wavelet transform, IDWT transform in Figure 7D) may be performed on these output features to obtain the adjusted component of the V component.

[0177] Example 13: In Example 4, with respect to the YUV signal enhancement module, the YUV signal enhancement module performs signal enhancement on the Y component, approximating the reconstructed Y component to the original signal before compression, the YUV signal enhancement module performs signal enhancement on the U component, approximating the reconstructed U component to the original signal before compression, and the YUV signal enhancement module performs signal enhancement on the V component, approximating the reconstructed V component to the original signal before compression. Based on this, the YUV signal enhancement module may perform feature enhancement processing on the initial Y component to obtain a target component corresponding to the Y component, and the target component is the Y component after signal restoration. The YUV signal enhancement module may perform feature enhancement processing on the adjusted U component to obtain a target component corresponding to the U component, and the target component is the U component after signal restoration. The YUV signal enhancement module may perform feature enhancement processing on the adjusted V component to obtain a target component corresponding to the V component, and the target component is the V component after signal restoration.

[0178] In one embodiment, feature enhancement processing may be performed on the initial Y component using at least one residual block network to obtain a target component corresponding to the Y component. Feature enhancement processing may be performed on the adjusted U component using at least one residual block network to obtain a target component corresponding to the U component. Feature enhancement processing may be performed on the adjusted V component using at least one residual block network to obtain a target component corresponding to the V component. Figure 7E is a structural diagram of a residual block cascade network used to enhance the Y, U, and V components, and the structure of this residual block cascade network is not limited.

[0179] In one embodiment, a U-Net network may be used to perform feature enhancement on the initial Y component to obtain a target component corresponding to the Y component. Alternatively, a U-Net network may be used to perform feature enhancement on the adjusted U component to obtain a target component corresponding to the U component. Furthermore, a U-Net network may be used to perform feature enhancement on the adjusted V component to obtain a target component corresponding to the V component. Figure 7F is a structural diagram showing the enhancement of the Y, U, and V components using a U-Net network. The U-Net network includes multiple downsampling network layers and multiple upsampling network layers, and the structure of this U-Net network is not limited.

[0180] Exemplary examples, each of the above embodiments may be implemented individually or in combination. For example, each embodiment in Examples 1 to 13 may be implemented individually, or at least two embodiments in Examples 1 to 13 may be implemented in combination.

[0181] For example, in the above-described embodiment, the contents of the encoding side are applicable to the decoding side, that is, the decoding side can process in the same way. The contents of the decoding side are applicable to the encoding side, that is, the encoding side can process in the same way.

[0182] Based on a similar patent application concept as described above, embodiments of the present invention further provide a decoding device applied to the decoding side, the decoding device may include a memory configured to store video data and a decoder configured to implement the decoding method in Embodiments 1 to 13, i.e., the processing flow on the decoding side.

[0183] For example, in one embodiment, the decoder is The bitstream corresponding to the current image block is decoded to obtain a reconstructed image block, the reconstructed image block includes a first initial component and a second initial component, the resolution of the first initial component is greater than or equal to the resolution of the second initial component, Based on the first initial component and the second initial component, a modified component corresponding to the second initial component is generated. The system is configured to perform feature enhancement processing on the adjusted component to obtain the restored target component corresponding to the second initial component.

[0184] For example, if the reconstructed image block is in YUV format, the first initial component is the luminance component, and the second initial component is the chromaticity U component and / or the chromaticity V component. Alternatively, if the reconstructed image block is in RGB format, the first initial component is the G component, and the second initial component is the R component and / or the B component.

[0185] Exemplary, the decoder is configured to further acquire a first image feature including image features that have not been processed by the neural network and / or image features acquired by the neural network, corresponding to the first initial component; acquire a second image feature including image features that have not been processed by the neural network and / or image features acquired by the neural network, corresponding to the second initial component; and generate the adjusted component based on the first and second image features.

[0186] For example, if the decoder includes a first feature map obtained by a neural network and a second feature map obtained by a neural network, the decoder is configured to input the first initial component into the first neural network to obtain the first feature map, input the second initial component into the second neural network to obtain the second feature map, perform an addition operation on the first and second feature maps to obtain the adjusted component, or perform a concatenation operation on the first and second feature maps to obtain the adjusted component.

[0187] Exemplary, if the decoder further includes a weight coefficient map obtained by a neural network for the first image feature and a second feature map obtained by a neural network for the second image feature, it is configured to input the first initial component into a first neural network to obtain a first feature map corresponding to the first initial component, perform guided filtering on the first feature map to obtain the weight coefficient map, input the second initial component into a second neural network to obtain the second feature map, and generate the adjusted component based on the weight coefficient map and the second feature map.

[0188] For example, the decoder is configured to further perform a convolution operation on the first feature map to obtain a convolved feature map, perform a weight mapping on the convolved feature map to obtain the weight coefficient map.

[0189] For example, the decoder is configured to perform a pooling operation on the convolutional feature map, a fully connected operation and a ReLU activation operation on the pooled feature map, a fully connected operation and a Sigmoid activation operation on the ReLU activation feature map, and generate the weight coefficient map based on the convolutional feature map and the Sigmoid activation feature map.

[0190] For example, the decoder is configured to perform a convolution operation and a sigmoid activation operation on the convolutional feature map to obtain the weight coefficient map.

[0191] Exemplary, the decoder is configured to further perform a multiplication operation on the weight coefficient map and the second feature map to obtain the multiplied feature map, perform a convolution operation on the first feature map to obtain the convolved feature map, perform an addition operation on the multiplied feature map and the convolved feature map to obtain the added feature map, and perform a concatenation operation on the added feature map and the convolved feature map to obtain the adjusted component.

[0192] Exemplary, the decoder is configured such that, if the first image feature includes K first feature maps obtained by the neural network and the second image feature includes K+1 second feature maps obtained by the neural network, then for the first first feature map and the first second feature map, the decoder inputs a first initial component into the neural network to obtain the first first feature map, inputs a second initial component into the neural network to obtain the first second feature map, for the i-th first feature map and the i-th second feature map, i is any integer between 2 and K, and K is a positive integer greater than 1, the decoder inputs the (i-1)th first feature map into the neural network to obtain the i-th first feature map, performs feature fusion on the (i-1)th first feature map and the (i-1)th second feature map to obtain the fused feature, inputs the fused feature into the neural network to obtain the i-th second feature map, obtains the last second feature map, and then generates the adjusted component based on the last second feature map.

[0193] Exemplary, the decoder is configured to further extract M-dimensional (where M is a positive integer) image features from the first initial component, concatenate the M-dimensional image features to obtain a concatenated multidimensional feature, and input the multidimensional feature into a first neural network to obtain the first feature map.

[0194] For example, the decoder is configured to perform a wavelet transform on the first initial component, obtain multiple frequency bands after the wavelet transform, input the multiple frequency bands or a portion of the multiple frequency bands into a first neural network to obtain the first feature map, perform a wavelet transform on the second initial component, obtain multiple frequency bands after the wavelet transform, input the multiple frequency bands or a portion of the multiple frequency bands into a second neural network to obtain the second feature map.

[0195] Exemplary, the decoder is configured to preprocess the first initial component and obtain the preprocessed first initial component, which is then input to the first neural network to obtain the first feature map. The decoder is configured to perform edge enhancement on the first initial component, obtain the edge-enhanced image features, perform multiscale feature extraction on the edge-enhanced image features, obtain multiscale features, and determine the preprocessed first initial component based on the multiscale features.

[0196] For example, the decoder is configured to further perform a convolution operation on the edge-enhanced image features to obtain the convolutional features, perform a downsampling operation on the convolutional features to obtain the downsampled features, perform a channel conversion on the downsampled features to obtain the channel conversion features, perform an upsampling operation on the channel conversion features to obtain the upsampled features, and generate the multiscale features based on the upsampled features and the channel conversion features.

[0197] For example, the decoder is configured to upsample the second initial component before generating a modified component corresponding to the second initial component based on the first and second initial components, thereby obtaining the upsampled second initial component. Here, the resolution of the upsampled second initial component is equal to the resolution of the first initial component.

[0198] Exemplary, the decoder is configured to further perform feature enhancement on the adjusted component and upsample the adjusted component before obtaining the restored target component corresponding to the second initial component, thereby obtaining the upsampled adjusted component. Here, the resolution of the upsampled adjusted component is equal to the resolution of the first initial component.

[0199] Exemplary, the decoder is configured to further perform feature enhancement on the adjusted component to obtain a restored target component corresponding to the second initial component, and then upsample the target component to obtain an upsampled target component. Here, the resolution of the upsampled target component is equal to the resolution of the first initial component.

[0200] Exemplary, the decoder is further configured to perform feature enhancement on the adjusted component using at least one residual block network to obtain a target component corresponding to the second initial component, or to perform feature enhancement on the adjusted component using a U-Net network to obtain a target component corresponding to the second initial component.

[0201] Based on a similar patent application concept as described above, the decoding device (also called a video decoder) according to an embodiment of the present invention, from a hardware perspective, can be specifically described in Figure 8 for the structure of its hardware architecture. The decoding device comprises a processor 811 and a machine-readable storage medium 812, the machine-readable storage medium 812 stores machine-executable instructions that can be executed by the processor 811, and the processor 811 is configured to carry out the methods of Embodiments 1 to 12 of the present invention by executing the machine-executable instructions.

[0202] For example, in one embodiment, the processor 811 executes machine-executable instructions, The bitstream corresponding to the current image block is decoded to obtain a reconstructed image block, the reconstructed image block includes a first initial component and a second initial component, the resolution of the first initial component is greater than or equal to the resolution of the second initial component, Based on the first initial component and the second initial component, a modified component corresponding to the second initial component is generated. The adjusted component is subjected to feature enhancement processing to obtain the restored target component corresponding to the second initial component.

[0203] Based on a patent application concept similar to the above method, embodiments of the present invention provide an electronic device. The electronic device comprises a processor and a device-readable storage medium, the device-readable storage medium storing device-executable instructions that can be executed by the processor, and the processor is configured to perform the decoding methods of embodiments 1 to 12 of the present invention by executing the device-executable instructions.

[0204] Based on a similar patent application concept as described above, embodiments of the present invention further provide a machine-readable storage medium in which several computer instructions are stored, and when the computer instructions are executed by a processor, the methods disclosed in the above examples of the present invention, such as the decoding methods in each of the above embodiments, are carried out.

[0205] Based on a similar patent application concept as described above, embodiments of the present invention further provide a computer program which, when executed by a processor, performs the decoding method disclosed in the above examples of the present invention.

[0206] Based on a similar patent application concept as described above, embodiments of the present invention further provide a decoding device applied to the decoding side, wherein the decoding device is A decoding module configured to decode a bitstream corresponding to the current image block to obtain a reconstructed image block, wherein the reconstructed image block includes a first initial component and a second initial component, and the resolution of the first initial component is greater than or equal to the resolution of the second initial component. A decision module configured to generate an adjusted component corresponding to the second initial component based on the first initial component and the second initial component, The system includes a processing module configured to perform feature enhancement processing on the adjusted component to obtain a restored target component corresponding to the second initial component.

[0207] For example, if the reconstructed image block is in YUV format, the first initial component is the luminance component, and the second initial component is the chromaticity U component and / or the chromaticity V component. Alternatively, if the reconstructed image block is in RGB format, the first initial component is the G component, and the second initial component is the R component and / or the B component.

[0208] Exemplary, the decision module is configured to generate an adjusted component corresponding to the second initial component based on the first initial component and the second initial component, by specifically acquiring a first image feature corresponding to the first initial component, which includes image features that have not been processed by a neural network and / or image features acquired by a neural network, and acquiring a second image feature corresponding to the second initial component, which includes image features that have not been processed by a neural network and / or image features acquired by a neural network, and then generating the adjusted component based on the first and second image features.

[0209] For example, if the first image feature includes an image feature that has not been processed by a neural network, the first image feature includes at least one of a texture feature, a subjective feature, a frequency domain feature, or a histogram feature, and if the second image feature includes an image feature that has not been processed by a neural network, the second image feature includes at least one of a texture feature, a subjective feature, a frequency domain feature, or a histogram feature.

[0210] For example, if the first image feature includes a first feature map obtained by a neural network, and the second image feature includes a second feature map obtained by a neural network, the decision module is configured to, when generating an adjusted component corresponding to the second initial component based on the first and second initial components, to input the first initial component into a first neural network to obtain the first feature map, input the second initial component into a second neural network to obtain the second feature map, perform an addition operation on the first and second feature maps to obtain the adjusted component, or perform a concatenation operation on the first and second feature maps to obtain the adjusted component.

[0211] For example, if the first image feature includes a weight coefficient map obtained by a neural network, and the second image feature includes a second feature map obtained by a neural network, the decision module is configured to, when generating an adjusted component corresponding to the second initial component based on the first and second initial components, to input the first initial component into a first neural network, obtain a first feature map corresponding to the first initial component, perform guided filtering on the first feature map, obtain the weight coefficient map, input the second initial component into a second neural network, obtain the second feature map, and generate the adjusted component based on the weight coefficient map and the second feature map.

[0212] For example, the decision module is configured to perform guided filtering on the first feature map and obtain the weight coefficient map, specifically by performing a convolution operation on the first feature map, obtaining the convolved feature map, performing weight mapping on the convolved feature map, and obtaining the weight coefficient map.

[0213] For example, the decision module is configured to perform weight mapping on the convolutional feature map and obtain the weight coefficient map by specifically performing a pooling operation on the convolutional feature map, performing a fully connected operation and a ReLU activation operation on the feature map after the pooling operation, performing a fully connected operation and a Sigmoid activation operation on the feature map after the ReLU activation operation, and generating the weight coefficient map based on the convolutional feature map and the feature map after the Sigmoid activation operation.

[0214] For example, the decision module is configured to perform a weight mapping on the convolved feature map and obtain the weight coefficient map by specifically performing a convolution operation and a sigmoid activation operation on the convolved feature map to obtain the weight coefficient map.

[0215] Exemplary, the decision module is configured to generate the adjusted component based on the weight coefficient map and the second feature map by specifically performing a multiplication operation on the weight coefficient map and the second feature map to obtain the multiplied feature map, performing a convolution operation on the first feature map to obtain the convolved feature map, performing an addition operation on the multiplied feature map and the convolved feature map to obtain the added feature map, and concatenating the added feature map and the convolved feature map to obtain the adjusted component.

[0216] For example, if the first image feature includes K first feature maps obtained by a neural network, and the second image feature includes K+1 second feature maps obtained by a neural network, the decision module, based on the first and second initial components, generates an adjusted component corresponding to the second initial component, specifically, for the first first feature map and the first second feature map, inputting the first initial component into the neural network to obtain the first first feature map, and inputting the second initial component into the neural network to obtain the first second feature map The system is configured to obtain a feature, and for the i-th first feature map and the i-th second feature map, i is any integer between 2 and K, and K is a positive integer greater than 1. The (i-1)-th first feature map is input into the neural network to obtain the i-th first feature map, feature fusion is performed on the (i-1)-th first feature map and the (i-1)-th second feature map to obtain the fused feature, the fused feature is input into the neural network to obtain the i-th second feature map, and after obtaining the last second feature map, the adjusted component is generated based on the last second feature map.

[0217] For example, the decision module is configured to perform feature fusion on a first feature map and a second feature map to obtain the fused features by, specifically, performing an addition operation on the first feature map and the second feature map to obtain the fused features, or performing a concatenation operation on the first feature map and the second feature map to obtain the fused features, or performing guided filtering on the first feature map to obtain a weight coefficient map corresponding to the first feature map, and then generating the fused features based on the weight coefficient map and the second feature map.

[0218] Exemplary, the decision module is configured to input the first initial component into a first neural network to obtain the first feature map, specifically by extracting M-dimensional (where M is a positive integer) image features from the first initial component, concatenating the M-dimensional image features to obtain a concatenated multidimensional feature, and inputting the multidimensional feature into the first neural network to obtain the first feature map.

[0219] Exemplary, the M-dimensional image features include at least one of the following: an image feature obtained by convolving the first initial component using one convolution kernel; an image feature obtained by cascading the first initial component using two convolution kernels; a first-order spatial domain feature obtained by processing the first initial component using the Sobel operator; a second-order spatial domain feature obtained by processing the first initial component using the Laplacian operator; and a frequency domain feature obtained by performing a Fourier transform on the first initial component.

[0220] Exemplarily, when the determination module inputs the first initial component into the first neural network to obtain the first feature map, specifically, wavelet transform is performed on the first initial component to obtain a plurality of frequency bands after wavelet transform, and the plurality of frequency bands or some of the plurality of frequency bands are input into the first neural network to obtain the first feature map. When the determination module inputs the second initial component into the second neural network to obtain the second feature map, specifically, wavelet transform is performed on the second initial component to obtain a plurality of frequency bands after wavelet transform, and the plurality of frequency bands or some of the plurality of frequency bands are input into the second neural network to obtain the second feature map.

[0221] Exemplarily, before the determination module inputs the first initial component into the first neural network to obtain the first feature map, the determination module is further configured to perform preprocessing on the first initial component to obtain the first initial component after preprocessing, and the first initial component after preprocessing is used to input into the first neural network to obtain the first feature map. When the determination module performs preprocessing on the first initial component to obtain the first initial component after preprocessing, specifically, edge enhancement is performed on the first initial component to obtain image features after edge enhancement, multi-scale feature extraction is performed on the image features after edge enhancement to obtain multi-scale features, and the first initial component after preprocessing is determined based on the multi-scale features.

[0222] Exemplarily, the determination module performs multi-scale feature extraction on the image features after edge enhancement. When obtaining multi-scale features, specifically, a convolution operation is performed on the image features after edge enhancement to obtain the features after convolution, a downsampling operation is performed on the features after convolution to obtain the features after downsampling, a channel conversion is performed on the features after downsampling to obtain the features after channel conversion, an upsampling operation is performed on the features after channel conversion to obtain the features after upsampling, and the multi-scale features are generated based on the features after upsampling and the features after channel conversion.

[0223] Exemplarily, the processing module is further configured to perform upsampling on the second initial component before generating the adjusted component corresponding to the second initial component based on the first initial component and the second initial component, so as to obtain the second initial component after upsampling. Here, the resolution of the second initial component after upsampling is equal to the resolution of the first initial component.

[0224] Exemplarily, the processing module is configured to perform upsampling on the adjusted component before obtaining the restored target component corresponding to the second initial component by performing feature enhancement processing on the adjusted component, so as to obtain the adjusted component after upsampling. Here, the resolution of the adjusted component after upsampling is equal to the resolution of the first initial component.

[0225] Exemplarily, the processing module is configured to perform upsampling on the target component after obtaining the restored target component corresponding to the second initial component by performing feature enhancement processing on the adjusted component, so as to obtain the target component after upsampling. Here, the resolution of the target component after upsampling is equal to the resolution of the first initial component.

[0226] For example, the processing module is configured to perform feature enhancement processing on the adjusted component and obtain the restored target component corresponding to the second initial component, specifically by performing feature enhancement processing on the adjusted component using at least one residual block network to obtain the target component, or by performing feature enhancement processing on the adjusted component using a U-Net network to obtain the target component.

[0227] Those skilled in the art will understand that embodiments of the present invention may be provided as methods, systems, or computer program products. The present invention can be provided as complete hardware embodiments, complete software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can be provided in the form of computer program products implemented on one or more computer-compatible storage media (including, but not limited to, disk memory, CD-ROM, optical memory, etc.) containing computer-compatible program code.

[0228] The above are merely embodiments of the present invention and are not intended to limit it. The present invention can be modified in various ways by those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of the claims. [Explanation of symbols]

[0229] Steps 201-203 811 Processor 812 Machine-readable storage medium

Claims

1. An image decoding method, A step of decoding a bitstream corresponding to the current image block to obtain a reconstructed image block, wherein the reconstructed image block includes an initial luminance component and an initial chromaticity component, and the initial chromaticity component includes a first chromaticity component and a second chromaticity component. The steps include: performing a Haar wavelet transform on the initial luminance component to obtain the Haar wavelet frequency domain features of the initial luminance component; performing a Haar wavelet transform on the first chromaticity component to obtain the Haar wavelet frequency domain features of the first chromaticity component; and performing a Haar wavelet transform on the second chromaticity component to obtain the Haar wavelet frequency domain features of the second chromaticity component. The steps include adjusting the frequency domain characteristics of the first chromaticity component based on the Haar wavelet frequency domain characteristics of the initial luminance component and the Haar wavelet frequency domain characteristics of the first chromaticity component, and obtaining the adjusted first chromaticity component, The process includes the step of adjusting the frequency domain characteristics of the second chromaticity component based on the Haar wavelet frequency domain characteristics of the initial luminance component and the Haar wavelet frequency domain characteristics of the second chromaticity component, and obtaining the adjusted second chromaticity component. An image decoding method characterized by the following:

2. The step of adjusting the frequency domain characteristics of the first chromaticity component based on the Haar wavelet frequency domain characteristics of the initial luminance component and the Haar wavelet frequency domain characteristics of the first chromaticity component, and obtaining the adjusted first chromaticity component, is as follows: The process includes the steps of: generating a frequency-domain fused feature by concatenating the Haar wavelet frequency-domain feature of the initial luminance component and the Haar wavelet frequency-domain feature of the first chromaticity component; and obtaining an adjusted first chromaticity component by replacing the frequency-domain feature corresponding to the first chromaticity component with the frequency-domain fused feature based on the first chromaticity component. The step of adjusting the frequency domain characteristics of the second chromaticity component based on the Haar wavelet frequency domain characteristics of the initial luminance component and the Haar wavelet frequency domain characteristics of the second chromaticity component, and obtaining the adjusted second chromaticity component, is as follows: The process includes the steps of: generating a frequency-domain fused feature by concatenating the Haar wavelet frequency-domain feature of the initial luminance component and the Haar wavelet frequency-domain feature of the second chromaticity component; and obtaining an adjusted second chromaticity component by replacing the frequency-domain feature corresponding to the second chromaticity component with the frequency-domain fused feature based on the second chromaticity component. The method according to feature 1.

3. The reconstructed image block is a reconstructed image block in YUV format, the initial luminance component is the Y component, the first chromaticity component is the U component, and the second chromaticity component is the V component. Here, the resolution of the initial luminance component is greater than or equal to the resolution of the first chromaticity component, and the resolution of the initial luminance component is greater than or equal to the resolution of the second chromaticity component. The method according to feature 1.

4. Performing a Haar wavelet transform on the initial luminance component to obtain the Haar wavelet frequency domain features of the initial luminance component is: This includes performing at least one Haar wavelet transform on the initial luminance component, obtaining a plurality of frequency bands after at least one Haar wavelet transform of the initial luminance component, and determining the Haar wavelet frequency domain features of the initial luminance component based on the plurality of frequency bands after at least one Haar wavelet transform of the initial luminance component. Performing a Haar wavelet transform on the first chromaticity component to obtain the Haar wavelet frequency domain features of the first chromaticity component is: This includes performing at least one Haar wavelet transform on the first chromaticity component, obtaining a plurality of frequency bands after at least one Haar wavelet transform of the first chromaticity component, and determining the Haar wavelet frequency domain features of the first chromaticity component based on the plurality of frequency bands after at least one Haar wavelet transform of the first chromaticity component. Performing a Haar wavelet transform on the second chromaticity component to obtain the Haar wavelet frequency domain features of the second chromaticity component is: The method includes performing at least one Haar wavelet transform on the second chromaticity component, obtaining a plurality of frequency bands after at least one Haar wavelet transform of the second chromaticity component, and determining the Haar wavelet frequency domain features of the second chromaticity component based on the plurality of frequency bands after at least one Haar wavelet transform of the second chromaticity component. The method according to feature 1.

5. Performing at least one Haar wavelet transform on the initial luminance component is: This includes performing a first Haar wavelet transform on the initial luminance component, obtaining multiple frequency bands after the first Haar wavelet transform of the initial luminance component, performing a second Haar wavelet transform on the multiple frequency bands after the first Haar wavelet transform, and obtaining multiple frequency bands after two Haar wavelet transforms of the initial luminance component. Performing at least one Haar wavelet transform on the first chromaticity component is: This includes performing a first Haar wavelet transform on the first chromaticity component, obtaining multiple frequency bands after the first Haar wavelet transform of the first chromaticity component, performing a second Haar wavelet transform on the multiple frequency bands after the first Haar wavelet transform, and obtaining multiple frequency bands after two Haar wavelet transforms of the first chromaticity component. Performing at least one Haar wavelet transform on the second chromaticity component is: This includes performing a first Haar wavelet transform on the second chromaticity component to obtain multiple frequency bands after the first Haar wavelet transform of the second chromaticity component, and performing a second Haar wavelet transform on the multiple frequency bands after the first Haar wavelet transform to obtain multiple frequency bands after two Haar wavelet transforms of the second chromaticity component. The method according to feature 4.

6. The above method further, The process includes the steps of performing feature enhancement on the adjusted first chromaticity component using a first residual block network to obtain the restored target chromaticity component corresponding to the first chromaticity component, and performing feature enhancement on the adjusted second chromaticity component using a second residual block network to obtain the restored target chromaticity component corresponding to the second chromaticity component. The method according to feature 1.

7. An image encoding method, A step of obtaining a reconstructed image block corresponding to the current image block, wherein the reconstructed image block includes an initial luminance component and an initial chromaticity component, and the initial chromaticity component includes a first chromaticity component and a second chromaticity component. The steps include: performing a Haar wavelet transform on the initial luminance component to obtain the Haar wavelet frequency domain features of the initial luminance component; performing a Haar wavelet transform on the first chromaticity component to obtain the Haar wavelet frequency domain features of the first chromaticity component; and performing a Haar wavelet transform on the second chromaticity component to obtain the Haar wavelet frequency domain features of the second chromaticity component. The steps include adjusting the frequency domain characteristics of the first chromaticity component based on the Haar wavelet frequency domain characteristics of the initial luminance component and the Haar wavelet frequency domain characteristics of the first chromaticity component, and obtaining the adjusted first chromaticity component, The process includes the step of adjusting the frequency domain characteristics of the second chromaticity component based on the Haar wavelet frequency domain characteristics of the initial luminance component and the Haar wavelet frequency domain characteristics of the second chromaticity component, and obtaining the adjusted second chromaticity component. An image encoding method characterized by the following.

8. An image decoding device, A decoding module configured to decode a bitstream corresponding to the current image block to obtain a reconstructed image block, wherein the reconstructed image block includes an initial luminance component and an initial chromaticity component, and the initial chromaticity component includes a first chromaticity component and a second chromaticity component, The system comprises a determination module configured to perform a Haar wavelet transform on the initial luminance component to obtain the Haar wavelet frequency domain features of the initial luminance component, perform a Haar wavelet transform on the first chromaticity component to obtain the Haar wavelet frequency domain features of the first chromaticity component, perform a Haar wavelet transform on the second chromaticity component to obtain the Haar wavelet frequency domain features of the second chromaticity component, adjust the frequency domain features of the first chromaticity component based on the Haar wavelet frequency domain features of the initial luminance component and the Haar wavelet frequency domain features of the first chromaticity component to obtain the adjusted first chromaticity component, adjust the frequency domain features of the second chromaticity component based on the Haar wavelet frequency domain features of the initial luminance component and the Haar wavelet frequency domain features of the second chromaticity component to obtain the adjusted second chromaticity component, An image decoding device characterized by the following features.

9. An image encoding device, An acquisition module configured to acquire a reconstructed image block corresponding to the current image block, wherein the reconstructed image block includes an initial luminance component and an initial chromaticity component, and the initial chromaticity component includes a first chromaticity component and a second chromaticity component, The system comprises a determination module configured to perform a Haar wavelet transform on the initial luminance component to obtain the Haar wavelet frequency domain features of the initial luminance component, perform a Haar wavelet transform on the first chromaticity component to obtain the Haar wavelet frequency domain features of the first chromaticity component, perform a Haar wavelet transform on the second chromaticity component to obtain the Haar wavelet frequency domain features of the second chromaticity component, adjust the frequency domain features of the first chromaticity component based on the Haar wavelet frequency domain features of the initial luminance component and the Haar wavelet frequency domain features of the first chromaticity component to obtain the adjusted first chromaticity component, adjust the frequency domain features of the second chromaticity component based on the Haar wavelet frequency domain features of the initial luminance component and the Haar wavelet frequency domain features of the second chromaticity component to obtain the adjusted second chromaticity component, An image coding device characterized by the following:

10. An image decoding device, Processor and The system comprises a machine-readable storage medium in which machine-executable instructions that can be executed by the processor are stored, The processor is configured to perform the decoding method described in any one of claims 1 to 6 by executing the machine-executable instructions. An image decoding device characterized by the following features.

11. An image encoding device, Processor and The system comprises a machine-readable storage medium in which machine-executable instructions that can be executed by the processor are stored, The processor is configured to perform the encoding method described in claim 7 by executing the machine-executable instructions. An image coding device characterized by the following features.

12. A machine-readable storage medium, Multiple computer instructions are stored in the aforementioned machine-readable storage medium. When the computer instruction is executed by the processor, the decoding method described in any one of claims 1 to 6 is performed, or the encoding method described in claim 7 is performed. A machine-readable storage medium characterized by the following features.

13. It is a computer program, When the computer program is executed by the processor, the method according to any one of claims 1 to 7 is carried out. A computer program characterized by the following features.