Image decoding, encoding method, device, and its device based on neural network
By generating block-specific decoding neural networks from control parameters, the method addresses low stability and high complexity in neural network-based image encoding and decoding, achieving improved performance and adjustable complexity.
Patent Information
- Application Number
- JP2025501682
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-14
- Filing Date
- 2023-07-12
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2043-07-12
AI Technical Summary
Existing image encoding and decoding methods based on neural networks suffer from low stability, low generalization ability, and high complexity, particularly in video encoding where high data density and low latency are required.
The method involves decoding control parameters from a bitstream to generate a decoding neural network for each block, allowing for variable and adjustable decoding neural networks, which can improve stability and generalization while reducing complexity.
This approach enhances decoding performance and stability, enabling better encoding and decoding efficiency with adjustable complexity and bit rate control, outperforming single neural network frameworks.
Smart Images

Figure 2025522098000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of encoding and decoding technologies, and particularly relates to an image decoding, encoding method, apparatus, and device thereof based on a neural network.
Background Art
[0002] In order to save space, all video images are encoded before transmission, and complete video encoding may include processes such as prediction, transformation, quantization, entropy encoding, filtering, etc. For the prediction process, the prediction process may include intra-frame prediction and inter-frame prediction. The inter-frame prediction utilizes the temporal correlation of the video to predict the current pixel using the pixels of adjacent encoded images, thereby effectively removing the temporal redundancy of the video. The intra-frame prediction utilizes the spatial correlation of the video to predict the current pixel using the pixels of the encoded blocks of the image in the current frame, thereby achieving the purpose of removing the spatial redundancy of the video.
[0003] With the rapid development of deep learning, deep learning has achieved good results in many high-level computer vision problems such as image classification and target detection. Deep learning has also gradually begun to be applied in the field of encoding and decoding, that is, it has become possible to encode and decode images using a neural network. Although the encoding and decoding methods based on neural networks show great performance potential, the encoding and decoding methods based on neural networks still have problems such as low stability, low generalization ability, and high complexity.
Summary of the Invention
[0004] In view of this, the present invention provides an image decoding, encoding method, apparatus, and device thereof based on a neural network, which improve the encoding performance and decoding performance and solve problems such as low stability, low generalization ability, and high complexity.
[0005] The present invention is an image decoding method based on a neural network applied to a decoding side, comprising: decoding control parameters corresponding to a current block from a bit stream; Tab and obtaining neural network information corresponding to a decoding processing unit from the control parameters, and generating a decoding neural network corresponding to the decoding processing unit based on the neural network information; determining input features corresponding to the decoding processing unit, and processing the input features based on the decoding neural network to obtain output features corresponding to the decoding processing unit. Before An image decoding method based on a neural network is provided.
[0006] In some embodiments, when the neural network information includes basic layer information, the step of generating a decoding neural network corresponding to the decoding processing unit based on the neural network information includes: Determining a basic layer corresponding to the decoding processing unit based on the basic layer information; Generating a decoding neural network corresponding to the decoding processing unit based on the basic layer.
[0007] In some embodiments, the step of determining a basic layer corresponding to the decoding processing unit based on the basic layer information includes: The basic layer information includes a basic layer default network usage flag bit, and when the basic layer default network usage flag bit indicates that the basic layer uses a default network, the step includes obtaining a basic layer of the default network structure.
[0008] In some embodiments, the step of determining a basic layer corresponding to the decoding processing unit based on the basic layer information includes: The basic layer information includes a basic layer pre-designed network usage flag bit and a basic layer pre-designed network index number, and when the basic layer pre-designed network usage flag bit indicates that the basic layer uses a pre-designed network, the step includes selecting, from a pre-designed neural network pool, a basic layer of a pre-designed network structure corresponding to the basic layer pre-designed network index number. The pre-designed neural network pool includes network layers of at least one pre-designed network structure.
[0009] In some embodiments, when the basic layer is a first feature decoding network, the step of processing the input features based on the decoding neural network to obtain output features corresponding to the decoding processing unit includes: Decoding a first bit stream of the current block that is input based on the first feature decoding network, to obtain coefficient hyperparameter feature information corresponding to the current block that is output.
[0010] In some embodiments, when the basic layer is a coefficient hyperparameter feature generation network, the step of processing the input features based on the decoding neural network to obtain output features corresponding to the decoding processing unit includes: Based on the coefficient hyperparameter feature generation network, processing the input coefficient hyperparameter feature coefficient reconstruction value to obtain the coefficient hyperparameter feature value corresponding to the current block to be output, wherein the coefficient hyperparameter feature coefficient reconstruction value input to the coefficient hyperparameter feature generation network is determined based on the coefficient hyperparameter feature information output by the first feature decoding network.
[0011] In some embodiments, the step of determining the coefficient hyperparameter feature coefficient reconstruction value based on the coefficient hyperparameter feature information is when the control parameter includes first enable information and the first enable information indicates enabling a first inverse quantization process, including the step of performing inverse quantization on the coefficient hyperparameter feature information to obtain a coefficient hyperparameter feature coefficient reconstruction value.
[0012] In some embodiments, when the basic layer is a second feature decoding network, the step of processing the input feature based on the decoding neural network to obtain the output feature corresponding to the decoding processing unit is including the step of decoding the second bitstream of the current block to be input based on the second feature decoding network to obtain the image feature information corresponding to the current block to be output.
[0013] In some embodiments, the step of decoding the second bitstream of the current block to be input based on the second feature decoding network to obtain the image feature information corresponding to the current block to be output is including the step of decoding the second bitstream based on the coefficient hyperparameter feature value input to the second feature decoding network to obtain the image feature information, wherein the coefficient hyperparameter feature value is the output feature of the coefficient hyperparameter feature generation network.
[0014] In some embodiments, when the basic layer is a conversion network, the step of processing the input feature based on the decoding neural network to obtain the output feature corresponding to the decoding processing unit is including the step of processing the input image feature reconstruction value based on the conversion network to obtain the image low-level feature value corresponding to the current block to be output. The image feature reconstruction value input to the conversion network is determined based on the image feature information output by the second feature decoding network. The image low-level feature value is used to obtain a reconstructed image block corresponding to the current block.
[0015] In some embodiments, the step of determining the image feature reconstruction value based on the image feature information is as follows: When the control parameter includes second enable information and the second enable information indicates enabling a second inverse quantization process, the step includes performing inverse quantization on the image feature information to obtain an image feature reconstruction value.
[0016] In some embodiments, when the neural network information includes enhancement layer information, the step of generating a decoding neural network corresponding to the decoding processing unit based on the neural network information is as follows: Determining an enhancement layer corresponding to the decoding processing unit based on the enhancement layer information; Generating a decoding neural network corresponding to the decoding processing unit based on the enhancement layer.
[0017] In some embodiments, the step of determining an enhancement layer corresponding to the decoding processing unit based on the enhancement layer information is as follows: The enhancement layer information includes an enhancement layer default network usage flag bit. When the enhancement layer default network usage flag bit indicates that the enhancement layer uses a default network, the step includes obtaining an enhancement layer with a default network structure.
[0018] In some embodiments, the step of determining an enhancement layer corresponding to the decoding processing unit based on the enhancement layer information is as follows: The enhancement layer information includes an enhancement layer pre-designed network usage flag bit and an enhancement layer pre-designed network index number. When the enhancement layer pre-designed network usage flag bit indicates that the enhancement layer uses a pre-designed network, the step includes selecting, from a pre-designed neural network pool, an enhancement layer with a pre-designed network structure corresponding to the enhancement layer pre-designed network index number. The pre-designed neural network pool includes network layers of at least one pre-designed network structure.
[0019] In some embodiments, when the enhancement layer is a conversion network, the step of determining the input features corresponding to the decoding processing unit includes: decoding the bitstream of the current block to obtain image feature information corresponding to the current block; and determining a reconstructed value of the image features corresponding to the current block based on the image feature information.
[0020] In some embodiments, the step of processing the input features based on the decoding neural network to obtain output features corresponding to the decoding processing unit includes: processing the reconstructed value of the input image features based on the conversion network to obtain a low-level image feature value corresponding to the current block output, where the low-level image feature value is used to obtain a reconstructed image block corresponding to the current block.
[0021] In some embodiments, the step of determining the reconstructed value of the image features corresponding to the current block based on the image feature information includes: when the control parameter includes second enable information and the second enable information indicates enabling a second inverse quantization process, performing inverse quantization on the image feature information to obtain the reconstructed value of the image features.
[0022] In some embodiments, when the enhancement layer is a quality improvement network, the step of processing the input features based on the decoding neural network to obtain output features corresponding to the decoding processing unit includes: performing an enhancement process on the low-level image feature value corresponding to the current block input based on the quality improvement network to obtain a reconstructed image block corresponding to the current block output.
[0023] In some embodiments, the step of performing an enhancement process on the low-level image feature value corresponding to the current block based on the quality improvement network to obtain a reconstructed image block corresponding to the current block includes, when the control parameter includes third enable information and the third enable information indicates enabling a quality improvement process, performing an enhancement process on the low-level image feature value corresponding to the current block based on the quality improvement network to obtain a reconstructed image block corresponding to the current block.
[0024] In some embodiments, the step of decoding control parameters corresponding to the current block from the bitstream is including the step of decoding the bitstream corresponding to the current block via a control parameter decoding network to obtain control parameters corresponding to the current block, where the control parameters include control parameters of a conversion network and control parameters of a quality improvement network.
[0025] In some embodiments, the step of determining the enhancement layer corresponding to the decoding processing unit based on the enhancement layer information is including, when the enhancement layer information includes network parameters for generating an enhancement layer, generating an enhancement layer corresponding to the decoding processing unit based on the network parameters, where the network parameters include at least one of the number of layers of a neural network, the number of filterings, the filtering size index, the filtering coefficient zero flag bit, and the filtering coefficient.
[0026] The present invention is an image encoding method based on a neural network applied to an encoding side, comprising: Pre obtaining control parameters corresponding to a current block, wherein the control parameters include neural network information corresponding to a decoding processing unit, and the neural network information is used to determine a decoding neural network corresponding to the decoding processing unit; encoding the control parameters corresponding to the current block into a bitstream. sent There is provided an image encoding method based on a neural network, including the above steps.
[0027] The present invention relates to a memory configured to store video data; decoding control parameters corresponding to a current block from a bitstream; Con obtaining neural network information corresponding to a decoding processing unit from the control parameters, and generating a decoding neural network corresponding to the decoding processing unit based on the neural network information; determining input features corresponding to the decoding processing unit, and processing the input features based on the decoding neural network to obtain output features corresponding to the decoding processing unit. trol There is provided an image decoding apparatus based on a neural network, including a decoder configured to perform the above steps.
[0028] The present invention relates to a memory configured to store video data; PreA step of obtaining a control parameter corresponding to a current block, where the control parameter includes neural network information corresponding to a decoding processing unit, and the neural network information is used to determine a decoding neural network corresponding to the decoding processing unit, and a step of encoding the control parameter corresponding to the current block in a bitstream, and an encoder configured to perform the steps, and an image encoding apparatus based on a neural network are provided. sent The present invention provides a decoding device including a processor and a machine-readable storage medium, where the machine-readable storage medium stores machine-executable instructions executable by the processor,
[0029] and the processor is used to execute the machine-executable instructions to implement the image decoding method based on the neural network as described above. The present invention provides an encoding device including a processor and a machine-readable storage medium, where the machine-readable storage medium stores machine-executable instructions executable by the processor,
[0030] and the processor is used to execute the machine-executable instructions to implement the image encoding method based on the neural network as described above. and the processor is used to execute the machine-executable instructions to implement the image encoding method based on the neural network as described above.
[0031] As can be seen from the above technical solutions, in the embodiments of the present invention, control parameters corresponding to the current block are decoded from the bitstream, neural network information corresponding to the decoding processing unit is obtained from the control parameters, a decoding neural network corresponding to the decoding processing unit is generated based on the neural network information, image decoding can be realized based on the decoding neural network, and the decoding performance is improved. Image encoding can be realized based on the encoding neural network corresponding to the encoding processing unit, and the encoding performance is improved. By using a neural network (for example, a decoding neural network, an encoding neural network, etc.) to encode and decode an image, transmitting neural network information in a bitstream, and generating a decoding neural network corresponding to the decoding processing unit based on the neural network information, problems such as low stability, low generalization, and high complexity can be solved, that is, high stability, high generalization, and low complexity. A solution for dynamically adjusting the complexity of encoding and decoding can be provided, and it has better encoding performance and decoding performance compared to the framework of a single neural network. Since each current block corresponds to control parameters, the neural network information obtained from the control parameters is the neural network information for the current block, and a decoding neural network is generated for each current block respectively. That is, the decoding neural networks of different current blocks may be the same or different. Thereby, the decoding neural network at the block level, that is, the decoding neural network is variable and adjustable.
Brief Description of the Drawings
[0032]
Figure 1
Figure 2A
Figure 2B
Figure 2C
Figure 3
Figure 4
Figure 5A
Figure 5B
Figure 5C
Figure 5D
Figure 5E
Figure 6A
Figure 6B
Figure 6C
Figure 6D
Figure 7A
Figure 7B
Figure 7C
Figure 8A
Figure 8B
Best Mode for Carrying Out the Invention
[0033] The terms used in the embodiments of the present invention are for illustrative purposes only for specific embodiments and are not intended to limit the present invention. The singular forms "a kind", "the foregoing", and "said" used in the embodiments and claims of the present invention are also intended to include the plural forms unless the context clearly dictates otherwise. Also, it should be understood that the term "and / or" used in the present invention means any or all possible combinations of one or more of the related listed items. The embodiments of the present invention may use terms such as first, second, third, etc. to explain various information, but it should be understood that this information is not limited to these terms. These terms are used only to distinguish the same type of information. For example, within the scope of the embodiments of the present invention, depending on the context, the first information may be referred to as the second information, and similarly, the second information may be referred to as the first information. Also, the word "when" used here may be interpreted as "and", "when", or "in response to a decision".
[0034] The embodiments of the present invention provide an image decoding and encoding method based on a neural network and may be related to the following concepts.
[0035] Neural Network (NN): A neural network refers to an artificial neural network rather than a biological neural network. A neural network is a computational model composed of a large number of nodes (or called neurons) connected to each other. In a neural network, a neuron processing unit can represent different objects, such as features, alphabets, concepts, or some meaningful abstract modes. The types of processing units in a neural network can be divided into three categories: input units, output units, and hidden units. Input units receive external signals and data, output units realize the output of processing results, and hidden units are units that are between the input units and output units and cannot be observed from the outside of the system. The connection weights between neurons reflect the connection strength between units, and the representation and processing of information are reflected in the connection relationships of the processing units. A neural network is an unprogrammed, brain-like information processing method. Its essence is to obtain a parallel distributed information processing function through the transformation and dynamic behavior of the neural network, and to imitate the information processing functions of the human nervous system at different degrees and levels. In the field of video processing, commonly used neural networks may include, but are not limited to, convolutional neural networks (CNNs), recurrent neural networks (RNNs), fully connected networks, etc.
[0036] Convolutional Neural Network (CNN): A convolutional neural network is a feed-forward neural network and one of the representative network structures in deep learning technology. The artificial neurons of a convolutional neural network can respond to the peripheral units within a certain coverage range and exhibit excellent performance in large-scale image processing. The basic structure of a convolutional neural network includes two layers. One is the feature extraction layer (also called the convolutional layer). The input of each neuron is connected to the local receptive field of the previous layer to extract the local features. Once the local features are extracted, their positional relationships with other features are also determined. The other is the feature mapping layer (also called the activation layer). Each computational layer of the neural network consists of multiple feature mappings. Each feature mapping is a plane, and the weights of all neurons in the plane are equal. For the feature mapping structure, functions such as the Sigmoid function, ReLU function, Leaky-ReLU function, PReLU function, and GDN function can be used as the activation functions of the convolutional network. Also, since the neurons in one mapping plane share weights, the number of free parameters of the network decreases.
[0037] Compared with image processing algorithms, one of the advantages of convolutional neural networks is that they can avoid complex preprocessing processes for images (such as the extraction of artificial features), directly input the original image, and perform end-to-end learning. Compared with general neural networks, one of the advantages of convolutional neural networks is that general neural networks adopt a fully connected method, that is, all neurons from the input layer to the hidden layer are connected, resulting in a huge number of parameters, which makes the training of the network time-consuming or difficult. In contrast, convolutional neural networks avoid this difficulty through methods such as local connection and weight sharing.
[0038] Deconvolution Layer: The deconvolution layer is also called the transposed convolution layer. The operation processes of the deconvolution layer and the convolution layer are similar. The main difference is that the deconvolution layer can make the output larger than the input through padding (of course, it can also be the same). If the stride is 1, it means that the output size is equal to the input size. If the stride is N, it means that the width of the output feature is N times the width of the input feature, and the height of the output feature is N times the height of the input feature.
[0039] Generalization Ability: The generalization ability may refer to the adaptability of a machine learning algorithm to new samples. The purpose of learning is to learn the rules behind the data. For data outside the training set with the same rules, the trained network should be able to give appropriate outputs, and this ability may be called the generalization ability.
[0040] Rate-Distortion Optimization Principle: To evaluate the coding efficiency, there are two metrics, bitrate and PSNR (Peak Signal to Noise Ratio). The smaller the bitstream, the higher the compression ratio, and the higher the PSNR, the better the quality of the reconstructed image. When making mode selection, the discriminant is essentially a comprehensive evaluation of both. For example, the cost corresponding to a mode: J(mode) = D + λ * R, where D represents Distortion, which can usually be evaluated using the SSE (Sum of the Squared Errors) metric. SSE refers to the sum of the mean squares of the differences between the reconstructed image block and the source image. To consider the cost, the SAD metric can also be used. SAD refers to the sum of the absolute values of the differences between the reconstructed image block and the source image. λ is the Lagrange multiplier, and R is the actual number of bits required for encoding the image block in this mode, including the total number of bits required for encoding mode information, motion information, residuals, etc. When making mode selection, using the rate-distortion principle to compare and determine the coding mode can usually guarantee optimal coding performance.
[0041] Video Coding Framework: FIG. 1 is a schematic diagram of the video coding framework on the encoding side. The encoding-side processing process of the embodiments of the present invention may be implemented by this video coding framework. Since the schematic diagram of the video decoding framework may be similar to FIG. 1, the description is omitted here. However, the decoding-side processing process of the embodiments of the present invention may also be implemented by the video decoding framework.
[0042] Exemplarily, as shown in FIG. 1, the video encoding framework may include modules such as prediction, transformation, quantization, entropy encoder, inverse quantization, inverse transformation, reconstruction, filtering, etc. On the encoding side, through the cooperation among these modules, the encoding-side processing process can be realized. Also, the video decoding framework may include modules such as prediction, transformation, quantization, entropy decoder, inverse quantization, inverse transformation, reconstruction, filtering, etc. On the decoding side, through the cooperation among these modules, the decoding-side processing process can be realized.
[0043] Regarding each module on the encoding side, a very large number of encoding tools have been proposed, and each tool usually has many modes. For different video sequences, the encoding tools that bring the optimal encoding performance are often different. Therefore, in the encoding process, usually, RDO (Rate-Distortion Optimize) is used to compare the encoding performance of different tools or modes and select the optimal mode. After determining the optimal tool or mode, the determination information of the tool or mode is transmitted by a method of encoding mark information into the bitstream. Such a method brings encoding complexity, but for different contents, the optimal mode combination can be adaptively selected to obtain the optimal encoding performance. The decoding side can obtain the relevant mode information by directly analyzing the mark information, and the influence of complexity is small.
[0044] Hereinafter, the structures on the encoding side and the decoding side will be briefly described. FIG. 2A is a schematic block diagram of an example of the encoding side for implementing an embodiment of the present invention. In FIG. 2A, the encoding side includes a prediction processing unit, a residual calculation unit, a conversion processing unit, a quantization unit, an encoding unit, an inverse quantization unit, an inverse conversion processing unit, a reconstruction unit, and a filter unit. In one example, the encoding side may further include a buffer and a decoded picture buffer (DPB). The buffer is used to buffer the reconstructed image blocks output by the reconstruction unit, and the decoded picture buffer is used to buffer the filtered image blocks output by the filter unit.
[0045] The input to the encoding side (also called the encoder) is an image block of an image (which may also be called the image to be encoded). The image block is also called the current block or the block to be encoded. The encoding side may further include a splitting unit (not shown) for splitting the image to be encoded into a plurality of image blocks. The encoding side is used to encode the image to be encoded block by block. For example, an encoding process is executed for each image block. The prediction processing unit receives or acquires an image block (which may also be called the current image block to be encoded of the current image to be encoded, and the current block, and the image block may be understood as the true value of the image block) and the reconstructed image data, and is used to predict the current block based on the relevant data in the reconstructed image data to obtain the predicted block of the current block. In one example, the prediction processing unit may include an inter-frame prediction unit, an intra-frame prediction unit, and a mode selection unit. The mode selection unit is used to select an intra-frame prediction mode or an inter-frame prediction mode. When the intra-frame prediction mode is selected, the prediction process is executed by the intra-frame prediction unit. When the inter-frame prediction mode is selected, the prediction process may be executed by the inter-frame prediction unit.
[0046] The residual calculation unit is used to calculate the residual between the true value of the image block and the predicted block of the image block to obtain a residual block. For example, the residual calculation unit may subtract the pixel value of the predicted block from the pixel value of the image block for each pixel.
[0047] The conversion processing unit is used to perform a conversion such as a discrete cosine transform (DCT) or a discrete sine transform (DST) on the residual block to obtain conversion coefficients in the conversion domain. The conversion coefficients may be called conversion residual coefficients, and the conversion residual coefficients can represent the residual block in the conversion domain.
[0048] The quantization unit is used to apply scalar quantization or vector quantization to quantize the conversion coefficients to obtain quantized conversion coefficients. The quantized conversion coefficients may be called quantized residual coefficients. The quantization process can reduce the bit depth for some or all of the conversion coefficients. For example, during quantization, an n-bit conversion coefficient may be truncated to an m-bit conversion coefficient, where n is greater than m. The degree of quantization can be changed by adjusting the quantization parameter (QP). For example, in the case of scalar quantization, fine quantization or coarse quantization can be achieved by applying different scales. A small quantization stride corresponds to fine quantization, and a large quantization stride corresponds to coarse quantization. The quantization parameter (QP) may indicate an appropriate quantization stride.
[0049] The symbolization unit encodes the quantized residual coefficients, outputs the encoded image data (i.e., the encoding result of the image block to be currently encoded) in the form of an encoded bitstream, and transmits the encoded bitstream to the decoder, or may transmit it to the decoder later or store it for use in search. The symbolization unit may further be used to encode other syntax elements of the current image block, such as encoding the prediction mode into the bitstream. The encoding algorithms include, but are not limited to, variable length coding (VLC) algorithm, context adaptive VLC (CAVLC) algorithm, arithmetic coding algorithm, context adaptive binary arithmetic coding (CABAC) algorithm, syntax-based context-adaptive binary arithmetic coding (SBAC) algorithm, probability interval partitioning entropy (PIPE) algorithm.
[0050] The inverse quantization unit is used to perform inverse quantization on the quantized coefficients to obtain the inverse quantized coefficients. The inverse quantization is the inverse application of the quantization unit. For example, based on or using the same quantization stride as the quantization unit, an inverse quantization scheme of the quantization scheme applied by the quantization unit may be applied. The inverse quantized coefficients may also be referred to as inverse quantized residual coefficients.
[0051] The inverse transformation processing unit is used to perform an inverse transformation on the above-mentioned inverse quantized coefficients. The inverse transformation is the inverse application of the above-mentioned transformation processing unit. For example, the inverse transformation may include an inverse discrete cosine transform (IDCT) or an inverse discrete sine transform (IDST) for obtaining an inverse transformation block in the pixel region (also called the sample region). The inverse transformation block may also be called an inverse transformation inverse quantized block or an inverse transformation residual block.
[0052] The reconstruction unit is used to add an inverse transformation block (i.e., an inverse transformation residual block) to a prediction block to obtain a reconstructed block in the sample region. The reconstruction unit may be an adder. For example, it adds the sample values (i.e., pixel values) of the residual block and the sample values of the prediction block. The reconstructed block output by the reconstruction unit may be used later to predict other image blocks, such as being used in the intra-frame prediction mode.
[0053] The filter unit (or abbreviated as "filter") is used to filter the reconstructed block to obtain a filtered block so as to perform pixel conversion smoothly or improve the quality of the image. The filter unit may be a loop filter unit intended to represent one or more loop filters. For example, the filter unit may be a deblocking filter, a sample-adaptive offset (SAO) filter, or other filters such as a bilateral filter, an adaptive loop filter (ALF), an edge sharpening or smoothing filter, a collaborative filter, etc. In one example, the filtered block output by the filtering unit may be used later to predict other image blocks, such as being used in the inter-frame prediction mode, but is not limited thereto.
[0054] Figure 2B is a schematic block diagram of an example of a decoder side (also called a decoder) for implementing an embodiment of the present invention. The decoder is used to receive, for example, encoded image data encoded by an encoder (i.e., an encoded bitstream, for example, including an encoded bitstream of an image block and syntax elements associated therewith) and obtain a decoded image. The decoder includes a decoding unit, an inverse quantization unit, an inverse transform processing unit, a prediction processing unit, a reconstruction unit, and a filter unit. In some embodiments, the decoder may perform a decoding process that is substantially the reverse of the encoding process described for the encoder of FIG. 2A. In one example, the decoder may further include a buffer and a decoded image buffer. The buffer is used to buffer the reconstructed image blocks output by the reconstruction unit, and the decoded image buffer is used to buffer the filtered image blocks output by the filter unit.
[0055] The decoding unit performs decoding on the encoded image data to obtain quantized coefficients and / or decoded encoding parameters (for example, the encoding parameters may include any one or all of inter-frame prediction parameters, intra-frame prediction parameters, filter parameters, and / or other syntax elements). The decoding unit is further used to transfer the decoded encoding parameters to the prediction processing unit so that the prediction processing unit can perform a prediction process based on the encoding parameters. The function of the inverse quantization unit may be the same as that of the inverse quantization unit of the encoder and is used to inverse-quantize the quantized coefficients decoded by the decoding unit.
[0056] The function of the inverse transformation processing unit may be the same as that of the inverse transformation processing unit of the encoder, and the function of the reconstruction unit (e.g., an adder) may be the same as that of the reconstruction unit of the encoder. It is used to perform an inverse transformation (e.g., inverse DCT, inverse integer transformation, or a conceptually similar inverse transformation process) on the quantized coefficients to obtain an inverse transformation block (also called an inverse transformation residual block), and the inverse transformation block is the residual block of the current image block in the pixel domain.
[0057] The prediction processing unit is used to receive or acquire encoded image data (e.g., the encoded bitstream of the current image block) and reconstructed image data. The prediction processing unit further receives or acquires, for example, prediction-related parameters and / or information regarding the selected prediction mode (i.e., decoded encoded parameters) from a decoding unit, and may predict the current image block based on the relevant data in the reconstructed image data and the decoded encoded parameters to obtain a prediction block of the current image block.
[0058] In one example, the prediction processing unit may include an inter-frame prediction unit, an intra-frame prediction unit, and a mode selection unit. The mode selection unit is used to select an intra-frame prediction mode or an inter-frame prediction mode. When the intra-frame prediction mode is selected, the prediction process is executed by the intra-frame prediction unit, and when the inter-frame prediction mode is selected, the prediction process is executed by the inter-frame prediction unit.
[0059] The reconstruction unit is used to add the inverse transformation block (i.e., the inverse transformation residual block) to the prediction block to obtain a reconstructed block in the sample domain. For example, the sample values of the inverse transformation residual block and the sample values of the prediction block may be added.
[0060] The filter unit is used to filter the reconstructed blocks to obtain filtered blocks, and the filtered blocks are decoded image blocks.
[0061] In addition, in the encoder and decoder of the embodiments of the present invention, the processing result of a certain process may be further processed, and the further processed result may be output to the next process. For example, after processes such as interpolation filtering, motion vector derivation, or filtering, further processing such as Clip or shift may be performed on the processing result of the corresponding process.
[0062] Based on the encoder and decoder, the embodiments of the present invention provide an implementable encoding / decoding method. As shown in FIG. 2C, FIG. 2C is a schematic flowchart of encoding and decoding provided by the embodiments of the present invention. The encoding and decoding method includes processes 1 to 5, and processes 1 to 5 may be executed by the above decoder and encoder. Process 1: Divide an image of one frame into one or more non-overlapping parallel encoding units. The one or more parallel encoding units do not depend on each other and can be encoded and decoded completely in parallel and independently of each other, like the parallel encoding unit 1 and parallel encoding unit 2 shown in FIG. 2C.
[0063] Process 2: For each parallel encoding unit, it may be further divided into one or more non-overlapping independent encoding units. Each independent encoding unit may not depend on each other, but may share some parallel encoding unit header information. For example, the width of the independent encoding unit is w_lcu and the height is h_lcu. When a parallel encoding unit is divided into one independent encoding unit, the size of the independent encoding unit is exactly the same as that of the parallel encoding unit. Otherwise, the width of the independent encoding unit must be greater than the height (except for the edge area).
[0064] Generally, the independent encoding unit may have a fixed size of w_lcu×h_lcu, where both w_lcu and h_lcu are powers of 2 (N≧0). For example, the size of the independent encoding unit can be 128×4, 64×4, 32×4, 16×4, 8×4, 32×2, 16×2, or 8×2, etc.
[0065] As one possible example, the independent encoding unit may have a fixed size of 128×4. When the size of the parallel encoding unit is 256×8, the parallel encoding unit may be evenly divided into four independent encoding units. When the size of the parallel encoding unit is 288×10, the parallel encoding unit may be divided such that the first and second rows each consist of two 128×4 and one 32×4 independent encoding units, and the third row consists of two 128×2 and one 32×2 independent encoding units. Note that the independent encoding unit may include three components of luminance Y, chrominance Cb, and chrominance Cr, three components of red (R), green (G), and blue (B), or three components of luminance Y, chrominance Co, and chrominance Cg, or it may include only any one of these components. When the independent encoding unit includes three components, the sizes of these three components may be exactly the same or different, specifically related to the input format of the image.
[0066]
[0067] When the sizes of the sub-encoding unit and the independent encoding unit are the same (i.e., when the independent encoding unit is divided into only one sub-encoding unit), the size may be any of the sizes described in Process 2. When the independent encoding unit is divided into a plurality of non-overlapping sub-encoding units, possible division examples include horizontal equal division (the height of the sub-encoding unit is the same as that of the independent encoding unit, but the width is different and may be 1 / 2, 1 / 4, 1 / 8, 1 / 16, etc.), vertical equal division (the width of the sub-encoding unit is the same as that of the independent encoding unit, but the height is different and may be 1 / 2, 1 / 4, 1 / 8, 1 / 16, etc.), horizontal and vertical equal division (quad-tree division), etc., and horizontal equal division is preferred.
[0068] The width of the sub-encoding unit is w_cu and the height is h_cu, and the width must be greater than the height (except for the edge region). Usually, the sub-encoding unit is of a fixed w_cu×h_cu, where both w_cu and h_cu are powers of 2 (N is 0 or more), for example, 16×4, 8×4, 16×2, 8×2, 8×1, 4×1, etc. For example, the sub-encoding unit is fixed at 16×4. When the size of the independent encoding unit is 64×4, the independent encoding unit is evenly divided into 4 sub-encoding units. When the size of the independent encoding unit is 72×4, it is divided into 4 sub-encoding units of 16x4 + 1 sub-encoding unit of 8×4. Note that the sub-encoding unit may include three components of luminance Y, chrominance Cb, and chrominance Cr (or three components of red R, green G, and blue B, or luminance Y, chrominance Co, and chrominance Cg), or may include only one of these components. When including three components, the sizes of these components may be exactly the same or different, specifically related to the input format of the image.
[0069] Note that Process 3 may be an optional step in the encoding and decoding method, and the encoder / decoder may perform encoding and decoding on the residual coefficients (or residual values) of the independent encoding unit obtained in Process 2.
[0070] Process 4: For the sub - encoding unit, it may be further divided into one or more non - overlapping prediction groups (PG, which may be abbreviated as Group). Each PG is encoded and decoded according to the selected prediction mode, and the predicted value of the PG is obtained to form the predicted value of the entire sub - encoding unit. Based on the predicted value and the original value of the sub - encoding unit, the residual value of the sub - encoding unit is obtained.
[0071] Process 5: Based on the residual value of the sub - encoding unit, the sub - encoding units are grouped to obtain one or more non - overlapping residual blocks (RB). The residual coefficients of each RB are encoded and decoded according to the selected mode to form a residual coefficient stream. Specifically, it can be divided into those that perform transformation on the residual coefficients and those that do not.
[0072] Here, the selected mode of the encoding and decoding method for the residual coefficient in Process 5 may include, but is not limited to, any of a semi-fixed-length coding mode, an exponential Golomb coding method, a Golomb-Rice coding method, a truncated unary coding method, a run-length coding method, a method of directly coding the original residual value, etc. For example, the encoder may directly code the coefficients in the RB. In another example, the encoder may perform a transformation (e.g., DCT, DST, Hadamard transformation, etc.) on the residual block and then code the transformed coefficients. As one possible example, when the RB is relatively small, the encoder may directly perform uniform quantization on each coefficient in the RB and then perform binary coding. When the RB is relatively large, it may be further divided into a plurality of coefficient groups (CGs), and after performing uniform quantization on each CG, binary coding may be performed. In some embodiments of the present invention, the coefficient group (CG) and the quantization group (QG) may be the same, and of course, the coefficient group and the quantization group may also be different.
[0073] Next, the encoding part of the residual coefficients in the semi-fixed length encoding method will be exemplarily described. First, the maximum value of the absolute value of the residuals within one RB block is defined as the modified maximum (MM). Next, the number of encoding bits for the residual coefficients within the RB block is determined (the number of encoding bits for the residual coefficients within the same RB block is the same). For example, if the critical limit (CL) of the current RB block is 2 and the current residual coefficient is 1, 2 bits are required to encode the residual coefficient 1, which is represented as 01. If the CL of the current RB block is 7, this represents encoding an 8-bit residual coefficient and a 1-bit sign bit. The determination of CL is to find the minimum M value that satisfies the condition that all the residuals of the current sub-block are within the range of [-2^(M-1), 2^(M-1)]. If both of the two boundary values of -2^(M-1) and 2^(M-1) exist, M is incremented by 1, that is, M+1 bits are required to encode all the residuals of the current RB block. If only one of the two boundary values of -2^(M-1) and 2^(M-1) exists, one Trailing bit is encoded to determine whether the boundary value is -2^(M-1) or 2^(M-1). If neither -2^(M-1) nor 2^(M-1) exists for all the residuals, there is no need to encode the Trailing bit. Also, in some special cases, the encoder may directly encode the original value of the image instead of the residual value.
[0074] With the rapid development of deep learning, neural networks can adaptively construct feature descriptions from training data and have higher flexibility and versatility. Therefore, deep learning has achieved remarkable results in many high-level computer vision problems such as image classification and object detection. Deep learning has also gradually begun to be applied in the field of encoding and decoding, that is, using neural networks to encode and decode images. For example, instead of deblocking filtering technology and sample adaptive offset technology, by using the convolutional neural network VRCNN (Variable Filter Size Convolutional Neural Network) to perform post-processing filtering on the image after intra-frame encoding, the subjective and objective quality of the reconstructed image can be significantly improved. Also, for example, applying neural networks to intra-frame prediction, proposing an intra-frame prediction mode based on block upsampling and downsampling, for the blocks of intra-frame prediction, first performing downsampling encoding, and then performing upsampling on the reconstructed pixels by the neural network, a performance improvement of up to 9.0% can be obtained in ultra-high-definition sequences. Obviously, neural networks can effectively avoid the limitations of artificially set modes, obtain neural networks that meet actual needs through data-driven, and significantly improve the encoding performance.
[0075] Exemplarily, the encoding and decoding methods based on neural networks show great performance potential. However, the encoding and decoding methods based on neural networks still have problems such as low stability, low generalization ability, and high complexity. First, neural networks are still evolving rapidly. New network structures emerge one after another. For a specific problem of a module with an encoder, and even for a general problem, it is not yet clear which network structure is optimal. Therefore, using a single fixed neural network for a module with an encoder has a high risk. Also, since the formation of a neural network depends greatly on training data, when the training data does not contain the characteristics of actual problems, the performance is likely to decline when dealing with such problems. In related technologies, one mode usually uses only one neural network. When the neural network lacks generalization ability, this mode will lead to a decline in encoding performance. Moreover, since the data density that needs to be processed in video encoding is high and low latency is required, video encoding standard technologies have very high requirements for complexity, especially the complexity of decoding. However, to obtain good encoding performance, generally, the number of parameters of the neural network for encoding is very large (for example, more than 1M), and the number of multiplications and additions generated by one application of the neural network averages more than 100K times per pixel. On the other hand, a simplified neural network with fewer network layers or parameters cannot obtain optimal encoding performance, but the encoding and decoding complexity brought about by the use of such a neural network can be significantly reduced. Also, many image encoding methods take the entire frame image as input (the cache overhead is very large) and cannot effectively control the output bit rate. However, in actual applications, it is usually necessary to obtain images with an arbitrary bit rate.
[0076] In response to the above discovery, embodiments of the present invention provide an image decoding method and an image encoding method based on a neural network, which can encode and decode an image using a neural network (for example, a decoding neural network, an encoding neural network, etc.). In this embodiment, the optimization concept is not only to focus on the encoding performance and decoding performance, but also to pay attention to the complexity (especially the parallelism of coefficient encoding and decoding) and application functions (the variability of the bit rate and the support for fine-tuning, that is, the controllability of the bit rate). Based on the above optimization concept, several possible methods are provided in the embodiments of the present invention. a) Fix the structure of the neural network, but do not limit the network parameters (that is, weight parameters) of the neural network, and the network parameters may be indexed in the ID manner. b) The structure of the neural network is flexible, and the related structure parameters may be transmitted by syntax encoding, and the network parameters may be indexed in the ID manner (some parameters may be transmitted by syntax encoding, and a certain bit rate cost will increase). c) The structure of the neural network is flexible (high-level syntax setting), the network parameters of some networks (for example, shallow networks) are fixed (saving the bit rate), and the network parameters of the remaining networks are transmitted by syntax encoding (maintaining the space for performance optimization).
[0077] Hereinafter, in connection with several specific embodiments, the decoding method and encoding method in the embodiments of the present invention will be described in detail.
[0078] Embodiment 1: Embodiments of the present invention provide an image decoding method based on a neural network. FIG. 3 is a schematic flowchart of the method. The method may be applied to a decoding side (also called a video decoder), and the method may include steps 301 to 303.
[0079] In step 301, control parameters and image information corresponding to the current block are decoded from the bitstream.
[0080] In step 302, neural network information corresponding to the decoding processing unit is obtained from the control parameters, and a decoding neural network corresponding to the decoding processing unit is generated based on the neural network information.
[0081] In step 303, input features corresponding to the decoding processing unit are determined based on the image information, and the input features are processed based on the decoding neural network to obtain output features corresponding to the decoding processing unit.
[0082] In one possible embodiment, when the neural network information includes basic layer information and enhancement layer information, based on the basic layer information, a basic layer corresponding to the decoding processing unit is determined, and based on the enhancement layer information, an enhancement layer corresponding to the decoding processing unit is determined. Then, based on the basic layer and the enhancement layer, a decoding neural network corresponding to the decoding processing unit is generated.
[0083] Exemplarily, for the process of determining the basic layer, the basic layer information includes a basic layer default network usage flag bit, and when the basic layer default network usage flag bit indicates that the basic layer uses the default network, the basic layer of the default network structure is obtained.
[0084] Exemplarily, for the process of determining the basic layer, the basic layer information includes a basic layer pre-designed network usage flag bit and a basic layer pre-designed network index number, and when the basic layer pre-designed network usage flag bit indicates that the basic layer uses the pre-designed network, the basic layer of the pre-designed network structure corresponding to the basic layer pre-designed network index number may be selected from the pre-designed neural network pool.
[0085] Here, the pre-designed neural network pool may include network layers of at least one pre-designed network structure.
[0086] Exemplarily, for the process of determining the enhancement layer, the enhancement layer information includes an enhancement layer default network usage flag bit, and when the enhancement layer default network usage flag bit indicates that the enhancement layer uses the default network, the enhancement layer of the default network structure is obtained.
[0087] Exemplarily, for the process of determining the enhancement layer, the enhancement layer information includes an enhancement layer pre-designed network usage flag bit and an enhancement layer pre-designed network index number, and when the enhancement layer pre-designed network usage flag bit indicates that the enhancement layer uses the pre-designed network, the enhancement layer of the pre-designed network structure corresponding to the enhancement layer pre-designed network index number may be selected from the pre-designed neural network pool.
[0088] Here, the pre-designed neural network pool includes network layers of at least one pre-designed network structure.
[0089] Exemplarily, for the process of determining the enhancement layer, when the enhancement layer information includes network parameters for generating the enhancement layer, an enhancement layer corresponding to the decoding processing unit is generated based on the network parameters. Here, the network parameters may include at least one of the number of layers of the neural network, the transposed convolution layer flag bit, the number of transposed convolution layers, the quantization stride of each transposed convolution layer, the number of channels of each transposed convolution layer, the size of the convolution kernel, the filtering number, the filtering size index, the filtering coefficient zero flag bit, the filtering coefficient, the activation layer flag bit, the activation layer type, etc., but are not limited thereto. Of course, the above are just some examples and are not limited in this regard.
[0090] In one possible embodiment, the image information may include coefficient hyperparameter feature information and image feature information. The step of determining the input features corresponding to the decoding processing unit based on the image information and processing the input features based on the decoding neural network to obtain the output features corresponding to the decoding processing unit, when executing the decoding process of generating the coefficient hyperparameter features, based on the coefficient hyperparameter feature information, determining a coefficient hyperparameter feature coefficient reconstruction value, and based on the decoding neural network, performing an inverse transformation process on the coefficient hyperparameter feature coefficient reconstruction value to obtain a coefficient hyperparameter feature value, where the coefficient hyperparameter feature value is used to decode the image feature information from the bitstream; and when executing the decoding process of inverse-transforming the image features, based on the image feature information, determining an image feature reconstruction value, and based on the decoding neural network, performing an inverse transformation process on the image feature reconstruction value to obtain a low-level image feature value, where the low-level image feature value is used to obtain the reconstructed image block corresponding to the current block, may be included, but is not limited thereto.
[0091] Exemplarily, the step of determining the coefficient hyperparameter feature coefficient reconstruction value based on the coefficient hyperparameter feature information may include, when the control parameter includes first enable information and the first enable information indicates enabling the first inverse quantization process, performing inverse quantization on the coefficient hyperparameter feature information to obtain a coefficient hyperparameter feature coefficient reconstruction value, but is not limited thereto. The step of determining the image feature reconstruction value based on the image feature information may include, when the control parameter includes second enable information and the second enable information indicates enabling the second inverse quantization process, performing inverse quantization on the image feature information to obtain an image feature reconstruction value, but is not limited thereto.
[0092] Exemplarily, when the control parameter includes the third enable information and the third enable information indicates enabling high-quality processing, when performing a high-quality decoding process, an image low-level feature value may be obtained, and based on a decoding neural network, enhancement processing may be performed on the image low-level feature value to obtain a reconstructed image block corresponding to the current block.
[0093] In one possible embodiment, the decoding device may include a control parameter decoding unit, a first feature decoding unit, a second feature decoding unit, a coefficient hyperparameter feature generation unit, and an image feature inverse transformation unit, and the image information may include coefficient hyperparameter feature information and image feature information. Here, the control parameter decoding unit can decode the control parameter from the bitstream, the first feature decoding unit can decode the coefficient hyperparameter feature information from the bitstream, and the second feature decoding unit can decode the image feature information from the bitstream. When the coefficient hyperparameter feature generation unit is a decoding processing unit, a coefficient hyperparameter feature coefficient reconstruction value can be determined based on the coefficient hyperparameter feature information, and the coefficient hyperparameter feature generation unit can perform inverse transformation processing on the coefficient hyperparameter feature coefficient reconstruction value based on a decoding neural network to obtain a coefficient hyperparameter feature value, and the coefficient hyperparameter feature value is used for the second feature decoding unit to decode the image feature information from the bitstream. When the image feature inverse transformation unit is a decoding processing unit, an image feature reconstruction value is determined based on the image feature information, and the image feature inverse transformation unit can perform inverse transformation processing on the image feature reconstruction value based on a decoding neural network to obtain an image low-level feature value, and the image low-level feature value is used to obtain a reconstructed image block corresponding to the current block.
[0094] Exemplarily, the decoding device may further include a first inverse quantization unit and a second inverse quantization unit. The control parameter may include first enable information of the first inverse quantization unit. When the first enable information indicates enabling the first inverse quantization unit, the first inverse quantization unit may obtain coefficient hyperparameter feature information from the first feature decoding unit, perform inverse quantization on the coefficient hyperparameter feature information to obtain a coefficient hyperparameter feature coefficient reconstruction value, and provide the coefficient hyperparameter feature coefficient reconstruction value to the coefficient hyperparameter feature generation unit. The control parameter may include second enable information of the second inverse quantization unit. When the second enable information indicates enabling the second inverse quantization unit, the second inverse quantization unit may obtain image feature information from the second feature decoding unit, perform inverse quantization on the image feature information to obtain an image feature reconstruction value, and provide the image feature reconstruction value to the image feature inverse transformation unit.
[0095] Exemplarily, the decoding device may further include a quality improvement unit. Here, the control parameter may include third enable information of the quality improvement unit. When the third enable information indicates enabling the quality improvement unit, if the quality improvement unit is a decoding processing unit, the quality improvement unit may obtain an image low-level feature value from the image feature inverse transformation unit, and perform an enhancement process on the image low-level feature value based on the decoding neural network to obtain a reconstructed image block corresponding to the current block.
[0096] Exemplarily, the above execution order is merely an exemplification for facilitating the explanation. In actual applications, the execution order between steps may be changed and is not limited to this execution order. Further, in other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in this specification, and the method may include more or fewer steps than those described in this specification. Also, a single step described in this specification may be decomposed and described as a plurality of steps in other embodiments, and a plurality of steps described in this specification may be combined and described as a single step in other embodiments.
[0097] As can be seen from the above technical solution, in the embodiment of the present invention, the control parameters corresponding to the current block are decoded from the bitstream, the neural network information corresponding to the decoding processing unit is obtained from the control parameters, the decoding neural network corresponding to the decoding processing unit is generated based on the neural network information, image decoding can be realized based on the decoding neural network, and the decoding performance is improved. Image encoding can be realized based on the encoding neural network corresponding to the encoding processing unit, and the encoding performance is improved. By using a deep learning network (for example, a decoding neural network, an encoding neural network, etc.) to encode and decode an image, transmitting neural network information in a bitstream, and generating a decoding neural network corresponding to the decoding processing unit based on the neural network information, problems such as low stability, low generalization, and high complexity can be solved, that is, high stability, high generalization, and low complexity. A solution for dynamically adjusting the complexity of encoding and decoding can be provided, and it has better encoding performance compared to the framework of a single deep learning network. Since each current block corresponds to control parameters, the neural network information obtained from the control parameters is the neural network information for the current block, and a decoding neural network is generated for each current block respectively. That is, the decoding neural networks of different current blocks may be the same or different. Thereby, the decoding neural network at the block level, that is, the decoding neural network is variable and adjustable.
[0098] Embodiment 2: The embodiment of the present invention provides an image encoding method based on a neural network. FIG. 4 is a schematic flowchart of the method. The method may be applied to the encoding side (also called a video encoder), and the method may include steps 401 to 403.
[0099] In step 401, input features corresponding to the encoding processing unit are determined based on the current block, the input features are processed based on the encoding neural network corresponding to the encoding processing unit to obtain output features corresponding to the encoding processing unit, and image information corresponding to the current block, such as coefficient hyperparameter feature information and image feature information, etc., is determined based on the output features.
[0100] In step 402, control parameters corresponding to the current block are obtained, and the control parameters may include neural network information corresponding to the decoding processing unit, and the neural network information is used to determine the decoding neural network corresponding to the decoding processing unit.
[0101] In step 403, the image information and control parameters corresponding to the current block are encoded into the bitstream.
[0102] In one possible embodiment, the neural network information includes basic layer information and enhancement layer information, and the decoding neural network includes a basic layer determined based on the basic layer information and an enhancement layer determined based on the enhancement layer information.
[0103] Exemplarily, the basic layer information includes a basic layer default network usage flag bit, and when the basic layer default network usage flag bit indicates that the basic layer uses the default network, the decoding neural network uses the basic layer of the default network structure.
[0104] Exemplarily, the base layer information includes a base layer pre-designed network usage flag bit and a base layer pre-designed network index number. And when the base layer pre-designed network usage flag bit indicates that the base layer uses the pre-designed network, the decoding neural network may use the base layer of the pre-designed network structure corresponding to the base layer pre-designed network index number selected from the pre-designed neural network pool.
[0105] Here, the pre-designed neural network pool may include network layers of at least one pre-designed network structure.
[0106] Exemplarily, the enhancement layer information includes an enhancement layer default network usage flag bit. And when the enhancement layer default network usage flag bit indicates that the enhancement layer uses the default network, the decoding neural network uses the enhancement layer of the default network structure.
[0107] Exemplarily, the enhancement layer information includes an enhancement layer pre-designed network usage flag bit and an enhancement layer pre-designed network index number. And when the enhancement layer pre-designed network usage flag bit indicates that the enhancement layer uses the pre-designed network, the decoding neural network may use the enhancement layer of the pre-designed network structure corresponding to the enhancement layer pre-designed network index number selected from the pre-designed neural network pool.
[0108] Here, the pre-designed neural network pool includes network layers of at least one pre-designed network structure.
[0109] Exemplarily, when the enhancement layer information includes network parameters for generating an enhancement layer, the decoding neural network may use the enhancement layer generated based on the network parameters. Here, the network parameters may include at least one of the number of layers of the neural network, the transposed convolution layer flag bit, the number of transposed convolution layers, the quantization stride of each transposed convolution layer, the number of channels of each transposed convolution layer, the size of the convolution kernel, the number of filterings, the filtering size index, the filtering coefficient zero flag bit, the filtering coefficient, the activation layer flag bit, and the activation layer type, but is not limited thereto. Of course, the above are only some examples of network parameters and are not limited thereto.
[0110] In one possible embodiment, the image information may include coefficient hyperparameter feature information and image feature information, determining the input features corresponding to the encoding processing unit based on the current block, processing the input features based on the encoding neural network corresponding to the encoding processing unit to obtain the output features corresponding to the encoding processing unit, and determining the image information corresponding to the current block based on the output features. The steps may include, in the encoding process of performing feature transformation, performing feature transformation on the current block based on the encoding neural network to obtain the image feature value corresponding to the current block, where the image feature value is used to determine the image feature information, and in the encoding process of performing coefficient hyperparameter feature transformation, performing coefficient hyperparameter feature transformation on the image feature value based on the encoding neural network to obtain the coefficient hyperparameter feature coefficient value, where the coefficient hyperparameter feature coefficient value is used to determine the coefficient hyperparameter feature information, but is not limited thereto.
[0111] Exemplarily, in the process of determining coefficient hyperparameter feature information based on coefficient hyperparameter feature coefficient values, quantization may be performed on the coefficient hyperparameter feature coefficient values to obtain coefficient hyperparameter feature coefficient quantization values, and the coefficient hyperparameter feature information may be determined based on the coefficient hyperparameter feature coefficient quantization values. Here, the control parameter may further include first enable information, and the first enable information is used to indicate that the first quantization process is enabled.
[0112] Exemplarily, in the process of determining image feature information based on image feature values, quantization may be performed on the image feature values to obtain image feature quantization values, and the image feature information may be determined based on the image feature quantization values. Here, the control parameter may further include second enable information, and the second enable information is used to indicate that the second quantization process is enabled.
[0113] Exemplarily, the step of obtaining the control parameter corresponding to the current block may include the step of determining neural network information used in the decoding process of the inverse transformation of the image features of the decoding device based on the network structure of the encoding neural network used in the encoding process of the feature transformation, where the neural network information is used to determine the decoding neural network corresponding to the decoding process of the inverse transformation of the image features of the decoding device, and / or the step of determining neural network information used in the decoding process of the generation of the coefficient hyperparameter features of the decoding device based on the network structure of the encoding neural network used in the encoding process of the coefficient hyperparameter feature transformation, where the neural network information is used to determine the decoding neural network corresponding to the decoding process of the generation of the coefficient hyperparameter features of the decoding device, but is not limited thereto.
[0114] In one possible embodiment, the encoding device may include a control parameter encoding unit, a first feature encoding unit, a second feature encoding unit, a feature conversion unit, and a coefficient hyperparameter feature conversion unit, and the image information may include coefficient hyperparameter feature information and image feature information. Here, the control parameter encoding unit encodes the control parameter into a bitstream, the first feature encoding unit encodes the coefficient hyperparameter feature information into the bitstream, the second feature encoding unit encodes the image feature information into the bitstream. When the feature conversion unit is an encoding processing unit, the feature conversion unit can perform feature conversion on the current block based on an encoding neural network to obtain an image feature value corresponding to the current block, and the image feature value is used to determine the image feature information. When the coefficient hyperparameter feature conversion unit is an encoding processing unit, the coefficient hyperparameter feature conversion unit can perform coefficient hyperparameter feature conversion on the image feature value based on an encoding neural network to obtain a coefficient hyperparameter feature coefficient value, and the coefficient hyperparameter feature coefficient value is used to determine the coefficient hyperparameter feature information.
[0115] Exemplarily, the encoding device may further include a first quantization unit and a second quantization unit. The first quantization unit obtains a coefficient hyperparameter feature coefficient value from the coefficient hyperparameter feature conversion unit, performs quantization on the coefficient hyperparameter feature coefficient value to obtain a coefficient hyperparameter feature coefficient quantization value, and may determine coefficient hyperparameter feature information based on the coefficient hyperparameter feature coefficient quantization value. Here, the control parameter may further include first enable information of the first quantization unit, and the first enable information is used to indicate that the first quantization unit is enabled. Further, the second quantization unit obtains an image feature value from the feature conversion unit, performs quantization on the image feature value to obtain an image feature quantization value, and may determine image feature information based on the image feature quantization value. Here, the control parameter may further include second enable information of the second quantization unit, and the second enable information is used to indicate that the second quantization unit is enabled.
[0116] Exemplarily, the step of obtaining the control parameter corresponding to the current block may include the step of determining neural network information corresponding to the image feature inverse conversion unit of the decoding device based on the network structure of the encoding neural network of the feature conversion unit, where the neural network information is used to determine the decoding neural network corresponding to the image feature inverse conversion unit of the decoding device, and / or the step of determining neural network information corresponding to the coefficient hyperparameter feature generation unit of the decoding device based on the network structure of the encoding neural network of the coefficient hyperparameter feature conversion unit, where the neural network information is used to determine the decoding neural network corresponding to the coefficient hyperparameter feature generation unit of the decoding device, but is not limited thereto.
[0117] In one possible embodiment, the image information may include coefficient hyperparameter feature information and image feature information. When performing the decoding process for generating the coefficient hyperparameter features, based on the coefficient hyperparameter feature information, a reconstructed value of the coefficient hyperparameter feature coefficients is determined, and based on the decoding neural network, an inverse transformation process is performed on the reconstructed value of the coefficient hyperparameter feature coefficients to obtain a coefficient hyperparameter feature value, and the coefficient hyperparameter feature value may be used to decode the image feature information from the bitstream. When performing the decoding process for the inverse transformation of the image features, based on the image feature information, a reconstructed value of the image features is determined, and based on the decoding neural network, an inverse transformation process is performed on the reconstructed value of the image features to obtain a low-level image feature value, and the low-level image feature value may be used to obtain a reconstructed image block corresponding to the current block.
[0118] Exemplarily, the step of determining the reconstructed value of the coefficient hyperparameter feature coefficients based on the coefficient hyperparameter feature information may include, but is not limited to, performing inverse quantization on the coefficient hyperparameter feature information to obtain the reconstructed value of the coefficient hyperparameter feature coefficients. The step of determining the reconstructed value of the image features based on the image feature information may include, but is not limited to, performing inverse quantization on the image feature information to obtain the reconstructed value of the image features.
[0119] Exemplarily, when performing the decoding process for quality improvement, a low-level image feature value is obtained, and based on the decoding neural network, an enhancement process is performed on the low-level image feature value to obtain a reconstructed image block corresponding to the current block.
[0120] Exemplarily, the step of obtaining the control parameter corresponding to the current block includes at least one of the following steps, but is not limited thereto: a step of determining neural network information used for the decoding process of generating the coefficient hyperparameter feature of the decoding device based on the network structure of the decoding neural network used in the encoding process of generating the coefficient hyperparameter feature of the encoding device, where the neural network information is used to determine the decoding neural network used in the decoding process of generating the coefficient hyperparameter feature of the decoding device; a step of determining neural network information used for the decoding process of inverse-transforming the image feature of the decoding device based on the network structure of the decoding neural network used in the encoding process of inverse-transforming the image feature of the encoding device, where the neural network information is used to determine the decoding neural network used in the decoding process of inverse-transforming the image feature of the decoding device; and a step of determining neural network information used for the decoding process of enhancing the quality of the decoding device based on the network structure of the decoding neural network used in the encoding process of enhancing the quality of the encoding device, where the neural network information is used to determine the decoding neural network used in the decoding process of enhancing the quality of the decoding device.
[0121] In one possible embodiment, the encoding device may include a first feature decoding unit, a second feature decoding unit, a coefficient hyperparameter feature generation unit, and an image feature inverse transformation unit, and the image information may include coefficient hyperparameter feature information and image feature information. Here, the first feature decoding unit can decode the coefficient hyperparameter feature information from the bitstream, the second feature decoding unit can decode the image feature information from the bitstream, and after determining the coefficient hyperparameter feature coefficient reconstruction value based on the coefficient hyperparameter feature information, the coefficient hyperparameter feature generation unit can perform an inverse transformation process on the coefficient hyperparameter feature coefficient reconstruction value based on the decoding neural network to obtain the coefficient hyperparameter feature value. The coefficient hyperparameter feature value is used for the second feature decoding unit to decode the image feature information from the bitstream and for the second feature encoding unit to encode the image feature information into the bitstream. After determining the image feature reconstruction value based on the image feature information, the image feature inverse transformation unit can perform an inverse transformation process on the image feature reconstruction value based on the decoding neural network to obtain the image low-level feature value. The image low-level feature value is used to obtain the reconstructed image block corresponding to the current block.
[0122] Exemplarily, the encoding device may further include a first inverse quantization unit and a second inverse quantization unit. The first inverse quantization unit can obtain the coefficient hyperparameter feature information from the first feature decoding unit and perform inverse quantization on the coefficient hyperparameter feature information to obtain the coefficient hyperparameter feature coefficient reconstruction value. The second inverse quantization unit can obtain the image feature information from the second feature decoding unit and perform inverse quantization on the image feature information to obtain the image feature reconstruction value.
[0123] Exemplarily, the encoding device may further include a quality improvement unit. The quality improvement unit may obtain the image low-level feature values from the image feature inverse conversion unit and perform an enhancement process on the image low-level feature values based on the decoding neural network to obtain a reconstructed image block.
[0124] Exemplarily, the step of obtaining the control parameter corresponding to the current block may include at least one of the following steps, but is not limited thereto: a step of determining the neural network information corresponding to the coefficient hyperparameter feature generation unit of the decoding device based on the network structure of the decoding neural network of the coefficient hyperparameter feature generation unit of the encoding device, where the neural network information is used to determine the decoding neural network corresponding to the coefficient hyperparameter feature generation unit of the decoding device; a step of determining the neural network information corresponding to the image feature inverse conversion unit of the decoding device based on the network structure of the decoding neural network of the image feature inverse conversion unit of the encoding device, where the neural network information is used to determine the decoding neural network corresponding to the image feature inverse conversion unit of the decoding device; and a step of determining the neural network information corresponding to the quality improvement unit of the decoding device based on the network structure of the decoding neural network of the quality improvement unit of the encoding device, where the neural network information is used to determine the decoding neural network corresponding to the quality improvement unit of the decoding device.
[0125] In one possible embodiment, before determining the input features corresponding to the encoding processing unit based on the current block, the current image may be divided into N non-overlapping image blocks, where N is a positive integer. For each image block, boundary padding may be performed to obtain the image block after boundary padding. When performing boundary padding on each image block, the padding value may not depend on the reconstructed pixel values of adjacent image blocks. Based on the image block after boundary padding, N current blocks may be generated.
[0126] In one possible embodiment, before determining the input features corresponding to the encoding processing unit based on the current block, the current image may be divided into a plurality of basic blocks, each basic block including at least one image block. For each image block, boundary padding may be performed to obtain the image block after boundary padding. When performing boundary padding on each image block, the padding value of the image block is not allowed to depend on the reconstructed pixel values of other image blocks within the same basic block, but is allowed to depend on the reconstructed pixel values of image blocks in different basic blocks. Based on the image block after boundary padding, a plurality of current blocks may be generated.
[0127] Exemplarily, the above execution order is merely an example for ease of explanation. In actual applications, the execution order between steps may be changed and is not limited to this execution order. Also, in other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in this specification. The method may include more or fewer steps than those described in this specification. Also, a single step described in this specification may be decomposed and described as a plurality of steps in other embodiments, and a plurality of steps described in this specification may be combined and described as a single step in other embodiments.
[0128] As can be seen from the above technical solutions, in the embodiments of the present invention, control parameters corresponding to the current block are decoded from the bitstream, neural network information corresponding to the decoding processing unit is obtained from the control parameters, and a decoding neural network corresponding to the decoding processing unit is generated based on the neural network information, and image decoding can be realized based on the decoding neural network, improving the decoding performance. Image encoding can be realized based on the encoding neural network corresponding to the encoding processing unit, improving the encoding performance. By using a deep learning network (for example, a decoding neural network, an encoding neural network, etc.) to encode and decode an image, transmitting neural network information in a bitstream, and generating a decoding neural network corresponding to the decoding processing unit based on the neural network information, problems such as low stability, low generalization, and high complexity can be solved, that is, high stability, high generalization, and low complexity. A solution can be provided to dynamically adjust the complexity of encoding and decoding, and it has better encoding performance compared to the framework of a single deep learning network. Since each current block corresponds to control parameters, the neural network information obtained from the control parameters is the neural network information for the current block, and a decoding neural network is generated for each current block respectively. That is, the decoding neural networks of different current blocks may be the same or different. Thus, the block-level decoding neural network, that is, the decoding neural network is variable and adjustable.
[0129] Embodiment 3: The embodiments of the present invention provide an image encoding method and an image decoding method based on a neural network with a variable and adjustable bit rate, which can realize the high parallelism of image blocks, and the bit rate is controllable and adjustable. FIG. 5A is a schematic diagram of the image encoding method and the image decoding method, showing the encoding process of the image encoding method and the decoding process of the image decoding method.
[0130] Exemplarily, for an image encoding method based on a neural network, the encoding process may include the following steps S11 to S15.
[0131] In step S11, the block division unit divides the current image (i.e., the original image) into N non-overlapping image blocks (i.e., original image blocks, which may be denoted as original image block 1, original image block 2, …, original image block N), where N is a positive integer.
[0132] In step S12, boundary padding is performed on each original image block to obtain an image block after boundary padding, and N current blocks are generated based on the image block after boundary padding. That is, boundary padding is performed on each of the N original image blocks to obtain N image blocks after boundary padding, and the N image blocks after boundary padding are used as the N current blocks.
[0133] Exemplarily, as shown in FIG. 5B, when performing boundary padding on each original image block, the padding value may not depend on the reconstructed pixel values of adjacent image blocks. Thereby, it is ensured that each original image block can be independently encoded in parallel, and the encoding performance can be improved.
[0134] In step S13, the encoding unit determines the encoding parameters of the current block based on the information of the encoded block. The encoding parameters are used to control the magnitude of the encoding bit rate of the current block (for example, parameters such as quantization stride), and there is no limitation on these encoding parameters.
[0135] In step S14, the control unit writes the control parameters that are required by the decoding side and cannot be derived into the bitstream.
[0136] In step S15, the padded image block (i.e., the current block) is input into an encoding unit based on a neural network. The encoding unit encodes the current block based on encoding parameters and outputs a bitstream of the current block. Exemplarily, when the encoding unit encodes the current block based on the encoding parameters, the encoding unit may encode the current block using a neural network.
[0137] Exemplarily, for an image decoding method based on a neural network, the decoding process may include the following steps S21 to S24.
[0138] In step S21, the decoding side decodes control parameters that the current block requires and cannot derive from the bitstream.
[0139] In step S22, based on the control parameters and the bitstream of the current block, a reconstruction image block corresponding to the current block is obtained by a decoding unit based on a neural network, that is, the current block is decoded to obtain a reconstruction image block, for example, a reconstruction image block 1 corresponding to the original image block 1, a reconstruction image block 2 corresponding to the original image block 2,..., a reconstruction image block N corresponding to the original image block N.
[0140] Exemplarily, when the decoding unit decodes the current block, the current block may be decoded using a neural network.
[0141] In step S23, based on the control parameters, it is determined whether to perform filtering processing on a certain current block. If filtering processing is to be performed, the filtering merging unit performs filtering processing based on the current block and information of at least one adjacent reconstruction image block to obtain a filtered image block.
[0142] In step S24, the filtered image blocks are merged to obtain a reconstructed image.
[0143] In one possible embodiment, when performing boundary padding on the original image block, the padding value may be a padding preset value. The padding preset value may be a default value agreed upon by encoding and decoding (such as 0, or 1<<(1-depth), where depth is the bit depth, for example 8, 10, 12, etc.), or a value transmitted to the decoding side by high-level syntax encoding, or may be obtained by performing processing such as mirroring or nearest neighbor copy based on the pixels of the current block. There is no limitation on the method for obtaining this padding value.
[0144] In one possible embodiment, when performing boundary padding on the original image block, the size of the padding extension of the surrounding blocks may be a default value agreed upon by encoding and decoding (such as 1, 2, 4, etc.), or a value related to the size of the current block, or a value transmitted to the decoding side by high-level syntax encoding. There is no limitation on the size of the padding extension of this surrounding block.
[0145] Example 4: The embodiments of the present invention provide an image encoding method and an image decoding method based on a neural network with a variable and adjustable bit rate, which can realize the high parallelism of image blocks, the bit rate is controllable and adjustable, and furthermore, information between adjacent blocks (such as reconstructed pixels of adjacent blocks, etc.) may be utilized. FIG. 5C is a schematic diagram of the image encoding method and the image decoding method, showing the encoding process of the image encoding method and the decoding process of the image decoding method.
[0146] Exemplarily, for the image encoding method based on a neural network, its encoding process may include the following steps S31 to S35.
[0147] In step S31, the block splitting unit splits the current image (i.e., the original image) into a plurality of basic blocks, each basic block includes at least one image block. Taking M basic blocks as an example, M is a positive integer, and the M basic blocks include a total of N non-overlapping image blocks (i.e., original image blocks, denoted as original image block 1, original image block 2, …, original image block N), and N is a positive integer.
[0148] In step S32, boundary padding is performed on each original image block to obtain an image block after boundary padding, and N current blocks are generated based on the image block after boundary padding, that is, boundary padding is performed on each of the N original image blocks to obtain N image blocks after boundary padding, and the N image blocks after boundary padding are used as the N current blocks.
[0149] Exemplarily, when performing boundary padding on each original image block, the padding value of the original image block is not dependent on the reconstructed pixel values of other original image blocks within the same basic block, and it is permitted to be dependent on the reconstructed pixel values of original image blocks within different basic blocks.
[0150] In this embodiment, the concept of a basic block is introduced. As shown in FIG. 5D, each basic block includes at least one image block (i.e., the original image block). The image blocks within a basic block do not refer to each other, but can refer to the reconstruction information of other image blocks located in different basic blocks. For example, as shown in FIG. 5D, for image block 1, since the image block on its left side (i.e., the adjacent block on the left side of image block 1) is located in the same basic block as image block 1, the reconstruction information of this image block cannot be used as the padding value of image block 1, and the image block is padded using a preset padding value. For image block 1, since the image block above it (i.e., the adjacent block above image block 1) is located in a different basic block from image block 1, the reconstruction information of this image block can be used as the padding value of image block 1. When using the reconstruction information of this image block as the padding value of image block 1, the reconstruction value of this image block may be used, or the reconstruction value before filtering of this image block may be used.
[0151] Obviously, by introducing basic blocks, parallelism can be guaranteed (each image block within a basic block can be encoded and decoded in parallel), and it is also helpful for performance improvement (the reconstruction information of image blocks in adjacent basic blocks can be utilized).
[0152] In step S33, the encoding unit determines the encoding parameters of the current block based on the information of the encoded block. The encoding parameters are used to control the magnitude of the encoding bit rate of the current block (for example, parameters such as quantization stride), and there is no limitation on these encoding parameters.
[0153] In step S34, the control unit writes the control parameters that are required by the decoding side and cannot be derived into the bitstream.
[0154] In step S35, the padded image block (i.e., the current block) is input into an encoding unit based on a neural network. The encoding unit encodes the current block based on encoding parameters and outputs a bitstream of the current block. Exemplarily, when the encoding unit encodes the current block based on the encoding parameters, the encoding unit may encode the current block using a neural network.
[0155] Exemplarily, for an image decoding method based on a neural network, the decoding process may include the following steps S41 to S44.
[0156] In step S41, the decoding side decodes control parameters that the current block requires and cannot derive from the bitstream.
[0157] In step S42, based on the control parameters and the bitstream of the current block, a reconstruction image block corresponding to the current block is obtained by a decoding unit based on a neural network, that is, the current block is decoded to obtain a reconstruction image block, for example, a reconstruction image block 1 corresponding to the original image block 1, a reconstruction image block 2 corresponding to the original image block 2,..., a reconstruction image block N corresponding to the original image block N.
[0158] In step S43, based on the control parameters, it is determined whether to perform filtering processing on a certain current block. If filtering processing is to be performed, the filtering merging unit performs filtering processing based on the current block and information of at least one adjacent reconstruction image block to obtain a filtered image block.
[0159] In step S44, the filtered image blocks are merged to obtain a reconstructed image.
[0160] In one possible embodiment, when performing boundary padding on the original image block, the padding value may be a padding preset value. The padding preset value may be a default value agreed upon by encoding and decoding (such as 0, or 1<<(1-depth), where depth is the bit depth, such as 8, 10, 12, etc.), or a value transmitted to the decoding side by high-level syntax encoding, or a value obtained by performing processing such as mirroring or nearest neighbor copy based on the pixels of the current block. The method for obtaining this padding value is not limited.
[0161] In one possible embodiment, when performing boundary padding on the original image block, the size of the padding extension of the surrounding blocks may be a default value agreed upon by encoding and decoding (such as 1, 2, 4, etc.), or a value related to the size of the current block, or a value transmitted to the decoding side by high-level syntax encoding. The size of the padding extension of this surrounding block is not limited.
[0162] In one possible embodiment, the number of image blocks included in the basic block may be a default value agreed upon by encoding and decoding (such as 1, 4, 16, etc.), or a value related to the size of the current image, or a value transmitted to the decoding side by high-level syntax encoding. The number of image blocks in this basic block is not limited and may be selected according to actual needs.
[0163] In one possible embodiment, in Examples 3 and 4, for the block splitting process (steps S11 and S31) of the original image, as shown in FIG. 5E, an image domain conversion may be performed on the original image, and block splitting may be performed on the image after the image domain conversion. Correspondingly, the filtered image blocks may be merged into one image, an image domain inverse conversion may be performed on this image, and a reconstructed image may be obtained based on the image after the image domain inverse conversion. Exemplarily, performing an image domain conversion on the original image may be a conversion process from an RGB domain image to a YUV domain image (the corresponding image domain inverse conversion is an inverse conversion process from a YUV domain image to an RGB domain image), or a process such as wavelet conversion or Fourier conversion may be introduced to generate an image in a new domain (the corresponding image domain inverse conversion is an inverse wavelet conversion or an inverse Fourier conversion process). The image domain conversion process can be realized using either a neural network or a non-neural network, and the image domain conversion process is not limited herein.
[0164] Example 5: The embodiments of the present invention provide an image decoding method based on a neural network, which may be applied to a decoding side (also referred to as a video decoder). FIG. 6A is a schematic structural diagram of the decoding side. The decoding side may include a control parameter decoding unit, a first feature decoding unit, a second feature decoding unit, a coefficient hyperparameter feature generation unit, an image feature inverse conversion unit, a first inverse quantization unit, a second inverse quantization unit, and a quality improvement unit. The first inverse quantization unit, the second inverse quantization unit, and the quality improvement unit are optional units, and in a specific scenario, it is also possible to turn off or skip the processes of these optional units.
[0165] In this embodiment, for each current block (i.e., an image block), the bitstream corresponding to the current block includes three parts: bitstream 0 (a bitstream including control parameters), bitstream 1 (a bitstream including coefficient hyperparameter feature information), and bitstream 2 (a bitstream including image feature information). The coefficient hyperparameter feature information and the image feature information may be collectively referred to as image information.
[0166] Exemplarily, the method for image decoding based on a neural network in this embodiment may include the following steps S51 to S58.
[0167] In step S51, decode bitstream 0 corresponding to the current block to obtain the control parameters corresponding to the current block. For example, the control parameter decoding unit may decode bitstream 0 corresponding to the current block to obtain the control parameters corresponding to the current block, that is, the control parameter decoding unit may decode the control parameters from bitstream 0. The control parameters may include the control parameters of the first inverse quantization unit, the control parameters of the second inverse quantization unit, the control parameters of the coefficient hyperparameter feature generation unit, the control parameters of the image feature inverse transformation unit, and the control parameters of the quality improvement unit. For the content of the control parameters, reference may be made to subsequent embodiments, and the description is omitted here.
[0168] In step S52, decode bitstream 1 corresponding to the current block to obtain the coefficient hyperparameter feature information corresponding to the current block. For example, the first feature decoding unit may decode bitstream 1 corresponding to the current block to obtain the coefficient hyperparameter feature information corresponding to the current block, that is, the first feature decoding unit may decode the coefficient hyperparameter feature information from bitstream 1.
[0169] In step S53, a coefficient hyperparameter feature coefficient reconstruction value is determined based on the coefficient hyperparameter feature information.
[0170] Exemplarily, when the control parameter includes first enable information corresponding to a first inverse quantization unit, the first enable information may indicate enabling the first inverse quantization unit (i.e., enabling the first inverse quantization unit to perform a first inverse quantization process), and the first enable information may also indicate not enabling the first inverse quantization unit. For example, when the first enable information is a first value, it indicates enabling the first inverse quantization unit, and when the first enable information is a second value, it indicates not enabling the first inverse quantization unit.
[0171] Exemplarily, when the first enable information indicates enabling the first inverse quantization unit, the coefficient hyperparameter feature information may be a coefficient hyperparameter feature coefficient quantization value C_q, and the first inverse quantization unit may perform inverse quantization on the coefficient hyperparameter feature coefficient quantization value C_q to obtain a coefficient hyperparameter feature coefficient reconstruction value C'. When the first enable information indicates not enabling the first inverse quantization unit, the coefficient hyperparameter feature information may be the coefficient hyperparameter feature coefficient reconstruction value C', that is, directly decode the coefficient hyperparameter feature coefficient reconstruction value C' from the bitstream 1.
[0172] In step S54, an inverse transformation process is performed on the coefficient hyperparameter feature coefficient reconstruction value C' to obtain a coefficient hyperparameter feature value P. For example, the coefficient hyperparameter feature generation unit performs an inverse transformation process on the coefficient hyperparameter feature coefficient reconstruction value C' to obtain a coefficient hyperparameter feature value P. For example, the coefficient hyperparameter feature generation unit performs an inverse transformation process on the coefficient hyperparameter feature coefficient reconstruction value C' based on a decoding neural network to obtain a coefficient hyperparameter feature value P.
[0173] In one possible embodiment, the coefficient hyperparameter feature generation unit may obtain neural network information 1 corresponding to the coefficient hyperparameter feature generation unit from the control parameter, and generate a decoding neural network 1 corresponding to the coefficient hyperparameter feature generation unit based on the neural network information 1.
[0174] Also, the coefficient hyperparameter feature generation unit may determine an input feature corresponding to the coefficient hyperparameter feature generation unit (for example, the coefficient hyperparameter feature coefficient reconstruction value C'), process the input feature based on the decoding neural network 1 (for example, perform an inverse transformation process on the coefficient hyperparameter feature coefficient reconstruction value C'), and obtain an output feature corresponding to the coefficient hyperparameter feature generation unit (for example, the coefficient hyperparameter feature value P).
[0175] Exemplarily, the neural network information 1 may include basic layer information and enhancement layer information. The coefficient hyperparameter feature generation unit may determine a basic layer corresponding to the coefficient hyperparameter feature generation unit based on the basic layer information, and determine an enhancement layer corresponding to the coefficient hyperparameter feature generation unit based on the enhancement layer information. The coefficient hyperparameter feature generation unit may also generate a decoding neural network 1 corresponding to the coefficient hyperparameter feature generation unit based on the basic layer and the enhancement layer. For example, the decoding neural network 1 may be obtained by combining the basic layer and the enhancement layer.
[0176] After obtaining the decoding neural network 1, the coefficient hyperparameter feature generation unit may perform an inverse transformation process on the coefficient hyperparameter feature coefficient reconstruction value C' via the decoding neural network 1 to obtain the coefficient hyperparameter feature value P. The inverse transformation process is not limited herein.
[0177] In step S55, the bitstream 2 corresponding to the current block is decoded to obtain the image feature information corresponding to the current block. For example, the second feature decoding unit may decode the bitstream 2 corresponding to the current block to obtain the image feature information corresponding to the current block, that is, the second feature decoding unit may decode the image feature information from the bitstream 2. When decoding the bitstream 2 corresponding to the current block, the second feature decoding unit may decode the bitstream 2 corresponding to the current block using the coefficient hyperparameter feature value P, and this decoding process is not limited.
[0178] In step S56, an image feature reconstruction value is determined based on the image feature information.
[0179] Exemplarily, when the control parameter includes the second enable information corresponding to the second inverse quantization unit, the second enable information may indicate enabling the second inverse quantization unit (that is, enabling the second inverse quantization unit to perform the second inverse quantization process), and the second enable information may also indicate not enabling the second inverse quantization unit. For example, when the second enable information is the first value, it indicates enabling the second inverse quantization unit, and when the second enable information is the second value, it indicates not enabling the second inverse quantization unit.
[0180] Exemplarily, when the second enable information indicates enabling the second inverse quantization unit, the image feature information may be the image feature quantization value F_q, and the second inverse quantization unit may obtain the image feature quantization value F_q and perform inverse quantization on the image feature quantization value F_q to obtain the image feature reconstruction value F'. When the second enable information indicates not enabling the second inverse quantization unit, the image feature information may also be the image feature reconstruction value F', that is, directly decode the image feature reconstruction value F' from the bitstream 2.
[0181] In step S57, an inverse transformation process is performed on the image feature reconstruction value F' to obtain the image low-level feature value LF. For example, the image feature inverse transformation unit performs an inverse transformation process on the image feature reconstruction value F' to obtain the image low-level feature value LF. For example, based on a decoding neural network, an inverse transformation process is performed on the image feature reconstruction value F' to obtain the image low-level feature value LF.
[0182] In one possible embodiment, the image feature inverse transformation unit may obtain the neural network information 2 corresponding to the image feature inverse transformation unit from the control parameter, and generate the decoding neural network 2 corresponding to the image feature inverse transformation unit based on the neural network information 2. The image feature inverse transformation unit determines the input feature corresponding to the image feature inverse transformation unit (for example, the image feature reconstruction value F'), processes the input feature based on the decoding neural network 2 (for example, performs an inverse transformation process on the image feature reconstruction value F'), and may obtain the output feature corresponding to the image feature inverse transformation unit (for example, the image low-level feature value LF).
[0183] Exemplarily, the neural network information 2 may include basic layer information and enhancement layer information. The image feature inverse transformation unit may determine the basic layer corresponding to the image feature inverse transformation unit based on the basic layer information, and determine the enhancement layer corresponding to the image feature inverse transformation unit based on the enhancement layer information. The image feature inverse transformation unit may generate the decoding neural network 2 corresponding to the image feature inverse transformation unit based on the basic layer and the enhancement layer. For example, the decoding neural network 2 may be obtained by combining the basic layer and the enhancement layer.
[0184] After obtaining the decoding neural network 2, the image feature inverse transformation unit may perform an inverse transformation process on the image feature reconstruction value F' through the decoding neural network 2 to obtain the image low-level feature value LF, and the inverse transformation process is not limited herein.
[0185] In step S58, based on the image low-level feature value LF, a reconstructed image block I corresponding to the current block is determined.
[0186] Exemplarily, when the control parameter includes third enable information corresponding to the high-quality unit, the third enable information may indicate enabling the high-quality unit (i.e., enabling the high-quality unit to perform high-quality processing), or the third enable information may indicate not enabling the high-quality unit. For example, when the third enable information is a first value, it may indicate enabling the high-quality unit, and when the third enable information is a second value, it may indicate not enabling the high-quality unit.
[0187] Exemplarily, when the third enable information indicates enabling the high-quality unit, the high-quality unit obtains the image low-level feature value LF, performs enhancement processing on the image low-level feature value LF, and obtains a reconstructed image block I corresponding to the current block. When the third enable information indicates not enabling the high-quality unit, the image low-level feature value LF is used as the reconstructed image block I corresponding to the current block.
[0188] Exemplarily, when the high-quality unit performs enhancement processing on the image low-level feature value LF, the high-quality unit may perform enhancement processing on the image low-level feature value LF based on the decoding neural network to obtain a reconstructed image block I corresponding to the current block.
[0189] In one possible embodiment, the high-quality unit may obtain neural network information 3 corresponding to the high-quality unit from the control parameter, and generate a decoding neural network 3 corresponding to the high-quality unit based on the neural network information 3.
[0190] The high-quality unit may determine input features corresponding to the high-quality unit (for example, the image low-level feature value LF), process the input features based on the decoding neural network 3 (for example, perform enhancement processing on the image low-level feature value LF), and obtain output features corresponding to the high-quality unit (for example, the reconstructed image block I corresponding to the current block).
[0191] Exemplarily, the neural network information 3 may include basic layer information and enhancement layer information. The high-quality unit may determine the basic layer corresponding to the high-quality unit based on the basic layer information and determine the enhancement layer corresponding to the high-quality unit based on the enhancement layer information. The high-quality unit may generate the decoding neural network 3 corresponding to the high-quality unit based on the basic layer and the enhancement layer. For example, the decoding neural network 3 may be obtained by combining the basic layer and the enhancement layer.
[0192] After obtaining the decoding neural network 3, the high-quality unit may perform enhancement processing on the image low-level feature value LF via the decoding neural network 3 to obtain the reconstructed image block I corresponding to the current block, and the enhancement processing process is not limited herein.
[0193] Example 6: The embodiments of the present invention provide an image decoding method based on a neural network, which may be applied to the decoding side. FIG. 6B is a schematic structural diagram of the decoding side. The decoding side may include a control parameter decoding unit, a first feature decoding unit, a second feature decoding unit, a coefficient hyperparameter feature generation unit, and an image feature inverse transformation unit. In this embodiment, for each current block (i.e., an image block), the bitstream corresponding to the current block includes three parts: bitstream 0 (a bitstream including control parameters), bitstream 1 (a bitstream including coefficient hyperparameter feature information), and bitstream 2 (a bitstream including image feature information).
[0194] Exemplarily, the image decoding method based on a neural network in this embodiment may include the following steps S61 to S66.
[0195] In step S61, the control parameter decoding unit decodes the bitstream 0 corresponding to the current block to obtain the control parameter corresponding to the current block.
[0196] In step S62, the first feature decoding unit decodes the bitstream 1 corresponding to the current block to obtain the coefficient hyperparameter feature information corresponding to the current block, and the coefficient hyperparameter feature information may be the coefficient hyperparameter feature coefficient reconstruction value C'.
[0197] In step S63, the coefficient hyperparameter feature generation unit performs an inverse transformation process on the coefficient hyperparameter feature coefficient reconstruction value C' to obtain the coefficient hyperparameter feature value P. For example, based on a decoding neural network, an inverse transformation process is performed on the coefficient hyperparameter feature coefficient reconstruction value C' to obtain the coefficient hyperparameter feature value P.
[0198] In step S64, the second feature decoding unit decodes the bitstream 2 corresponding to the current block to obtain the image feature information corresponding to the current block, and the image feature information may be the image feature reconstruction value F'. Exemplarily, the second feature decoding unit may obtain the coefficient hyperparameter feature value P, and use the coefficient hyperparameter feature value P to decode the bitstream 2 corresponding to the current block to obtain the image feature reconstruction value F'.
[0199] In step S65, the image feature inverse transformation unit performs an inverse transformation process on the image feature reconstruction value F' to obtain the image low-level feature value LF. For example, based on a decoding neural network, an inverse transformation process is performed on the image feature reconstruction value F' to obtain the image low-level feature value LF.
[0200] In step S66, based on the image low-level feature value LF, a reconstructed image block I corresponding to the current block is determined.
[0201] For example, the image low-level feature value LF may be directly used as the reconstructed image block I. Or, when the decoding device further includes a quality improvement unit, the quality improvement unit performs enhancement processing on the image low-level feature value LF to obtain a reconstructed image block I corresponding to the current block. For example, based on a decoding neural network, enhancement processing is performed on the image low-level feature value LF to obtain the reconstructed image block I.
[0202] Example 7: For Example 5 and Example 6, the first feature decoding unit may decode the bitstream 1 corresponding to the current block to obtain coefficient hyperparameter feature information corresponding to the current block. For example, the first feature decoding unit includes at least one coefficient decoding module. In one possible embodiment, the coefficient decoding module performs coefficient decoding using an entropy decoding method, that is, uses the entropy decoding method to decode the bitstream 1 corresponding to the current block to obtain coefficient hyperparameter feature information corresponding to the current block.
[0203] Exemplarily, the entropy decoding method may include, but is not limited to, entropy decoding methods such as CAVLC (Context-Adaptive Varialbe Length Coding) or CABAC (Context-based Adaptive Binary Arithmetic Coding), and is not limited thereto.
[0204] Exemplarily, when performing coefficient decoding using an entropy decoding method, as the probability model for entropy decoding, a preset probability model may be used. The preset probability model may be set according to actual needs, and there is no limitation on this preset probability model. For example, based on the preset probability model, the coefficient decoding module may perform coefficient decoding using the entropy decoding method.
[0205] Example 8: Regarding Example 5, the first inverse quantization unit may perform inverse quantization on the coefficient hyperparameter feature coefficient quantization value C_q (i.e., the coefficient hyperparameter feature information) to obtain the coefficient hyperparameter feature coefficient reconstruction value C'. For example, the first inverse quantization unit may not exist, or when the first inverse quantization unit exists, the first inverse quantization unit may be selectively skipped based on control parameters (e.g., high-level syntax, e.g., the first enable information, etc.), or it may be determined to enable the first inverse quantization unit based on the control parameters.
[0206] Exemplarily, when the first inverse quantization unit does not exist, the coefficient hyperparameter feature coefficient reconstruction value C' is the same as the coefficient hyperparameter feature coefficient quantization value C_q, that is, there is no need to perform inverse quantization on the coefficient hyperparameter feature coefficient quantization value C_q. When selectively skipping the first inverse quantization unit based on the control parameters, the coefficient hyperparameter feature coefficient reconstruction value C' is the same as the coefficient hyperparameter feature coefficient quantization value C_q, that is, there is no need to perform inverse quantization on the coefficient hyperparameter feature coefficient quantization value C_q. Although it is determined to enable the first inverse quantization unit based on the control parameters, when the step parameter qstep corresponding to the coefficient hyperparameter feature coefficient quantization value C_q is 1, the coefficient hyperparameter feature coefficient reconstruction value C' is the same as the coefficient hyperparameter feature coefficient quantization value C_q, that is, there is no need to perform inverse quantization on the coefficient hyperparameter feature coefficient quantization value C_q.
[0207] Exemplarily, it is determined to enable the first inverse quantization unit based on a control parameter, and when the step parameter qstep corresponding to the coefficient hyperparameter feature coefficient quantization value C_q is not 1, the first inverse quantization unit may perform inverse quantization on the coefficient hyperparameter feature coefficient quantization value C_q based on the control parameter (e.g., a quantization-related parameter) to obtain a coefficient hyperparameter feature coefficient reconstruction value C'. For example, the first inverse quantization unit performs the following processing. Obtain a quantization-related parameter corresponding to the coefficient hyperparameter feature coefficient quantization value C_q from the control parameter (the control parameter is included in the bitstream, and the control parameter may include the quantization-related parameter), for example, the step parameter qstep or the quantization parameter qp. Based on the step parameter qstep or the quantization parameter qp, determine a multiplication factor mult and a shift factor shift corresponding to the coefficient hyperparameter feature coefficient quantization value C_q. Assuming that the coefficient hyperparameter feature coefficient quantization value C_q is Coff_hyper and the coefficient hyperparameter feature coefficient reconstruction value C' is Coff_hyper_rec, then Coff_hyper_rec = (Coff_hyper * mult) << shift. As described above, when performing inverse quantization on the coefficient hyperparameter feature coefficient quantization value C_q, the coefficient hyperparameter feature coefficient reconstruction value C' may be obtained using the above formula.
[0208] Regarding the quantization-related parameter (e.g., step parameter qstep) corresponding to the coefficient hyperparameter feature quantization value C_q, it may include: 1) using the same step parameter qstep for each coefficient hyperparameter feature quantization value of each feature channel; 2) using different step parameters qstep for the coefficient hyperparameter feature quantization values of each feature channel, but using the same step parameter qstep for each coefficient hyperparameter feature quantization value within a feature channel; 3) using different step parameters qstep for each coefficient hyperparameter feature quantization value of each feature channel. In the above process, the step parameter qstep may be referred to as the quantization stride.
[0209] Example 9: For Example 5 and Example 6, the coefficient hyperparameter feature generation unit may perform an inverse transformation process on the coefficient hyperparameter feature coefficient reconstruction value C’ based on the decoding neural network to obtain the coefficient hyperparameter feature value P. In one possible embodiment, as shown in FIG. 6C, the coefficient hyperparameter feature generation unit may include a decoding neural network 1. The decoding neural network 1 may include a basic layer and an enhancement layer. The coefficient hyperparameter feature coefficient reconstruction value C’ is the input feature of the decoding neural network 1, and the coefficient hyperparameter feature value P is the output feature of the decoding neural network 1. The decoding neural network 1 is used to perform an inverse transformation process on the coefficient hyperparameter feature coefficient reconstruction value C’.
[0210] In this embodiment, the decoding neural network 1 is divided into a basic layer and an enhancement layer. The basic layer may include at least one network layer, or may not include a network layer, that is, the basic layer may be empty. The enhancement layer may include at least one network layer, or may not include a network layer, that is, the enhancement layer may be empty. For the multiple network layers in the decoding neural network 1, according to actual needs, the multiple network layers may be divided into a basic layer and an enhancement layer. For example, the first M1 network layers are used as the basic layer, and the remaining network layers are used as the enhancement layer, or the first M2 network layers are used as the enhancement layer, and the remaining network layers are used as the basic layer, or the last M3 network layers are used as the basic layer, and the remaining network layers are used as the enhancement layer, or the last M4 network layers are used as the enhancement layer, and the remaining network layers are used as the basic layer, or the network layers with odd numbers are used as the basic layer, and the remaining network layers are used as the enhancement layer, or the network layers with even numbers are used as the basic layer, and the remaining network layers are used as the enhancement layer. Of course, the above are just some examples, and the division method is not limited.
[0211] For example, a network layer with a fixed network structure can be used as the basic layer, and a network layer with a non-fixed network structure can be used as the enhancement layer. For example, for a certain network layer in the decoding neural network 1, when decoding multiple image blocks, if the network layer uses the same network structure, the network layer is regarded as a network layer with a fixed network structure and used as the basic layer. Also, for example, for a certain network layer in the decoding neural network 1, when decoding multiple image blocks, if the network layer uses different network structures, the network layer is regarded as a network layer with a non-fixed network structure and used as the enhancement layer.
[0212] Exemplarily, for the decoding neural network 1, the size of the output feature may be larger than the size of the input feature, or the size of the output feature may be equal to the size of the input feature, or the size of the output feature may be smaller than the size of the input feature.
[0213] For the decoding neural network 1, the basic layer and the enhancement layer include at least one transposed convolutional layer. For example, the basic layer includes at least one transposed convolutional layer, and the enhancement layer may include at least one transposed convolutional layer or may not include a transposed convolutional layer. Or, the enhancement layer includes at least one transposed convolutional layer, and the basic layer may include at least one transposed convolutional layer or may not include a transposed convolutional layer.
[0214] Exemplarily, the decoding neural network 1 may include, but is not limited to, a transposed convolutional layer, an activation layer, etc., and is not limited thereto. For example, the decoding neural network 1 sequentially includes one transposed convolutional layer with a stride of 2, one activation layer, one transposed convolutional layer with a stride of 1, and one activation layer. All of the above network layers may be used as basic layers. That is, the basic layer includes one transposed convolutional layer with a stride of 2, one activation layer, one transposed convolutional layer with a stride of 1, and one activation layer. In this case, the enhancement layer is empty. Of course, some of the network layers may be used as the enhancement layer, and is not limited thereto. Also, for example, the decoding neural network 1 sequentially includes one transposed convolutional layer with a stride of 2 and one transposed convolutional layer with a stride of 1. All of the above network layers may be used as basic layers. That is, the basic layer includes one transposed convolutional layer with a stride of 2 and one transposed convolutional layer with a stride of 1. In this case, the enhancement layer is empty. Of course, some of the network layers may be used as the enhancement layer, and is not limited thereto. Also, for example, the decoding neural network 1 sequentially includes one transposed convolutional layer with a stride of 2, one activation layer, one transposed convolutional layer with a stride of 2, one activation layer, one transposed convolutional layer with a stride of 1, and one activation layer. All of the above network layers may be used as basic layers. That is, the basic layer includes one transposed convolutional layer with a stride of 2, one activation layer, one transposed convolutional layer with a stride of 2, one activation layer, one transposed convolutional layer with a stride of 1, and one activation layer. In this case, the enhancement layer is empty. Of course, some of the network layers may be used as the enhancement layer, and is not limited thereto. Of course, the above are just some examples and are not limited thereto.
[0215] In one possible implementation, for the coefficient hyperparameter feature generation unit, a network layer of the default network structure (the network layer of the default network structure may be composed of at least one network layer) may be set, and all network parameters related to the network layer of the default network structure are fixed. For example, the network parameters may include at least one of the number of layers of the neural network, the transposed convolution layer flag bit, the number of transposed convolution layers, the quantization stride of each transposed convolution layer, the number of channels of each transposed convolution layer, the size of the convolution kernel, the number of filterings, the filtering size index, the filtering coefficient zero flag bit, the filtering coefficient, the number of activation layers, the activation layer flag bit, and the activation layer type, but are not limited thereto, that is, all the above network parameters are fixed. For example, in the network layer of the default network structure, the number of transposed convolution layers is fixed, the number of activation layers is fixed, the number of channels of each transposed convolution layer is fixed, the size of the convolution kernel is fixed, the filtering coefficient is fixed, etc. For example, the number of channels of the transposed convolution layer is 4, 8, 16, 32, 64, 128, or 256, etc., and the size of the convolution kernel is 1*1, 3*3, or 5×5, etc. Obviously, all network parameters in the network layer of the default network structure are fixed, and since these network parameters are known, the network layer of the default network structure can be directly obtained.
[0216] In one possible embodiment, for the coefficient hyperparameter feature generation unit, a pre-designed neural network pool may be set, and the pre-designed neural network pool may include at least one network layer of a pre-designed network structure (the network layer of the pre-designed network structure may be composed of at least one network layer), and the network parameters related to the network layer of the pre-designed network structure may all be set according to actual needs. For example, the network parameters may include at least one of the number of layers of the neural network, the transposed convolution layer flag bit, the number of transposed convolution layers, the quantization stride of each transposed convolution layer, the number of channels of each transposed convolution layer, the size of the convolution kernel, the filtering number, the filtering size index, the filtering coefficient zero flag bit, the filtering coefficient, the number of activation layers, the activation layer flag bit, the activation layer type, but are not limited thereto. That is, the above network parameters may all be set according to actual needs.
[0217] For example, the pre-designed neural network pool may include the network layer of the pre-designed network structure s1, the network layer of the pre-designed network structure s2, and the network layer of the pre-designed network structure s3. Here, for the network layer of the pre-designed network structure s1, network parameters such as the number of transposed convolution layers, the number of activation layers, the number of channels of each transposed convolution layer, the size of the convolution kernel, and the filtering coefficient may be set in advance. After all the network parameter settings are completed, the network layer of the pre-designed network structure s1 can be obtained. Similarly, the network layer of the pre-designed network structure s2 and the network layer of the pre-designed network structure s3 can be obtained, which will not be repeated here.
[0218] In one possible embodiment, based on network parameters, for the coefficient hyperparameter feature generation unit, a network layer of a variable network structure (the network layer of the variable network structure may be composed of at least one network layer) may be dynamically generated. The network parameters related to the network layer of the variable network structure are not preset but are dynamically generated by the encoding side. For example, the encoding side may send the network parameters corresponding to the coefficient hyperparameter feature generation unit to the decoding side. The network parameters may include at least one of the number of layers of the neural network, the transposed convolution layer flag bit, the number of transposed convolution layers, the quantization stride of each transposed convolution layer, the number of channels of each transposed convolution layer, the size of the convolution kernel, the number of filterings, the filtering size index, the filtering coefficient zero flag bit, the filtering coefficient, the number of activation layers, the activation layer flag bit, the activation layer type, etc., but are not limited thereto. The decoding side dynamically generates a network layer of a variable network structure based on the above network parameters. For example, the encoding side encodes network parameters such as the number of transposed convolution layers, the number of activation layers, the number of channels of each transposed convolution layer, the size of the convolution kernel, the filtering coefficient, etc. into the bitstream, and the decoding side analyzes the network parameters from the bitstream and may generate a network layer of a variable network structure based on these network parameters. The generation process is not limited herein.
[0219] In one possible embodiment, the decoding neural network 1 may be divided into a basic layer and an enhancement layer. Regarding the combination method of the basic layer and the enhancement layer, the following methods may be included, but are not limited thereto. Method 1: The basic layer uses the network layer of the default network structure, the enhancement layer uses the network layer of the default network structure, and the basic layer and the enhancement layer are combined to form the decoding neural network 1. Method 2: The basic layer uses the network layer of the default network structure, the enhancement layer uses the network layer of the pre-designed network structure, and the basic layer and the enhancement layer are combined to form the decoding neural network 1. Method 3: The basic layer uses the network layer of the default network structure, the enhancement layer uses the network layer of the variable network structure, and the basic layer and the enhancement layer are combined to form the decoding neural network 1. Method 4: The basic layer uses the network layer of the pre-designed network structure, the enhancement layer uses the network layer of the default network structure, and the basic layer and the enhancement layer are combined to form the decoding neural network 1. Method 5: The basic layer uses the network layer of the pre-designed network structure, the enhancement layer uses the network layer of the pre-designed network structure, and the basic layer and the enhancement layer are combined to form the decoding neural network 1. Method 6: The basic layer uses the network layer of the pre-designed network structure, the enhancement layer uses the network layer of the variable network structure, and the basic layer and the enhancement layer are combined to form the decoding neural network 1.
[0220] In one possible embodiment, the control parameter may include neural network information 1 corresponding to the coefficient hyperparameter feature generation unit, and the coefficient hyperparameter feature generation unit may analyze the neural network information 1 from the control parameter and generate a decoding neural network 1 based on the neural network information 1. For example, the neural network information 1 may include basic layer information and enhancement layer information. Based on the basic layer information, a basic layer may be determined, and based on the enhancement layer information, an enhancement layer may be determined. The basic layer and the enhancement layer may be combined to obtain the decoding neural network 1. For example, the neural network information 1 corresponding to the coefficient hyperparameter feature generation unit may be obtained using the following cases.
[0221] Case 1: The basic layer information includes a basic layer default network usage flag bit, and the basic layer default network usage flag bit indicates that the basic layer uses the default network. In this case, the coefficient hyperparameter feature generation unit knows, based on the basic layer information, that the basic layer uses the network layer of the default network structure, and thus obtains the basic layer of the default network structure (i.e., the network layer of the default network structure). The enhancement layer information includes an enhancement layer default network usage flag bit, and the enhancement layer default network usage flag bit indicates that the enhancement layer uses the default network. In this case, the coefficient hyperparameter feature generation unit knows, based on the enhancement layer information, that the enhancement layer uses the network layer of the default network structure, and thus obtains the enhancement layer of the default network structure (i.e., the network layer of the default network structure). Based on this, the basic layer of the default network structure and the enhancement layer of the default network structure may be combined to obtain the decoding neural network 1.
[0222] Case 2: The basic layer information includes a basic layer default network usage flag bit, and the basic layer default network usage flag bit indicates that the basic layer uses the default network. In this case, the coefficient hyperparameter feature generation unit knows, based on the basic layer information, that the basic layer uses a network layer with a default network structure, and thus obtains the basic layer of the default network structure (i.e., the network layer of the default network structure). The enhancement layer information includes an enhancement layer pre-designed network usage flag bit and an enhancement layer pre-designed network index number, and the enhancement layer pre-designed network usage flag bit indicates that the enhancement layer uses the pre-designed network. In this case, the coefficient hyperparameter feature generation unit knows, based on the enhancement layer information, that the enhancement layer uses a network layer with a pre-designed network structure, and thus selects, from the pre-designed neural network pool, the enhancement layer of the pre-designed network structure corresponding to the enhancement layer pre-designed network index number (for example, when the enhancement layer pre-designed network index number is 0, the network layer of the pre-designed network structure s1 is used as the enhancement layer, and when the enhancement layer pre-designed network index number is 1, the network layer of the pre-designed network structure s2 is used as the enhancement layer). Based on this, a decoding neural network 1 may be obtained by combining the basic layer of the default network structure and the enhancement layer of the pre-designed network structure.
[0223] Case 3: The basic layer information includes a basic layer default network usage flag bit, and the basic layer default network usage flag bit indicates that the basic layer uses the default network. In this case, the coefficient hyperparameter feature generation unit knows, based on the basic layer information, that the basic layer uses a network layer with a default network structure. Therefore, the basic layer of the default network structure (i.e., the network layer of the default network structure) is obtained. The enhanced layer information includes network parameters for generating an enhanced layer. In this case, the coefficient hyperparameter feature generation unit may analyze the network parameters from the control parameters and generate an enhanced layer with a variable network structure (i.e., the network layer of the variable network structure) based on the network parameters. Here, the network parameters may include at least one of the number of layers of the neural network, the transposed convolution layer flag bit, the number of transposed convolution layers, the quantization stride of each transposed convolution layer, the number of channels of each transposed convolution layer, the size of the convolution kernel, the number of filterings, the filtering size index, the filtering coefficient zero flag bit, the filtering coefficient, the number of activation layers, the activation layer flag bit, and the activation layer type, but are not limited thereto. The coefficient hyperparameter feature generation unit may generate an enhanced layer with a variable network structure based on the above network parameters. For example, the coefficient hyperparameter feature generation unit analyzes network parameters such as the number of transposed convolution layers, the quantization stride (stride) of each transposed convolution layer, the number of channels of each transposed convolution layer, the size of the convolution kernel, and the activation layer type from the control parameters, and the coefficient hyperparameter feature generation unit generates an enhanced layer with a variable network structure based on these network parameters. Based on this, the coefficient hyperparameter feature generation unit may combine the basic layer of the default network structure and the enhanced layer of the variable network structure to obtain the decoding neural network 1.
[0224] Case 4: The basic layer information includes a basic layer pre - design network usage flag bit and a basic layer pre - design network index number. The basic layer pre - design network usage flag bit indicates that the basic layer uses the pre - design network. In this case, the coefficient hyper - parameter feature generation unit knows, based on the basic layer information, that the basic layer uses the network layer of the pre - design network structure. Therefore, it selects the basic layer of the pre - design network structure corresponding to the basic layer pre - design network index number from the pre - design neural network pool (for example, when the basic layer pre - design network index number is 0, the network layer of the pre - design network structure s1 is used as the basic layer; when the basic layer pre - design network index number is 1, the network layer of the pre - design network structure s2 is used as the basic layer). The enhancement layer information includes an enhancement layer default network usage flag bit, and the enhancement layer default network usage flag bit indicates that the enhancement layer uses the default network. In this case, the coefficient hyper - parameter feature generation unit knows, based on the enhancement layer information, that the enhancement layer uses the network layer of the default network structure. Therefore, it obtains the enhancement layer of the default network structure (that is, the network layer of the default network structure). Based on this, a decoding neural network 1 may be obtained by combining the basic layer of the pre - design network structure and the enhancement layer of the default network structure.
[0225] Case 5: The basic layer information includes a basic layer pre-designed network usage flag bit and a basic layer pre-designed network index number, and the basic layer pre-designed network usage flag bit indicates that the basic layer uses the pre-designed network. In this case, the coefficient hyperparameter feature generation unit knows, based on the basic layer information, that the basic layer uses the network layer of the pre-designed network structure. Therefore, it may select, from the pre-designed neural network pool, the basic layer of the pre-designed network structure corresponding to the basic layer pre-designed network index number (for example, the network layer of the pre-designed network structure s1, etc.). The enhancement layer information includes an enhancement layer pre-designed network usage flag bit and an enhancement layer pre-designed network index number, and the enhancement layer pre-designed network usage flag bit indicates that the enhancement layer uses the pre-designed network. In this case, the coefficient hyperparameter feature generation unit knows, based on the enhancement layer information, that the enhancement layer uses the network layer of the pre-designed network structure. Therefore, it may select, from the pre-designed neural network pool, the enhancement layer of the pre-designed network structure corresponding to the enhancement layer pre-designed network index number (for example, the network layer of the pre-designed network structure s1, etc.). Based on this, the basic layer of the pre-designed network structure and the enhancement layer of the pre-designed network structure may be combined to obtain the decoding neural network 1.
[0226] Case 6: The basic layer information includes a basic layer pre - design network usage flag bit and a basic layer pre - design network index number, and the basic layer pre - design network usage flag bit indicates that the basic layer uses the pre - design network. In this case, the coefficient hyper - parameter feature generation unit knows, based on the basic layer information, that the basic layer uses the network layer of the pre - design network structure. Therefore, it may select, from the pre - design neural network pool, the basic layer of the pre - design network structure corresponding to the basic layer pre - design network index number (for example, the network layer of the pre - design network structure s1, etc.). The enhancement layer information includes network parameters for generating the enhancement layer. In this case, the coefficient hyper - parameter feature generation unit may analyze the network parameters from the control parameters and generate an enhancement layer of a variable network structure (i.e., a network layer of a variable network structure) based on the network parameters. For example, the coefficient hyper - parameter feature generation unit may analyze network parameters such as the number of de - convolutional layers, the quantization stride of each de - convolutional layer, the number of channels of each de - convolutional layer, the size of the convolutional kernel, and the activation layer type from the control parameters, and the coefficient hyper - parameter feature generation unit generates an enhancement layer of a variable network structure based on these network parameters. Based on this, the coefficient hyper - parameter feature generation unit may combine the basic layer of the pre - design network structure and the enhancement layer of the variable network structure to obtain the decoding neural network 1.
[0227] In one possible embodiment, for the above - mentioned network structures of the basic layer and the enhancement layer, both may be determined by the decoded control parameters. The control parameters may include neural network information 1, and the neural network information 1 may include basic layer information and enhancement layer information. An example of the neural network information 1 for the coefficient hyper - parameter feature generation unit is shown in Table 1. In Table 1, u(n) represents an n - bit fixed - length coding method, and ae(v) represents a variable - length coding method.
Table 1
[0228] In Table 1, hyper_basic_layer_use_default_para_flag is the basic layer default network usage flag bit for the coefficient hyperparameter feature generation unit. hyper_basic_layer_use_default_para_flag is a binary variable. A value of 1 indicates that the basic layer of the coefficient hyperparameter feature generation unit uses the default network, and a value of 0 indicates that the basic layer of the coefficient hyperparameter feature generation unit does not use the default network. The value of HyperBasicLayerUseDefaultParaFlag may be equal to the value of hyper_basic_layer_use_default_para_flag.
[0229] In Table 1, hyper_basic_layer_use_predesigned_para_flag is the basic layer pre-designed network usage flag bit for the coefficient hyperparameter feature generation unit. hyper_basic_layer_use_predesigned_para_flag is a binary variable. When the value of this binary variable is 1, it indicates that the basic layer of the coefficient hyperparameter feature generation unit uses the pre-designed network. When the value of this binary variable is 0, it indicates that the basic layer of the coefficient hyperparameter feature generation unit does not use the pre-designed network. The value of HyperBasicLayerUsePredesignedParaFlag may be equal to the value of hyper_basic_layer_use_predesigned_para_flag.
[0230] In Table 1, hyper_basic_id is the basic layer pre - design network index number for the coefficient hyper - parameter feature generation unit. It may be a 32 - bit unsigned integer and indicates the index number in the pre - design neural network pool of the neural network used for the basic layer.
[0231] In Table 1, hyper_enhance_layer_use_default_para_flag is the enhancement layer default network usage flag bit for the coefficient hyper - parameter feature generation unit. hyper_enhance_layer_use_default_para_flag is a binary variable. When the value of this binary variable is 1, it indicates that the enhancement layer of the coefficient hyper - parameter feature generation unit uses the default network. When the value of this binary variable is 0, it indicates that the enhancement layer of the coefficient hyper - parameter feature generation unit does not use the default network. The value of HyperEnhanceLayerUseDefaultParaFlag may be equal to the value of hyper_enhance_layer_use_default_para_flag.
[0232] In Table 1, hyper_enhance_layer_use_predesigned_para_flag is the enhancement layer pre - design network usage flag bit for the coefficient hyper - parameter feature generation unit. hyper_enhance_layer_use_predesigned_para_flag is a binary variable. When the value of this binary variable is 1, it indicates that the enhancement layer of the coefficient hyper - parameter feature generation unit uses the pre - design network. When the value of this binary variable is 0, it indicates that the enhancement layer of the coefficient hyper - parameter feature generation unit does not use the pre - design network. The value of HyperEnhanceLayerUsePredesignedParaFlag may be equal to the value of hyper_enhance_layer_use_predesigned_para_flag.
[0233] In Table 1, hyper_enhance_id is the enhancement layer pre - design network index number for the coefficient hyper - parameter feature generation unit, which may be a 32 - bit unsigned integer, and indicates the index number in the pre - design neural network pool of the neural network used in the enhancement layer.
[0234] In the above process, the range of hyper_basic_id is [id_min, id_max], where id_min is preferably 0 and id_max is preferably 2^32 - 1. The segment [a, b] is a reserved segment for the later expansion of the pre - design neural network pool. For the pre - design neural network pool of the basic layer, the pre - design neural network pool may include multiple basic layer pre - design networks, such as two, three, four, etc., or dozens of basic layer pre - design networks, or even more basic layer pre - design networks, and is not limited thereto. Obviously, the fact that id_max is preferably 2^32 - 1 is only an example, and in some cases, the value of id_max may be dynamically adjusted.
[0235] In the above process, the range of hyper_enhance_id is [id_min, id_max], where id_min is preferably 0 and id_max is preferably 2^32 - 1. The segment [a, b] is a reserved segment for the later expansion of the pre - design neural network pool. For the pre - design neural network pool of the enhancement layer, the pre - design neural network pool may include multiple enhancement layer pre - design networks, such as two, three, four, etc., or dozens of enhancement layer pre - design networks, or even more enhancement layer pre - design networks, and is not limited thereto. Obviously, the fact that id_max is preferably 2^32 - 1 is only an example, and in some cases, the value of id_max may be dynamically adjusted.
[0236] In one possible embodiment, the network structure of the reinforcement layer may be determined by the decoded control parameters. The control parameters may include neural network information 1. The neural network information 1 may include reinforcement layer information. An example of the neural network information 1 for the coefficient hyperparameter feature generation unit is shown in Table 2 and Table 3. [Table 2] [Table 3]
[0237] In Table 2, layer_num indicates the number of layers of the neural network and is used to indicate the number of network layers of the neural network. When included in a network structure with an activation layer, the number of layers is not counted extra, and the value of LayerNum is equal to layer_num.
[0238] In Table 3, deconv_layer_flag indicates the transposed convolution layer flag bit, and deconv_layer_flag is a binary variable. When the value of the binary variable is 1, it indicates that the current layer is a transposed convolution layer network. When the value of the binary variable is 0, it indicates that the current layer is not a transposed convolution layer network, and the value of DeconvLayerFlag is equal to the value of deconv_layer_flag.
[0239] In Table 3, stride_num indicates the quantization stride of the transposed convolution layer.
[0240] In Table 3, filter_num indicates the number of filterings, that is, the number of filterings of the current layer.
[0241] In Table 3, filter_size_index indicates the filtering size index, that is, the value of the current filtering size index.
[0242] In Table 3, filter_coeff_zero_flag[i][j] indicates the filtering coefficient zero flag bit and is a binary variable. When the value of the binary variable is 1, it indicates that the current filtering coefficient is 0. When the value of the binary variable is 0, it indicates that the current filtering coefficient is not 0, and the value of FilterCoeffZeroFlag[i][j] is equal to the value of filter_coeff_zero_flag[i][j].
[0243] In Table 3, filter_coeff[i][j] indicates the filtering coefficient, that is, the current filtering coefficient value.
[0244] In Table 3, activation_layer_flag indicates the activation layer flag bit, and activation_layer_flag is a binary variable. When the value of the binary variable is 1, it indicates that the current layer is an activation layer. When the value of the binary variable is 0, it indicates that the current layer is not an activation layer, and the value of ActivationLayerFlag is equal to the value of activation_layer_flag.
[0245] In Table 3, activation_layer_type indicates the activation layer type, that is, the specific type of the activation layer of the current layer.
[0246] Example 10: For Example 5 and Example 6, the second feature decoding unit may decode the bitstream 2 corresponding to the current block to obtain the image feature information corresponding to the current block. For example, the second feature decoding unit may include at least one coefficient decoding module and one probability model acquisition module. In one possible embodiment, the coefficient decoding module may perform coefficient decoding using the entropy decoding method, that is, use the entropy decoding method to decode the bitstream 2 corresponding to the current block to obtain the image feature information corresponding to the current block.
[0247] Exemplarily, the entropy decoding method may include, but is not limited to, CAVLC or CABAC, etc., and is not limited thereto.
[0248] Exemplarily, the method of using the features generated by the coefficient hyperparameter feature generation unit in the second feature decoding unit may include the following.
[0249] Method 1: The probability model acquisition module is used to acquire the probability model of entropy decoding. For example, the probability model acquisition module acquires the coefficient hyperparameter feature value P from the coefficient hyperparameter feature generation unit. Based on this, the coefficient decoding module can acquire the coefficient hyperparameter feature value P from the probability model acquisition module, and based on the coefficient hyperparameter feature value P, the coefficient decoding module can perform coefficient decoding using the entropy decoding method.
[0250] Method 2: The coefficient analysis process (for example, the CABAC or CAVCL decoding process) does not depend on the features generated by the coefficient hyperparameter feature generation unit and can directly analyze to obtain the coefficient value (in this way, the analysis throughput or rate can be guaranteed). Based on the features generated by the coefficient hyperparameter feature generation unit, the analyzed coefficient value is converted to obtain the image feature quantization value F_q. As an example, when the coefficient value obtained in the coefficient analysis process is 0 and the corresponding feature value generated by the coefficient hyperparameter feature generation unit is u, F_q = u; when the coefficient value obtained in the coefficient analysis process is 1 and the corresponding feature value generated by the coefficient hyperparameter feature generation unit is u, F_q = u + x, where x is the corresponding coefficient variance.
[0251] Example 11: For Example 5, the second inverse quantization unit may perform inverse quantization on the image feature quantization value F_q (i.e., the image feature information) to obtain the image feature reconstruction value F'. For example, the second inverse quantization unit may not exist, or when the second inverse quantization unit exists, the second inverse quantization unit may be selectively skipped based on control parameters (e.g., high-level syntax, e.g., the second enable information, etc.), or it may be determined to enable the second inverse quantization unit based on the control parameters. Exemplarily, when the second inverse quantization unit does not exist, the image feature reconstruction value F' is the same as the image feature quantization value F_q, that is, it is not necessary to perform inverse quantization on the image feature quantization value F_q. When the second inverse quantization unit is selectively skipped based on the control parameters, the image feature reconstruction value F' is the same as the image feature quantization value F_q, that is, it is not necessary to perform inverse quantization on the image feature quantization value F_q. Although it is determined to enable the second inverse quantization unit based on the control parameters, when the step parameter qstep corresponding to the image feature quantization value F_q is 1, the image feature reconstruction value F' is the same as the image feature quantization value F_q, that is, it is not necessary to perform inverse quantization on the image feature quantization value F_q.
[0252] Exemplarily, it is determined to enable the second inverse quantization unit based on control parameters, and when the step parameter qstep corresponding to the image feature quantization value Fq is not 1, the second inverse quantization unit may perform inverse quantization on the image feature quantization value Fq based on the control parameters (for example, quantization-related parameters) to obtain an image feature reconstruction value F'. For example, the second inverse quantization unit performs the following processing. Obtain the quantization-related parameter (the control parameter may be included in the bitstream and the control parameter may include the quantization-related parameter) corresponding to the image feature quantization value Fq from the control parameters, for example, the step parameter qstep or the quantization parameter qp. Based on the step parameter qstep or the quantization parameter qp, determine the multiplication factor mult and the shift factor shift corresponding to the image feature quantization value Fq. Assuming that the image feature quantization value Fq is Coff_hyper and the image feature reconstruction value F' is Coff_hyper_rec, then Coff_hyper_rec = (Coff_hyper * mult) << shift, that is, when performing inverse quantization, the image feature reconstruction value F' may be obtained using the above formula.
[0253] Regarding the quantization-related parameter (for example, the step parameter qstep) corresponding to the image feature quantization value Fq, it includes: 1) all the image feature quantization values of each feature channel use the same step parameter qstep; 2) the image feature quantization values of each feature channel use different step parameters qstep, but each image feature quantization value within the feature channel uses the same step parameter qstep; 3) all the image feature quantization values of each feature channel use different step parameters qstep.
[0254] Example 12: For Example 5 and Example 6, the image feature inverse transformation unit may perform an inverse transformation process on the image feature reconstruction value F' based on the decoding neural network to obtain the image low-level feature value LF. In one possible embodiment, as shown in FIG. 6D, the image feature inverse transformation unit may include a decoding neural network 2. The decoding neural network 2 may include a basic layer and an enhancement layer. The image feature reconstruction value F' is the input feature of the decoding neural network 2, and the image low-level feature value LF is the output feature of the decoding neural network 2. The decoding neural network 2 is used to perform an inverse transformation process on the image feature reconstruction value F'.
[0255] In this embodiment, the decoding neural network 2 may be divided into a basic layer and an enhancement layer. The basic layer may include at least one network layer, or the basic layer may not include a network layer, that is, the basic layer may be empty. The enhancement layer may include at least one network layer, or the enhancement layer may not include a network layer, that is, the enhancement layer may be empty. For the multiple network layers in the decoding neural network 2, according to actual needs, the multiple network layers may be divided into a basic layer and an enhancement layer. For example, the network layer with a fixed network structure may be used as the basic layer, and the network layer with a non-fixed network structure may be used as the enhancement layer.
[0256] Exemplarily, for the decoding neural network 2, the size of the output feature may be larger than the size of the input feature, or the size of the output feature may be equal to the size of the input feature, or the size of the output feature may be smaller than the size of the input feature.
[0257] For the decoding neural network 2, the basic layer and the enhancement layer include at least one transposed convolutional layer. For example, the basic layer includes at least one transposed convolutional layer, and the enhancement layer may or may not include at least one transposed convolutional layer. Alternatively, the enhancement layer includes at least one transposed convolutional layer, and the basic layer may or may not include at least one transposed convolutional layer.
[0258] For the decoding neural network 2, the basic layer and the enhancement layer include at least one residual structure layer. For example, the basic layer includes at least one residual structure layer, and the enhancement layer may or may not include a residual structure layer. Alternatively, the enhancement layer includes at least one residual structure layer, and the basic layer may or may not include a residual structure layer.
[0259] Exemplarily, for the decoding neural network 2, it may include, but is not limited to, a transposed convolutional layer, an activation layer, etc., and is not limited thereto. For example, the decoding neural network 2 sequentially includes one transposed convolutional layer with a stride of 2, one activation layer, one transposed convolutional layer with a stride of 1, and one activation layer. All of the above network layers are used as the basic layer, and the enhancement layer may be empty. Also, for example, the decoding neural network 2 sequentially includes one transposed convolutional layer with a stride of 2 and one transposed convolutional layer with a stride of 1. All of the above network layers are used as the basic layer, and the enhancement layer may be empty. Also, for example, the decoding neural network 2 sequentially includes one transposed convolutional layer with a stride of 2, one activation layer, one transposed convolutional layer with a stride of 2, one activation layer, one transposed convolutional layer with a stride of 1, and one activation layer. All of the above network layers are used as the basic layer, and the enhancement layer may be empty.
[0260] Exemplarily, when there is no quality improvement unit, the number of output features (number of filters) of the last network layer of the image feature inverse transformation unit is 1 or 3. Specifically, when the output is only the value of one channel (for example, a grayscale image), the number of output features of the last network layer of the image feature inverse transformation unit is 1, and when the output is only the values of three channels (for example, RGB or YUV format), the number of output features of the last network layer of the image feature inverse transformation unit is 3.
[0261] Exemplarily, when there is a quality improvement unit, the number of output features (number of filters) of the last network layer of the image feature inverse transformation unit may be 1 or 3, or may be other values, and is not limited thereto.
[0262] In one possible embodiment, for the image feature inverse transformation unit, the network layers of the default network structure may be set, and all network parameters related to the network layers of the default network structure are fixed. For example, in the network layers of the default network structure, the number of transposed convolution layers is fixed, the number of activation layers is fixed, the number of channels of each transposed convolution layer is fixed, the size of the convolution kernel is fixed, the filtering coefficient is fixed, and so on. Obviously, all network parameters in the network layers of the default network structure are fixed, and since these network parameters are known, the network layers of the default network structure can be directly obtained.
[0263] In one possible embodiment, for the image feature inverse transformation unit, a pre-designed neural network pool may be set up. The pre-designed neural network pool includes network layers of at least one pre-designed network structure, and the network parameters related to the network layers of the pre-designed network structure may all be set according to actual needs. For example, the pre-designed neural network pool includes network layers of pre-designed network structure t1, network layers of pre-designed network structure t2, and network layers of pre-designed network structure t3. Here, for the network layers of pre-designed network structure t1, network parameters such as the number of deconvolution layers, the number of activation layers, the number of channels of each deconvolution layer, the size of the convolution kernel, and the filtering coefficient may be set in advance. After all the network parameter settings are completed, the network layers of pre-designed network structure t1 can be obtained, and the same applies hereinafter.
[0264] In one possible embodiment, based on the network parameters, for the image feature inverse transformation unit, network layers of a variable network structure may be dynamically generated. The network parameters related to the network layers of the variable network structure are not pre-set but are dynamically generated by the encoding side. For example, the encoding side may encode network parameters such as the number of deconvolution layers, the number of activation layers, the number of channels of each deconvolution layer, the size of the convolution kernel, and the filtering coefficient into the bitstream. Therefore, the decoding side may analyze the above network parameters from the bitstream and generate network layers of the variable network structure based on these network parameters.
[0265] In one possible embodiment, the control parameter may include neural network information 2 corresponding to the image feature inverse transformation unit, and the image feature inverse transformation unit may analyze the neural network information 2 from the control parameter and generate a decoding neural network 2 based on the neural network information 2. For example, the neural network information 2 may include basic layer information and enhancement layer information. The basic layer may be determined based on the basic layer information, and the enhancement layer may be determined based on the enhancement layer information. The basic layer and the enhancement layer may be combined to obtain the decoding neural network 2. For example, the neural network information 2 may be obtained using the following cases.
[0266] Case 1: The basic layer information includes a basic layer default network usage flag bit, and the basic layer default network usage flag bit indicates that the basic layer uses the default network. In this case, based on the basic layer information, the image feature inverse transformation unit knows that the basic layer uses the network layer of the default network structure, and thus obtains the basic layer of the default network structure. The enhancement layer information includes an enhancement layer default network usage flag bit, and the enhancement layer default network usage flag bit indicates that the enhancement layer uses the default network. In this case, based on the enhancement layer information, it is known that the enhancement layer uses the network layer of the default network structure, and thus the enhancement layer of the default network structure is obtained. The basic layer of the default network structure and the enhancement layer of the default network structure are combined to obtain the decoding neural network 2.
[0267] Case 2: The basic layer information includes a basic layer default network usage flag bit, and the basic layer default network usage flag bit indicates that the basic layer uses the default network. In this case, based on the basic layer information, the image feature inverse transformation unit knows that the basic layer uses the network layer of the default network structure, and thus obtains the basic layer of the default network structure. The enhancement layer information includes an enhancement layer pre-designed network usage flag bit and an enhancement layer pre-designed network index number, and the enhancement layer pre-designed network usage flag bit indicates that the enhancement layer uses the pre-designed network. In this case, based on the enhancement layer information, it is known that the enhancement layer uses the network layer of the pre-designed network structure, and thus an enhancement layer of the pre-designed network structure corresponding to the enhancement layer pre-designed network index number is selected from the pre-designed neural network pool. The basic layer of the default network structure and the enhancement layer of the pre-designed network structure may be combined to obtain the decoding neural network 2.
[0268] Case 3: The basic layer information includes a basic layer default network usage flag bit, and the basic layer default network usage flag bit indicates that the basic layer uses the default network. The image feature inverse transformation unit knows that the basic layer uses the network layer of the default network structure based on the basic layer information, and thus obtains the basic layer of the default network structure. The enhancement layer information includes network parameters for generating the enhancement layer. In this case, the network parameters are analyzed from the control parameters, and an enhancement layer of a variable network structure is generated based on the network parameters. The basic layer of the default network structure and the enhancement layer of the variable network structure are combined to obtain the decoding neural network 2.
[0269] Case 4: The basic layer information includes a basic layer pre-designed network usage flag bit and a basic layer pre-designed network index number, and the basic layer pre-designed network usage flag bit indicates that the basic layer uses the pre-designed network. In this case, based on the basic layer information, the image feature inverse transformation unit knows that the basic layer uses the network layer of the pre-designed network structure. Therefore, it selects the basic layer of the pre-designed network structure corresponding to the basic layer pre-designed network index number from the pre-designed neural network pool. The enhancement layer information includes an enhancement layer default network usage flag bit, and the enhancement layer default network usage flag bit indicates that the enhancement layer uses the default network. In this case, based on the enhancement layer information, it knows that the enhancement layer uses the network layer of the default network structure. Therefore, it obtains the enhancement layer of the default network structure. The basic layer of the pre-designed network structure and the enhancement layer of the default network structure may be combined to obtain the decoding neural network 2.
[0270] Case 5: The basic layer information includes a basic layer pre-designed network usage flag bit and a basic layer pre-designed network index number, and the basic layer pre-designed network usage flag bit indicates that the basic layer uses the pre-designed network. In this case, based on the basic layer information, the image feature inverse conversion unit knows that the basic layer uses the network layer of the pre-designed network structure. Therefore, it may select the basic layer of the pre-designed network structure corresponding to the basic layer pre-designed network index number from the pre-designed neural network pool. The enhancement layer information includes an enhancement layer pre-designed network usage flag bit and an enhancement layer pre-designed network index number, and the enhancement layer pre-designed network usage flag bit indicates that the enhancement layer uses the pre-designed network. In this case, based on the enhancement layer information, it knows that the enhancement layer uses the network layer of the pre-designed network structure. Therefore, it may select the enhancement layer of the pre-designed network structure corresponding to the enhancement layer pre-designed network index number from the pre-designed neural network pool. Combining the basic layer of the pre-designed network structure and the enhancement layer of the pre-designed network structure, a decoding neural network 2 may be obtained.
[0271] Case 6: The basic layer information includes a basic layer pre-designed network usage flag bit and a basic layer pre-designed network index number, and the basic layer pre-designed network usage flag bit indicates that the basic layer uses the pre-designed network. In this case, the image feature inverse conversion unit knows, based on the basic layer information, that the basic layer uses the network layer of the pre-designed network structure. Therefore, it may select the basic layer of the pre-designed network structure corresponding to the basic layer pre-designed network index number from the pre-designed neural network pool. The enhancement layer information includes network parameters for generating the enhancement layer. In this case, the network parameters are analyzed from the control parameters, and an enhancement layer with a variable network structure is generated based on the network parameters. The basic layer of the pre-designed network structure and the enhancement layer of the variable network structure are combined to obtain the decoding neural network 2.
[0272] In one possible embodiment, for the above network structures of the basic layer and the enhancement layer, both may be determined by the decoded control parameters. The control parameters may include neural network information 2, and the neural network information 2 may include basic layer information and enhancement layer information. The neural network information 2 for the image feature inverse conversion unit is similar to Tables 1, 2, and 3, but the relevant information is for the image feature inverse conversion unit instead of the coefficient hyperparameter feature generation unit, and will not be repeated here.
[0273]
[0274] For example, the high-quality unit may not exist, or if the high-quality unit exists, the high-quality unit may be selectively skipped based on control parameters (e.g., high-level syntax, e.g., the third enable information, etc.), or it may be determined to enable the high-quality unit based on the control parameters. Exemplarily, when it is determined to enable the high-quality unit based on the control parameters, the high-quality unit may be used to remove problems of image quality degradation such as blocking artifacts between blocks and quantization distortion. For example, the high-quality unit may perform an enhancement process on the image low-level feature value LF based on a decoded neural network to obtain a reconstructed image block I corresponding to the current block.
[0275] In one possible embodiment, the high-quality unit may include a decoded neural network 3. The decoded neural network 3 may include a basic layer and an enhancement layer. The image low-level feature value LF is an input feature of the decoded neural network 3, and the reconstructed image block I is an output feature of the decoded neural network 3. The decoded neural network 3 is used to perform an enhancement process on the image low-level feature value LF.
[0276] In this embodiment, the decoded neural network 3 may be divided into a basic layer and an enhancement layer. The basic layer may include at least one network layer, or the basic layer may not include a network layer, that is, the basic layer may be empty. The enhancement layer may include at least one network layer, or the enhancement layer may not include a network layer, that is, the enhancement layer may be empty. For the multiple network layers in the decoded neural network 3, according to actual needs, the multiple network layers may be divided into a basic layer and an enhancement layer. For example, a network layer with a fixed network structure may be used as the basic layer, and a network layer with a non-fixed network structure may be used as the enhancement layer.
[0277] Exemplarily, for the decoding neural network 3, the size of the output feature may be larger than the size of the input feature, or the size of the output feature may be equal to the size of the input feature, or the size of the output feature may be smaller than the size of the input feature.
[0278] For the decoding neural network 3, the basic layer and the enhancement layer include at least one transposed convolutional layer. For example, the basic layer includes at least one transposed convolutional layer, and the enhancement layer may or may not include at least one transposed convolutional layer. Or, the enhancement layer includes at least one transposed convolutional layer, and the basic layer may or may not include at least one transposed convolutional layer.
[0279] For the decoding neural network 3, the basic layer and the enhancement layer include at least one residual structure layer. For example, the basic layer includes at least one residual structure layer, and the enhancement layer may or may not include a residual structure layer. Or, the enhancement layer includes at least one residual structure layer, and the basic layer may or may not include a residual structure layer.
[0280] Exemplarily, for the decoding neural network 3, it may include a transposed convolutional layer, an activation layer, etc., but is not limited thereto, and is not limited in this regard. For example, the decoding neural network 3 sequentially includes one transposed convolutional layer with a stride of 2, one activation layer, one transposed convolutional layer with a stride of 1, and one activation layer. All the above network layers are used as the basic layer, and the enhancement layer may be empty. Also, for example, the decoding neural network 3 sequentially includes one transposed convolutional layer with a stride of 2 and one transposed convolutional layer with a stride of 1. All the above network layers are used as the basic layer, and the enhancement layer may be empty. Also, for example, the decoding neural network 3 sequentially includes one transposed convolutional layer with a stride of 2, one activation layer, one transposed convolutional layer with a stride of 2, one activation layer, one transposed convolutional layer with a stride of 1, and one activation layer. All the above network layers are used as the basic layer, and the enhancement layer may be empty.
[0281] Exemplarily, the number of output feature numbers (filter numbers) of the last network layer of the high-quality unit is 1 or 3. Specifically, when the output is only the value of one channel (for example, a grayscale image), the number of output features of the last network layer is 1, and when the output is only the values of three channels (for example, RGB or YUV format), the number of output features of the last network layer is 3.
[0282] In one possible embodiment, for the high-quality unit, the network layers of the default network structure may be set, and all network parameters related to the network layers of the default network structure are fixed. For the high-quality unit, a pre-designed neural network pool may be set, and the pre-designed neural network pool includes at least one network layer of the pre-designed network structure, and the network parameters related to the network layers of the pre-designed network structure may all be set according to actual needs. Based on the network parameters, for the high-quality unit, the network layers of the variable network structure may be dynamically generated, and the network parameters related to the network layers of the variable network structure are not preset, but are dynamically generated by the encoding side.
[0283] In one possible embodiment, the control parameter may include neural network information 3 corresponding to the high-quality unit, and the high-quality unit may analyze the neural network information 3 from the control parameter and generate a decoding neural network 3 based on the neural network information 3. For example, the neural network information 3 may include basic layer information and enhancement layer information, determine the basic layer based on the basic layer information, determine the enhancement layer based on the enhancement layer information, and combine the basic layer and the enhancement layer to obtain the decoding neural network 3. For example, the high-quality unit may obtain the neural network information 3 using the following cases.
[0284] Case 1: The base layer information includes a base layer default network usage flag bit, and the base layer default network usage flag bit indicates that the base layer uses the default network. In this case, based on the base layer information, the high-quality unit knows that the base layer uses the network layer of the default network structure, and thus obtains the base layer of the default network structure. The enhancement layer information includes an enhancement layer default network usage flag bit, and the enhancement layer default network usage flag bit indicates that the enhancement layer uses the default network. In this case, based on the enhancement layer information, the high-quality unit knows that the enhancement layer uses the network layer of the default network structure, and thus obtains the enhancement layer of the default network structure. Combining the base layer of the default network structure and the enhancement layer of the default network structure, a decoding neural network 3 is obtained.
[0285] Case 2: The base layer information includes a base layer default network usage flag bit, and the base layer default network usage flag bit indicates that the base layer uses the default network. In this case, based on the base layer information, the high-quality unit knows that the base layer uses the network layer of the default network structure, and thus obtains the base layer of the default network structure. The enhancement layer information includes an enhancement layer pre-designed network usage flag bit and an enhancement layer pre-designed network index number, and the enhancement layer pre-designed network usage flag bit indicates that the enhancement layer uses the pre-designed network. In this case, based on the enhancement layer information, the high-quality unit knows that the enhancement layer uses the network layer of the pre-designed network structure, and thus selects the enhancement layer of the pre-designed network structure corresponding to the enhancement layer pre-designed network index number from the pre-designed neural network pool. Combining the base layer of the default network structure and the enhancement layer of the pre-designed network structure may also obtain a decoding neural network 3.
[0286] Case 3: The base layer information includes a base layer default network usage flag bit, and the base layer default network usage flag bit indicates that the base layer uses the default network. The high-quality unit knows, based on the base layer information, that the base layer uses the network layer of the default network structure, and thus obtains the base layer of the default network structure. The enhancement layer information includes network parameters for generating the enhancement layer. In this case, the network parameters are parsed from the control parameters, and an enhancement layer of a variable network structure is generated based on the network parameters. The base layer of the default network structure and the enhancement layer of the variable network structure are combined to obtain the decoding neural network 3.
[0287] Case 4: The base layer information includes a base layer pre-designed network usage flag bit and a base layer pre-designed network index number, and the base layer pre-designed network usage flag bit indicates that the base layer uses the pre-designed network. In this case, the high-quality unit knows, based on the base layer information, that the base layer uses the network layer of the pre-designed network structure, and thus selects the base layer of the pre-designed network structure corresponding to the base layer pre-designed network index number from the pre-designed neural network pool. The enhancement layer information includes an enhancement layer default network usage flag bit, and the enhancement layer default network usage flag bit indicates that the enhancement layer uses the default network. In this case, based on the enhancement layer information, it is known that the enhancement layer uses the network layer of the default network structure, and thus the enhancement layer of the default network structure is obtained. The base layer of the pre-designed network structure and the enhancement layer of the default network structure may be combined to obtain the decoding neural network 3.
[0288] Case 5: The base layer information includes a base layer pre-designed network usage flag bit and a base layer pre-designed network index number, and the base layer pre-designed network usage flag bit indicates that the base layer uses the pre-designed network. In this case, based on the base layer information, the high-quality unit knows that the base layer uses the network layer of the pre-designed network structure. Therefore, it may select the base layer of the pre-designed network structure corresponding to the base layer pre-designed network index number from the pre-designed neural network pool. The enhancement layer information includes an enhancement layer pre-designed network usage flag bit and an enhancement layer pre-designed network index number, and the enhancement layer pre-designed network usage flag bit indicates that the enhancement layer uses the pre-designed network. In this case, based on the enhancement layer information, it knows that the enhancement layer uses the network layer of the pre-designed network structure. Therefore, it may select the enhancement layer of the pre-designed network structure corresponding to the enhancement layer pre-designed network index number from the pre-designed neural network pool. Combining the base layer of the pre-designed network structure and the enhancement layer of the pre-designed network structure, a decoding neural network 3 may be obtained.
[0289] Case 6: The basic layer information includes a basic layer pre-designed network usage flag bit and a basic layer pre-designed network index number. The basic layer pre-designed network usage flag bit indicates that the basic layer uses the pre-designed network. In this case, the high-quality unit knows, based on the basic layer information, that the basic layer uses the network layer of the pre-designed network structure. Therefore, the basic layer of the pre-designed network structure corresponding to the basic layer pre-designed network index number may be selected from the pre-designed neural network pool. The enhancement layer information includes network parameters for generating the enhancement layer. In this case, the network parameters are analyzed from the control parameters, and an enhancement layer with a variable network structure is generated based on the network parameters. The basic layer of the pre-designed network structure and the enhancement layer of the variable network structure are combined to obtain a decoding neural network 3.
[0290] In one possible embodiment, for the above network structures of the basic layer and the enhancement layer, both may be determined by the decoded control parameters. The control parameters may include neural network information 3. The neural network information 3 may include basic layer information and enhancement layer information. The neural network information 3 for the high-quality unit is similar to Tables 1, 2, and 3, but the relevant information is for the high-quality unit rather than the coefficient hyperparameter feature generation unit, and will not be repeated here.
[0291] Embodiment 14: The embodiments of the present invention provide an image encoding method based on a neural network, which may be applied to the encoding side (also called a video encoder). FIG. 7A is a schematic structural diagram of the encoding side. The encoding side may include a control parameter encoding unit, a feature conversion unit, a coefficient hyperparameter feature conversion unit, a first quantization unit, a second quantization unit, a first feature encoding unit, a second feature encoding unit, a first feature decoding unit, a second feature decoding unit, a coefficient hyperparameter feature generation unit, an image feature inverse conversion unit, a first inverse quantization unit, a second inverse quantization unit, and a quality improvement unit.
[0292] Exemplarily, the first quantization unit, the second quantization unit, the first inverse quantization unit, the second inverse quantization unit, and the quality improvement unit are optional units. In a specific scenario, it is also possible to turn off or skip the processes of these optional units.
[0293] In this embodiment, for each current block (i.e., an image block), the bitstream corresponding to the current block includes three parts: bitstream 0 (the bitstream including control parameters), bitstream 1 (the bitstream including coefficient hyperparameter feature information), and bitstream 2 (the bitstream including image feature information). The coefficient hyperparameter feature information and the image feature information may be collectively referred to as image information.
[0294] Exemplarily, the image encoding method based on the neural network in this embodiment may include the following steps S71 to S79.
[0295] In step S71, perform feature transformation on the current block I to obtain an image feature value F corresponding to the current block I. For example, the feature transformation unit may perform feature transformation on the current block I to obtain an image feature value F corresponding to the current block I. Exemplarily, the feature transformation unit may perform feature transformation on the current block I based on an encoding neural network to obtain an image feature value F corresponding to the current block I. Here, the current block I is an input feature of the encoding neural network, and the image feature value F is an output feature of the encoding neural network.
[0296] In step S72, determine image feature information based on the image feature value F.
[0297] Exemplarily, when enabling the second quantization unit, the second quantization unit may obtain the image feature value F from the feature transformation unit, perform quantization on the image feature value F to obtain an image feature quantization value F_q, and determine image feature information based on the image feature quantization value F_q, that is, the image feature information may be the image feature quantization value F_q. In this case, the control parameter encoding unit may encode the second enable information of the second quantization unit into the bitstream 0, that is, the control parameter includes the second enable information of the second quantization unit, and the second enable information is used to indicate that the second quantization unit is enabled. The control parameter encoding unit may encode quantization-related parameters corresponding to the second quantization unit, such as a step parameter qstep or a quantization parameter qp, etc., into the bitstream 0.
[0298] Exemplarily, when the second quantization unit is not enabled, determine image feature information based on the image feature value F, that is, the image feature information may be the image feature value F. In this case, the control parameter encoding unit may encode the second enable information of the second quantization unit into the bitstream 0, and the second enable information is used to indicate that the second quantization unit is not enabled.
[0299] In step S73, coefficient hyperparameter feature transformation is performed on the image feature value F to obtain a coefficient hyperparameter feature coefficient value C. For example, the coefficient hyperparameter feature transformation unit performs coefficient hyperparameter feature transformation on the image feature value F to obtain a coefficient hyperparameter feature coefficient value C. For example, based on an encoding neural network, coefficient hyperparameter feature transformation is performed on the image feature value F to obtain a coefficient hyperparameter feature coefficient value C. Here, the image feature value F is an input feature of the encoding neural network, and the coefficient hyperparameter feature coefficient value C is an output feature of the encoding neural network.
[0300] In step S74, coefficient hyperparameter feature information is determined based on the coefficient hyperparameter feature coefficient value C.
[0301] Exemplarily, when enabling the first quantization unit, the first quantization unit obtains the coefficient hyperparameter feature coefficient value C from the coefficient hyperparameter feature transformation unit, performs quantization on the coefficient hyperparameter feature coefficient value C to obtain a coefficient hyperparameter feature coefficient quantization value C_q, and may determine the coefficient hyperparameter feature information based on the coefficient hyperparameter feature coefficient quantization value C_q. That is, the coefficient hyperparameter feature information may be the coefficient hyperparameter feature coefficient quantization value C_q. In this case, the control parameter encoding unit may encode the first enable information of the first quantization unit into the bitstream 0. That is, the control parameter includes the first enable information of the first quantization unit, and the first enable information is used to indicate that the first quantization unit is enabled. The control parameter encoding unit may encode quantization-related parameters corresponding to the first quantization unit into the bitstream 0.
[0302] Exemplarily, when the first quantization unit is not enabled, coefficient hyperparameter feature information is determined based on the coefficient hyperparameter feature coefficient value C, that is, the coefficient hyperparameter feature information may be the coefficient hyperparameter feature coefficient value C. In this case, the control parameter encoding unit may encode the first enable information of the first quantization unit into the bitstream 0, and the first enable information is used to indicate that the first quantization unit is not enabled.
[0303] In step S75, coefficient hyperparameter feature information (for example, the coefficient hyperparameter feature coefficient quantization value C_q or the coefficient hyperparameter feature coefficient value C) is encoded to obtain a bitstream 1. For example, the first feature encoding unit may encode the coefficient hyperparameter feature information into the bitstream corresponding to the current block, and for ease of distinction, the bitstream including the coefficient hyperparameter feature information is denoted as the bitstream 1.
[0304] In step S76, the bitstream 1 corresponding to the current block is decoded to obtain coefficient hyperparameter feature information (for example, the coefficient hyperparameter feature coefficient quantization value C_q or the coefficient hyperparameter feature coefficient value C), and a coefficient hyperparameter feature coefficient reconstruction value is determined based on the coefficient hyperparameter feature information.
[0305] Exemplarily, the first feature decoding unit may decode the bitstream 1 corresponding to the current block to obtain coefficient hyperparameter feature information.
[0306] Exemplarily, when enabling the first quantization unit, the coefficient hyperparameter feature information is the coefficient hyperparameter feature coefficient quantization value C_q. In this case, the first inverse quantization unit may perform inverse quantization on the coefficient hyperparameter feature coefficient quantization value C_q to obtain the coefficient hyperparameter feature coefficient reconstruction value C'. Obviously, the coefficient hyperparameter feature coefficient reconstruction value C' may be the same as the coefficient hyperparameter feature coefficient value C. When the first quantization unit is not enabled, the coefficient hyperparameter feature information is the coefficient hyperparameter feature coefficient value C. In this case, the coefficient hyperparameter feature coefficient value C may be used as the coefficient hyperparameter feature coefficient reconstruction value C'. As described above, the coefficient hyperparameter feature coefficient reconstruction value C' can be obtained.
[0307] As can be seen from the above, since the coefficient hyperparameter feature coefficient reconstruction value C' is the same as the coefficient hyperparameter feature coefficient value C, the process of step S76 can be omitted, and the coefficient hyperparameter feature coefficient value C may be directly used as the coefficient hyperparameter feature coefficient reconstruction value C'. In this case, the encoding-side structure shown in FIG. 7A may be improved to obtain the encoding-side structure shown in FIG. 7B.
[0308] In step S77, inverse transformation processing is performed on the coefficient hyperparameter feature coefficient reconstruction value C' to obtain the coefficient hyperparameter feature value P. For example, the coefficient hyperparameter feature generation unit performs inverse transformation processing on the coefficient hyperparameter feature coefficient reconstruction value C' to obtain the coefficient hyperparameter feature value P.
[0309] In step S78, based on the coefficient hyperparameter feature value P, image feature information (e.g., image feature quantization value F_q or image feature value F) is encoded to obtain bitstream 2. For example, the second feature encoding unit may encode the image feature information into the bitstream corresponding to the current block. For ease of distinction, the bitstream including the image feature information is denoted as bitstream 2.
[0310] In step S79, the control parameter encoding unit acquires the control parameter corresponding to the current block. The control parameter may include neural network information. The control parameter corresponding to the current block is encoded into the bit stream, and the bit stream including the control parameter is denoted as bit stream 0.
[0311] In one possible embodiment, the feature transformation unit may perform feature transformation on the current block I using an encoding neural network. The control parameter encoding unit may determine neural network information 2 corresponding to the image feature inverse transformation unit of the decoding device based on the network structure of the encoding neural network. The neural network information 2 is used to determine the decoding neural network 2 corresponding to the image feature inverse transformation unit, and the neural network information 2 corresponding to the image feature inverse transformation unit is encoded into bit stream 0.
[0312] In one possible embodiment, the coefficient hyperparameter feature transformation unit may perform coefficient hyperparameter feature transformation on the image feature value F using an encoding neural network. The control parameter encoding unit may determine neural network information 1 corresponding to the coefficient hyperparameter feature generation unit of the decoding device based on the network structure of the encoding neural network. The neural network information 1 is used to determine the decoding neural network 1 corresponding to the coefficient hyperparameter feature generation unit, and the neural network information 1 corresponding to the coefficient hyperparameter feature generation unit is encoded into bit stream 0.
[0313] In one possible embodiment, as shown in FIGS. 7A and 7B, further, the following steps S80 to S86 may be included.
[0314] In step S80, decrypt the bitstream 1 corresponding to the current block to obtain the coefficient hyperparameter feature information corresponding to the current block. For example, the first feature decryption unit may decrypt the bitstream 1 corresponding to the current block to obtain the coefficient hyperparameter feature information corresponding to the current block.
[0315] In step S81, determine the coefficient hyperparameter feature coefficient reconstruction value based on the coefficient hyperparameter feature information.
[0316] For example, when enabling the first quantization unit, the coefficient hyperparameter feature information is the coefficient hyperparameter feature coefficient quantization value C_q. In this case, the first inverse quantization unit may perform inverse quantization on the coefficient hyperparameter feature coefficient quantization value C_q to obtain the coefficient hyperparameter feature coefficient reconstruction value C'. Or, when the first quantization unit is not enabled, the coefficient hyperparameter feature information is the coefficient hyperparameter feature coefficient value C. In this case, the coefficient hyperparameter feature coefficient value C may be used as the coefficient hyperparameter feature coefficient reconstruction value C'.
[0317] In step S82, perform inverse transformation processing on the coefficient hyperparameter feature coefficient reconstruction value C' to obtain the coefficient hyperparameter feature value P. For example, the coefficient hyperparameter feature generation unit may perform inverse transformation processing on the coefficient hyperparameter feature coefficient reconstruction value C' to obtain the coefficient hyperparameter feature value P. For example, the coefficient hyperparameter feature generation unit may perform inverse transformation processing on the coefficient hyperparameter feature coefficient reconstruction value C' based on the decryption neural network to obtain the coefficient hyperparameter feature value P.
[0318] In step S83, decode the bitstream 2 corresponding to the current block to obtain the image feature information corresponding to the current block. For example, the second feature decoding unit may decode the bitstream 2 corresponding to the current block to obtain the image feature information corresponding to the current block. When decoding the bitstream 2 corresponding to the current block, the second feature decoding unit may decode the bitstream 2 corresponding to the current block using the coefficient hyperparameter feature value P.
[0319] In step S84, determine the image feature reconstruction value based on the image feature information.
[0320] For example, when enabling the second quantization unit, the image feature information is the image feature quantization value F_q, and the second inverse quantization unit may perform inverse quantization on the image feature quantization value F_q to obtain the image feature reconstruction value F'. When the second quantization unit is not enabled, the image feature information is the image feature value F, and the image feature value F may be used as the image feature reconstruction value F'.
[0321] In step S85, perform inverse transformation processing on the image feature reconstruction value F' to obtain the image low-level feature value LF. For example, the image feature inverse transformation unit may perform inverse transformation processing on the image feature reconstruction value F' to obtain the image low-level feature value LF. For example, based on the decoded neural network, perform inverse transformation processing on the image feature reconstruction value F' to obtain the image low-level feature value LF.
[0322] In step S86, based on the image low-level feature value LF, a reconstructed image block I corresponding to the current block is determined. Exemplarily, when enabling the high-quality unit, the high-quality unit performs enhancement processing on the image low-level feature value LF to obtain the reconstructed image block I corresponding to the current block. For example, based on a decoding neural network, enhancement processing is performed on the image low-level feature value LF to obtain the reconstructed image block I corresponding to the current block. When the high-quality unit is not enabled, the image low-level feature value LF is used as the reconstructed image block I.
[0323] Exemplarily, for steps S80 to S86, reference may be made to Embodiment 5, and details are not repeated here.
[0324] In one possible embodiment, the coefficient hyperparameter feature generation unit may perform inverse transformation processing on the coefficient hyperparameter feature coefficient reconstruction value C' using a decoding neural network to obtain the coefficient hyperparameter feature value P. The control parameter encoding unit may determine neural network information 1 corresponding to the coefficient hyperparameter feature generation unit of the decoding device based on the network structure of the decoding neural network. The neural network information 1 is used to determine the decoding neural network 1 corresponding to the coefficient hyperparameter feature generation unit, and the neural network information 1 is encoded into the bitstream 0.
[0325] In one possible embodiment, the image feature inverse transformation unit may perform inverse transformation processing on the image feature reconstruction value F' using a decoding neural network to obtain the image low-level feature value LF. The control parameter encoding unit may determine neural network information 2 corresponding to the image feature inverse transformation unit of the decoding device based on the network structure of the decoding neural network. The neural network information 2 is used to determine the decoding neural network 2 corresponding to the image feature inverse transformation unit, and the neural network information 2 is encoded into the bitstream 0.
[0326] In one possible embodiment, the quality improvement unit may perform enhancement processing on the image low-level feature value LF using a decoding neural network to obtain a reconstructed image block I corresponding to the current block, and the control parameter encoding unit may determine neural network information 3 corresponding to the quality improvement unit of the decoding device based on the network structure of the decoding neural network, where the neural network information 3 is used to determine the decoding neural network 3 corresponding to the quality improvement unit, and the neural network information 3 is encoded into the bitstream 0.
[0327] Example 15: An embodiment of the present invention provides an image encoding method based on a neural network, which may be applied to the encoding side. FIG. 7C is a schematic structural diagram of the encoding side. The encoding side may include a control parameter encoding unit, a feature conversion unit, a coefficient hyperparameter feature conversion unit, a first feature encoding unit, a second feature encoding unit, a first feature decoding unit, a second feature decoding unit, a coefficient hyperparameter feature generation unit, an image feature inverse conversion unit, and a quality improvement unit.
[0328] Exemplarily, the image encoding method based on the neural network in this embodiment may include the following steps S91 to S99.
[0329] In step S91, perform feature conversion on the current block I to obtain an image feature value F corresponding to the current block I. For example, the feature conversion unit may perform feature conversion on the current block I based on an encoding neural network to obtain an image feature value F corresponding to the current block I.
[0330] In step S92, coefficient hyperparameter feature transformation is performed on the image feature value F to obtain a coefficient hyperparameter feature coefficient value C. For example, the coefficient hyperparameter feature transformation unit performs coefficient hyperparameter feature transformation on the image feature value F based on an encoding neural network to obtain a coefficient hyperparameter feature coefficient value C.
[0331] In step S93, the coefficient hyperparameter feature coefficient value C is encoded to obtain a bitstream 1. For example, the first feature encoding unit may encode the coefficient hyperparameter feature coefficient value C into the bitstream corresponding to the current block to obtain a bitstream 1.
[0332] In step S94, a coefficient hyperparameter feature coefficient reconstruction value C' is determined based on the coefficient hyperparameter feature coefficient value C. For example, the coefficient hyperparameter feature coefficient reconstruction value C' is the same as the coefficient hyperparameter feature coefficient value C, and inverse transformation processing is performed on the coefficient hyperparameter feature coefficient reconstruction value C' to obtain a coefficient hyperparameter feature value P. For example, the coefficient hyperparameter feature generation unit performs inverse transformation processing on the coefficient hyperparameter feature coefficient reconstruction value C' to obtain a coefficient hyperparameter feature value P.
[0333] In step S95, based on the coefficient hyperparameter feature value P, the image feature value F is encoded to obtain a bitstream 2. For example, the second feature encoding unit encodes the image feature value F into the bitstream corresponding to the current block to obtain a bitstream 2.
[0334] In step S96, the control parameter encoding unit acquires the control parameter corresponding to the current block. The control parameter may include neural network information, encodes the control parameter corresponding to the current block into the bitstream, and denotes the bitstream including the control parameter as bitstream 0.
[0335] In step S97, bitstream 1 corresponding to the current block is decoded to obtain the coefficient hyperparameter feature coefficient value C corresponding to the current block. For example, the first feature decoding unit may decode bitstream 1 corresponding to the current block to obtain the coefficient hyperparameter feature coefficient value C corresponding to the current block.
[0336] Also, a coefficient hyperparameter feature coefficient reconstruction value C' is determined based on the coefficient hyperparameter feature coefficient value C.
[0337] In step S98, an inverse transformation process is performed on the coefficient hyperparameter feature coefficient reconstruction value C' to obtain the coefficient hyperparameter feature value P. For example, the coefficient hyperparameter feature generation unit performs an inverse transformation process on the coefficient hyperparameter feature coefficient reconstruction value C' based on the decoded neural network to obtain the coefficient hyperparameter feature value P. Bitstream 2 corresponding to the current block is decoded to obtain the image feature value F corresponding to the current block, and the image feature value F is used as the image feature reconstruction value F'. For example, the second feature decoding unit decodes bitstream 2 corresponding to the current block using the coefficient hyperparameter feature value P.
[0338] In step S99, an inverse transformation process is performed on the image feature reconstruction value F' to obtain the image low-level feature value LF. For example, the image feature inverse transformation unit performs an inverse transformation process on the image feature reconstruction value F' based on the decoded neural network to obtain the image low-level feature value LF. Based on the image low-level feature value LF, the reconstructed image block I corresponding to the current block is determined. For example, the high-quality enhancement unit performs an enhancement process on the image low-level feature value LF based on the decoded neural network to obtain the reconstructed image block I corresponding to the current block.
[0339] Example 16: For Example 14 and Example 15, the first feature encoding unit may encode the coefficient hyperparameter feature information into the bitstream 1 corresponding to the current block. The encoding process of the first feature encoding unit corresponds to the decoding process of the first feature decoding unit. Referring to Example 7, for example, the first feature encoding unit may use an entropy encoding method (such as an entropy encoding method like CAVLC or CABAC) to encode the coefficient hyperparameter feature information, which will not be repeated here. The first feature decoding unit may decode the bitstream 1 corresponding to the current block to obtain the coefficient hyperparameter feature information corresponding to the current block. For this process, reference may be made to Example 7.
[0340] For Example 14 and Example 15, the second feature encoding unit may encode the image feature information based on the coefficient hyperparameter feature value P to obtain the bitstream 2. The encoding process of the second feature encoding unit corresponds to the decoding process of the second feature decoding unit. Referring to Example 10, for example, the second feature encoding unit may use an entropy encoding method (such as an entropy encoding method like CAVLC or CABAC) to encode the image feature information, which will not be repeated here. The second feature decoding unit may decode the bitstream 2 corresponding to the current block to obtain the image feature information corresponding to the current block. For this process, reference may be made to Example 10.
[0341] For Embodiment 14 and Embodiment 15, the first quantization unit may perform quantization on the coefficient hyperparameter feature coefficient value C to obtain the coefficient hyperparameter feature coefficient quantization value C_q. The quantization process of the first quantization unit corresponds to the inverse quantization process of the first inverse quantization unit. Referring to Embodiment 8, for example, the first quantization unit performs quantization on the coefficient hyperparameter feature coefficient value C based on quantization-related parameters, which will not be repeated here. Regarding the quantization-related parameters (for example, the step parameter qstep, also called the quantization stride), 1) each feature value of each feature channel uses the same quantization stride, 2) each feature channel uses different quantization strides, but each feature value within the feature channel uses the same quantization stride, 3) each feature value of each feature channel uses different quantization strides. The first inverse quantization unit on the encoding side may perform inverse quantization on the coefficient hyperparameter feature coefficient quantization value C_q to obtain the coefficient hyperparameter feature coefficient reconstruction value C'. The processing process of the first inverse quantization unit on the encoding side may refer to Embodiment 10.
[0342] The second quantization unit may perform quantization on the image feature value F to obtain the image feature quantization value F_q. The quantization process of the second quantization unit corresponds to the inverse quantization process of the second inverse quantization unit. Referring to Embodiment 11, for example, the second quantization unit performs quantization on the image feature value F based on quantization-related parameters, which will not be repeated here. Regarding the quantization stride, 1) each feature value of each feature channel uses the same quantization stride, 2) each feature channel uses different quantization strides, but each feature value within the feature channel uses the same quantization stride, 3) each feature value of each feature channel uses different quantization strides. The second inverse quantization unit on the encoding side may perform inverse quantization on the image feature quantization value F_q to obtain the image feature reconstruction value F'. The processing process of the second inverse quantization unit on the encoding side may refer to Embodiment 11.
[0343] Example 17: For Example 14 and Example 15, the feature transformation unit may perform feature transformation on the current block I based on the encoding neural network to obtain the image feature value F corresponding to the current block I. Based on this, based on the network structure of the encoding neural network, the neural network information 2 corresponding to the image feature inverse transformation unit of the decoding device is determined, and the neural network information 2 corresponding to the image feature inverse transformation unit may be encoded into the bitstream 0. For the method by which the decoding device generates the decoding neural network 2 corresponding to the image feature inverse transformation unit based on the neural network information 2, reference may be made to Example 12, which will not be repeated here.
[0344] In one possible embodiment, the feature transformation unit may include an encoding neural network 2. The encoding neural network 2 may include a basic layer and an enhancement layer. The encoding neural network 2 may be divided into a basic layer and an enhancement layer. The basic layer may include at least one network layer, or the basic layer may not include a network layer, that is, the basic layer may be empty. The enhancement layer may include at least one network layer, or the enhancement layer may not include a network layer, that is, the enhancement layer may be empty. For the multiple network layers in the encoding neural network 2, according to actual needs, the multiple network layers may be divided into a basic layer and an enhancement layer. For example, the network layer with a fixed network structure may be used as the basic layer, and the network layer with a non-fixed network structure may be used as the enhancement layer.
[0345] Exemplarily, for the encoding neural network 2, the size of the output feature may be smaller than the size of the input feature, or the size of the output feature may be equal to the size of the input feature, or the size of the output feature may be larger than the size of the input feature.
[0346] Exemplarily, for the encoded neural network 2, the basic layer and the enhancement layer include at least one convolutional layer. For example, the basic layer includes at least one convolutional layer, and the enhancement layer may or may not include at least one convolutional layer. Or, the enhancement layer includes at least one convolutional layer, and the basic layer may or may not include at least one convolutional layer.
[0347] For the encoded neural network 2, the basic layer and the enhancement layer include at least one residual structure layer. For example, the basic layer includes at least one residual structure layer, and the enhancement layer may or may not include a residual structure layer. Or, the enhancement layer includes at least one residual structure layer, and the basic layer may or may not include a residual structure layer.
[0348] For the encoded neural network 2, it may include convolutional layers, activation layers, etc., but is not limited thereto, and is not limited in this regard. For example, the encoded neural network 2 sequentially includes one convolutional layer with a stride of 2, one activation layer, one convolutional layer with a stride of 1, and one activation layer, and all the above network layers are used as the basic layer. Also, for example, the encoded neural network 2 sequentially includes one convolutional layer with a stride of 2 and one convolutional layer with a stride of 1, and all the above network layers are used as the basic layer. Also, for example, the encoded neural network 2 sequentially includes one convolutional layer with a stride of 2, one activation layer, one convolutional layer with a stride of 2, one activation layer, one convolutional layer with a stride of 1, and one activation layer, and all the above network layers are used as the basic layer.
[0349] Exemplarily, the number of input channels of the first network layer of the feature transformation unit is 1 or 3. When the input image block contains only one channel (e.g., a grayscale image), the number of input channels of the first network layer of the feature transformation unit is 1, and when the input image block contains three channels (e.g., in RGB or YUV format), the number of input channels of the first network layer of the feature transformation unit is 3.
[0350] Note that the network structure of the encoding neural network 2 of the feature transformation unit and the network structure of the decoding neural network 2 of the image feature inverse transformation unit (see Example 12) may be symmetric, and the network parameters of the encoding neural network 2 of the feature transformation unit and the network parameters of the decoding neural network 2 of the image feature inverse transformation unit may be the same or different.
[0351] In one possible embodiment, for the feature transformation unit, network layers of a default network structure may be set, and all network parameters regarding the network layers of the default network structure are fixed. For example, in the network layers of the default network structure, the number of transposed convolution layers is fixed, the number of activation layers is fixed, the number of channels of each transposed convolution layer is fixed, the size of the convolution kernel is fixed, the filtering coefficient is fixed, etc. Note that the network layers of the default network structure set for the feature transformation unit and the network layers of the default network structure set for the image feature inverse transformation unit (see Example 12) may have a symmetric structure.
[0352] In one possible embodiment, for the feature conversion unit, a pre-designed neural network pool (corresponding to the pre-designed neural network pool of the decoding device) may be set. The pre-designed neural network pool includes network layers of at least one pre-designed network structure, and the network parameters related to the network layers of the pre-designed network structure may all be set according to actual needs. For example, the pre-designed neural network pool may include a network layer of a pre-designed network structure t1’, a network layer of a pre-designed network structure t2’, a network layer of a pre-designed network structure t3’, etc. Note that the network layer of the pre-designed network structure t1’ and the network layer of the pre-designed network structure t1 (see Embodiment 12) may be a symmetric structure. The network layer of the pre-designed network structure t2’ and the network layer of the pre-designed network structure t2 may be a symmetric structure. The network layer of the pre-designed network structure t3’ and the network layer of the pre-designed network structure t3 may be a symmetric structure.
[0353] In one possible embodiment, based on the network parameters, for the feature conversion unit, a network layer of a variable network structure may be dynamically generated. The network parameters related to the network layer of the variable network structure are not preset but are dynamically generated by the encoding side.
[0354] In one possible embodiment, the encoding side may determine neural network information 2 corresponding to the image feature inverse conversion unit based on the network structure of the encoding neural network 2, and encode the neural network information 2 corresponding to the image feature inverse conversion unit into the bitstream 0. For example, the encoding side may determine the neural network information 2 using the following cases.
[0355] Case 1: When the basic layer of the encoding neural network 2 uses the network layer of the default network structure and the enhancement layer uses the network layer of the default network structure, on the encoding side, the basic layer information and the enhancement layer information are encoded into the bitstream 0, and the basic layer information and the enhancement layer information constitute the neural network information 2 corresponding to the image feature inverse transformation unit. The basic layer information includes a basic layer default network usage flag bit, and the basic layer default network usage flag bit indicates that the basic layer uses the default network. The enhancement layer information includes an enhancement layer default network usage flag bit, and the enhancement layer default network usage flag bit indicates that the enhancement layer uses the default network.
[0356] Case 2: When the basic layer of the encoding neural network 2 uses the network layer of the default network structure and the enhancement layer uses the network layer of the pre-designed network structure, on the encoding side, the basic layer information and the enhancement layer information are encoded into the bitstream 0, and the basic layer information and the enhancement layer information constitute the neural network information 2 corresponding to the image feature inverse transformation unit. The basic layer information includes a basic layer default network usage flag bit, and the basic layer default network usage flag bit indicates that the basic layer uses the default network. The enhancement layer information includes an enhancement layer pre-designed network usage flag bit and an enhancement layer pre-designed network index number, and the enhancement layer pre-designed network usage flag bit indicates that the enhancement layer uses the pre-designed network, and the enhancement layer pre-designed network index number indicates the corresponding index in the pre-designed neural network pool of the network layer of the pre-designed network structure. For example, when the enhancement layer of the encoding neural network 2 uses the network layer of the pre-designed network structure t1’, the enhancement layer pre-designed network index number indicates the network layer of the first pre-designed network structure t1’ in the pre-designed neural network pool.
[0357] Case 3: When the basic layer of the encoded neural network 2 uses the network layer of the default network structure and the enhancement layer uses the enhancement layer of the variable network structure, on the encoding side, the basic layer information and the enhancement layer information are encoded into the bitstream 0, and the basic layer information and the enhancement layer information constitute the neural network information 2 corresponding to the image feature inverse transformation unit. The basic layer information includes a basic layer default network usage flag bit, and the basic layer default network usage flag bit indicates that the basic layer uses the default network. The enhancement layer information includes network parameters for generating the enhancement layer, and the network parameters may include, but are not limited to, at least one of the number of layers of the neural network, the transposed convolution layer flag bit, the number of transposed convolution layers, the quantization stride of each transposed convolution layer, the number of channels of each transposed convolution layer, the size of the convolution kernel, the number of filterings, the filtering size index, the filtering coefficient zero flag bit, the filtering coefficient, the activation layer flag bit, and the activation layer type. Note that the network parameters in the enhancement layer information may be different from the network parameters used in the enhancement layer of the encoded neural network 2, that is, for each network parameter used in the enhancement layer, a network parameter symmetric to the network parameter may be generated, and the generation process is not limited, and this symmetric network parameter is transmitted to the decoding side as the enhancement layer information.
[0358] Case 4: When the basic layer of the encoded neural network 2 uses the network layer of the pre-designed network structure and the enhancement layer uses the network layer of the default network structure, the basic layer information and the enhancement layer information are encoded into the bitstream 0, and the basic layer information and the enhancement layer information constitute the neural network information 2 corresponding to the image feature inverse conversion unit. The basic layer information includes a basic layer pre-designed network usage flag bit and a basic layer pre-designed network index number. The basic layer pre-designed network usage flag bit indicates that the basic layer uses the pre-designed network, and the basic layer pre-designed network index number indicates the corresponding index in the pre-designed neural network pool of the network layer of the pre-designed network structure. The enhancement layer information includes an enhancement layer default network usage flag bit, and the enhancement layer default network usage flag bit indicates that the enhancement layer uses the default network.
[0359] Case 5: When the basic layer of the encoded neural network 2 uses the network layer of the pre-designed network structure and the enhancement layer uses the network layer of the pre-designed network structure, on the encoding side, the basic layer information and the enhancement layer information are encoded into the bit stream 0, and the basic layer information and the enhancement layer information constitute the neural network information 2 corresponding to the image feature inverse transformation unit. The basic layer information may include a basic layer pre-designed network usage flag bit and a basic layer pre-designed network index number. The basic layer pre-designed network usage flag bit is used to indicate that the basic layer uses the pre-designed network, and the basic layer pre-designed network index number is used to indicate the corresponding index in the pre-designed neural network pool of the network layer of the pre-designed network structure. The enhancement layer information may include an enhancement layer pre-designed network usage flag bit and an enhancement layer pre-designed network index number. The enhancement layer pre-designed network usage flag bit is used to indicate that the enhancement layer uses the pre-designed network, and the enhancement layer pre-designed network index number is used to indicate the corresponding index in the pre-designed neural network pool of the network layer of the pre-designed network structure.
[0360] Case 6: When the basic layer of the encoded neural network 2 uses the network layer of the pre-designed network structure and the enhancement layer uses the enhancement layer of the variable network structure, the basic layer information and the enhancement layer information are encoded into the bitstream 0, and the basic layer information and the enhancement layer information constitute the neural network information 2 corresponding to the image feature inverse transformation unit. The basic layer information includes the basic layer pre-designed network usage flag bit and the basic layer pre-designed network index number. The basic layer pre-designed network usage flag bit indicates that the basic layer uses the pre-designed network, and the basic layer pre-designed network index number indicates the corresponding index in the pre-designed neural network pool of the network layer of the pre-designed network structure. The enhancement layer information includes the network parameters for generating the enhancement layer, and the network parameters may include, but are not limited to, at least one of the number of layers of the neural network, the transposed convolution layer flag bit, the number of transposed convolution layers, the quantization stride of each transposed convolution layer, the number of channels of each transposed convolution layer, the size of the convolution kernel, the number of filterings, the filtering size index, the filtering coefficient zero flag bit, the filtering coefficient, the activation layer flag bit, and the activation layer type. The network parameters in the enhancement layer information may be different from the network parameters used in the enhancement layer of the encoded neural network 2, that is, for each network parameter used in the enhancement layer, a network parameter symmetric to the network parameter may be generated, and the generation process is not limited, and this symmetric network parameter is transmitted to the decoding side as the enhancement layer information.
[0361] Example 18: For Example 14 and Example 15, the image feature inverse transformation unit on the encoding side may perform an inverse transformation process on the image feature reconstruction value F' based on the decoding neural network to obtain the image low-level feature value LF. Based on this, based on the network structure of the decoding neural network, the neural network information 2 corresponding to the image feature inverse transformation unit of the decoding device is determined, and the neural network information 2 corresponding to the image feature inverse transformation unit may be encoded into the bitstream 0.
[0362] In one possible embodiment, the image feature inverse transformation unit on the encoding side may include a decoding neural network 2. The decoding neural network 2 may include a basic layer and an enhancement layer. The network structure of the decoding neural network 2 on the encoding side is the same as the network structure of the decoding neural network 2 on the decoding side. Reference may be made to Example 12, and the description is omitted here.
[0363] In one possible embodiment, the encoding side may set the network layer of the default network structure for the image feature inverse transformation unit. The network layer of this default network structure is the same as the network layer of the default network structure on the decoding side. The encoding side may set a pre-designed neural network pool for the image feature inverse transformation unit. The pre-designed neural network pool may include at least one network layer of the pre-designed network structure. This pre-designed neural network pool is the same as the pre-designed neural network pool on the decoding side. The encoding side may dynamically generate the network layer of the variable network structure for the image feature inverse transformation unit based on the network parameters. The encoding side may determine the neural network information 2 corresponding to the image feature inverse transformation unit based on the network structure of the decoding neural network 2, and encode the neural network information 2 into the bitstream 0.
[0364] For example, when the basic layer of the decoding neural network 2 uses the network layer of the default network structure and the enhancement layer uses the network layer of the default network structure, the basic layer information includes a basic layer default network usage flag bit, and the basic layer default network usage flag bit indicates that the basic layer uses the default network. The enhancement layer information includes an enhancement layer default network usage flag bit, and the enhancement layer default network usage flag bit indicates that the enhancement layer uses the default network. Also, for example, when the basic layer of the decoding neural network 2 uses the network layer of the default network structure and the enhancement layer uses the network layer of the pre-designed network structure, the basic layer information includes a basic layer default network usage flag bit, and the basic layer default network usage flag bit indicates that the basic layer uses the default network. The enhancement layer information includes an enhancement layer pre-designed network usage flag bit and an enhancement layer pre-designed network index number, and the enhancement layer pre-designed network usage flag bit indicates that the enhancement layer uses the pre-designed network, and the enhancement layer pre-designed network index number indicates the corresponding index in the pre-designed neural network pool of the network layer of the pre-designed network structure. Also, for example, when the basic layer of the decoding neural network 2 uses the network layer of the default network structure and the enhancement layer uses the enhancement layer of the variable network structure, the basic layer information includes a basic layer default network usage flag bit, and the basic layer default network usage flag bit indicates that the basic layer uses the default network. The enhancement layer information includes network parameters for generating the enhancement layer, and the network parameters in the enhancement layer information may be the same as the network parameters used for the enhancement layer of the decoding neural network 2.
[0365] For example, when the basic layer of the decoding neural network 2 uses a network layer with a pre-designed network structure and the enhancement layer uses a network layer with a default network structure, the basic layer information includes a basic layer pre-designed network usage flag bit and a basic layer pre-designed network index number. The basic layer pre-designed network usage flag bit indicates that the basic layer uses a pre-designed network, and the basic layer pre-designed network index number indicates the corresponding index in the pre-designed neural network pool of the network layer with the pre-designed network structure. The enhancement layer information includes an enhancement layer default network usage flag bit, and the enhancement layer default network usage flag bit indicates that the enhancement layer uses a default network. Also, for example, when the basic layer of the decoding neural network 2 uses a network layer with a pre-designed network structure and the enhancement layer uses a network layer with a pre-designed network structure, the basic layer information includes a basic layer pre-designed network usage flag bit and a basic layer pre-designed network index number. The basic layer pre-designed network usage flag bit indicates that the basic layer uses a pre-designed network, and the basic layer pre-designed network index number indicates the corresponding index in the pre-designed neural network pool of the network layer with the pre-designed network structure. The enhancement layer information includes an enhancement layer pre-designed network usage flag bit and an enhancement layer pre-designed network index number. The enhancement layer pre-designed network usage flag bit indicates that the enhancement layer uses a pre-designed network, and the enhancement layer pre-designed network index number indicates the corresponding index in the pre-designed neural network pool of the network layer with the pre-designed network structure.Also, for example, when the basic layer of the decoding neural network 2 uses a network layer with a pre-designed network structure and the enhancement layer uses an enhancement layer with a variable network structure, the basic layer information includes a basic layer pre-designed network usage flag bit and a basic layer pre-designed network index number. The basic layer pre-designed network usage flag bit indicates that the basic layer uses the pre-designed network, and the basic layer pre-designed network index number indicates the corresponding index in the pre-designed neural network pool of the network layer with the pre-designed network structure. The enhancement layer information includes network parameters, and the network parameters in the enhancement layer information are the same as the network parameters used in the enhancement layer of the decoding neural network 2.
[0366] Example 19: For Example 14 and Example 15, the coefficient hyperparameter feature conversion unit may perform coefficient hyperparameter feature conversion on the image feature value F based on the encoding neural network to obtain a coefficient hyperparameter feature coefficient value C. Based on this, based on the network structure of the encoding neural network, neural network information 1 corresponding to the coefficient hyperparameter feature generation unit of the decoding device is determined, and neural network information 1 corresponding to the coefficient hyperparameter feature generation unit may be encoded into the bitstream 0. For the method by which the decoding device generates the decoding neural network 1 corresponding to the coefficient hyperparameter feature generation unit based on the neural network information 1, reference may be made to Example 9, which will not be repeated here.
[0367] In one possible embodiment, the coefficient hyperparameter feature conversion unit may include an encoding neural network 1, the encoding neural network 1 may include a basic layer and an enhancement layer, the encoding neural network 1 may be divided into a basic layer and an enhancement layer, the basic layer may include at least one network layer, the basic layer may not include a network layer, that is, the basic layer may be empty. The enhancement layer may include at least one network layer, the enhancement layer may not include a network layer, that is, the enhancement layer may be empty. It should be noted that the network structure of the encoding neural network 1 of the coefficient hyperparameter feature conversion unit and the network structure of the decoding neural network 1 of the coefficient hyperparameter feature generation unit (see Embodiment 9) may be symmetric, and the network parameters of the encoding neural network 1 of the coefficient hyperparameter feature conversion unit and the network parameters of the decoding neural network 1 of the coefficient hyperparameter feature generation unit may be the same or different. The description of the network structure of this encoding neural network 1 is omitted.
[0368] In one possible embodiment, for the coefficient hyperparameter feature conversion unit, the network layers of the default network structure may be set, and all network parameters related to the network layers of the default network structure are fixed. Note that the network layers of the default network structure set for the coefficient hyperparameter feature conversion unit and the network layers of the default network structure set for the coefficient hyperparameter feature generation unit may have a symmetric structure. For the coefficient hyperparameter feature conversion unit, a pre-designed neural network pool (corresponding to the pre-designed neural network pool of the decoding device) may be set. The pre-designed neural network pool includes at least one network layer of the pre-designed network structure, and the network parameters related to the network layer of the pre-designed network structure may all be set according to actual needs. Based on the network parameters, for the coefficient hyperparameter feature conversion unit, network layers of a variable network structure may be dynamically generated, and the network parameters related to the network layers of the variable network structure are dynamically generated by the encoding side.
[0369] In one possible embodiment, the encoding side may determine neural network information 1 corresponding to the coefficient hyperparameter feature generation unit based on the network structure of the encoding neural network 1, and encode the neural network information 1 into the bitstream 0. Here, the neural network information 1 may include basic layer information and enhanced layer information. For the method of encoding the neural network information 1, it is similar to the method of encoding the neural network information 2, and reference may be made to Embodiment 17, which will not be repeated here. For example, when the basic layer of the encoding neural network 1 uses the network layer of the default network structure and the enhanced layer uses the network layer of the default network structure, the basic layer information includes a basic layer default network usage flag bit, and the basic layer default network usage flag bit indicates that the basic layer uses the default network. The enhanced layer information includes an enhanced layer default network usage flag bit, and the enhanced layer default network usage flag bit indicates that the enhanced layer uses the default network.
[0370] Example 20: For Example 14 and Example 15, the coefficient hyperparameter feature generation unit on the encoding side may perform an inverse transformation process on the coefficient hyperparameter feature coefficient reconstruction value C' based on the decoding neural network to obtain the coefficient hyperparameter feature value P. Based on this, neural network information 1 corresponding to the coefficient hyperparameter feature generation unit of the decoding device may be determined based on the network structure of the decoding neural network, and the neural network information 1 corresponding to the coefficient hyperparameter feature generation unit may be encoded into the bitstream 0.
[0371] In one possible embodiment, the coefficient hyperparameter feature generation unit on the encoding side may include a decoding neural network 1. The decoding neural network 1 may include a basic layer and an enhancement layer. The network structure of the decoding neural network 1 on the encoding side is the same as the network structure of the decoding neural network 1 on the decoding side. Reference may be made to Example 9, and the description is omitted here.
[0372] In one possible embodiment, the encoding side may set the network layer of the default network structure for the coefficient hyperparameter feature generation unit, and the network layer of this default network structure is the same as the network layer of the default network structure on the decoding side. The encoding side may set a pre-designed neural network pool for the coefficient hyperparameter feature generation unit, and the pre-designed neural network pool may include at least one network layer of the pre-designed network structure, and this pre-designed neural network pool is the same as the pre-designed neural network pool on the decoding side. The encoding side may dynamically generate the network layer of the variable network structure for the coefficient hyperparameter feature generation unit based on the network parameters. The encoding side may determine the neural network information 1 corresponding to the coefficient hyperparameter feature generation unit based on the network structure of the decoding neural network 1, and encode the neural network information 1 into the bitstream 0. Here, the neural network information 1 may include basic layer information and enhanced layer information. Regarding the method of encoding the neural network information 1, it is similar to the method of encoding the neural network information 2, and reference may be made to Example 18, which will not be repeated here. For example, when the basic layer of the decoding neural network 1 (used on the encoding side) uses the network layer of the default network structure and the enhanced layer uses the enhanced layer of the variable network structure, the basic layer information includes the basic layer default network usage flag bit, and the basic layer default network usage flag bit indicates that the basic layer uses the default network. The enhanced layer information includes the network parameters for generating the enhanced layer, and the network parameters in the enhanced layer information may be the same as the network parameters used in the enhanced layer of the decoding neural network 1 (i.e., the enhanced layer used on the encoding side).
[0373] Example 21: For Example 14 and Example 15, the high-quality unit on the encoding side may perform enhancement processing on the image low-level feature value LF based on the decoding neural network to obtain a reconstructed image block I corresponding to the current block. Based on this, the neural network information 3 corresponding to the high-quality unit of the decoding device is determined based on the network structure of the decoding neural network, and the neural network information 3 corresponding to the high-quality unit may be encoded into the bitstream 0. For example, the high-quality unit on the encoding side may include a decoding neural network 3, the decoding neural network 3 may include a basic layer and an enhancement layer, the network structure of the decoding neural network 3 on the encoding side is the same as the network structure of the decoding neural network 3 on the decoding side, reference may be made to Example 13, and the description is omitted here.
[0374] Exemplarily, for the quality improvement unit, the encoding side may set the network layer of the default network structure, and the network layer of this default network structure is the same as the network layer of the default network structure on the decoding side. For the quality improvement unit, the encoding side may set a pre-designed neural network pool, and the pre-designed neural network pool may include at least one network layer of the pre-designed network structure, and this pre-designed neural network pool is the same as the pre-designed neural network pool on the decoding side. Based on the network parameters, the encoding side may dynamically generate the network layer of the variable network structure for the quality improvement unit. Based on the network structure of the decoding neural network 3, the encoding side determines the neural network information 3 corresponding to the quality improvement unit and encodes the neural network information 3 into the bitstream 0. Here, the neural network information 3 may include basic layer information and enhancement layer information. Regarding the method of encoding the neural network information 3, it is similar to the method of encoding the neural network information 2, and reference may be made to Embodiment 18, which will not be repeated here. For example, when the basic layer of the decoding neural network 3 uses the network layer of the default network structure and the enhancement layer uses the enhancement layer of the variable network structure, the basic layer information includes a basic layer default network usage flag bit, and the basic layer default network usage flag bit indicates that the basic layer uses the default network. The enhancement layer information includes the network parameters for generating the enhancement layer, and the network parameters in the enhancement layer information may be the same as the network parameters used in the enhancement layer of the decoding neural network 3.
[0375] In one possible embodiment, for the network parameters in Examples 1 to 21, fixed-point network parameters may be used. For example, the filtering weights in the network parameters may be represented with bit widths of 4, 8, 16, 32, or 64 bits. For the feature values output from the network, bit width limitation may be performed. For example, it may be limited to be represented with bit widths of 4, 8, 16, 32, or 64 bits.
[0376] Exemplarily, the bit width of the network parameters may be limited to 8 bits, and the magnitude of its value may be limited to be between [-127, 127]. The bit width of the feature values output from the network may be limited to 8 bits, and the magnitude of its value may be limited to be between [-127, 127].
[0377] Example 22: Regarding the related syntax table of the image header, Table 4 provides syntax information regarding the image header, i.e., image-level syntax. In Table 4, u(n) is used to represent an n-bit fixed-length coding method. [Table 4]
[0378] In Table 4, pic_width represents the width of the image, pic_height represents the height of the image, pic_format represents the image format, such as image formats like RGB444, YUV444, YUV420, YUV422, etc. bu_width represents the width of the basic block, bu_height represents the height of the basic block, block_width represents the width of the image block, block_height represents the height of the image block, bit_depth represents the image bit depth, pic_qp represents the quantization parameter in the current image, lossless_flag represents the flag indicating whether to perform lossless coding mark on the current image, feature_map_max_bit_depth represents the maximum bit depth of the feature map, and the maximum bit depth of the feature map is used to limit the maximum and minimum values of the input or output feature map of the network.
[0379] Exemplarily, each of the above embodiments may be implemented alone or in combination. For example, each of the embodiments 1 to 22 may be implemented alone, or at least two of the embodiments 1 to 22 may be implemented in combination.
[0380] Exemplarily, in each of the above embodiments, the content on the encoding side may be applied to the decoding side, that is, it may be processed in the same manner by the decoding side, and the content on the decoding side may be applied to the encoding side, that is, it may be processed in the same manner by the encoding side.
[0381] Based on the same concept as the above method, embodiments of the present invention further provide an image decoding device based on a neural network. The device is applied to the decoding side and includes a memory configured to store video data, and a decoder configured to implement the decoding method in the above embodiments 1 to 22, that is, the processing process on the decoding side.
[0382] For example, in one possible embodiment, the decoder is The step of decoding control parameters and image information corresponding to the current block from the bit stream, The step of obtaining neural network information corresponding to the decoding processing unit from the control parameters, and generating a decoding neural network corresponding to the decoding processing unit based on the neural network information, The step of determining input features corresponding to the decoding processing unit based on the image information, and processing the input features based on the decoding neural network to obtain output features corresponding to the decoding processing unit, are configured to be implemented.
[0383] Based on the same concept as the above method, an embodiment of the present invention further provides an image encoding device based on a neural network. The device is applied to the encoding side, and the device includes a memory configured to store video data, and an encoder configured to implement the encoding methods in the above Embodiments 1 to 22, that is, the processing process on the encoding side.
[0384] For example, in one possible embodiment, the encoder is The step of determining input features corresponding to the encoding processing unit based on the current block, processing the input features based on the encoding neural network corresponding to the encoding processing unit to obtain output features corresponding to the encoding processing unit, and determining image information corresponding to the current block based on the output features, The step of obtaining control parameters corresponding to the current block, where the control parameters include neural network information corresponding to the decoding processing unit, and the neural network information is used to determine the decoding neural network corresponding to the decoding processing unit, The step of encoding the image information and control parameters corresponding to the current block into a bit stream, are configured to be implemented.
[0385] Based on the same concept as the above method, for the decoding device (also called a video decoder) provided by an embodiment of the present invention, from a hardware perspective, its schematic hardware architecture diagram may specifically refer to FIG. 8A. It includes a processor 811 and a machine-readable storage medium 812. The machine-readable storage medium 812 stores machine-executable instructions executable by the processor 811, and the processor 811 is used to execute the machine-executable instructions to implement the decoding methods of Embodiments 1 to 22 of the present invention.
[0386] Based on the same concept as the above method, for the encoding device (also called a video encoder) provided by an embodiment of the present invention, from a hardware perspective, its schematic hardware architecture diagram may specifically refer to FIG. 8B. It includes a processor 821 and a machine-readable storage medium 822. The machine-readable storage medium 822 stores machine-executable instructions executable by the processor 821, and the processor 821 is used to execute the machine-executable instructions to implement the encoding methods of Embodiments 1 to 22 of the present invention.
[0387] Based on the same concept as the above method, an embodiment of the present invention further provides a machine-readable storage medium storing several computer instructions. When the computer instructions are executed by a processor, the methods disclosed in the above examples of the present invention, such as the decoding method or the encoding method in each of the above embodiments, can be implemented.
[0388] Based on the same concept as the above method, an embodiment of the present invention further provides a computer application program. When the computer application program is executed by a processor, the decoding method or the encoding method disclosed in the above examples of the present invention can be implemented.
[0389] Based on the same concept as the above method, an embodiment of the present invention further provides an image decoding device based on a neural network. The device is applied to the decoding side. The device includes a decoding module for decoding control parameters and image information corresponding to the current block from a bitstream, an acquisition module for obtaining neural network information corresponding to a decoding processing unit from the control parameters, and generating a decoding neural network corresponding to the decoding processing unit based on the neural network information, and a processing module for determining input features corresponding to the decoding processing unit based on the image information, and processing the input features based on the decoding neural network to obtain output features corresponding to the decoding processing unit.
[0390] Exemplarily, when the neural network information includes basic layer information and enhancement layer information, when generating a decoding neural network corresponding to the decoding processing unit based on the neural network information, the acquisition module specifically determines a basic layer corresponding to the decoding processing unit based on the basic layer information, determines an enhancement layer corresponding to the decoding processing unit based on the enhancement layer information, and is used to generate a decoding neural network corresponding to the decoding processing unit based on the basic layer and the enhancement layer.
[0391] Exemplarily, when the acquisition module determines a basic layer corresponding to the decoding processing unit based on the basic layer information, specifically, the basic layer information includes a basic layer default network usage flag bit, and when the basic layer default network usage flag bit indicates that the basic layer uses a default network, it is used to obtain the basic layer of the default network structure.
[0392] Exemplarily, when the acquisition module determines the base layer corresponding to the decoding processing unit based on the base layer information, specifically, the base layer information includes a base layer pre-designed network usage flag bit and a base layer pre-designed network index number, and when the base layer pre-designed network usage flag bit indicates that the base layer uses a pre-designed network, it is used to select the base layer of the pre-designed network structure corresponding to the base layer pre-designed network index number from the pre-designed neural network pool, and the pre-designed neural network pool includes network layers of at least one pre-designed network structure.
[0393] Exemplarily, when the acquisition module determines the enhancement layer corresponding to the decoding processing unit based on the enhancement layer information, specifically, the enhancement layer information includes an enhancement layer default network usage flag bit, and when the enhancement layer default network usage flag bit indicates that the enhancement layer uses a default network, it is used to obtain the enhancement layer of the default network structure.
[0394] Exemplarily, when the acquisition module determines the enhancement layer corresponding to the decoding processing unit based on the enhancement layer information, specifically, the enhancement layer information includes an enhancement layer pre-designed network usage flag bit and an enhancement layer pre-designed network index number, and when the enhancement layer pre-designed network usage flag bit indicates that the enhancement layer uses a pre-designed network, it is used to select the enhancement layer of the pre-designed network structure corresponding to the enhancement layer pre-designed network index number from the pre-designed neural network pool, and the pre-designed neural network pool includes network layers of at least one pre-designed network structure.
[0395] Exemplarily, when the acquisition module determines the enhancement layer corresponding to the decoding processing unit based on the enhancement layer information, specifically, when the enhancement layer information includes network parameters for generating an enhancement layer, an enhancement layer corresponding to the decoding processing unit is generated based on the network parameters. The network parameters include at least one of the number of layers of a neural network, a transposed convolution layer flag bit, the number of transposed convolution layers, the quantization stride of each transposed convolution layer, the number of channels of each transposed convolution layer, the size of a convolution kernel, the number of filterings, a filtering size index, a filtering coefficient zero flag bit, a filtering coefficient, an activation layer flag bit, and an activation layer type.
[0396] Exemplarily, the image information includes coefficient hyperparameter feature information and image feature information. When the processing module determines the input features corresponding to the decoding processing unit based on the image information and processes the input features based on the decoding neural network to obtain the output features corresponding to the decoding processing unit, specifically, when executing a decoding process for generating coefficient hyperparameter feature coefficients, a coefficient hyperparameter feature coefficient reconstruction value is determined based on the coefficient hyperparameter feature information, and an inverse transformation process is performed on the coefficient hyperparameter feature coefficient reconstruction value based on the decoding neural network to obtain a coefficient hyperparameter feature value. When executing a decoding process for inverse-transforming image features, an image feature reconstruction value is determined based on the image feature information, and an inverse transformation process is performed on the image feature reconstruction value based on the decoding neural network to obtain a low-level image feature value. The coefficient hyperparameter feature value is used for decoding the image feature information from a bitstream, and the low-level image feature value is used for obtaining a reconstructed image block corresponding to the current block.
[0397] Exemplarily, when the processing module determines the coefficient hyperparameter feature coefficient reconstruction value based on the coefficient hyperparameter feature information, specifically, when the control parameter includes the first enable information and the first enable information indicates enabling the first inverse quantization process, it is used to perform inverse quantization on the coefficient hyperparameter feature information to obtain the coefficient hyperparameter feature coefficient reconstruction value.
[0398] Exemplarily, when the processing module determines the image feature reconstruction value based on the image feature information, specifically, when the control parameter includes the second enable information and the second enable information indicates enabling the second inverse quantization process, it is used to perform inverse quantization on the image feature information to obtain the image feature reconstruction value.
[0399] Exemplarily, when the control parameter includes the third enable information and the third enable information indicates enabling the high-quality processing, when the processing module further executes the high-quality decoding process, it is used to obtain the low-level image feature value and perform enhancement processing on the low-level image feature value based on the decoding neural network to obtain the reconstructed image block corresponding to the current block.
[0400] Based on the same concept as the above method, an embodiment of the present invention further provides an image encoding device based on a neural network. The device is applied to the encoding side. The device determines input features corresponding to an encoding processing unit based on a current block, processes the input features based on an encoding neural network corresponding to the encoding processing unit to obtain output features corresponding to the encoding processing unit, and has a processing module for determining image information corresponding to the current block based on the output features, and an acquisition module for acquiring control parameters corresponding to the current block. The control parameters include neural network information corresponding to a decoding processing unit, and the neural network information is used to determine a decoding neural network corresponding to the decoding processing unit. The device further includes an encoding module for encoding the image information and the control parameters corresponding to the current block into a bitstream.
[0401] Exemplarily, the neural network information includes base layer information and enhancement layer information, and the decoding neural network includes a base layer determined based on the base layer information and an enhancement layer determined based on the enhancement layer information.
[0402] Exemplarily, the base layer information includes a base layer default network usage flag bit, and when the base layer default network usage flag bit indicates that the base layer uses a default network, the decoding neural network uses a base layer of the default network structure.
[0403] Exemplarily, the basic layer information includes a basic layer pre-designed network usage flag bit and a basic layer pre-designed network index number, and when the basic layer pre-designed network usage flag bit indicates that the basic layer uses a pre-designed network, the decoding neural network uses the basic layer of the pre-designed network structure corresponding to the basic layer pre-designed network index number selected from the pre-designed neural network pool, and the pre-designed neural network pool includes network layers of at least one pre-designed network structure.
[0404] Exemplarily, the enhancement layer information includes an enhancement layer default network usage flag bit, and when the enhancement layer default network usage flag bit indicates that the enhancement layer uses a default network, the decoding neural network uses the enhancement layer of the default network structure.
[0405] Exemplarily, the enhancement layer information includes an enhancement layer pre-designed network usage flag bit and an enhancement layer pre-designed network index number, and when the enhancement layer pre-designed network usage flag bit indicates that the enhancement layer uses a pre-designed network, the decoding neural network uses the enhancement layer of the pre-designed network structure corresponding to the enhancement layer pre-designed network index number selected from the pre-designed neural network pool, and the pre-designed neural network pool includes network layers of at least one pre-designed network structure.
[0406] Exemplarily, when the enhancement layer information includes network parameters for generating an enhancement layer, the decoding neural network uses the enhancement layer generated based on the network parameters, and the network parameters include at least one of the number of layers of the neural network, the transposed convolution layer flag bit, the number of transposed convolution layers, the quantization stride of each transposed convolution layer, the number of channels of each transposed convolution layer, the size of the convolution kernel, the number of filterings, the filtering size index, the filtering coefficient zero flag bit, the filtering coefficient, the activation layer flag bit, and the activation layer type.
[0407] Exemplarily, the processing module further divides the current image into N non-overlapping image blocks, where N is a positive integer, performs boundary padding on each image block to obtain an image block after boundary padding, and when performing boundary padding on each image block, the padding value does not depend on the reconstructed pixel values of adjacent image blocks, and is used to generate N current blocks based on the image block after boundary padding.
[0408] Exemplarily, the processing module further divides the current image into a plurality of basic blocks, each basic block includes at least one image block, performs boundary padding on each image block to obtain an image block after boundary padding, and when performing boundary padding on each image block, the padding value of the image block is allowed not to depend on the reconstructed pixel values of other image blocks within the same basic block and to depend on the reconstructed pixel values of image blocks within different basic blocks, and is used to generate a plurality of current blocks based on the image block after boundary padding.
[0409] It should be understood by those skilled in the art that embodiments of the present invention can be provided as a method, a system, or a computer program product. The present invention may adopt the form of embodiments implemented by hardware, embodiments implemented by software, or embodiments in combination of software and hardware. Embodiments of the present invention may also adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0410] The above are only embodiments of the present invention and do not limit the present invention. For those skilled in the art, various modifications and changes are possible to the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle scope of the present invention shall all be included within the scope of the claims of the present invention.
Explanation of Reference Numerals
[0411] 811 Processor 812 Machine-readable Storage Medium 821 Processor 822 Machine-readable Storage Medium
Claims
1. Decoding control parameters and image information corresponding to a current block from a bitstream; Obtaining neural network information corresponding to a decoding processing unit from the control parameters, and generating a decoding neural network corresponding to the decoding processing unit based on the neural network information; Determining input features corresponding to the decoding processing unit based on the image information, and processing the input features based on the decoding neural network to obtain output features corresponding to the decoding processing unit, A method for image decoding based on a neural network, characterized by the above.
2. When the neural network information includes basic layer information and enhancement layer information, the step of generating a decoding neural network corresponding to the decoding processing unit based on the neural network information includes: Determining a basic layer corresponding to the decoding processing unit based on the basic layer information; Determining an enhancement layer corresponding to the decoding processing unit based on the enhancement layer information; Generating a decoding neural network corresponding to the decoding processing unit based on the basic layer and the enhancement layer. The method according to claim 1, characterized by the above.
3. The step of determining a basic layer corresponding to the decoding processing unit based on the basic layer information includes: The basic layer information includes a basic layer default network usage flag bit, and when the basic layer default network usage flag bit indicates that the basic layer uses a default network, obtaining a basic layer of the default network structure. The method according to claim 2, characterized by the above.
4. The step of determining a basic layer corresponding to the decoding processing unit based on the basic layer information includes: The basic layer information includes a basic layer pre-designed network usage flag bit and a basic layer pre-designed network index number, and when the basic layer pre-designed network usage flag bit indicates that the basic layer uses a pre-designed network, selecting a basic layer of the pre-designed network structure corresponding to the basic layer pre-designed network index number from a pre-designed neural network pool. The pre-designed neural network pool includes network layers of at least one pre-designed network structure. The method according to claim 2, characterized in that.
5. Based on the reinforcement layer information, the step of determining the reinforcement layer corresponding to the decoding processing unit is: The reinforcement layer information includes a reinforcement layer default network usage flag bit, and when the reinforcement layer default network usage flag bit indicates that the reinforcement layer uses the default network, it includes the step of obtaining the reinforcement layer of the default network structure. The method according to claim 2, characterized in that.
6. Based on the reinforcement layer information, the step of determining the reinforcement layer corresponding to the decoding processing unit is: The reinforcement layer information includes a reinforcement layer pre-designed network usage flag bit and a reinforcement layer pre-designed network index number, and when the reinforcement layer pre-designed network usage flag bit indicates that the reinforcement layer uses the pre-designed network, it includes the step of selecting, from the pre-designed neural network pool, the reinforcement layer of the pre-designed network structure corresponding to the reinforcement layer pre-designed network index number. The pre-designed neural network pool includes network layers of at least one pre-designed network structure. The method according to claim 2, characterized in that.
7. Based on the reinforcement layer information, the step of determining the reinforcement layer corresponding to the decoding processing unit is: When the reinforcement layer information includes network parameters for generating the reinforcement layer, it includes the step of generating the reinforcement layer corresponding to the decoding processing unit based on the network parameters. The network parameters are: At least one of the number of layers of the neural network, the transposed convolution layer flag bit, the number of transposed convolution layers, the quantization stride of each transposed convolution layer, the number of channels of each transposed convolution layer, the size of the convolution kernel, the number of filterings, the filtering size index, the filtering coefficient zero flag bit, the filtering coefficient, the activation layer flag bit, and the activation layer type. The method according to claim 2, characterized in that.
8. The image information includes coefficient hyperparameter feature information and image feature information. The step of determining input features corresponding to the decoding processing unit based on the image information and processing the input features based on the decoding neural network to obtain output features corresponding to the decoding processing unit is When executing the decoding process of coefficient hyperparameter feature generation, based on the coefficient hyperparameter feature information, determining a coefficient hyperparameter feature coefficient reconstruction value, and based on the decoding neural network, performing an inverse transformation process on the coefficient hyperparameter feature coefficient reconstruction value to obtain a coefficient hyperparameter feature value, where the coefficient hyperparameter feature value is used to decode the image feature information from the bitstream, and the step When executing the decoding process of image feature inverse transformation, based on the image feature information, determining an image feature reconstruction value, and based on the decoding neural network, performing an inverse transformation process on the image feature reconstruction value to obtain an image low-level feature value, where the image low-level feature value is used to obtain a reconstructed image block corresponding to the current block, and the step The method according to any one of claims 1 to 7, characterized by the above.
9. The step of determining a coefficient hyperparameter feature coefficient reconstruction value based on the coefficient hyperparameter feature information is When the control parameter includes first enable information and the first enable information indicates enabling a first inverse quantization process, the step includes performing inverse quantization on the coefficient hyperparameter feature information to obtain a coefficient hyperparameter feature coefficient reconstruction value. The step of determining an image feature reconstruction value based on the image feature information is When the control parameter includes second enable information and the second enable information indicates enabling a second inverse quantization process, the step includes performing inverse quantization on the image feature information to obtain an image feature reconstruction value. The method according to claim 8, characterized by the above.
10. When the control parameter includes third enable information and the third enable information indicates enabling high-quality processing, when executing a high-quality decoding process, the method further includes obtaining the low-level image feature value and performing enhancement processing on the low-level image feature value based on the decoding neural network to obtain a reconstructed image block corresponding to the current block. The method according to claim 8, characterized in that. **Claim 11** Determining input features corresponding to an encoding processing unit based on a current block, processing the input features based on an encoding neural network corresponding to the encoding processing unit to obtain output features corresponding to the encoding processing unit, and determining image information corresponding to the current block based on the output features. Obtaining a control parameter corresponding to the current block, where the control parameter includes neural network information corresponding to a decoding processing unit, and the neural network information is used to determine a decoding neural network corresponding to the decoding processing unit. Encoding the image information and the control parameter corresponding to the current block into a bitstream. A neural network-based image encoding method, characterized in that. **Claim 12** The neural network information includes basic layer information and enhancement layer information, and the decoding neural network includes a basic layer determined based on the basic layer information and an enhancement layer determined based on the enhancement layer information. The method according to claim 11, characterized in that. **Claim 13** The basic layer information includes a basic layer default network usage flag bit, and when the basic layer default network usage flag bit indicates that the basic layer uses a default network, the decoding neural network uses a basic layer with a default network structure. The method according to claim 12, characterized in that. **Claim 14** The basic layer information includes a basic layer pre-designed network usage flag bit and a basic layer pre-designed network index number, and when the basic layer pre-designed network usage flag bit indicates that the basic layer uses the pre-designed network, the decoding neural network uses the basic layer of the pre-designed network structure corresponding to the basic layer pre-designed network index number selected from the pre-designed neural network pool, The pre-designed neural network pool includes network layers of at least one pre-designed network structure, The method according to claim 12, characterized in that.
15. The enhancement layer information includes an enhancement layer default network usage flag bit, and when the enhancement layer default network usage flag bit indicates that the enhancement layer uses the default network, the decoding neural network uses the enhancement layer of the default network structure, The method according to claim 12, characterized in that.
16. The enhancement layer information includes an enhancement layer pre-designed network usage flag bit and an enhancement layer pre-designed network index number, and when the enhancement layer pre-designed network usage flag bit indicates that the enhancement layer uses the pre-designed network, the decoding neural network uses the enhancement layer of the pre-designed network structure corresponding to the enhancement layer pre-designed network index number selected from the pre-designed neural network pool, The pre-designed neural network pool includes network layers of at least one pre-designed network structure, The method according to claim 12, characterized in that.
17. When the enhancement layer information includes network parameters for generating an enhancement layer, the decoding neural network uses the enhancement layer generated based on the network parameters, The network parameters are, including at least one of the number of layers of the neural network, the transposed convolution layer flag bit, the number of transposed convolution layers, the quantization stride of each transposed convolution layer, the number of channels of each transposed convolution layer, the size of the convolution kernel, the number of filterings, the filtering size index, the filtering coefficient zero flag bit, the filtering coefficient, the activation layer flag bit, and the activation layer type The method according to claim 12, characterized in that.
18. Before determining the input features corresponding to the encoding processing unit based on the current block, further dividing the current image into N non-overlapping image blocks, where N is a positive integer, the step of performing boundary padding on each image block to obtain a boundary-padded image block, and when performing boundary padding on each image block, the padding value does not depend on the reconstructed pixel values of adjacent image blocks, the step of generating N current blocks based on the boundary-padded image blocks, including The method according to any one of claims 11 to 17, characterized in that.
19. Before determining the input features corresponding to the encoding processing unit based on the current block, further dividing the current image into a plurality of basic blocks, where each basic block includes at least one image block, the step of performing boundary padding on each image block to obtain a boundary-padded image block, and when performing boundary padding on each image block, the padding value of the image block does not depend on the reconstructed pixel values of other image blocks within the same basic block, and is permitted to depend on the reconstructed pixel values of image blocks within different basic blocks, the step of generating a plurality of current blocks based on the boundary-padded image blocks, including The method according to any one of claims 11 to 17, characterized in that.
20. a memory configured to store video data, decoding the control parameters and image information corresponding to the current block from the bitstream A step of obtaining neural network information corresponding to the decoding processing unit from the control parameter and generating a decoding neural network corresponding to the decoding processing unit based on the neural network information; A step of determining input features corresponding to the decoding processing unit based on the image information, processing the input features based on the decoding neural network, and obtaining output features corresponding to the decoding processing unit, and a decoder configured to perform the steps; An image decoding apparatus based on a neural network, characterized in that.
21. A memory configured to store video data; A step of determining input features corresponding to the encoding processing unit based on the current block, processing the input features based on the encoding neural network corresponding to the encoding processing unit to obtain output features corresponding to the encoding processing unit, and determining image information corresponding to the current block based on the output features; A step of obtaining a control parameter corresponding to the current block, the control parameter including neural network information corresponding to the decoding processing unit, and the neural network information being used to determine the decoding neural network corresponding to the decoding processing unit; A step of encoding the image information and the control parameter corresponding to the current block into a bitstream, and an encoder configured to perform the steps; An image encoding apparatus based on a neural network, characterized in that.
22. A decoding device including a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions executable by the processor; The processor is used to execute the machine-executable instructions to implement the method according to any one of Claims 1 to 10; A decoding device, characterized in that.
23. An encoding device including a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions executable by the processor; The processor is used to execute the machine-executable instructions to implement the method according to any one of Claims 11 to 19; An encoding device, characterized in that.
Citation Information
Patent Citations
Data compression using conditional entropy models
US20200027247A1
Methods And Apparatuses For Learned Image Compression
US20200160565A1
Method and apparatus for variable rate compression with a conditional autoencoder
US20200304147A1
Techniques for signaling neural network topology and parameters in coded video stream
WO2022146522A1
Progressive data compression using artificial neural networks
WO2022159897A1