Decoding method, device and equipment

By optimizing the activation function and adjusting the network structure, the problems of poor decoding performance and high complexity of neural networks are solved, and the effect of improving decoding performance and image quality at low complexity is achieved.

CN120343280APending Publication Date: 2025-07-18HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411384621.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing neural network-based encoding and decoding methods have problems with poor decoding performance and high complexity.

Method used

By optimizing the activation function, the target activation function of upper and lower clamps is used to nonlinearly adjust the features, reducing the complexity of the activation function, and combining the optimization of the synthetic transform network and the mean hyperparameter decoding network, the robustness and decoding performance of the neural network are improved.

Benefits of technology

While maintaining low complexity, it effectively ensures the quality of reconstructed image blocks, improves decoding performance, reduces the phenomenon of neural network decoding garbled code on high-frequency images, improves decoding performance and reduces complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343280A_ABST
    Figure CN120343280A_ABST
Patent Text Reader

Abstract

The invention provides a decoding method, device and equipment thereof. The decoding method comprises the following steps: decoding a first code stream corresponding to a current image block to obtain a coefficient hyper-parameter feature; determining a probability distribution parameter based on the coefficient hyper-parameter feature, and decoding a second code stream corresponding to the current image block based on the probability distribution parameter to obtain a residual feature; determining a target mean value feature based on the coefficient hyper-parameter feature; determining a reconstruction feature based on the target mean feature and the residual feature; inputting the reconstruction features into a synthesis transformation network to obtain a reconstruction image block corresponding to the current image block; wherein the synthetic transformation network comprises a nonlinear residual network layer, and the nonlinear residual network layer comprises an activation layer and a jump connection structure. Through the technical scheme of the invention, the decoding performance can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of coding and decoding technologies, and in particular, to a decoding method, apparatus, and device thereof. Background Art

[0002] For the purpose of saving space, video images are transmitted after being encoded. A complete video encoding process may include prediction, transformation, quantization, entropy encoding, filtering, etc. For the prediction process, the prediction process may include intra-frame prediction and inter-frame prediction. Inter-frame prediction refers to using the correlation in the video time domain and the pixels of adjacent encoded images to predict the current pixel, so as to effectively remove the redundancy in the video time domain. Intra-frame prediction refers to using the correlation in the video spatial domain and the pixels of the encoded blocks in the current frame image to predict the current pixel, so as to remove the redundancy in the video spatial domain.

[0003] With the rapid development of deep learning, deep learning has achieved success in many high-level computer vision problems. For example, in the fields of image classification, object detection, etc., deep learning has gradually been applied in the field of coding and decoding, that is, a neural network can be used to encode and decode images. Although the coding and decoding method based on a neural network shows great performance potential, there are still problems such as poor decoding performance and high complexity in the coding and decoding method based on a neural network. Summary of the Invention

[0004] In view of this, this application provides a decoding method, apparatus, and device thereof, which can improve decoding performance and reduce complexity.

[0005] This application provides a decoding method, which is applied to a decoding end. The method includes:

[0006] Decoding a first bitstream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block;

[0007] Determining probability distribution parameters based on the coefficient hyperparameter feature, and decoding a second bitstream corresponding to the current image block based on the probability distribution parameters to obtain a residual feature corresponding to the current image block;

[0008] Determining a target mean feature corresponding to the current image block based on the coefficient hyperparameter feature;

[0009] Determining a reconstructed feature corresponding to the current image block based on the target mean feature and the residual feature;

[0010] Input the reconstructed feature into a synthesis transformation network to obtain a reconstructed image block corresponding to the current image block; wherein, the synthesis transformation network includes a non-linear residual network layer, and the non-linear residual network layer includes at least an activation layer and a skip connection structure; the output feature of the activation layer is not greater than an upper threshold value, and / or, the output feature of the activation layer is not less than a lower threshold value.

[0011] The present application provides a decoding method, which is applied to a decoding end, and the method includes:

[0012] Decode a first bitstream corresponding to the current image block to obtain a coefficient hyperparameter feature corresponding to the current image block;

[0013] Input the coefficient hyperparameter feature into a probability hyperparameter decoding network to obtain probability distribution parameters, and decode a second bitstream corresponding to the current image block based on the probability distribution parameters to obtain a residual feature corresponding to the current image block;

[0014] Input the coefficient hyperparameter feature into a mean hyperparameter decoding network to obtain an initial mean feature corresponding to the current image block, and determine a target mean feature corresponding to the current image block based on the initial mean feature;

[0015] Determine a reconstructed feature corresponding to the current image block based on the target mean feature and the residual feature;

[0016] Determine a reconstructed image block corresponding to the current image block based on the reconstructed feature;

[0017] Wherein, the probability hyperparameter decoding network includes an activation layer, and the output feature of the activation layer is not greater than an upper threshold value, and / or, the output feature of the activation layer is not less than a lower threshold value;

[0018] And / or, the mean hyperparameter decoding network includes an activation layer, and the output feature of the activation layer is not greater than an upper threshold value, and / or, the output feature of the activation layer is not less than a lower threshold value.

[0019] The present application provides a decoding method, which is applied to a decoding end, and the method includes:

[0020] Decode a first bitstream corresponding to the current image block to obtain a coefficient hyperparameter feature corresponding to the current image block;

[0021] Input the coefficient hyperparameter feature into a hyperparameter decoding network to obtain a reference feature corresponding to the current image block;

[0022] Input the reference feature and the already obtained reconstructed feature of the current image block into a context model to obtain probability distribution parameters and a target mean feature corresponding to the current image block;

[0023] Decode the second bitstream corresponding to the current image block based on the probability distribution parameter to obtain the residual feature corresponding to the current image block, and determine the reconstruction feature based on the target mean feature and the residual feature;

[0024] Determine the reconstructed image block corresponding to the current image block based on the reconstruction feature;

[0025] Wherein, the hyperparameter decoding network includes an activation layer, and the output feature of the activation layer is not greater than the upper limit threshold, and / or the output feature of the activation layer is not less than the lower limit threshold;

[0026] And / or, the context model includes an activation layer, and the output feature of the activation layer is not greater than the upper limit threshold, and / or the output feature of the activation layer is not less than the lower limit threshold.

[0027] The present application provides a decoding method, which is applied to a decoding end, and the method includes:

[0028] Decode the first bitstream corresponding to the current image block to obtain the coefficient hyperparameter feature corresponding to the current image block;

[0029] Determine the probability distribution parameter based on the coefficient hyperparameter feature, and decode the second bitstream corresponding to the current image block based on the probability distribution parameter to obtain the residual feature corresponding to the current image block;

[0030] Input the coefficient hyperparameter feature into the mean hyperparameter decoding network to obtain the initial mean feature corresponding to the current image block, and determine the target mean feature corresponding to the current image block based on the initial mean feature; wherein, the mean hyperparameter decoding network sequentially includes a processing layer, an upsampling layer for performing an upsampling operation, and a cropping layer; wherein, the feature output by the upsampling layer is obtained as the initial mean feature after passing through the cropping layer;

[0031] Determine the reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature;

[0032] Determine the reconstructed image block corresponding to the current image block based on the reconstruction feature.

[0033] The present application provides a decoding device, which is applied to a decoding end, and the device includes:

[0034] A decoding module, configured to decode the first bitstream corresponding to the current image block to obtain the coefficient hyperparameter feature corresponding to the current image block; determine the probability distribution parameter based on the coefficient hyperparameter feature, and decode the second bitstream corresponding to the current image block based on the probability distribution parameter to obtain the residual feature corresponding to the current image block;

[0035] A determination module, configured to determine a target mean feature corresponding to the current image block based on the coefficient hyperparameter feature; and determine a reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature;

[0036] An acquisition module, configured to input the reconstruction feature into a synthesis transformation network to obtain a reconstructed image block corresponding to the current image block; wherein, the synthesis transformation network includes a non-linear residual network layer, and the non-linear residual network layer at least includes an activation layer and a skip connection structure; wherein, the output feature of the activation layer is not greater than an upper threshold value, and / or, the output feature of the activation layer is not less than a lower threshold value.

[0037] The present application provides a decoding device, which is applied to a decoding end, and the device includes:

[0038] A decoding module, configured to decode a first bitstream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block; input the coefficient hyperparameter feature into a probability hyperparameter decoding network to obtain probability distribution parameters, and decode a second bitstream corresponding to the current image block based on the probability distribution parameters to obtain a residual feature corresponding to the current image block;

[0039] A determination module, configured to input the coefficient hyperparameter feature into a mean hyperparameter decoding network to obtain an initial mean feature corresponding to the current image block, and determine a target mean feature corresponding to the current image block based on the initial mean feature; determine a reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature;

[0040] An acquisition module, configured to determine a reconstructed image block corresponding to the current image block based on the reconstruction feature;

[0041] Wherein, the probability hyperparameter decoding network includes an activation layer, and the output feature of the activation layer is not greater than an upper threshold value, and / or, the output feature of the activation layer is not less than a lower threshold value;

[0042] And / or, the mean hyperparameter decoding network includes an activation layer, and the output feature of the activation layer is not greater than an upper threshold value, and / or, the output feature of the activation layer is not less than a lower threshold value.

[0043] The present application provides a decoding device, which is applied to a decoding end, and the device includes:

[0044] A decoding module, configured to decode a first bitstream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block;

[0045] A determination module, configured to input the coefficient hyperparameter feature to a hyperparameter decoding network to obtain a reference feature corresponding to the current image block; input the reference feature and the obtained reconstruction feature of the current image block to a context model to obtain a probability distribution parameter and a target mean feature corresponding to the current image block;

[0046] The decoding module is further configured to decode the second bitstream corresponding to the current image block based on the probability distribution parameter to obtain a residual feature corresponding to the current image block;

[0047] The determination module is further configured to determine the reconstruction feature based on the target mean feature and the residual feature;

[0048] An acquisition module, configured to determine a reconstructed image block corresponding to the current image block based on the reconstruction feature;

[0049] Wherein, the hyperparameter decoding network includes an activation layer, and the output feature of the activation layer is not greater than an upper threshold value, and / or, the output feature of the activation layer is not less than a lower threshold value;

[0050] And / or, the context model includes an activation layer, and the output feature of the activation layer is not greater than an upper threshold value, and / or, the output feature of the activation layer is not less than a lower threshold value.

[0051] The present application provides a decoding device, which is applied to a decoding end. The device includes:

[0052] A decoding module, configured to decode a first bitstream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block; determine a probability distribution parameter based on the coefficient hyperparameter feature, and decode a second bitstream corresponding to the current image block based on the probability distribution parameter to obtain a residual feature corresponding to the current image block;

[0053] A determination module, configured to input the coefficient hyperparameter feature to a mean hyperparameter decoding network to obtain an initial mean feature corresponding to the current image block, and determine a target mean feature corresponding to the current image block based on the initial mean feature; wherein, the mean hyperparameter decoding network sequentially includes a processing layer, an upsampling layer for performing an upsampling operation, and a cropping layer; wherein, the feature output by the upsampling layer obtains the initial mean feature after passing through the cropping layer; determine a reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature;

[0054] An acquisition module, configured to determine a reconstructed image block corresponding to the current image block based on the reconstruction feature.

[0055] This application provides a decoding device, which includes: a processor and a machine-readable storage medium storing machine-executable instructions that can be executed by the processor;

[0056] The processor is configured to execute the machine-executable instructions to implement the above decoding method.

[0057] This application provides an electronic device, which includes: a processor and a machine-readable storage medium storing machine-executable instructions that can be executed by the processor;

[0058] The processor is configured to execute the machine-executable instructions to implement the above decoding method.

[0059] This application provides a machine-readable storage medium storing a number of computer instructions, which can implement the above decoding method when executed by a processor.

[0060] This application provides a computer program, which implements the above decoding method when executed by a processor.

[0061] As can be seen from the above technical solutions, in the embodiments of this application, an end-to-end video image compression method is proposed, which can implement the decoding of video images based on a neural network. By optimizing the activation function and using a target activation function with upper and lower clamping to perform non-linear adjustment on features, the complexity of the activation function is reduced, and then the complexity of the neural network is reduced. The target activation function can improve the robustness of the neural network and reduce the phenomenon of decoding garbled codes on high-frequency images. This enables the neural network to effectively guarantee the quality of the reconstructed image blocks while maintaining low complexity, improve the quality of the reconstructed image, achieve the purpose of improving the decoding performance, and reduce the complexity. By optimizing the mean hyperparameter decoding network, the complexity of the neural network is greatly reduced while maintaining the performance of the neural network unchanged, achieving the purpose of improving the decoding performance and reducing the complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 is a schematic diagram of a three-dimensional feature matrix in an embodiment of this application;

[0063] Figure 2 is a schematic flowchart of a decoding method in an embodiment of this application;

[0064] Figure 3 is a schematic flowchart of an encoding method in an embodiment of this application;

[0065] Figures 4A - 4D is a schematic diagram of an encoding and decoding framework in an embodiment of this application;

[0066] Figure 5A It is a schematic diagram of a probability hyperparameter decoding network in an implementation manner of the present application;

[0067] Figure 5B It is a schematic diagram of a Relu activation function in an implementation manner of the present application;

[0068] Figure 5C It is a schematic diagram of a probability hyperparameter decoding network in an implementation manner of the present application;

[0069] Figure 5D It is a schematic diagram of a Relu6 activation function in an implementation manner of the present application;

[0070] Figure 6A and Figure 6B It is a schematic diagram of a mean hyperparameter decoding network in an implementation manner of the present application;

[0071] Figure 6C and Figure 6D It is a schematic diagram of a context model in an implementation manner of the present application;

[0072] Figures 7A - 7I It is a schematic diagram of a synthesis transformation network in an implementation manner of the present application;

[0073] Figure 8A It is a schematic diagram of a ResAU layer in an implementation manner of the present application;

[0074] Figure 8B It is a schematic diagram of a LeakyRelu activation function in an implementation manner of the present application;

[0075] Figures 8C - 8H It is a schematic diagram of a ResAU layer in an implementation manner of the present application;

[0076] Figures 9A - 9C It is a schematic diagram of an upsampling operation in an implementation manner of the present application;

[0077] Figure 9D It is a schematic diagram of a mean hyperparameter decoding network in an implementation manner of the present application;

[0078] Figures 10A - 10F It is a schematic diagram of a convolution operation in an implementation manner of the present application;

[0079] Figure 11A It is a hardware structure diagram of a decoding end device in an implementation manner of the present application;

[0080] Figure 11B It is a hardware structure diagram of an encoding end device in an implementation manner of the present application. Specific implementation manners

[0081] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the embodiments of the present application and the claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to any or all possible combinations of one or more of the associated listed items. It should be understood that although the terms first, second, third, etc. may be used in the embodiments of the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information, depending on the context. In addition, the word "if" can be interpreted as "when", or "while", or "in response to determining".

[0082] In the embodiments of the present application, a decoding method, apparatus, and device are proposed, which may involve the following concepts:

[0083] Entropy Encoding: Entropy encoding is an encoding that does not lose any information according to the entropy principle during the encoding process. The information entropy is the average amount of information of the information source (a measure of uncertainty). The encoding methods of entropy encoding may include, but are not limited to, Shannon encoding, Huffman encoding, and arithmetic coding.

[0084] Neural Network (NN): A neural network refers to an artificial neural network, which is an operation model composed of a large number of nodes (or called neurons) connected to each other. In a neural network, the neuron processing unit can represent different objects, such as features, letters, concepts, or some meaningful abstract patterns. The types of processing units in a neural network can be divided into three categories: input units, output units, and hidden units. The input units receive signals and data from the external world; the output units implement the output of the processing results; the hidden units are located between the input and output units and cannot be observed from the outside of the system. The connection weights between neurons reflect the connection strength between units, and the representation and processing of information are reflected in the connection relationship of the processing units. A neural network is a non-programmed, brain-like style of information processing method. The essence of a neural network is to obtain a parallel distributed information processing function through the transformation and dynamic behavior of the neural network and to imitate the information processing function of the human brain nervous system to varying degrees and levels. In the field of video processing, common neural networks may include, but are not limited to, Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), fully connected network, etc.

[0085] Convolutional Neural Network (CNN): A convolutional neural network is a feedforward neural network and one of the most representative network structures in deep learning technology. The artificial neurons of a convolutional neural network can respond to the surrounding units within a certain coverage range and perform excellently in large-scale image processing. The basic structure of a convolutional neural network can include two layers. One is the feature extraction layer (also known as the convolutional layer), where the input of each neuron is connected to the local receptive field of the previous layer and extracts the features of this local area. Once the local features are extracted, the positional relationship between it and other features is also determined. The other is the feature mapping layer (also known as the activation layer). Each computational layer of the neural network consists of multiple feature maps, and each feature map is a plane where the weights of all neurons on the plane are equal. The feature mapping structure can use functions such as the Sigmoid function, ReLU function, Leaky-ReLU function, PReLU function, GDN function, etc. as the activation function of the convolutional network. In addition, since the neurons on a mapping surface share weights, the number of free parameters of the network is reduced.

[0086] Exemplarily, one of the advantages of a convolutional neural network compared to image processing algorithms is that it avoids the complex preprocessing process of images (such as extracting artificial features, etc.) and can directly input the original image for end-to-end learning. One of the advantages of a convolutional neural network compared to ordinary neural networks is that ordinary neural networks all use a fully connected method, that is, all neurons from the input layer to the hidden layer are fully connected, which will lead to a huge number of parameters, making the network training time-consuming or even difficult to train. However, a convolutional neural network avoids this difficulty through methods such as local connection and weight sharing.

[0087] Deconvolution layer: The deconvolution layer is also known as the transposed convolution layer. The working process of the deconvolution layer is very similar to that of the convolutional layer. The main difference is that the deconvolution layer will make the output larger than the input (it can also be the same) through padding. If the stride is 1, it means the output size is equal to the input size; if the stride is N, it means the width of the output feature is N times the width of the input feature, and the height of the output feature is N times the height of the input feature.

[0088] Generalization Ability: Generalization ability can refer to the adaptability of machine learning algorithms to fresh samples. The purpose of learning is to learn the rules hidden behind the data pairs, so that for data outside the learning set with the same rules, the trained network can also give appropriate outputs, and this ability can be called generalization ability.

[0089] Feature: The feature involved in this application is a three-dimensional feature matrix or tensor of C*W*H. See Figure 1 As shown, it is a schematic diagram of a three-dimensional feature matrix. In the three-dimensional feature matrix, C represents the number of channels, H represents the feature height, and W represents the feature width. The three-dimensional feature matrix can be the input of a neural network or the output of a neural network.

[0090] Rate-Distortion Optimized: There are two major indicators for evaluating coding efficiency: bit rate and PSNR (Peak Signal to Noise Ratio). The smaller the bitstream, the larger the compression ratio, and the larger the PSNR, the better the quality of the reconstructed image. When selecting a mode, the discrimination formula is essentially a comprehensive evaluation of the two. For example, the cost corresponding to a mode: J(mode) = D + λ*R, where D represents Distortion, and usually the SSE index can be used to measure it. SSE refers to the mean square sum of the differences between the reconstructed image block and the source image. To consider the cost, the SAD index can also be used. SAD refers to the sum of the absolute values of the differences between the reconstructed image block and the source image; λ is the Lagrange multiplier, and R is the actual number of bits required for encoding the image block in this mode, including the total number of bits required for encoding mode information, motion information, residuals, etc. When selecting a mode, if the rate-distortion principle is used to make a comparison decision on the coding mode, the best coding performance can usually be guaranteed.

[0091] For each module at the encoding end, a very large number of coding tools have been proposed, and each tool often has multiple modes. For different video sequences, the coding tools that can obtain the optimal coding performance are often different. Therefore, during the encoding process, RDO (Rate-Distortion Opitimize) is usually used to compare the coding performance of different tools or modes to select the best mode. After determining the optimal tool or mode, the decision information of the tool or mode is transmitted by encoding marker information in the bitstream. Although this method brings a relatively high coding complexity, it can adaptively select the optimal mode combination for different contents and obtain the optimal coding performance. The decoding end can obtain the relevant mode information by directly parsing the flag information, and the impact on complexity is relatively small.

[0092] In the end-to-end general framework of image encoding and decoding, it mainly includes the feature main information part and the hyperprior side information part. The feature main information part includes the analysis network, quantization, normal entropy encoding, normal entropy decoding, and synthesis network. The hyperprior side information part includes the hyperprior analysis network, quantization, factorized entropy encoding, factorized entropy decoding, and hyperprior synthesis network. The image components are compressed and encoded and reconstructed and restored by the analysis network and synthesis network of the feature main information part respectively. The hyperprior side information part is mainly used to model the probability of the feature main information and guide the entropy encoding and decoding of the feature main information. In the end-to-end general framework of image encoding and decoding, there are problems such as the mean hyperparameter decoding network first performing upsampling and then convolution processing, with relatively high complexity, and the hardware complexity of the LeakyRelu function in the non-linear residual network layer of the synthesis transformation network being relatively high.

[0093] In view of the above findings, in the embodiments of the present application, a feature clipping and network simplification technology for video image decoding is proposed. For the non-linear residual network layer of the synthesis transformation network, an activation function with upper and lower clamping (used to replace the LeakyRelu function) can be used to non-linearly adjust the features, thereby solving the problem of relatively high hardware complexity of the LeakyRelu function. For the mean hyperparameter decoding network, convolution processing can be performed first and then upsampling, thereby solving the problem of relatively high complexity.

[0094] The decoding method and encoding method in the embodiments of the present application will be described in detail below in conjunction with several specific embodiments.

[0095] Embodiment 1: A decoding method is proposed in the embodiments of the present application. Refer to Figure 2 As shown in the flowchart of the decoding method, this method can be applied to the decoding end (also referred to as a video decoder), and this method may include:

[0096] Step 201: Decode the first bitstream corresponding to the current image block to obtain the coefficient hyperparameter feature corresponding to the current image block.

[0097] Step 202: Determine the probability distribution parameter based on the coefficient hyperparameter feature, and decode the second bitstream corresponding to the current image block based on the probability distribution parameter to obtain the residual feature corresponding to the current image block.

[0098] Step 203: Determine the target mean feature corresponding to the current image block based on the coefficient hyperparameter feature.

[0099] Step 204: Determine the reconstructed feature corresponding to the current image block based on the target mean feature and the residual feature.

[0100] Step 205: Input the reconstructed feature into the synthesis transformation network to obtain the reconstructed image patch corresponding to the current image patch. Exemplarily, the synthesis transformation network includes a non-linear residual network layer, and the non-linear residual network layer can at least include an activation layer and a skip connection structure; the output feature of the activation layer is not greater than the upper threshold, and / or the output feature of the activation layer is not less than the lower threshold.

[0101] Exemplarily, the non-linear residual network layer can be a ResAU network layer, or it can be other types of network layers, which is not limited herein, as long as it can implement non-linear residual processing. Hereinafter, the ResAU network layer will be taken as an example for illustration.

[0102] Exemplarily, the output feature of the activation layer of the non-linear residual network layer can be not greater than the upper threshold and not less than the lower threshold, that is, both the upper threshold and the lower threshold of the activation layer are restricted. In this case, the activation layer can use a target activation function with clamping to non-linearly adjust the feature, and the target activation function corresponds to the upper threshold and the lower threshold, and the output feature of the activation layer is not greater than the upper threshold and not less than the lower threshold.

[0103] Alternatively, the output feature of the activation layer of the non-linear residual network layer can be not less than the lower threshold, that is, only the lower threshold of the activation layer is restricted, and the upper threshold of the activation layer is not restricted. In this case, the activation layer can use the Relu activation function to non-linearly adjust the feature, and the Relu activation function corresponds to the lower threshold, and the output feature of the activation layer is not less than the lower threshold.

[0104] Alternatively, the output feature of the activation layer of the non-linear residual network layer can be not greater than the upper threshold, that is, only the upper threshold of the activation layer is restricted, and the lower threshold of the activation layer is not restricted. In this case, the activation layer can use a certain activation function to non-linearly adjust the feature, and the activation function corresponds to the upper threshold, and the output feature of the activation layer is not greater than the upper threshold.

[0105] Exemplarily, the non-linear residual network layer can further include a processing layer. For example, the non-linear residual network layer sequentially includes an activation layer, a processing layer, and a skip connection structure. Among them, the processing layer can include but is not limited to at least one of the following: one or more ordinary convolution layers, one or more transposed convolution layers, one or more deformable convolution layers, one or more depthwise separable convolution layers, one or more grouped convolution layers, one or more dilated convolution layers, one or more global processing unit layers. The global processing unit layer is used to implement the global processing function, which can be a transformer layer or other types of network layers, as long as it can implement the global processing function. Hereinafter, the transformer layer will be taken as an example for illustration.

[0106] For example, when there is an activation layer and the processing layer includes a network layer, the network layer can be located in front of the activation layer or behind the activation layer. When there is an activation layer and the processing layer includes multiple network layers, all the network layers can be located in front of the activation layer, or all of them can be located behind the activation layer. Alternatively, some of the network layers can be located in front of the activation layer and the remaining network layers can be located behind the activation layer.

[0107] For another example, when there are multiple activation layers (taking the first activation layer and the second activation layer as an example) and the processing layer includes a network layer, the network layer can be located in front of the first activation layer, between the first activation layer and the second activation layer, or behind the second activation layer. When the processing layer includes multiple network layers, all the network layers can be located in front of the first activation layer, between the first activation layer and the second activation layer, or behind the second activation layer. Alternatively, some of the network layers can be located in front of the first activation layer and the remaining network layers can be located between the first activation layer and the second activation layer. Or, some of the network layers can be located in front of the first activation layer and the remaining network layers can be located behind the second activation layer. Or, some of the network layers can be located in front of the first activation layer, some can be located between the first activation layer and the second activation layer, and the remaining network layers can be located behind the second activation layer. Or, some of the network layers can be located between the first activation layer and the second activation layer and the remaining network layers can be located behind the second activation layer.

[0108] Exemplarily, determining the probability distribution parameters based on the coefficient hyperparameter feature may include, but is not limited to: inputting the coefficient hyperparameter feature into a probability hyperparameter decoding network to obtain the probability distribution parameters. Exemplarily, the probability hyperparameter decoding network may at least include an activation layer, and the activation layer uses a target activation function with upper and lower clamping (such as the Relu6 activation function) to non-linearly adjust the feature. Wherein, the target activation function can correspond to an upper limit threshold and a lower limit threshold, and the output feature of the activation layer is not greater than the upper limit threshold and not less than the lower limit threshold.

[0109] Exemplarily, determining the target mean feature corresponding to the current image block based on the coefficient hyperparameter feature may include, but is not limited to: inputting the coefficient hyperparameter feature into a mean hyperparameter decoding network to obtain the initial mean feature corresponding to the current image block, and determining the target mean feature based on the initial mean feature. Exemplarily, the mean hyperparameter decoding network includes an activation layer, and the activation layer uses a target activation function with upper and lower clamping to non-linearly adjust the feature; wherein, the target activation function corresponds to an upper limit threshold and a lower limit threshold, and the output feature of the activation layer is not greater than the upper limit threshold and not less than the lower limit threshold.

[0110] Exemplarily, determining the target mean feature based on the initial mean feature may include, but is not limited to: using the initial mean feature as the target mean feature; or, inputting the initial mean feature and the obtained reconstruction feature of the current image block into a context model to obtain the target mean feature. Exemplarily, the context model includes an activation layer, and the activation layer non-linearly adjusts the feature using a target activation function with upper and lower clamping; wherein the target activation function corresponds to an upper threshold and a lower threshold, and the output feature of the activation layer is not greater than the upper threshold and not less than the lower threshold.

[0111] Exemplarily, determining the probability distribution parameter based on the coefficient hyperparameter feature and determining the target mean feature corresponding to the current image block based on the coefficient hyperparameter feature may include, but is not limited to: inputting the coefficient hyperparameter feature into a hyperparameter decoding network to obtain a reference feature corresponding to the current image block; inputting the reference feature and the obtained reconstruction feature of the current image block into a context model to obtain the probability distribution parameter and the target mean feature corresponding to the current image block. Exemplarily, the probability hyperparameter decoding network includes an activation layer, and the activation layer non-linearly adjusts the feature using a target activation function with upper and lower clamping; and / or, the context model includes an activation layer, and the activation layer non-linearly adjusts the feature using a target activation function with upper and lower clamping; the target activation function corresponds to an upper threshold and a lower threshold, and the output feature of the activation layer is not greater than the upper threshold and not less than the lower threshold.

[0112] Exemplarily, the upper threshold corresponding to the target activation function is a configured fixed upper threshold, and the lower threshold corresponding to the target activation function is a configured fixed lower threshold; or, the upper threshold corresponding to the target activation function is an adaptively learned upper threshold, and the lower threshold corresponding to the target activation function is an adaptively learned lower threshold.

[0113] Exemplarily, when adaptively learning the upper threshold, the same upper threshold is learned for all channels of the activation layer, or, the upper threshold is learned separately for each channel of the activation layer, and the upper thresholds corresponding to different channels are the same or different.

[0114] Exemplarily, when adaptively learning the lower threshold, the same lower threshold is learned for all channels of the activation layer, or, the lower threshold is learned separately for each channel of the activation layer, and the lower thresholds corresponding to different channels are the same or different.

[0115] An embodiment of the present application proposes a decoding method, which can be applied to the decoding end, and the method may include:

[0116] Step S11: Decode the first bitstream corresponding to the current image block to obtain the coefficient hyperparameter feature corresponding to the current image block.

[0117] Step S12: Input the coefficient hyperparameter features into the probability hyperparameter decoding network to obtain probability distribution parameters, and decode the second bitstream corresponding to the current image block based on the probability distribution parameters to obtain the residual features corresponding to the current image block.

[0118] Step S13: Input the coefficient hyperparameter features into the mean hyperparameter decoding network to obtain the initial mean features corresponding to the current image block, and determine the target mean features corresponding to the current image block based on the initial mean features.

[0119] Step S14: Determine the reconstruction features corresponding to the current image block based on the target mean features and the residual features.

[0120] Step S15: Determine the reconstructed image block corresponding to the current image block based on the reconstruction features.

[0121] In a possible implementation manner, the probability hyperparameter decoding network may include an activation layer, and the output features of the activation layer are not greater than the upper threshold, and / or, the output features of the activation layer are not less than the lower threshold. For example, the output features of the activation layer may not be greater than the upper threshold, and the output features of the activation layer are not less than the lower threshold, that is, both the upper threshold and the lower threshold of the activation layer are restricted. In this case, the activation layer may use a target activation function with upper and lower clamping to non-linearly adjust the features, and the target activation function corresponds to the upper threshold and the lower threshold, and the output features of the activation layer are not greater than the upper threshold, and the output features of the activation layer are not less than the lower threshold. Or, the output features of the activation layer may not be less than the lower threshold, that is, only the lower threshold of the activation layer is restricted, and the upper threshold of the activation layer is not restricted. In this case, the activation layer may use the Relu activation function to non-linearly adjust the features, and the Relu activation function corresponds to the lower threshold, and the output features of the activation layer are not less than the lower threshold. Or, the output features of the activation layer may not be greater than the upper threshold, that is, only the upper threshold of the activation layer is restricted, and the lower threshold of the activation layer is not restricted. In this case, the activation layer may use a certain activation function to non-linearly adjust the features, and the activation function corresponds to the upper threshold, and the output features of the activation layer are not greater than the upper threshold.

[0122] In a possible implementation, the mean hyperparameter decoding network may include an activation layer, and the output features of the activation layer are not greater than an upper threshold value, and / or, the output features of the activation layer are not less than a lower threshold value. For example, the output features of the activation layer may not be greater than the upper threshold value, and the output features of the activation layer are not less than the lower threshold value, that is, both the upper threshold value and the lower threshold value of the activation layer are restricted. In this case, the activation layer may use a target activation function with upper and lower clamping to nonlinearly adjust the features, and the target activation function corresponds to the upper threshold value and the lower threshold value, the output features of the activation layer are not greater than the upper threshold value, and the output features of the activation layer are not less than the lower threshold value. Or, the output features of the activation layer may not be less than the lower threshold value, that is, only the lower threshold value of the activation layer is restricted, and the upper threshold value of the activation layer is not restricted. In this case, the activation layer may use a Relu activation function to nonlinearly adjust the features, and the Relu activation function corresponds to the lower threshold value, and the output features of the activation layer are not less than the lower threshold value. Or, the output features of the activation layer may not be greater than the upper threshold value, that is, only the upper threshold value of the activation layer is restricted, and the lower threshold value of the activation layer is not restricted. In this case, the activation layer may use a certain activation function to nonlinearly adjust the features, and the activation function corresponds to the upper threshold value, and the output features of the activation layer are not greater than the upper threshold value.

[0123] Exemplarily, the upper threshold value corresponding to the target activation function is a configured fixed upper threshold value, and the lower threshold value corresponding to the target activation function is a configured fixed lower threshold value; or, the upper threshold value corresponding to the target activation function is an adaptively learned upper threshold value, and the lower threshold value corresponding to the target activation function is an adaptively learned lower threshold value.

[0124] Exemplarily, when adaptively learning the upper threshold value, the same upper threshold value is learned for all channels of the activation layer, or, the upper threshold value is learned separately for each channel of the activation layer, and the upper threshold values corresponding to different channels are the same or different.

[0125] Exemplarily, when adaptively learning the lower threshold value, the same lower threshold value is learned for all channels of the activation layer, or, the lower threshold value is learned separately for each channel of the activation layer, and the lower threshold values corresponding to different channels are the same or different.

[0126] An embodiment of the present application proposes a decoding method, which can be applied to a decoding end, and the method may include:

[0127] Step S21, decoding the first bitstream corresponding to the current image block to obtain the coefficient hyperparameter features corresponding to the current image block.

[0128] Step S22, inputting the coefficient hyperparameter features into the hyperparameter decoding network to obtain the reference features corresponding to the current image block.

[0129] Step S23: Input the reference feature and the obtained reconstruction feature of the current image block into the context model to obtain the probability distribution parameter and the target mean feature corresponding to the current image block.

[0130] Step S24: Decode the second bitstream corresponding to the current image block based on the probability distribution parameter to obtain the residual feature corresponding to the current image block, and determine the reconstruction feature based on the target mean feature and the residual feature.

[0131] Step S25: Determine the reconstructed image block corresponding to the current image block based on the reconstruction feature.

[0132] In a possible implementation, the hyperparameter decoding network may include an activation layer, and the output feature of the activation layer is not greater than the upper threshold, and / or the output feature of the activation layer is not less than the lower threshold. For example, the output feature of the activation layer may not be greater than the upper threshold, and the output feature of the activation layer is not less than the lower threshold, that is, both the upper threshold and the lower threshold of the activation layer are restricted. In this case, the activation layer may use a target activation function with upper and lower clamping to non-linearly adjust the feature, and the target activation function corresponds to the upper threshold and the lower threshold, the output feature of the activation layer is not greater than the upper threshold, and the output feature of the activation layer is not less than the lower threshold. Or, the output feature of the activation layer may not be less than the lower threshold, that is, only the lower threshold of the activation layer is restricted, and the upper threshold of the activation layer is not restricted. In this case, the activation layer may use the Relu activation function to non-linearly adjust the feature, and the Relu activation function corresponds to the lower threshold, and the output feature of the activation layer is not less than the lower threshold. Or, the output feature of the activation layer may not be greater than the upper threshold, that is, only the upper threshold of the activation layer is restricted, and the lower threshold of the activation layer is not restricted. In this case, the activation layer may use a certain activation function to non-linearly adjust the feature, and the activation function corresponds to the upper threshold, and the output feature of the activation layer is not greater than the upper threshold.

[0133] In a possible implementation, the context model may include an activation layer, and the output features of the activation layer are not greater than an upper threshold value, and / or, the output features of the activation layer are not less than a lower threshold value. For example, the output features of the activation layer may not be greater than the upper threshold value, and the output features of the activation layer are not less than the lower threshold value, that is, both the upper threshold value and the lower threshold value of the activation layer are restricted. In this case, the activation layer may use a target activation function with upper and lower clamping to non-linearly adjust the features, and the target activation function corresponds to the upper threshold value and the lower threshold value, and the output features of the activation layer are not greater than the upper threshold value, and the output features of the activation layer are not less than the lower threshold value. Or, the output features of the activation layer may not be less than the lower threshold value, that is, only the lower threshold value of the activation layer is restricted, and the upper threshold value of the activation layer is not restricted. In this case, the activation layer may use the Relu activation function to non-linearly adjust the features, and the Relu activation function corresponds to the lower threshold value, and the output features of the activation layer are not less than the lower threshold value. Or, the output features of the activation layer may not be greater than the upper threshold value, that is, only the upper threshold value of the activation layer is restricted, and the lower threshold value of the activation layer is not restricted. In this case, the activation layer may use a certain activation function to non-linearly adjust the features, and the activation function corresponds to the upper threshold value, and the output features of the activation layer are not greater than the upper threshold value.

[0134] Exemplarily, the upper threshold value corresponding to the target activation function is a configured fixed upper threshold value, and the lower threshold value corresponding to the target activation function is a configured fixed lower threshold value; or, the upper threshold value corresponding to the target activation function is an adaptively learned upper threshold value, and the lower threshold value corresponding to the target activation function is an adaptively learned lower threshold value.

[0135] Exemplarily, when adaptively learning the upper threshold value, the same upper threshold value is learned for all channels of the activation layer, or, the upper threshold value is separately learned for each channel of the activation layer, and the upper threshold values corresponding to different channels are the same or different.

[0136] Exemplarily, when adaptively learning the lower threshold value, the same lower threshold value is learned for all channels of the activation layer, or, the lower threshold value is separately learned for each channel of the activation layer, and the lower threshold values corresponding to different channels are the same or different.

[0137] In an embodiment of the present application, a decoding method is proposed. This method can be applied to the decoding end, and this method may include:

[0138] Step S31: Decode the first bitstream corresponding to the current image block to obtain the coefficient hyperparameter features corresponding to the current image block.

[0139] Step S32: Determine the probability distribution parameters based on the coefficient hyperparameter features, and decode the second bitstream corresponding to the current image block based on the probability distribution parameters to obtain the residual features corresponding to the current image block.

[0140] Step S33: Input the coefficient hyperparameter feature into the mean hyperparameter decoding network to obtain the initial mean feature corresponding to the current image block, and determine the target mean feature corresponding to the current image block based on the initial mean feature. Exemplarily, the mean hyperparameter decoding network may sequentially include a processing layer, an upsampling layer for performing an upsampling operation, and a cropping layer. Among them, the feature output by the upsampling layer is obtained as the initial mean feature after passing through the cropping layer, that is, the upsampling layer is located at the end of the mean hyperparameter decoding network. After the feature output by the upsampling layer passes through the cropping layer, the output feature of the mean hyperparameter decoding network is obtained.

[0141] Exemplarily, the cropping layer is used to crop and align the pixel positions after upsampling. The cropping layer can be a Crop layer or other network layers, as long as it can implement the cropping and alignment function. In the following, the Crop layer is taken as an example for illustration.

[0142] Step S34: Determine the reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature.

[0143] Step S35: Determine the reconstructed image block corresponding to the current image block based on the reconstruction feature.

[0144] Exemplarily, the processing layer of the mean hyperparameter decoding network does not include an upsampling layer for performing an upsampling operation, that is, the upsampling layer for performing an upsampling operation is deployed behind the processing layer, so that the processing layer does not perform an upsampling operation. The upsampling layer can be located behind the processing layer, and the upsampling layer is located at the end of the mean hyperparameter decoding network, that is, after the feature output by the upsampling layer passes through the Crop layer, the output feature of the mean hyperparameter decoding network is obtained. Or, an additional convolutional layer (such as a 1*1 convolutional layer or a 3*3 convolutional layer, and this convolutional layer only performs a simple convolution operation) can be deployed between the upsampling layer and the Crop layer. After the feature output by the upsampling layer passes through a convolutional layer and a Crop layer, the output feature (initial mean feature) of the mean hyperparameter decoding network is obtained. For the convenience of description, in the following embodiments, it is taken as an example that the mean hyperparameter decoding network may sequentially include a processing layer, an upsampling layer for performing an upsampling operation, and a Crop layer.

[0145] Exemplarily, based on the mean hyperparameter decoding network, the coefficient hyperparameter feature can be processed by the processing layer to obtain a processed feature; the processed feature can be upsampled by the upsampling layer to obtain an upsampled feature; the upsampled feature can be cropped and aligned by the Crop layer to obtain an initial mean feature.

[0146] Exemplarily, upsampling the processed features through an upsampling layer to obtain upsampled features may include, but is not limited to: upsampling the processed features in a pixel rearrangement manner to obtain upsampled features; or, upsampling the processed features in a transposed convolution manner to obtain upsampled features; or, upsampling the processed features in a nearest neighbor interpolation manner to obtain upsampled features.

[0147] Exemplarily, the processing layer may include at least one of the following: one or more ordinary convolution layers, one or more transposed convolution layers, one or more deformable convolution layers, one or more depthwise separable convolution layers, one or more grouped convolution layers, one or more dilated convolution layers, one or more global processing unit layers, one or more activation layers using the Relu activation function, one or more activation layers using a target activation function with upper and lower clamping, one or more activation layers using the LeakyRelu activation function; there are arbitrary skip connections between the layers of the processing layer.

[0148] Exemplarily, the above execution order is only an example given for convenience of description. In actual applications, the execution order between steps can also be changed, and no limitation is imposed on this execution order. Moreover, in other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in this specification, and the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; multiple steps described in this specification may also be combined into a single step for description in other embodiments.

[0149] As can be seen from the above technical solutions, in the embodiments of the present application, an end-to-end video image compression method is proposed, which can implement the decoding of video images based on a neural network. By optimizing the activation function and using a target activation function with upper and lower clamping to perform non-linear adjustment on the features, the complexity of the activation function is reduced, and then the complexity of the neural network is reduced. The target activation function can improve the robustness of the neural network and reduce the phenomenon of decoding garbled codes on high-frequency images. This enables the neural network to effectively ensure the quality of the reconstructed image blocks while maintaining low complexity, improve the quality of the reconstructed image, achieve the purpose of improving the decoding performance, and reduce the complexity. By optimizing the mean hyperparameter decoding network, the complexity of the neural network is greatly reduced while maintaining the performance of the neural network unchanged, achieving the purpose of improving the decoding performance and reducing the complexity.

[0150] Embodiment 2: In the embodiments of the present application, an encoding method is proposed. Refer to Figure 3 As shown, it is a schematic flowchart of the encoding method. This method can be applied to an encoding end (also referred to as a video encoder), and this method may include:

[0151] Step 301: The encoding end decodes the first bitstream corresponding to the current image block to obtain the coefficient hyperparameter feature corresponding to the current image block.

[0152] Step 302: The encoding end determines the probability distribution parameter based on the coefficient hyperparameter feature, and decodes the second bitstream corresponding to the current image block based on the probability distribution parameter to obtain the residual feature corresponding to the current image block.

[0153] Step 303: The encoding end determines the target mean feature corresponding to the current image block based on the coefficient hyperparameter feature.

[0154] Step 304: The encoding end determines the reconstructed feature corresponding to the current image block based on the target mean feature and the residual feature.

[0155] Step 305: The encoding end inputs the reconstructed feature into the synthesis transformation network to obtain the reconstructed image block corresponding to the current image block. The synthesis transformation network includes a non-linear residual network layer, and the non-linear residual network layer can at least include an activation layer and a skip connection structure; the output feature of the activation layer is not greater than the upper limit threshold, and / or, the output feature of the activation layer is not less than the lower limit threshold.

[0156] Exemplarily, the output feature of the activation layer of the non-linear residual network layer can be not greater than the upper limit threshold, and the output feature of the activation layer is not less than the lower limit threshold, that is, both the upper limit threshold and the lower limit threshold of the activation layer are restricted. In this case, the activation layer can use a target activation function with upper and lower clamping to non-linearly adjust the feature, and the target activation function corresponds to the upper limit threshold and the lower limit threshold, and the output feature of the activation layer is not greater than the upper limit threshold, and the output feature of the activation layer is not less than the lower limit threshold.

[0157] Alternatively, the output feature of the activation layer of the non-linear residual network layer can be not less than the lower limit threshold, that is, only the lower limit threshold of the activation layer is restricted, and the upper limit threshold of the activation layer is not restricted. In this case, the activation layer can use the Relu activation function to non-linearly adjust the feature, and the Relu activation function corresponds to the lower limit threshold, and the output feature of the activation layer is not less than the lower limit threshold.

[0158] Alternatively, the output feature of the activation layer of the non-linear residual network layer can be not greater than the upper limit threshold, that is, only the upper limit threshold of the activation layer is restricted, and the lower limit threshold of the activation layer is not restricted. In this case, the activation layer can use a certain activation function to non-linearly adjust the feature, and the activation function corresponds to the upper limit threshold, and the output feature of the activation layer is not greater than the upper limit threshold.

[0159] In an embodiment of the present application, an encoding method is proposed, which is applied to an encoding end and includes: decoding a first bitstream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block. Inputting the coefficient hyperparameter feature into a probability hyperparameter decoding network to obtain probability distribution parameters, and decoding a second bitstream corresponding to the current image block based on the probability distribution parameters to obtain a residual feature corresponding to the current image block. Inputting the coefficient hyperparameter feature into a mean hyperparameter decoding network to obtain an initial mean feature corresponding to the current image block, and determining a target mean feature corresponding to the current image block based on the initial mean feature. Determining a reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature. Determining a reconstructed image block corresponding to the current image block based on the reconstruction feature.

[0160] In a possible implementation manner, the probability hyperparameter decoding network may include an activation layer, and the output feature of the activation layer is not greater than an upper threshold, and / or the output feature of the activation layer is not less than a lower threshold. And / or, the mean hyperparameter decoding network may include an activation layer, and the output feature of the activation layer is not greater than an upper threshold, and / or the output feature of the activation layer is not less than a lower threshold.

[0161] In an embodiment of the present application, an encoding method is proposed, which can be applied to an encoding end, and the method includes: decoding a first bitstream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block. Inputting the coefficient hyperparameter feature into a hyperparameter decoding network to obtain a reference feature corresponding to the current image block. Inputting the reference feature and the obtained reconstruction feature of the current image block into a context model to obtain probability distribution parameters and a target mean feature corresponding to the current image block. Decoding a second bitstream corresponding to the current image block based on the probability distribution parameters to obtain a residual feature corresponding to the current image block, and determining a reconstruction feature based on the target mean feature and the residual feature. Determining a reconstructed image block corresponding to the current image block based on the reconstruction feature.

[0162] In a possible implementation manner, the hyperparameter decoding network may include an activation layer, and the output feature of the activation layer is not greater than an upper threshold, and / or the output feature of the activation layer is not less than a lower threshold. And / or, the context model may include an activation layer, and the output feature of the activation layer is not greater than an upper threshold, and / or the output feature of the activation layer is not less than a lower threshold.

[0163] In an embodiment of the present application, a coding method is proposed. This method can be applied to the coding end and may include: decoding a first bitstream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block. Determining probability distribution parameters based on the coefficient hyperparameter feature, and decoding a second bitstream corresponding to the current image block based on the probability distribution parameters to obtain a residual feature corresponding to the current image block. Inputting the coefficient hyperparameter feature into a mean hyperparameter decoding network to obtain an initial mean feature corresponding to the current image block, and determining a target mean feature corresponding to the current image block based on the initial mean feature; wherein, the mean hyperparameter decoding network may sequentially include a processing layer, an upsampling layer for performing an upsampling operation, and a cropping layer. Determining a reconstructed feature corresponding to the current image block based on the target mean feature and the residual feature. Determining a reconstructed image block corresponding to the current image block based on the reconstructed feature. Among them, for the mean hyperparameter decoding network, the feature output by the upsampling layer is obtained as the initial mean feature after passing through the cropping layer, that is, the upsampling layer is located at the end of the mean hyperparameter decoding network. In this way, after the feature output by the upsampling layer passes through the cropping layer, the output feature of the mean hyperparameter decoding network (i.e., the initial mean feature) is obtained.

[0164] Exemplarily, the processing process at the coding end is similar to that at the decoding end. The same parts will not be repeated here. The processing process at the decoding end can be applied to the coding end, that is, the coding end adopts the same processing method as the decoding end.

[0165] Exemplarily, the above execution order is only an example for convenient description. In actual applications, the execution order between steps can also be changed, and no limitation is imposed on this execution order. Moreover, in other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in this specification. The steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; multiple steps described in this specification may also be combined into a single step for description in other embodiments.

[0166] As can be seen from the above technical solutions, in the embodiments of the present application, an end-to-end video image compression method is proposed, which can implement the encoding of video images based on a neural network. By optimizing the activation function and using a target activation function with upper and lower clamping to non-linearly adjust the features, the complexity of the activation function is reduced, and then the complexity of the neural network is reduced. The target activation function can improve the robustness of the neural network and reduce the phenomenon of encoding garbled characters in the neural network for high-frequency images. This enables the neural network to effectively ensure the quality of the reconstructed image blocks while maintaining a low complexity, improve the quality of the reconstructed images, achieve the purpose of improving the encoding performance, and reduce the complexity. By optimizing the mean hyperparameter encoding network, the complexity of the neural network is greatly reduced while maintaining the performance of the neural network unchanged, achieving the purpose of improving the encoding performance and reducing the complexity.

[0167] Embodiment 3: For Embodiment 1 and Embodiment 2, regarding the processing procedure at the encoding end, reference can be made to Figure 4A as shown. Of course, Figure 4A this is only an example of the processing procedure at the encoding end, and the processing procedure at this encoding end is not limited.

[0168] After the encoding end obtains the current image block x (the current image block x can be the original image block x, that is, the input image block), it can analyze and transform the current image block x through an analysis transformation network (i.e., a neural network) to obtain the image feature y corresponding to the current image block x. Among them, performing feature transformation on the current image block x through the analysis transformation network means: transforming the current image block x into the image feature y in the latent domain, so as to facilitate all subsequent processes to be operated in the latent domain.

[0169] Exemplarily, an image can be divided into 1 image block or multiple image blocks. If the image is divided into 1 image block, then the current image block x can also be the image, that is, the encoding and decoding process for the image block can also be directly applied to the image.

[0170] After the encoding end obtains the image feature y, it performs coefficient hyperparameter feature transformation on the image feature y to obtain the coefficient hyperparameter feature z. For example, the image feature y can be input to a hyperparameter encoding network (i.e., a neural network), and the hyperparameter encoding network performs coefficient hyperparameter feature transformation on the image feature y to obtain the coefficient hyperparameter feature z. Among them, the hyperparameter encoding network can be a pre-trained neural network, and the training process of this hyperparameter encoding network is not limited, as long as it can perform coefficient hyperparameter feature transformation on the image feature y. Among them, the image feature y in the latent domain obtains the hyperprior latent information z after passing through the hyperparameter encoding network.

[0171] After the encoding end obtains the coefficient hyperparameter feature z, it can quantize the coefficient hyperparameter feature z to obtain the hyperparameter quantization feature corresponding to the coefficient hyperparameter feature z, that isFigure 4A The Q operation in Figure 4A is a quantization process. After obtaining the hyperparameter quantization feature corresponding to the coefficient hyperparameter feature z, the hyperparameter quantization feature is encoded to obtain Bitstream#1 (i.e., the first bitstream) corresponding to the current image block, that is, Figure 4A The AE operation in Figure 4A represents an encoding process, such as an entropy encoding process. Alternatively, the encoding end can also directly encode the coefficient hyperparameter feature z to obtain Bitstream#1 corresponding to the current image block. Among them, the hyperparameter quantization feature or the coefficient hyperparameter feature z carried in Bitstream#1 is mainly used to obtain the parameters of the mean and probability distribution models.

[0172] After obtaining Bitstream#1 corresponding to the current image block, the encoding end can send Bitstream#1 corresponding to the current image block to the decoding end. For the processing process of the decoding end for Bitstream#1 corresponding to the current image block, see the subsequent embodiments.

[0173] After the encoding end obtains Bitstream#1 corresponding to the current image block, it can also decode Bitstream#1 to obtain the hyperparameter quantization feature, that is, Figure 4A The AD in Figure 4A represents a decoding process. Then, the hyperparameter quantization feature is dequantized to obtain the coefficient hyperparameter feature z_hat. The coefficient hyperparameter feature z_hat and the coefficient hyperparameter feature z can be the same or different. Figure 4A The IQ operation in Figure 4A is a dequantization process. Alternatively, after the encoding end obtains Bitstream#1 corresponding to the current image block, it can decode Bitstream#1 to obtain the coefficient hyperparameter feature z_hat, without involving the dequantization process of the coefficient hyperparameter feature z_hat.

[0174] For the encoding process of Bitstream#1, an encoding method with a fixed probability density model can be adopted. For the decoding process of Bitstream#1, a decoding method with a fixed probability density model can be adopted. There is no restriction on this encoding and decoding process.

[0175] After the encoding end obtains the coefficient hyperparameter feature z_hat, it can input the coefficient hyperparameter feature z_hat to the mean hyperparameter decoding network. The mean hyperparameter decoding network processes based on the coefficient hyperparameter feature z_hat to obtain the initial mean feature m (the initial mean feature m is an intermediate parameter for the mean feature). There is no restriction on the processing process of this mean hyperparameter decoding network.

[0176] After obtaining the initial mean feature m, the encoding end can use the initial mean feature m as the target mean feature mu (i.e., the predicted value mu). Alternatively, the encoding end can input the initial mean feature m and the decoded reconstruction feature y_hat (i.e., the obtained reconstruction feature of the current image block, and the determination process of the reconstruction feature y_hat is described in the subsequent embodiments) to the context model, and the context model performs a context-based prediction process to obtain the target mean feature mu (i.e., the mean mu) corresponding to the current image block. For example, for the prediction process of the context model, the input data of the context model includes the initial mean feature m and the decoded reconstruction feature y_hat, and the two are jointly input to obtain a more accurate target mean feature mu, which is used to subtract the original feature to obtain the residual r_hat and add to the decoded residual to obtain the reconstruction feature y_hat.

[0177] The mean hyperparameter decoding network and the context model are optional neural networks, that is, it is also possible not to have the mean hyperparameter decoding network and the context model, that is, there is no need to determine the target mean feature mu through the mean hyperparameter decoding network and the context model.

[0178] After obtaining the image feature y, the encoding end can determine the residual feature r based on the image feature y and the target mean feature mu, such as taking the difference between the image feature y and the target mean feature mu as the residual feature r. Then, the residual feature r is processed to obtain the image feature s, and this feature processing process is not limited and can be any feature processing method. In this case, it is necessary to deploy the mean hyperparameter decoding network and the context model to provide the target mean feature mu. Alternatively, after obtaining the image feature y, the encoding end can process the image feature y to obtain the image feature s, and this feature processing process is not limited and can be any feature processing method. In this case, there is no need to deploy the mean hyperparameter decoding network and the context model.

[0179] After obtaining the image feature s, the encoding end can quantize the image feature s to obtain the image quantization feature corresponding to the image feature s, that is Figure 4A The Q operation in is the quantization process. After obtaining the image quantization feature corresponding to the image feature s, the encoding end can encode the image quantization feature to obtain the Bitstream#2 (i.e., the second bitstream) corresponding to the current image block, that is Figure 4A The AE operation in represents the encoding process, such as the entropy encoding process. Alternatively, the encoding end can also directly encode the image feature s to obtain the Bitstream#2 corresponding to the current image block without involving the quantization process of the image feature s.

[0180] Alternatively, after obtaining the residual feature r, the encoding end may not perform feature processing on the residual feature r, but directly quantize the residual feature r to obtain the image quantization feature corresponding to the residual feature r, and the encoding end encodes the image quantization feature to obtain the Bitstream#2 (i.e., the second bitstream) corresponding to the current image block. Alternatively, after obtaining the residual feature r, the encoding end directly encodes the residual feature r to obtain the Bitstream#2 corresponding to the current image block without involving the quantization process.

[0181] After obtaining the Bitstream#2 corresponding to the current image block, the encoding end may send the Bitstream#2 corresponding to the current image block to the decoding end. For the processing process of the decoding end for the Bitstream#2 corresponding to the current image block, see the subsequent embodiments.

[0182] After obtaining the Bitstream#2 corresponding to the current image block, the encoding end may also decode the Bitstream#2 to obtain the image quantization feature, that is, Figure 4A the AD in represents the decoding process. Then, the encoding end may perform inverse quantization on the image quantization feature to obtain the image feature s'. The image feature s' may be the same as or different from the image feature s. Figure 4A the IQ operation in is the inverse quantization process. Alternatively, after obtaining the Bitstream#2 corresponding to the current image block, the encoding end may also decode the Bitstream#2 to obtain the image feature s' without involving the inverse quantization process of the image quantization feature. After obtaining the image feature s', the encoding end may perform feature recovery (i.e., the inverse process of feature processing) on the image feature s'. There is no limitation on this feature recovery, and it may be any feature recovery method to obtain the residual feature r_hat. The residual feature r_hat may be the same as or different from the residual feature r.

[0183] Alternatively, after obtaining the Bitstream#2 corresponding to the current image block, the encoding end may also decode the Bitstream#2 to obtain the image quantization feature, and then the encoding end may perform inverse quantization on the image quantization feature to obtain the residual feature r_hat. Alternatively, after obtaining the Bitstream#2 corresponding to the current image block, the encoding end may also decode the Bitstream#2 to obtain the residual feature r_hat without involving the inverse quantization process of the image quantization feature.

[0184] After obtaining the residual feature r_hat, the encoding end determines the image feature y_hat (i.e., the reconstructed feature) based on the residual feature r_hat and the target mean feature mu. The image feature y_hat may or may not be the same as the image feature y. For example, the sum of the residual feature r_hat and the target mean feature mu can be used as the reconstructed feature y_hat. In this case, a mean hyperparameter decoding network and a context model need to be deployed, and the target mean feature mu is provided by the mean hyperparameter decoding network and the context model. Alternatively, after the encoding end obtains Bitstream#2 corresponding to the current image block and performs operations such as decoding and inverse quantization, the reconstructed feature y_hat can be directly obtained. In this case, the mean hyperparameter decoding network and the context model do not need to be deployed.

[0185] After the encoding end obtains the reconstructed feature y_hat, it can perform a synthesis transformation on the reconstructed feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the reconstructed feature y_hat is input into a synthesis transformation network, and the synthesis transformation network performs a synthesis transformation on the reconstructed feature y_hat to obtain the reconstructed image block x_hat. Thus, the image reconstruction process is completed.

[0186] Exemplarily, when the encoding end encodes the feature to obtain Bitstream#2 corresponding to the current image block, it needs to first determine a probability distribution model and then encode the feature based on this probability distribution model. In addition, when the encoding end decodes Bitstream#2, it also needs to first determine a probability distribution model and then decode Bitstream#2 based on this probability distribution model.

[0187] To obtain the probability distribution model, continue to refer to Figure 4A As shown, after the encoding end obtains the coefficient hyperparameter feature z_hat, it can perform an inverse transformation of the coefficient hyperparameter feature on z_hat to obtain the probability distribution parameters. For example, the coefficient hyperparameter feature z_hat is input into a probability hyperparameter decoding network, and the probability hyperparameter decoding network performs an inverse transformation of the coefficient hyperparameter feature on z_hat to obtain the probability distribution parameter p. After obtaining the probability distribution parameters, a probability distribution model can be generated based on the probability distribution parameters. Among them, the probability hyperparameter decoding network can be a trained neural network, and the training process of this probability hyperparameter decoding network is not limited as long as it can perform an inverse transformation of the coefficient hyperparameter feature on z_hat.

[0188] In a possible implementation manner, the above processing process of the encoding end can be executed by a deep learning model or a neural network model, so as to implement an end-to-end image compression and encoding process, and this encoding process is not limited.

[0189] Example 4: For Examples 1 and 2, regarding the processing procedure at the decoding end, reference can be made to Figure 4B as shown. Of course, Figure 4B this is only an example of the processing procedure at the decoding end, and the processing procedure at the decoding end is not limited thereto.

[0190] After the decoding end obtains Bitstream#1 corresponding to the current image block, it can decode Bitstream#1 to obtain the hyperparameter quantization feature, that is, Figure 4B the AD in represents the decoding process. The hyperparameter quantization feature is dequantized to obtain the coefficient hyperparameter feature z_hat, Figure 4B the IQ operation in is the dequantization process. Or, after the decoding end obtains Bitstream#1 corresponding to the current image block, it can decode Bitstream#1 to obtain the coefficient hyperparameter feature z_hat without involving the dequantization process.

[0191] Regarding the decoding process of Bitstream#1, the decoding method using a fixed probability density model can be adopted, and this is not limited.

[0192] The image can be divided into 1 image block or multiple image blocks. If the image is divided into 1 image block, the current image block x can also be the image, that is, the decoding process for the image block can also be directly applied to the image.

[0193] After the decoding end obtains the coefficient hyperparameter feature z_hat, it can input the coefficient hyperparameter feature z_hat into the mean hyperparameter decoding network. The mean hyperparameter decoding network processes based on the coefficient hyperparameter feature z_hat to obtain the initial mean feature m (the initial mean feature m is the intermediate parameter for the mean feature), and the processing procedure of the mean hyperparameter decoding network is not limited thereto.

[0194] After the decoding end obtains the initial mean feature m, it can use the initial mean feature m as the target mean feature mu (i.e., the predicted value mu). Or, the decoding end can input the initial mean feature m and the decoded reconstruction feature y_hat (i.e., the obtained reconstruction feature of the current image block, and the determination process of the reconstruction feature y_hat is described in the subsequent examples) into the context model. The context model performs the context-based prediction process to obtain the target mean feature mu (i.e., the mean mu) corresponding to the current image block. For example, for the prediction process of the context model, the input data of the context model includes the initial mean feature m and the decoded reconstruction feature y_hat, and the two are jointly input to obtain a more accurate target mean feature mu. The target mean feature mu is used to subtract from the original feature to obtain the residual r_hat and add to the decoded residual to obtain the reconstruction feature y_hat.

[0195] The mean hyperparameter decoding network and the context model are optional neural networks, that is, it is also possible not to have the mean hyperparameter decoding network and the context model, that is, there is no need to determine the target mean feature mu through the mean hyperparameter decoding network and the context model.

[0196] After the decoding end obtains Bitstream#2 corresponding to the current image block, it can also decode Bitstream#2 to obtain the image quantization feature, that is, Figure 4B the AD in represents the decoding process. Then, the decoding end can inverse-quantize the image quantization feature to obtain the image feature s'. Figure 4B the IQ operation in is the inverse-quantization process. Alternatively, after the decoding end obtains Bitstream#2 corresponding to the current image block, it can also decode Bitstream#2 to obtain the image feature s', without involving the inverse-quantization process of the image quantization feature. After the decoding end obtains the image feature s', it can perform feature restoration on the image feature s'. There is no limit to this feature restoration, and it can be any feature restoration method to obtain the residual feature r_hat.

[0197] Alternatively, after the decoding end obtains Bitstream#2 corresponding to the current image block, the decoding end can also decode Bitstream#2 to obtain the image quantization feature. Then, the decoding end can inverse-quantize the image quantization feature to obtain the residual feature r_hat. Alternatively, after the decoding end obtains Bitstream#2 corresponding to the current image block, the decoding end can also decode Bitstream#2 to obtain the residual feature r_hat, without involving the inverse-quantization process of the image quantization feature.

[0198] After the decoding end obtains the residual feature r_hat, it determines the image feature y_hat (i.e., the reconstructed feature) based on the residual feature r_hat and the target mean feature mu. The image feature y_hat may be the same as or different from the image feature y. For example, the sum of the residual feature r_hat and the target mean feature mu can be used as the reconstructed feature y_hat. In this case, it is necessary to deploy the mean hyperparameter decoding network and the context model, and the mean hyperparameter decoding network and the context model provide the target mean feature mu. Alternatively, after the decoding end obtains Bitstream#2 corresponding to the current image block and performs operations such as decoding and inverse-quantization, it can directly obtain the reconstructed feature y_hat. In this case, it is not necessary to deploy the mean hyperparameter decoding network and the context model.

[0199] After the decoding end obtains the reconstructed feature y_hat, it can perform a synthesis transformation on the reconstructed feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the reconstructed feature y_hat is input to a synthesis transformation network, and the synthesis transformation network performs a synthesis transformation on the reconstructed feature y_hat to obtain the reconstructed image block x_hat. Thus, the image reconstruction process is completed.

[0200] Exemplarily, when the decoding end decodes Bitstream#2, it needs to first determine a probability distribution model, and then decode Bitstream#2 based on this probability distribution model. To obtain the probability distribution model, continue to refer to Figure 4B As shown, after the decoding end obtains the coefficient hyperparameter feature z_hat, it can also perform an inverse transformation of the coefficient hyperparameter feature on z_hat to obtain probability distribution parameters. For example, the decoding end inputs the coefficient hyperparameter feature z_hat to a probability hyperparameter decoding network, and the probability hyperparameter decoding network performs an inverse transformation of the coefficient hyperparameter feature on z_hat to obtain the probability distribution parameter p. After obtaining the probability distribution parameters, the decoding end can generate a probability distribution model based on the probability distribution parameters.

[0201] Among them, the probability hyperparameter decoding network can be a trained neural network. There is no limitation on the training process of this probability hyperparameter decoding network, as long as it can perform an inverse transformation of the coefficient hyperparameter feature on z_hat.

[0202] In a possible implementation manner, the above processing process of the decoding end can be executed by a deep learning model or a neural network model, so as to implement an end-to-end image compression and decoding process. There is no limitation on this decoding process.

[0203] Example 5: For Examples 1 and 2, regarding the processing process of the encoding end, refer to Figure 4C As shown, of course, Figure 4C This is just an example of the processing process of the encoding end. There is no limitation on the processing process of the encoding end.

[0204] After the encoding end obtains the current image block x, it can perform an analysis transformation on the current image block x through an analysis transformation network to obtain the image feature y corresponding to the current image block x. Performing a feature transformation on the current image block x through the analysis transformation network means: transforming the current image block x into the image feature y in the latent domain, so as to facilitate all subsequent processes to be operated in the latent domain.

[0205] Exemplarily, an image can be divided into 1 image block or multiple image blocks. If the image is divided into 1 image block, then the current image block x can also be the image, that is, the encoding and decoding process of the image block can also be directly applied to the image.

[0206] After the encoding end obtains the image feature y, it performs a coefficient hyperparameter feature transformation on the image feature y to obtain a coefficient hyperparameter feature z. For example, the image feature y is input into a hyperparameter encoding network, and the hyperparameter encoding network performs a coefficient hyperparameter feature transformation on the image feature y to obtain the coefficient hyperparameter feature z. The hyperparameter encoding network is a pre-trained neural network that can perform a coefficient hyperparameter feature transformation on the image feature y. The image feature y in the latent domain passes through the hyperparameter encoding network to obtain the hyperprior latent information z.

[0207] After the encoding end obtains the coefficient hyperparameter feature z, it can quantize the coefficient hyperparameter feature z to obtain a hyperparameter quantization feature corresponding to the coefficient hyperparameter feature z, and encode the hyperparameter quantization feature to obtain Bitstream#1 (i.e., the first bitstream) corresponding to the current image block. Alternatively, it can also directly encode the coefficient hyperparameter feature z to obtain Bitstream#1 corresponding to the current image block.

[0208] After obtaining Bitstream#1 corresponding to the current image block, the encoding end can send Bitstream#1 corresponding to the current image block to the decoding end. For the processing process of the decoding end for Bitstream#1 corresponding to the current image block, see the subsequent embodiments.

[0209] After the encoding end obtains Bitstream#1 corresponding to the current image block, it can also decode Bitstream#1 to obtain a hyperparameter quantization feature, and dequantize the hyperparameter quantization feature to obtain the coefficient hyperparameter feature z_hat. Alternatively, after the encoding end obtains Bitstream#1 corresponding to the current image block, it can decode Bitstream#1 to obtain the coefficient hyperparameter feature z_hat.

[0210] For the encoding process of Bitstream#1, an encoding method with a fixed probability density model can be adopted. For the decoding process of Bitstream#1, a decoding method with a fixed probability density model can be adopted. There is no limitation on this encoding and decoding process.

[0211] After the encoding end obtains the coefficient hyperparameter feature z_hat, it can input the coefficient hyperparameter feature z_hat into a hyperparameter decoding network, and the hyperparameter decoding network processes it based on the coefficient hyperparameter feature z_hat to obtain a reference feature m (the reference feature m is an intermediate parameter for the mean feature and probability distribution parameters). There is no limitation on the processing process of this hyperparameter decoding network.

[0212] After the encoding end obtains the reference feature m, it can input the reference feature m and the decoded reconstructed feature ŷ (i.e., the obtained reconstructed feature of the current image block. For the determination process of the reconstructed feature ŷ, refer to the subsequent embodiments) into the context model. The context model performs a context-based prediction process to obtain the target mean feature μ (i.e., the predicted value μ, i.e., the mean μ) and the probability distribution parameter p corresponding to the current image block. For example, for the prediction process of the context model, the input data of the context model includes the reference feature m and the decoded reconstructed feature ŷ. The two are jointly input to obtain a more accurate target mean feature μ and probability distribution parameter p. The target mean feature μ is used to subtract the original feature to obtain the residual r̂ and add it to the decoded residual to obtain the reconstructed feature ŷ. The probability distribution parameter p is used to encode Bitstream#2.

[0213] After the encoding end obtains the image feature y, it can determine the residual feature r based on the image feature y and the target mean feature μ. For example, the difference between the image feature y and the target mean feature μ is used as the residual feature r. Then, the residual feature r is processed to obtain the image feature s. There is no limitation on this feature processing process, and it can be any feature processing method. Alternatively, after the encoding end obtains the image feature y, it can process the image feature y to obtain the image feature s. There is no limitation on this feature processing process, and it can be any feature processing method. After the encoding end obtains the image feature s, it can quantize the image feature s to obtain the image quantization feature corresponding to the image feature s, and encode the image quantization feature to obtain Bitstream#2 (i.e., the second bitstream) corresponding to the current image block. Or, after the encoding end obtains the image feature s, it can also directly encode the image feature s to obtain Bitstream#2 corresponding to the current image block without involving the quantization process of the image feature s.

[0214] Alternatively, after the encoding end obtains the residual feature r, it can directly quantize the residual feature r without processing the residual feature r to obtain the image quantization feature corresponding to the residual feature r. The encoding end encodes the image quantization feature to obtain Bitstream#2 (i.e., the second bitstream) corresponding to the current image block. Or, after the encoding end obtains the residual feature r, it directly encodes the residual feature r to obtain Bitstream#2 corresponding to the current image block without involving the quantization process.

[0215] Exemplarily, when the encoding end encodes the feature to obtain Bitstream#2 corresponding to the current image block, it can generate a probability distribution model based on the probability distribution parameter p, and then encode the feature based on the probability distribution model.

[0216] After obtaining Bitstream#2 corresponding to the current image block, the encoder may send Bitstream#2 corresponding to the current image block to the decoder. For the processing of Bitstream#2 corresponding to the current image block by the decoder, refer to the subsequent embodiments.

[0217] After obtaining Bitstream#2 corresponding to the current image block, the encoder can also decode Bitstream#2 to obtain image quantization features, and dequantize the image quantization features to obtain image features s'. Alternatively, after obtaining Bitstream#2 corresponding to the current image block, the encoder can also decode Bitstream#2 to obtain image features s' without involving the dequantization process of image quantization features. After obtaining image features s', the encoder can perform feature recovery on image features s', and there is no restriction on this feature recovery, which can be any feature recovery method, to obtain residual features r_hat.

[0218] Alternatively, after obtaining Bitstream#2 corresponding to the current image block, the encoder can also decode Bitstream#2 to obtain image quantization features, and then the encoder can dequantize the image quantization features to obtain the residual feature r_hat. Alternatively, after obtaining Bitstream#2 corresponding to the current image block, the encoder can also decode Bitstream#2 to obtain the residual feature r_hat without involving the dequantization process of the image quantization features.

[0219] Exemplarily, when the encoder decodes Bitstream#2, it may generate a probability distribution model based on the probability distribution parameter p, and then decode Bitstream#2 based on the probability distribution model, and there is no restriction on this decoding process.

[0220] After obtaining the residual feature r_hat, the encoder determines the image feature y_hat (ie, the reconstruction feature) based on the residual feature r_hat and the target mean feature mu, such as taking the sum of the residual feature r_hat and the target mean feature mu as the reconstruction feature y_hat.

[0221] After obtaining the reconstructed feature y_hat, the encoding end can perform a synthetic transformation on the reconstructed feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the reconstructed feature y_hat is input into the synthetic transformation network, and the synthetic transformation network performs a synthetic transformation on the reconstructed feature y_hat to obtain the reconstructed image block x_hat. At this point, the image reconstruction process is completed.

[0222] In a possible implementation, the processing process of the above encoding end can be executed by a deep learning model or a neural network model, so as to realize the end-to-end image compression and encoding process, and the encoding process is not limited thereto.

[0223] Example 6: For Examples 1 and 2, regarding the processing process of the decoding end, reference can be made to Figure 4D as shown. Of course, Figure 4D it is only an example of the processing process of the decoding end, and the processing process of the decoding end is not limited thereto.

[0224] After the decoding end obtains Bitstream#1 corresponding to the current image block, it can also decode Bitstream#1 to obtain the hyperparameter quantization feature, and perform inverse quantization on the hyperparameter quantization feature to obtain the coefficient hyperparameter feature z_hat. Alternatively, after the decoding end obtains Bitstream#1 corresponding to the current image block, it can decode Bitstream#1 to obtain the coefficient hyperparameter feature z_hat.

[0225] Regarding the decoding process of Bitstream#1, a decoding method using a fixed probability density model can be adopted, and this is not limited.

[0226] The image can be divided into 1 image block or multiple image blocks. If the image is divided into 1 image block, the current image block x can also be the image, that is, the decoding process for the image block can also be directly used for the image.

[0227] After the decoding end obtains the coefficient hyperparameter feature z_hat, it can input the coefficient hyperparameter feature z_hat into the hyperparameter decoding network, and the hyperparameter decoding network processes based on the coefficient hyperparameter feature z_hat to obtain the reference feature m (the reference feature m is an intermediate parameter for the mean feature and the probability distribution parameter), and the processing process of the hyperparameter decoding network is not limited thereto.

[0228] After the decoding end obtains the reference feature m, it can input the reference feature m and the decoded reconstruction feature y_hat (that is, the obtained reconstruction feature of the current image block, and the determination process of the reconstruction feature y_hat is described in the subsequent examples) into the context model, and the context model executes the context-based prediction process to obtain the target mean feature mu (that is, the predicted value mu, that is, the mean mu) and the probability distribution parameter p corresponding to the current image block. For example, for the prediction process of the context model, the input data of the context model includes the reference feature m and the decoded reconstruction feature y_hat, and the two are jointly input to obtain more accurate target mean feature mu and probability distribution parameter p. The target mean feature mu is used to subtract from the original feature to obtain the residual r_hat and add to the decoded residual to obtain the reconstruction feature y_hat, and the probability distribution parameter p is used to decode Bitstream#2.

[0229] After obtaining Bitstream#2 corresponding to the current image block, the decoding end can also decode Bitstream#2 to obtain image quantization features, and dequantize the image quantization features to obtain image features s'. Alternatively, after obtaining Bitstream#2 corresponding to the current image block, the decoding end can also decode Bitstream#2 to obtain image features s' without involving the dequantization process of image quantization features. After obtaining image features s', the decoding end can perform feature recovery on image features s', and there is no restriction on this feature recovery, which can be any feature recovery method, to obtain residual features r_hat.

[0230] Alternatively, after obtaining Bitstream#2 corresponding to the current image block, the decoding end may also decode Bitstream#2 to obtain image quantization features, and then, the decoding end may dequantize the image quantization features to obtain residual features r_hat. Alternatively, after obtaining Bitstream#2 corresponding to the current image block, the decoding end may also decode Bitstream#2 to obtain residual features r_hat without involving the dequantization process of image quantization features.

[0231] Exemplarily, when the decoding end decodes Bitstream#2, it can generate a probability distribution model based on the probability distribution parameter p, and then decode Bitstream#2 based on the probability distribution model, and there is no restriction on this decoding process.

[0232] After obtaining the residual feature r_hat, the decoding end determines the image feature y_hat (ie, the reconstruction feature) based on the residual feature r_hat and the target mean feature mu, such as taking the sum of the residual feature r_hat and the target mean feature mu as the reconstruction feature y_hat.

[0233] After obtaining the reconstructed feature y_hat, the decoding end can perform a synthetic transformation on the reconstructed feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the reconstructed feature y_hat is input into the synthetic transformation network, and the synthetic transformation network performs a synthetic transformation on the reconstructed feature y_hat to obtain the reconstructed image block x_hat. At this point, the image reconstruction process is completed.

[0234] In a possible implementation, the processing process at the decoding end may be performed by a deep learning model or a neural network model, thereby realizing an end-to-end image compression and decoding process, without any limitation on the decoding process.

[0235] Embodiment 7: In Embodiments 1 - 4, for the encoding end and the decoding end, after obtaining the first bitstream (Bitstream#1) corresponding to the current image block, the first bitstream corresponding to the current image block can be decoded to obtain the coefficient hyperparameter feature z_hat corresponding to the current image block. The probability distribution parameter p can be determined based on the coefficient hyperparameter feature z_hat. For example, the coefficient hyperparameter feature z_hat can be input into the probability hyperparameter decoding network to obtain the probability distribution parameter p. The second bitstream (Bitstream#2) corresponding to the current image block can be decoded based on the probability distribution parameter p to obtain the residual feature r_hat corresponding to the current image block.

[0236] The target mean feature mu corresponding to the current image block can be determined based on the coefficient hyperparameter feature z_hat. For example, the coefficient hyperparameter feature z_hat can be input into the mean hyperparameter decoding network to obtain the initial mean feature m corresponding to the current image block, and the initial mean feature m can be used as the target mean feature mu. Alternatively, the coefficient hyperparameter feature z_hat can be input into the mean hyperparameter decoding network to obtain the initial mean feature m corresponding to the current image block, and the initial mean feature m and the obtained reconstruction feature y_hat of the current image block can be input into the context model to obtain the target mean feature mu.

[0237] The reconstruction feature y_hat corresponding to the current image block is determined based on the target mean feature mu and the residual feature r_hat, and the reconstructed image block x_hat corresponding to the current image block is determined based on the reconstruction feature y_hat. For example, the reconstruction feature y_hat can be input into the synthesis transformation network to obtain the reconstructed image block x_hat corresponding to the current image block.

[0238] In Embodiments 1, 2, 5, and 6, for the encoding end and the decoding end, after obtaining the first bitstream (Bitstream#1) corresponding to the current image block, the first bitstream corresponding to the current image block can be decoded to obtain the coefficient hyperparameter feature z_hat corresponding to the current image block. The probability distribution parameter p and the target mean feature mu corresponding to the current image block can be determined based on the coefficient hyperparameter feature z_hat. For example, the coefficient hyperparameter feature z_hat can be input into the hyperparameter decoding network to obtain the reference feature m corresponding to the current image block. Then, the reference feature m and the obtained reconstruction feature y_hat of the current image block can be input into the context model to obtain the probability distribution parameter p and the target mean feature mu.

[0239] The second bitstream (Bitstream#2) corresponding to the current image block can be decoded based on the probability distribution parameter p to obtain the residual feature r_hat corresponding to the current image block. The reconstructed feature y_hat corresponding to the current image block is determined based on the target mean feature mu and the residual feature r_hat, and the reconstructed image block x_hat corresponding to the current image block is determined based on the reconstructed feature y_hat. For example, the reconstructed feature y_hat can be input into the synthesis transformation network to obtain the reconstructed image block x_hat corresponding to the current image block.

[0240] Embodiment 8: In Embodiments 1-7, for the encoding end and the decoding end, the coefficient hyperparameter feature z_hat can be input into the probability hyperparameter decoding network to obtain the probability distribution parameter p. Among them, the probability hyperparameter decoding network can be a trained neural network for performing the inverse transformation of the coefficient hyperparameter feature on the coefficient hyperparameter feature z_hat. For example, the structure of the probability hyperparameter decoding network can be referred to Figure 5A as shown. Of course, Figure 5A it is only an example of the probability hyperparameter decoding network, and the structure of this probability hyperparameter decoding network is not limited. Subsequently, Figure 5A the probability hyperparameter decoding network is taken as an example.

[0241] Exemplarily, the probability hyperparameter decoding network can sequentially include: a convolutional layer (such as a convolutional layer with a size of 1*1, an input channel number of c, and an output channel number of c), a Relu activation layer, a convolutional layer (such as a convolutional layer with a size of 3*3, an input channel number of c, and an output channel number of c), a Relu activation layer, a convolutional layer (such as a convolutional layer with a size of 1*1, an input channel number of c, and an output channel number of c), a Pixshuffle (upsampling) layer, and a Crop layer. The Crop layer is used to crop and align the pixel positions after upsampling. The Crop layer can be located behind the Pixshuffle layer to ensure the alignment of the feature map size.

[0242] Refer to Figure 5A as shown. The input data of the probability hyperparameter decoding network is the coefficient hyperparameter feature z_hat. After being processed by the probability hyperparameter decoding network, the probability distribution parameter p can be obtained and the probability distribution parameter p can be output.

[0243] Refer to Figure 5A as shown. The probability hyperparameter decoding network can include an activation layer, and the activation layer can adopt the Relu activation function. The Relu activation function can be referred to Figure 5B as shown, that is, the Relu activation function corresponds to a lower limit threshold. As the input feature decreases, by restricting the lower limit of the output feature, the activation layer will not output output features with very small values.

[0244] In this embodiment, the activation layer of the probability hyperparameter decoding network can also be optimized so that the activation layer adopts a target activation function with upper and lower clamping, that is, the activation layer uses the target activation function with upper and lower clamping to perform non-linear adjustment on the features. For example, the target activation function can be the Relu6 activation function, or other activation functions, as long as the upper and lower clamping are achieved simultaneously. The Relu6 activation function will be used as an example for subsequent description. Refer to Figure 5C As shown, it is a schematic structural diagram of the probability hyperparameter decoding network. The Relu activation layer is replaced by the Relu6 activation layer, that is, the Relu6 activation function is used for processing. Of course, the activation layer of the probability hyperparameter decoding network can also be optimized so that the activation layer adopts a certain activation function, and the activation function corresponds to an upper limit threshold, that is, the upper limit of the activation layer is restricted, and the lower limit of the activation layer is not restricted.

[0245] The Relu6 activation function (that is, the target activation function with upper and lower clamping) can be referred to Figure 5D As shown, the Relu6 activation function corresponds to an upper limit threshold and a lower limit threshold. As the input feature decreases, the lower limit of the output feature will be restricted, so that the output feature of the activation layer is not less than the lower limit threshold ( Figure 5D 0 is used as an example in this case). As the input feature increases, the upper limit of the output feature will be restricted, so that the output feature of the activation layer is not greater than the upper limit threshold ( Figure 5D 6 is used as an example in this case). By restricting the lower limit and the upper limit, the Relu6 activation function makes the output feature located in a specified interval (between the lower limit threshold and the upper limit threshold), so that the activation layer will not output output features with very large or very small values.

[0246] Exemplarily, the upper limit threshold corresponding to the Relu6 activation function can be a configured fixed upper limit threshold, that is, a certain fixed upper limit threshold is defaulted as the upper limit threshold corresponding to the Relu6 activation function. For example, the fixed upper limit threshold in 5D is 6. The lower limit threshold corresponding to the Relu6 activation function can be a configured fixed lower limit threshold, that is, a certain fixed lower limit threshold is defaulted as the lower limit threshold corresponding to the Relu6 activation function. For example, the fixed lower limit threshold in 5D is 0. Of course, both the fixed upper limit threshold and the fixed lower limit threshold can be configured according to experience, and the fixed upper limit threshold and the fixed lower limit threshold are not restricted in this embodiment.

[0247] Exemplarily, the upper threshold corresponding to the Relu6 activation function can be an adaptively learned upper threshold, that is, the upper threshold corresponding to the Relu6 activation function is determined through the training process instead of a fixed upper threshold. In this embodiment, the learning process of this upper threshold is not limited. Moreover, the lower threshold corresponding to the Relu6 activation function can be an adaptively learned lower threshold, that is, the lower threshold corresponding to the Relu6 activation function is determined through the training process instead of a fixed lower threshold.

[0248] Alternatively, the upper threshold corresponding to the Relu6 activation function can be an adaptively learned upper threshold instead of a fixed upper threshold, and moreover, the lower threshold corresponding to the Relu6 activation function can be a fixed lower threshold.

[0249] Alternatively, the lower threshold corresponding to the Relu6 activation function can be an adaptively learned lower threshold instead of a fixed lower threshold, and moreover, the upper threshold corresponding to the Relu6 activation function can be a fixed upper threshold.

[0250] Exemplarily, when adaptively learning the upper threshold corresponding to the Relu6 activation function, the same upper threshold can be learned for all channels of the activation layer. That is to say, only one upper threshold needs to be adaptively learned, and this upper threshold serves as the upper threshold for all channels of the activation layer. Alternatively, when adaptively learning the upper threshold corresponding to the Relu6 activation function, a separate upper threshold can be learned for each channel of the activation layer. That is to say, assuming the activation layer corresponds to K channels (i.e., the number of channels of the input feature, and K is a positive integer), then K upper thresholds need to be adaptively learned. The K upper thresholds correspond one-to-one with the K channels, each channel corresponds to a separate upper threshold, and the upper thresholds corresponding to different channels can be the same or different.

[0251] Exemplarily, when adaptively learning the lower threshold corresponding to the Relu6 activation function, the same lower threshold can be learned for all channels of the activation layer. That is to say, only one lower threshold needs to be adaptively learned, and this lower threshold serves as the lower threshold for all channels of the activation layer. Alternatively, when adaptively learning the lower threshold corresponding to the Relu6 activation function, a separate lower threshold can be learned for each channel of the activation layer. That is to say, assuming the activation layer corresponds to K channels (i.e., the number of channels of the input feature, and K is a positive integer), then K lower thresholds need to be adaptively learned. The K lower thresholds correspond one-to-one with the K channels, each channel corresponds to a separate lower threshold, and the lower thresholds corresponding to different channels can be the same or different.

[0252] Embodiment 9: In Embodiments 1 - 7, for the encoding end and the decoding end, the coefficient hyperparameter feature z_hat can be input into the mean hyperparameter decoding network to obtain the initial mean feature m corresponding to the current image block. For example, the structure of the mean hyperparameter decoding network can be referred to Figure 6A as shown. Of course, Figure 6A this is just an example of the mean hyperparameter decoding network, and the structure of this mean hyperparameter decoding network is not limited. Subsequently, Figure 6A the mean hyperparameter decoding network is taken as an example.

[0253] Exemplarily, the mean hyperparameter decoding network can successively include: a convolutional layer (such as a convolutional layer with a size of 1*1, an input channel number of c, and an output channel number of c), a transposed convolutional layer TConv (such as a convolutional layer with a size of 4*4, an input channel number of c, an output channel number of c, and a convolutional stride of s2), a Crop layer, a Relu activation layer, a convolutional layer (such as a convolutional layer with a size of 3*3, an input channel number of c, and an output channel number of c), a Relu activation layer, and a convolutional layer (such as a convolutional layer with a size of 3*3, an input channel number of c, and an output channel number of c). Refer to Figure 6A as shown. The input data of the mean hyperparameter decoding network is the coefficient hyperparameter feature z_hat. After being processed by the mean hyperparameter decoding network, the initial mean feature m can be obtained and output.

[0254] Refer to Figure 6A as shown. The mean hyperparameter decoding network can include an activation layer, and the activation layer uses the Relu activation function. The Relu activation function can be referred to Figure 5B as shown. On this basis, the activation layer of the mean hyperparameter decoding network can be optimized so that the activation layer uses a target activation function with upper and lower clamping, that is, the target activation function with upper and lower clamping is used to perform non - linear adjustment on the feature. For example, this target activation function can be the Relu6 activation function. Refer to Figure 6B as shown, which is a schematic diagram of the structure of the mean hyperparameter decoding network. The Relu activation layer is replaced by the Relu6 activation layer, that is, the Relu6 activation function is used for processing. Of course, the activation layer of the mean hyperparameter decoding network can also be optimized so that the activation layer uses a certain activation function, and this activation function corresponds to an upper limit threshold, that is, the upper limit of the activation layer is restricted, while the lower limit of the activation layer is not restricted.

[0255] The Relu6 activation function can be referred to Figure 5DAs shown, the Relu6 activation function corresponds to an upper threshold and a lower threshold. The upper threshold corresponding to the Relu6 activation function can be a configured fixed upper threshold, and the lower threshold corresponding to the Relu6 activation function can be a configured fixed lower threshold. Alternatively, the upper threshold corresponding to the Relu6 activation function can be an adaptively learned upper threshold, and / or the lower threshold corresponding to the Relu6 activation function can be an adaptively learned lower threshold.

[0256] Exemplarily, when adaptively learning the upper threshold corresponding to the Relu6 activation function, the same upper threshold can be learned for all channels of the activation layer, or, a separate upper threshold can be learned for each channel of the activation layer.

[0257] Exemplarily, when adaptively learning the lower threshold corresponding to the Relu6 activation function, the same lower threshold can be learned for all channels of the activation layer, or, a separate lower threshold can be learned for each channel of the activation layer.

[0258] Embodiment 10: In Embodiments 1 - 7, for the encoding end and the decoding end, the initial mean feature m and the obtained reconstructed feature y_hat of the current image block can be input to the context model to obtain the target mean feature mu. For example, the structure of the context model can be referred to Figure 6C as shown. Of course, Figure 6C this is only an example of the context model, and the structure of this context model is not limited. Subsequently, the context model of Figure 6C will be used as an example for illustration.

[0259] Exemplarily, the context model can include multiple parameter fusion networks. For each parameter fusion network, the parameter fusion network can sequentially include: a convolutional layer (such as a convolutional layer with a size of 1*1), a Relu activation layer, a convolutional layer (such as a convolutional layer with a size of 1*1), a Relu activation layer, a convolutional layer (such as a convolutional layer with a size of 1*1). Refer to Figure 6C as shown. The input data of the context model is the initial mean feature m and the reconstructed feature y_hat. After being processed by the context model, the target mean feature mu can be obtained, and the target mean feature mu is output. The processing process of this context model is not limited.

[0260] Refer to Figure 6C as shown. The parameter fusion network of the context model can include an activation layer, and the activation layer uses the Relu activation function. The Relu activation function can be referred to Figure 5BAs shown below. On this basis, the activation layer of the parameter fusion network can be optimized so that the activation layer adopts a target activation function with upper and lower clamping, that is, the feature is non-linearly adjusted by using the target activation function with upper and lower clamping. For example, the target activation function can be the Relu6 activation function. See Figure 6D As shown in Figure 6D , it is a schematic structural diagram of the parameter fusion network of the context model, and the Relu activation layer is replaced by the Relu6 activation layer. Of course, the activation layer of the context model can also be optimized so that the activation layer adopts an activation function, and the activation function corresponds to an upper limit threshold, that is, the upper limit of the activation layer is restricted, while the lower limit of the activation layer is not restricted.

[0261] The Relu6 activation function can be seen in Figure 5D As shown in Figure 5D , the Relu6 activation function corresponds to an upper limit threshold and a lower limit threshold. The upper limit threshold corresponding to the Relu6 activation function can be a configured fixed upper limit threshold, and the lower limit threshold corresponding to the Relu6 activation function can be a configured fixed lower limit threshold. Or, the upper limit threshold corresponding to the Relu6 activation function can be an upper limit threshold for adaptive learning, and / or, the lower limit threshold corresponding to the Relu6 activation function can be a lower limit threshold for adaptive learning.

[0262] Exemplarily, when adaptively learning the upper limit threshold corresponding to the Relu6 activation function, the same upper limit threshold can be learned for all channels of the activation layer, or, a separate upper limit threshold can be learned for each channel of the activation layer.

[0263] Exemplarily, when adaptively learning the lower limit threshold corresponding to the Relu6 activation function, the same lower limit threshold can be learned for all channels of the activation layer, or, a separate lower limit threshold can be learned for each channel of the activation layer.

[0264] In Embodiment 8 - Embodiment 10, the Relu activation function in the probability hyperparameter decoding network, the mean hyperparameter decoding network, and the context model can be replaced by the Relu6 activation function. If there is an activation layer in the probability hyperparameter decoding network that uses the LeakyRelu activation function, the LeakyRelu activation function can also be replaced by the Relu6 activation function, and this process is not restricted. And / or, if there is an activation layer in the mean hyperparameter decoding network that uses the LeakyRelu activation function, the LeakyRelu activation function can also be replaced by the Relu6 activation function, and this process is not restricted. And / or, if there is an activation layer in the context model that uses the LeakyRelu activation function, the LeakyRelu activation function can also be replaced by the Relu6 activation function, and this process is not restricted.

[0265] Embodiment 11: In Embodiments 1 - 7, for the encoding end and the decoding end, the coefficient hyperparameter feature z_hat can be input into the hyperparameter decoding network to obtain the reference feature m, and the reference feature m and the obtained reconstruction feature y_hat of the current image block can be input into the context model to obtain the probability distribution parameter p and the target mean feature mu.

[0266] The structure of the context model can be referred to Figure 6C as shown. The context model can include multiple parameter fusion networks, and the parameter fusion network can sequentially include: a convolutional layer, a Relu activation layer, a convolutional layer, a Relu activation layer, and a convolutional layer. Regarding the structure of the hyperparameter decoding network, no limitation is made in this embodiment, and the hyperparameter decoding network can include at least one Relu activation layer.

[0267] Regarding the context model and the hyperparameter decoding network, the Relu activation function is adopted for the activation layer, and the Relu activation function can be referred to Figure 5B as shown. On this basis, the activation layer of the context model and / or the hyperparameter decoding network can be optimized so that the activation layer adopts a target activation function with upper and lower clamping, that is, the target activation function with upper and lower clamping is used to perform non-linear adjustment on the feature. For example, the target activation function can be the Relu6 activation function, that is, the Relu activation layer is replaced with the Relu6 activation layer. The activation layer of the context model and / or the hyperparameter decoding network can also be optimized so that the activation layer adopts a certain activation function, and the activation function corresponds to an upper limit threshold, that is, the upper limit of the activation layer is restricted, while the lower limit of the activation layer is not restricted.

[0268] The Relu6 activation function can be referred to Figure 5D as shown. The Relu6 activation function corresponds to an upper limit threshold and a lower limit threshold. The upper limit threshold corresponding to the Relu6 activation function can be a configured fixed upper limit threshold, and the lower limit threshold corresponding to the Relu6 activation function can be a configured fixed lower limit threshold. Or, the upper limit threshold corresponding to the Relu6 activation function can be an adaptively learned upper limit threshold, and / or, the lower limit threshold corresponding to the Relu6 activation function can be an adaptively learned lower limit threshold.

[0269] Exemplarily, when adaptively learning the upper limit threshold corresponding to the Relu6 activation function, the same upper limit threshold can be learned for all channels of the activation layer, or, a separate upper limit threshold can be learned for each channel of the activation layer.

[0270] Exemplarily, when adaptively learning the lower limit threshold corresponding to the Relu6 activation function, the same lower limit threshold can be learned for all channels of the activation layer, or, a separate lower limit threshold can be learned for each channel of the activation layer.

[0271] Embodiment 12: In Embodiments 1 - 7, for the encoding end and the decoding end, the reconstructed feature y_hat can be input into the synthesis transformation network to obtain the reconstructed image block x_hat corresponding to the current image block. Among them, the synthesis transformation network can be a trained neural network for performing synthesis transformation on the reconstructed feature y_hat. For example, the synthesis transformation network can correspond to three branches, namely, a high - complexity synthesis transformation network, a medium - complexity synthesis transformation network, and a low - complexity synthesis transformation network, and the image qualities output by the three branches are different. At least one of the high - complexity synthesis transformation network, the medium - complexity synthesis transformation network, and the low - complexity synthesis transformation network can be deployed according to the computing power. For example, for the decoding end with higher computing power, the high - complexity synthesis transformation network, the medium - complexity synthesis transformation network, and the low - complexity synthesis transformation network can be deployed simultaneously. If a reconstructed image block x_hat with higher image quality needs to be output, the reconstructed feature y_hat can be input into the high - complexity synthesis transformation network; if a reconstructed image block x_hat with lower image quality needs to be output, the reconstructed feature y_hat can be input into the low - complexity synthesis transformation network; if a reconstructed image block x_hat with moderate image quality needs to be output, the reconstructed feature y_hat can be input into the medium - complexity synthesis transformation network. Another example is that for the decoding end with lower computing power, only the low - complexity synthesis transformation network can be deployed, and the reconstructed feature y_hat is input into the low - complexity synthesis transformation network. Taking the deployment of three branches as an example, see Figure 7A As shown, it is a schematic structural diagram of the synthesis transformation network, Figure 7A which shows the three branches of the synthesis transformation network.

[0272] For example, the structure of the low - complexity synthesis transformation network can be seen in Figure 7B As shown. Of course, Figure 7B this is only an example of the low - complexity synthesis transformation network, and there is no limitation in this regard. Subsequently, Figure 7BTake the following as an example. The low-complexity synthesis transformation network may sequentially include: an LRB layer, a convolutional layer (Conv, such as a convolutional layer with a size of 2*2, an input channel number of c, an output channel number of c, and a convolutional stride of 2), a Pixshuffle layer (such as a 2-fold upsampling layer), a ResAU layer, a Crop layer, a convolutional layer (Conv, such as a convolutional layer with a size of 2*2, an input channel number of c1, an output channel number of c2, and a convolutional stride of s2), a Pixshuffle layer (such as a 2-fold upsampling layer), a Crop layer, a ResAU layer, a convolutional layer (such as a convolutional layer with a size of 3*3, an input channel number of c, and an output channel number of c), a ResAU layer, a convolutional layer (such as a convolutional layer with a size of 1*1, an input channel number of c, and an output channel number of 16*cout), a Pixshuffle layer (such as a 4-fold upsampling layer), and a Crop layer. The Crop layer is used to crop and align the upsampled pixel positions. The Crop layer can be located behind the Pixshuffle layer to ensure the alignment of the feature map sizes.

[0273] See Figure 7B As shown, the input data of the low-complexity synthesis transformation network is the reconstructed feature y_hat. After being processed by the low-complexity synthesis transformation network, the reconstructed image patch x_hat can be obtained and the reconstructed image patch x_hat is output.

[0274] For example, the structure of the medium-complexity synthesis transformation network can be seen in Figure 7C As shown. Of course, Figure 7C This is just an example of the medium-complexity synthesis transformation network and is not limited thereto. Subsequently, take Figure 7C as an example. The medium-complexity synthesis transformation network may sequentially include: an LRB layer, a transposed convolutional layer (TConv, such as a convolutional layer with a size of 4*4, an input channel number of c1, an output channel number of c2, and a convolutional stride of S2), a ResAU layer, a Crop layer, a transposed convolutional layer (TConv, such as a convolutional layer with a size of 4*4, an input channel number of c2, an output channel number of c3, and a convolutional stride of S2), a Crop layer, a ResAU layer, a convolutional layer (Conv, such as a convolutional layer with a size of 3*3, an input channel number of c, and an output channel number of c), a ResAU layer, a convolutional layer (Conv, such as a convolutional layer with a size of 3*3, an input channel number of c, and an output channel number of c), a Pixshuffle layer (such as a 4-fold upsampling layer), and a Crop layer. The input data of the medium-complexity synthesis transformation network is the reconstructed feature y_hat. After being processed by the medium-complexity synthesis transformation network, the reconstructed image patch x_hat can be obtained and the reconstructed image patch x_hat is output.

[0275] For example, the structure of the high-complexity synthesis transformation network can be seen in Figure 7D As shown. Of course, Figure 7DThis is just an example of a high - complexity synthesis transformation network, without any limitations. Subsequently, taking Figure 7D as an example. The high - complexity synthesis transformation network may sequentially include: an RB layer, a transposed convolution layer (TConv, such as a convolution layer with a size of 3*3, an input channel number of c, an output channel number of c, and a convolution stride of S2), a ResAU layer, a Crop layer, a transposed convolution layer (TConv, such as a convolution layer with a size of 3*3, an input channel number of c, an output channel number of c, and a convolution stride of S2), a CAB layer, a Crop layer, a ResAU layer, a convolution layer (Conv, such as a convolution layer with a size of 1*1, an input channel number of c, an output channel number of c), a Pixshuffle layer (such as a 2 - fold upsampling layer), a TAM layer, a Crop layer, a ResAU layer, a transposed convolution layer (TConv, such as a convolution layer with a size of 3*3, an input channel number of c2, an output channel number of c3, and a convolution stride of S2), and a Crop layer.

[0276] See Figure 7D As shown, the input data of the high - complexity synthesis transformation network is the reconstructed feature y_hat. After being processed by the high - complexity synthesis transformation network, a reconstructed image patch x_hat can be obtained and the reconstructed image patch x_hat is output.

[0277] In Figure 7D , the CAB layer is a convolution - base attention block. The structure of the CAB layer can be seen in Figure 7E As shown. Of course, Figure 7E this is just an example, without any limitations.

[0278] In Figure 7D , the TAM layer is a Transformer - based attention block. The structure of the TAM layer can be seen in Figure 7F As shown. Of course, Figure 7F this is just an example, without any limitations. Or, the structure of the TAM layer can be seen in Figure 7G As shown. Of course, Figure 7G this is just an example, without any limitations.

[0279] In Figure 7B and Figure 7C , the LRB layer is a lightweight residual block. The structure of the LRB layer can be seen in Figure 7H As shown. Of course, Figure 7H these are just examples, without any limitations. In Figure 7DIn it, the RB layer is a residual block (Residual Block), and the structure of the RB layer can be seen in Figure 7I as shown. Of course, Figure 7I this is just an example and there is no limitation on this.

[0280] In the above synthesis transformation network, an activation layer using the Relu activation function (such as the Relu activation functions in the LRB layer and the RB layer) may be involved. On this basis, the Relu activation function in the synthesis transformation network can be optimized, and the Relu activation function can be replaced with a target activation function with upper and lower clamping (such as the Relu6 activation function), that is, the target activation function with upper and lower clamping is used to perform non-linear adjustment on the features. Among them, the Relu6 activation function corresponds to an upper threshold and a lower threshold. Of course, the activation layer of the synthesis transformation network can also be optimized so that the activation layer uses an activation function, and the activation function corresponds to an upper threshold, that is, the upper limit of the activation layer is restricted, and the lower limit of the activation layer is not restricted.

[0281] The upper threshold corresponding to the Relu6 activation function can be a configured fixed upper threshold, and the lower threshold corresponding to the Relu6 activation function can be a configured fixed lower threshold. Or, the upper threshold corresponding to the Relu6 activation function can be an upper threshold for adaptive learning, and / or the lower threshold corresponding to the Relu6 activation function can be a lower threshold for adaptive learning.

[0282] Embodiment 13: In Embodiment 12, for the high-complexity synthesis transformation network, the medium-complexity synthesis transformation network, and the low-complexity synthesis transformation network, all may include a ResAU layer, and the structure of the ResAU layer can be seen in Figure 8A as shown. The ResAU layer may sequentially include a LeakyRelu activation layer (that is, an activation layer using the LeakyRelu activation function), a grouped convolution layer (Group-Conv, such as a convolution layer with a size of 3*3, an input channel number of c, an output channel number of c, a convolution stride of s1, and a grouped number g of 4), and a convolution layer (Conv, such as a convolution layer with a size of 1*1, an input channel number of c, an output channel number of c, and a convolution stride of s1). In addition, the ResAU layer may further include a skip connection structure, that is, the input feature of the ResAU layer is connected to the output feature of the convolution layer, and operations such as feature addition, feature multiplication, and feature splicing can be performed. For example, assuming that the input feature of the ResAU layer is feature X1, the LeakyRelu activation layer processes feature X1 to obtain feature X2, the grouped convolution layer Group-Conv processes feature X2 to obtain feature X3, and the convolution layer Conv processes feature X3 to obtain feature X4. Through the skip connection structure, feature X1 and feature X4 can be added to obtain the output feature of the ResAU layer.

[0283] From Figure 8A It can be seen that the input of the ResAU layer has a residual connection (i.e., skip connection structure). The input features of the ResAU layer sequentially pass through the LeakyRelu function (LeakyRelu activation layer), group convolution (Group-Conv in the grouped convolution layer), and 1x1 conventional convolution (Conv in the convolution layer). Its output is added to the residual connection to obtain the output features of the ResAU layer.

[0284] The LeakyRelu activation layer of the ResAU layer uses the LeakyRelu activation function. The LeakyRelu activation function can be referred to Figure 8B as shown. That is, the LeakyRelu activation function does not correspond to a lower threshold or an upper threshold. Obviously, as the input features increase, the upper limit of the output features will not be restricted, resulting in the activation layer possibly outputting output features with very large values. As the input features decrease, the lower limit of the output features will not be restricted, resulting in the activation layer possibly outputting output features with very small values. The above LeakyRelu activation function has a relatively high complexity for the hardware, resulting in a relatively high complexity at the decoding end.

[0285] On this basis, in one implementation manner of this embodiment, the activation layer of the ResAU layer can be optimized so that the activation layer uses a target activation function with upper and lower clamping, that is, the activation layer uses a target activation function with upper and lower clamping to perform non-linear adjustment on the features. For example, the target activation function can be the Relu6 activation function, or other activation functions, as long as upper and lower clamping are achieved simultaneously. Hereinafter, the Relu6 activation function will be used as an example for description. Refer to Figure 8C as shown. It is a schematic diagram of the structure of the ResAU layer. The LeakyRelu activation layer is replaced by the Relu6 activation layer, that is, the Relu6 activation function is used for processing.

[0286] The Relu6 activation function (i.e., the target activation function with upper and lower clamping) can be referred to Figure 5D as shown. The Relu6 activation function corresponds to an upper threshold and a lower threshold. As the input features decrease, the lower limit of the output features will be restricted, so that the output features of the activation layer are not less than the lower threshold. As the input features increase, the upper limit of the output features will be restricted, so that the output features of the activation layer are not greater than the upper threshold. The Relu6 activation function restricts the lower and upper limits, so that the output features are between the lower threshold and the upper threshold, so that the activation layer will not output output features with very large or very small values.

[0287] The upper threshold corresponding to the Relu6 activation function can be a configured fixed upper threshold, that is, by default, a certain fixed upper threshold is used as the upper threshold corresponding to the Relu6 activation function. The lower threshold corresponding to the Relu6 activation function can be a configured fixed lower threshold, that is, by default, a certain fixed lower threshold is used as the lower threshold corresponding to the Relu6 activation function.

[0288] The upper threshold corresponding to the Relu6 activation function can be an adaptively learned upper threshold, that is, the upper threshold corresponding to the Relu6 activation function is determined through the training process, rather than a fixed upper threshold. The lower threshold corresponding to the Relu6 activation function can be an adaptively learned lower threshold, that is, the lower threshold corresponding to the Relu6 activation function is determined through the training process, rather than a fixed lower threshold. Or, the upper threshold corresponding to the Relu6 activation function can be an adaptively learned upper threshold, and the lower threshold corresponding to the Relu6 activation function can be a fixed lower threshold. Or, the lower threshold corresponding to the Relu6 activation function can be an adaptively learned lower threshold, and the upper threshold corresponding to the Relu6 activation function can be a fixed upper threshold.

[0289] When adaptively learning the upper threshold corresponding to the Relu6 activation function, the same upper threshold can be learned for all channels of the activation layer. Only one upper threshold needs to be adaptively learned, and this upper threshold is used as the upper threshold for all channels of the activation layer. Or, when adaptively learning the upper threshold corresponding to the Relu6 activation function, a separate upper threshold can be learned for each channel of the activation layer. Assuming that the activation layer corresponds to K channels, then K upper thresholds need to be adaptively learned. The K upper thresholds correspond one-to-one with the K channels, each channel corresponds to a separate upper threshold, and the upper thresholds corresponding to different channels can be the same or different.

[0290] When adaptively learning the lower threshold corresponding to the Relu6 activation function, the same lower threshold can be learned for all channels of the activation layer. Only one lower threshold needs to be adaptively learned, and this lower threshold is used as the lower threshold for all channels of the activation layer. Or, when adaptively learning the lower threshold corresponding to the Relu6 activation function, a separate lower threshold can be learned for each channel of the activation layer. Assuming that the activation layer corresponds to K channels, then K lower thresholds need to be adaptively learned. The K lower thresholds correspond one-to-one with the K channels, each channel corresponds to a separate lower threshold, and the lower thresholds corresponding to different channels can be the same or different.

[0291] When adapting and learning the upper and lower threshold values corresponding to the Relu6 activation function, the Relu6 activation function can also be referred to as the ClipRelu6 activation function. When the ClipRelu6 activation function is extended to the channel level, that is, when the upper and lower threshold values are learned separately for each channel, the Relu6 activation function can also be referred to as the chs-wise ClipRelu6 activation function.

[0292] In another implementation manner of this embodiment, the activation layer of the ResAU layer is optimized so that the activation layer uses the Relu activation function, that is, the activation layer uses the Relu activation function to perform non-linear adjustment on the features. Refer to Figure 8D As shown, it is a schematic structural diagram of the ResAU layer, and the LeakyRelu activation layer is replaced by the Relu activation layer, that is, the Relu activation function is used for processing.

[0293] The Relu activation function can be referred to Figure 5B As shown, the Relu activation function can correspond to a lower threshold value. As the input features decrease, the lower limit of the output features will be restricted, so that the output features of the activation layer are not less than the lower threshold value. By restricting the lower limit, the Relu activation function enables the activation layer not to output output features with very small values.

[0294] The lower threshold value corresponding to the Relu activation function can be a configured fixed lower threshold value, that is, a certain fixed lower threshold value is defaulted as the lower threshold value corresponding to the Relu activation function. Or, the lower threshold value corresponding to the Relu activation function can be an adaptively learned lower threshold value, that is, the lower threshold value corresponding to the Relu activation function is determined through the training process, rather than a fixed lower threshold value.

[0295] When adapting and learning the lower threshold value corresponding to the Relu activation function, the same lower threshold value can be learned for all channels of the activation layer, and only one lower threshold value needs to be adaptively learned, and this lower threshold value is used as the lower threshold value for all channels of the activation layer. Or, when adapting and learning the lower threshold value corresponding to the Relu activation function, a separate lower threshold value can be learned for each channel of the activation layer. Assuming that the activation layer corresponds to K channels, then K lower threshold values need to be adaptively learned. The K lower threshold values correspond one by one to the K channels, and each channel corresponds to a separate lower threshold value, and the lower threshold values corresponding to different channels can be the same or different.

[0296] In another implementation of this embodiment, the activation layer of the ResAU layer is optimized so that the activation layer uses a certain activation function, that is, the activation layer uses this activation function to perform non-linear adjustment on the features. This activation function can correspond to an upper limit threshold. As the input features increase, the upper limit of the output features will be restricted, so that the output features of the activation layer are not greater than this upper limit threshold. The upper limit threshold corresponding to this activation function can be a configured fixed upper limit threshold, that is, by default, a certain fixed upper limit threshold is used as the upper limit threshold corresponding to this activation function. Or, the upper limit threshold corresponding to this activation function can be an adaptively learned upper limit threshold, that is, the upper limit threshold corresponding to this activation function is determined through the training process.

[0297] When adaptively learning the upper limit threshold corresponding to this activation function, the same upper limit threshold can be learned for all channels of the activation layer. Only one upper limit threshold needs to be adaptively learned, and this upper limit threshold is used as the upper limit threshold for all channels of the activation layer. Or, when adaptively learning the upper limit threshold corresponding to the Relu activation function, a separate upper limit threshold can be learned for each channel of the activation layer. Assuming that the activation layer corresponds to K channels, then K upper limit thresholds need to be adaptively learned. The K upper limit thresholds correspond one-to-one with the K channels, and each channel corresponds to a separate upper limit threshold, and the upper limit thresholds corresponding to different channels can be the same or different.

[0298] See Figure 8C and Figure 8D As shown, the ResAU layer (i.e., the ResAU network layer) can include an activation layer (using the Relu6 activation function or the Relu activation function), a processing layer, and a skip connection structure. The skip connection structure (i.e., the residual connection) is used to connect the input features of the ResAU layer with the output features of the processing layer, and can perform operations such as feature addition, feature multiplication, and feature splicing. Of course, it can also be a skip connection structure in other positions, such as the skip connection structure is used to connect the output features of the activation layer with the output features of the processing layer. The position of this skip connection structure is not limited and can be a skip connection structure at any position.

[0299] Exemplarily, the ResAU layer at least includes an activation layer and a skip connection structure. In addition to the activation layer and the skip connection structure, the ResAU layer may include a processing layer or may not include a processing layer. When including a processing layer, the ResAU layer can sequentially include an activation layer, a processing layer, and a skip connection structure. In Figure 8C and Figure 8D , taking the case of including a processing layer as an example. In addition to Figure 8C and Figure 8D 's network structure, the network structure of the ResAU layer can also be seen in Figure 8E , Figure 8F , Figure 8G , Figure 8H as shown.

[0300] In Figure 8E , the ResAU layer may include a convolutional layer (such as a convolutional layer with a size of 3*3, an input channel number of c, an output channel number of c, and a convolutional stride of s1), an activation layer (using the Relu6 activation function or the Relu activation function), and a skip connection structure. In Figure 8F , the ResAU layer may include a convolutional layer (such as a convolutional layer with a size of 1*1, an input channel number of c, an output channel number of c, and a convolutional stride of s1), an activation layer (using the Relu6 activation function or the Relu activation function), a convolutional layer (such as a convolutional layer with a size of 3*3, an input channel number of c, an output channel number of c, and a convolutional stride of s1), and a skip connection structure. In Figure 8G , the ResAU layer may include a convolutional layer (such as a convolutional layer with a size of 3*3, an input channel number of c, an output channel number of c, and a convolutional stride of s1), an activation layer (using the Relu6 activation function or the Relu activation function), a convolutional layer (such as a convolutional layer with a size of 1*1, an input channel number of c, an output channel number of c, and a convolutional stride of s1), a convolutional layer (such as a convolutional layer with a size of 3*3, an input channel number of c, an output channel number of c, and a convolutional stride of s1), and a skip connection structure. In Figure 8H , the ResAU layer may include a convolutional layer (such as a convolutional layer with a size of 3*3, an input channel number of c, an output channel number of c, and a convolutional stride of s1), an activation layer (using the Relu6 activation function or the Relu activation function), a convolutional layer (such as a convolutional layer with a size of 1*1, an input channel number of c, an output channel number of c, and a convolutional stride of s1), an activation layer (using the Relu6 activation function or the Relu activation function), a convolutional layer (such as a convolutional layer with a size of 3*3, an input channel number of c, an output channel number of c, and a convolutional stride of s1), and a skip connection structure. Of course, the above are just a few examples and are not limited to this.

[0301] In a possible implementation, the processing layer (also referred to as a processing unit) for the ResAU layer may include, but is not limited to, at least one of the following: one or more ordinary convolutional layers, one or more transposed convolutional layers, one or more deformable convolutional layers, one or more depthwise separable convolutional layers, one or more grouped convolutional layers, one or more dilated convolutional layers, one or more transformer layers. For each convolutional layer in the processing layer (such as an ordinary convolutional layer, a transposed convolutional layer, a deformable convolutional layer, a depthwise separable convolutional layer, a grouped convolutional layer, a dilated convolutional layer, a transformer layer, etc.), a quantized convolution form may also be adopted, that is, the parameters within the convolutional layer are quantized, and the input features are quantized, and then the convolution operation is performed. That is to say, for each convolutional layer in the processing layer, this convolutional layer is a quantized convolutional layer (or an integerized convolutional layer, and the integerized bit width can be 2, 4, 8, 16, etc.), that is, the parameters within the convolutional layer are integerized parameters.

[0302] For each convolutional layer in the processing layer, these convolutional layers can be repeatedly stacked, and there is no restriction on the stacking form of these convolutional layers. For example, first stack two ordinary convolutional layers, then stack two grouped convolutional layers, then deploy a dilated convolutional layer, then deploy an ordinary convolutional layer, etc. Of course, this is just an example, and there is no restriction on the structure of this processing layer.

[0303] In a possible implementation, for Figure 4A and Figure 4B the encoding framework and decoding framework shown, a Relu6 activation function, or a ClipRelu6 activation function, or a chs-wiseClipRelu6 activation function can be connected after any layer of the probability hyperparameter decoding network. A Relu6 activation function, or a ClipRelu6 activation function, or a chs-wise ClipRelu6 activation function can be connected after any layer of the mean hyperparameter decoding network. A Relu6 activation function, or a ClipRelu6 activation function, or a chs-wise ClipRelu6 activation function can be connected after any layer of the context model. A Relu6 activation function, or a ClipRelu6 activation function, or a chs-wise ClipRelu6 activation function can be connected after any layer of the synthesis transformation network (such as a high-complexity synthesis transformation network, a medium-complexity synthesis transformation network, a low-complexity synthesis transformation network). For Figure 4C and Figure 4DThe shown encoding framework and decoding framework can connect the Relu6 activation function, or the ClipRelu6 activation function, or the chs-wise ClipRelu6 activation function behind any layer of the hyperparameter decoding network. The Relu6 activation function, or the ClipRelu6 activation function, or the chs-wise ClipRelu6 activation function can be connected behind any layer of the context model. The Relu6 activation function, or the ClipRelu6 activation function, or the chs-wise ClipRelu6 activation function can be connected behind any layer of the synthesis transformation network (such as the high-complexity synthesis transformation network, the medium-complexity synthesis transformation network, the low-complexity synthesis transformation network).

[0304] In a possible implementation manner, for Embodiments 1 - 13, by adopting the Relu6 activation function, the ClipRelu6 activation function, or the chs-wise ClipRelu6 activation function, the complexity of the model can be reduced.

[0305] Embodiment 14: In Embodiments 1 - 7, for the encoding end and the decoding end, the coefficient hyperparameter feature z_hat can be input to the mean hyperparameter decoding network to obtain the initial mean feature m corresponding to the current image block, that is, the input of the mean hyperparameter decoding network is the decoding information of the first bitstream (such as the coefficient hyperparameter feature z_hat), and the output of the mean hyperparameter decoding network is the mean mu (i.e., the initial mean feature m serves as the target mean feature mu), or the output of the mean hyperparameter decoding network is the latent feature m of the mean (i.e., the initial mean feature m needs to be output to the context model, and the context model outputs the target mean feature mu).

[0306] In a possible implementation manner, the mean hyperparameter decoding network can first perform upsampling and then convolution processing. This form of the mean hyperparameter decoding network has a high complexity, resulting in a high complexity at the decoding end. In view of the above discovery, in this embodiment, the mean hyperparameter decoding network can be simplified. The mean hyperparameter decoding network can sequentially include a processing layer, an upsampling layer, and a Crop layer. The upsampling layer is behind the processing layer and in front of the Crop layer. Since the upsampling layer is deployed behind the processing layer, the mean hyperparameter decoding network can first perform processing (such as convolution processing, etc.) and then perform upsampling. This form of the mean hyperparameter decoding network has a low complexity, making the complexity at the decoding end low.

[0307] Based on this mean hyperparameter decoding network, the coefficient hyperparameter feature z_hat can be input into the processing layer of the mean hyperparameter decoding network. The processing layer processes the coefficient hyperparameter feature z_hat (such as convolution processing and / or activation processing, etc.) to obtain a processed feature, and inputs the processed feature into the upsampling layer of the mean hyperparameter decoding network. The upsampling layer performs upsampling processing on the processed feature to obtain an upsampled feature, and inputs the upsampled feature into the Crop layer of the mean hyperparameter decoding network. The Crop layer performs cropping and alignment processing on the upsampled feature to obtain an initial mean feature m, and the initial mean feature m is used as the output feature of the mean hyperparameter decoding network. Exemplarily, the Crop layer is used to crop and align the pixel positions after upsampling, and can be located after the upsampling layer to ensure that the feature map sizes are aligned. There is no limit to the processing process of this Crop layer.

[0308] Exemplarily, the upsampling layer (which can also be referred to as a sampling layer or a Pixshuffle layer) can be located behind the processing layer (all network layers other than the upsampling layer and the Crop layer are classified into the processing layer), that is, the upsampling layer is located at the end of the mean hyperparameter decoding network. The form of the upsampling layer can include but is not limited to pixel rearrangement, pixshuffle, upshuffle, transposed convolution, nearest neighbor interpolation. There is no limit to the implementation method of this upsampling layer as long as it can perform upsampling on the input feature.

[0309] For example, the pixel rearrangement method can be used to perform upsampling processing on the processed feature (i.e., the output feature of the processing layer) to obtain an upsampled feature. Pixel rearrangement is used to rearrange the channel domain information to the spatial domain to achieve the upsampling effect.

[0310] For example, the pixshuffle method can be used to perform upsampling processing on the processed feature (i.e., the output feature of the processing layer) to obtain an upsampled feature. Pixshuffle is a way of spatial operation. By rearranging the feature points, the three-dimensional representation [4C, H, W] is rearranged to obtain [C, 2H, 2W], achieving the effect of magnifying the spatial resolution. See Figure 9A As shown, it is a schematic diagram of using the pixshuffle method to perform upsampling processing on the processed feature to obtain an upsampled feature.

[0311] For example, the upshuffle method can be used to perform upsampling processing on the processed feature (i.e., the output feature of the processing layer) to obtain an upsampled feature. Upshuffle rearranges the feature points to rearrange the three-dimensional representation [4C, H, W] to obtain [C, 2H, 2W], achieving the effect of magnifying the spatial resolution. The upshuffle has a different arrangement criterion compared to pixshuffle. See Figure 9BAs shown, it is a schematic diagram of upsampling the processed features in the way of upshuffle.

[0312] For example, the processed features can be upsampled by using transposed convolution to obtain the upsampled features.

[0313] For example, the processed features can be upsampled by using nearest neighbor interpolation (nearest neighbor interpolation can also be called nearest neighbor upsampling) to obtain the upsampled features. Nearest neighbor interpolation plays the role of magnifying the spatial resolution of the features. Nearest neighbor interpolation rearranges the three-dimensional representation [C, H, W] to obtain [C, 2H, 2W], and the extra pixel points are obtained by copying the points with the nearest spatial position as the current interpolation points. See Figure 9C As shown, it is a schematic diagram of using the nearest neighbor interpolation method.

[0314] Of course, the above are only several examples of the processing methods of the upsampling layer, and there is no limitation thereto.

[0315] In a possible implementation manner, for the processing layer (also called a processing unit or a processing link) of the mean hyperparameter decoding network, it includes but is not limited to at least one of the following: one or more ordinary convolutional layers, one or more transposed convolutional layers, one or more deformable convolutional layers, one or more depthwise separable convolutional layers, one or more grouped convolutional layers, one or more dilated convolutional layers, one or more transformer layers, one or more activation layers using the Relu activation function, one or more activation layers using the target activation function with upper and lower clamping, and one or more activation layers using the LeakyRelu activation function.

[0316] Exemplarily, for each convolutional layer in the processing layer (such as an ordinary convolutional layer, a transposed convolutional layer, a deformable convolutional layer, a depthwise separable convolutional layer, a grouped convolutional layer, a dilated convolutional layer, a transformer layer, etc.), the form of quantized convolution can also be adopted, that is, the parameters in the convolutional layer are quantized, and the input features are quantized, and then the convolution operation is performed. When quantizing the parameters in the convolutional layer, the bit width can be 2, 4, 8, 16, etc., and there is no limitation thereto. For example, for each convolutional layer in the processing layer, the convolutional layer is a quantized convolutional layer (or an integerized convolutional layer, and the integerized bit width can be 2, 4, 8, 16, etc.), that is, the parameters in the convolutional layer are integerized parameters. For example, the parameters in the ordinary convolutional layer are quantized parameters, the parameters in the transposed convolutional layer are quantized parameters, the parameters in the grouped convolutional layer are quantized parameters, and so on.

[0317] For each convolutional layer in the processing layer, these convolutional layers can be repeatedly stacked without any restrictions on the stacking form. For example, first stack two ordinary convolutional layers, then stack two grouped convolutional layers, then deploy a dilated convolutional layer, and then deploy an ordinary convolutional layer, etc. Of course, this is just an example and there are no restrictions on the structure of this processing layer.

[0318] For each layer in the processing layer, there are arbitrary skip connections (i.e., skip connection structures, which can also be called residual connections) between the layers of the processing layer. The skip connection structure is used to connect the input features of a certain network layer (such as any network layer) with the output features of another network layer (such as any network layer), and operations such as feature addition, feature multiplication, and feature concatenation can be performed.

[0319] If the processing layer includes a convolutional layer and an activation layer, the convolutional layer and the activation layer can be interleaved, and the convolutional layer or the activation layer can be repeatedly stacked. Any input-output node (i.e., any network layer) from front to back can be added through a residual connection (i.e., there are arbitrary skip connections between the layers of the processing layer and can be added through the skip connection). For example, the input features of the first network layer are added to the output features of the fifth network layer, the input features of the second network layer are added to the output features of the fifth network layer, the input features of the third network layer are added to the output features of the sixth network layer, and so on.

[0320] In a possible implementation, the structure of the mean hyperparameter decoding network (i.e., the optimized mean hyperparameter decoding network) can be seen in Figure 9D as shown. Of course, Figure 9D this is just an example of the mean hyperparameter decoding network and there are no restrictions on the structure of this mean hyperparameter decoding network. Subsequently, the mean hyperparameter decoding network of Figure 9D is taken as an example. The mean hyperparameter decoding network can sequentially include a processing layer, an upsampling layer, and a Crop layer, and the processing layer can sequentially include a convolutional layer (Conv, such as a convolutional layer with a size of 3*3, an input channel number of c, and an output channel number of c), a Relu activation layer (using the Relu activation function), a convolutional layer (Conv, such as a convolutional layer with a size of 3*3, an input channel number of c, and an output channel number of c), a Relu activation layer (using the Relu activation function), a grouped convolutional layer (gConv, such as a convolutional layer with a size of 3*3, an input channel number of c, and an output channel number of 4c, and a grouping number of 4), and a convolutional layer (such as a convolutional layer with a size of 1*1, an input channel number of 4c, and an output channel number of c).

[0321] For Figure 9DThe Relu activation layer in it can also be replaced by a Relu6 activation layer (using the Relu6 activation function), a ClipRelu6 activation layer (using the ClipRelu6 activation function), or a chs-wise ClipRelu6 activation layer (using the chs-wise ClipRelu6 activation function). For example, if any Relu activation layer is replaced by a Relu6 activation layer, or two Relu activation layers are replaced by a Relu6 activation layer.

[0322] Example 15: In Examples 1 - 14, it may involve ordinary convolutional layers, transposed convolutional layers, deformable convolutional layers, depthwise separable convolutional layers, grouped convolutional layers, dilated convolutional layers, transformer layers, etc.

[0323] Regarding the ordinary convolutional layer, the ordinary convolutional layer can also be called a conventional convolutional layer - Convolution (Conv). The ordinary convolutional layer starts with a small weight matrix, that is, a convolutional kernel (kernel), and gradually "scans" it on the two-dimensional input data. While the convolutional kernel "slides", it calculates the product of the weight matrix and the data matrix obtained by scanning, and then summarizes the results into an output pixel. See Figure 10A As shown in the figure, it is a schematic diagram of the ordinary convolutional layer. Figure 10A The convolutional kernel matrix, input data, and output data are shown in the figure, and the dotted white part is the padding value of the input data. Regarding the ordinary convolutional layer, in the form of a sliding window, the convolutional kernel operates from left to right and from top to bottom, and the stride parameter determines the step size of each slide.

[0324] Regarding the transposed convolutional layer - Transposed Convolution (Tconv), for some tasks, it is necessary to upsample the data. Conventional convolutional operations can only keep the resolution of the output feature map equal to or approximate to that of the input feature map and cannot achieve the purpose of upsampling. Therefore, the transposed convolution method can be used. Different from the conventional convolution, the stride parameter does not represent the step size of sliding, but represents the number of 0s filled between the input data. For example, when stride = 1, the transposed convolution is the same as the conventional convolution. When stride = 2, see Figure 10B As shown in the figure, it is a schematic diagram of the transposed convolutional layer. It can insert into the input feature map, and (stride - 1) 0s will be inserted between each feature point in the spatial domain. After expanding the input spatial domain resolution, it will be operated with the weights of the convolution.

[0325] Regarding the dilated convolutional layer (Dconv), the dilated convolutional layer can also be called the atrous convolutional layer or the dilated convolutional layer, which means inserting holes between the points of the normal convolutional kernel. It is relative to the normal discrete convolution. For a normal convolution with a stride of 2 and a padding of 1, see Figure 10CAs shown in the figure, it is a schematic diagram of the dilated convolutional layer, which inserts dilation into the convolution.

[0326] For the grouped convolutional layer - Group-Convolution (gconv), given the input x with a size of [C, H, W], the grouped convolution operation splits x by channels into the number of groups. Each group is separately convolved to obtain the input of the output group, and then the outputs of each group are concatenated by channels. See Figure 10D As shown in the figure, it is a schematic diagram of the grouped convolutional layer.

[0327] For the depthwise separable convolutional layer, the depthwise separable convolutional layer can be divided into two stages. The first stage is spatial separability, and the second stage is depthwise separability. Spatial separability is expressed as: the input [C, H, W] is divided into C groups, and the data of each group is [1, H, W]. Each group is separately convolved and then concatenated to obtain the output. See Figure 10E As shown in the figure, it is a schematic diagram of the spatial separability of the depthwise separable convolutional layer. The second stage is the depthwise separable stage, which is used to operate on the output of the first stage with a convolutional kernel of size 1x1. See Figure 10E As shown in the figure, it is a schematic diagram of the depthwise separability of the depthwise separable convolutional layer.

[0328] For the transformer layer, the transformer layer is a Transformer-based attention block. The structure of the transformer layer can be seen in Figure 7F Or Figure 7G As shown in the figure, it will not be elaborated here.

[0329] For each convolutional layer in the processing layer (such as the ordinary convolutional layer, transposed convolutional layer, deformable convolutional layer, depthwise separable convolutional layer, grouped convolutional layer, dilated convolutional layer, transformer layer, etc.), it is also possible to adopt the form of quantized convolution (quantized convolution can also be called integerized convolution), that is, to quantize the parameters within the convolutional layer. This process can also be called quantized convolution. Quantized convolution is not a specific form of convolution. All the above operation forms can be quantized, which means that the weights inside the convolution are weights quantized to a certain bit width. The weight data of the original convolutional kernel is represented in floating point by float. The weight data is scaled, offset, and integerized to ensure that the data can be represented by 4 bits, 8 bits, or 16 bits for the above floating point types.

[0330] Exemplarily, each of the above embodiments can be implemented alone or in combination. For example, each of Embodiment 1 to Embodiment 15 can be implemented alone, and at least two of Embodiment 1 to Embodiment 15 can be implemented in combination.

[0331] Exemplarily, in each of the above embodiments, the content at the encoding end can also be applied to the decoding end, that is, the decoding end can process in the same way, and the content at the decoding end can also be applied to the encoding end, that is, the encoding end can process in the same way.

[0332] Based on the same inventive concept as the above method, an embodiment of the present application further provides a decoding device, which is applied to the decoding end, and the device includes: a memory configured to store video data; a decoder configured to implement the decoding methods in Embodiment 1 to Embodiment 15 of the above, that is, the processing flow at the decoding end.

[0333] Based on the same inventive concept as the above method, an embodiment of the present application further provides an encoding device, which is applied to the encoding end, and the device includes: a memory configured to store video data; an encoder configured to implement the encoding methods in Embodiment 1 to Embodiment 15 of the above, that is, the processing flow at the encoding end.

[0334] Based on the same inventive concept as the above method, the decoding end device (which can also be called a video decoder) provided in the embodiment of the present application, in terms of the hardware level, the schematic diagram of its hardware architecture can be specifically referred to Figure 11A as shown. It includes: a processor 111 and a machine-readable storage medium 112, and the machine-readable storage medium 112 stores machine-executable instructions that can be executed by the processor 111; the processor 111 is used to execute the machine-executable instructions to implement the decoding methods in Embodiment 1 to 15 of the present application above.

[0335] Based on the same inventive concept as the above method, the encoding end device (which can also be called a video encoder) provided in the embodiment of the present application, in terms of the hardware level, the schematic diagram of its hardware architecture can be specifically referred to Figure 11B as shown. It includes: a processor 113 and a machine-readable storage medium 114, and the machine-readable storage medium 114 stores machine-executable instructions that can be executed by the processor 113; the processor 113 is used to execute the machine-executable instructions to implement the encoding methods in Embodiment 1 to 15 of the present application above.

[0336] Based on the same inventive concept as the above method, an embodiment of the present application provides an electronic device. It includes: a processor and a machine-readable storage medium, and the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the decoding method or the encoding method in Embodiment 1 to 15 of the present application above.

[0337] Based on the same application concept as the above method, an embodiment of the present application further provides a machine-readable storage medium, on which a number of computer instructions are stored. When the computer instructions are executed by a processor, the methods disclosed in the above examples of the present application can be implemented, such as the decoding method or the encoding method in the above embodiments.

[0338] Based on the same application concept as the above method, an embodiment of the present application further provides a computer program, which when executed by a processor, can implement the decoding method or the encoding method disclosed in the above examples of the present application.

[0339] Based on the same application concept as the above method, an embodiment of the present application further proposes a decoding device, which can be applied to a decoding end (also called a video decoder). The decoding device includes: a decoding module, configured to decode a first bitstream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block; determine probability distribution parameters based on the coefficient hyperparameter feature, and decode a second bitstream corresponding to the current image block based on the probability distribution parameters to obtain a residual feature corresponding to the current image block; a determination module, configured to determine a target mean feature corresponding to the current image block based on the coefficient hyperparameter feature; determine a reconstructed feature corresponding to the current image block based on the target mean feature and the residual feature; an acquisition module, configured to input the reconstructed feature into a synthesis transformation network to obtain a reconstructed image block corresponding to the current image block; wherein the synthesis transformation network includes a non-linear residual network layer, and the non-linear residual network layer at least includes an activation layer and a skip connection structure; wherein, the output feature of the activation layer is not greater than an upper threshold, and / or, the output feature of the activation layer is not less than a lower threshold.

[0340] Exemplarily, the activation layer uses a target activation function with upper and lower clamping to non-linearly adjust the feature. The target activation function corresponds to the upper threshold and the lower threshold, and the output feature of the activation layer is not greater than the upper threshold, and the output feature of the activation layer is not less than the lower threshold; or, the activation layer uses a Relu activation function to non-linearly adjust the feature. The Relu activation function corresponds to the lower threshold, and the output feature of the activation layer is not less than the lower threshold.

[0341] Exemplarily, the non-linear residual network layer sequentially includes an activation layer, a processing layer and a skip connection structure. The processing layer includes at least one of the following: one or more ordinary convolutional layers, one or more transposed convolutional layers, one or more deformable convolutional layers, one or more depthwise separable convolutional layers, one or more grouped convolutional layers, one or more dilated convolutional layers, one or more global processing unit layers.

[0342] Exemplarily, when determining the probability distribution parameters based on the coefficient hyperparameter features, the decoding module is specifically configured to: input the coefficient hyperparameter features into a probability hyperparameter decoding network to obtain the probability distribution parameters; wherein, the probability hyperparameter decoding network includes an activation layer, and the activation layer uses a target activation function with upper and lower clamping to nonlinearly adjust the features; wherein, the target activation function corresponds to an upper limit threshold and a lower limit threshold, and the output features of the activation layer are not greater than the upper limit threshold and not less than the lower limit threshold.

[0343] Exemplarily, when determining the target mean features corresponding to the current image block based on the coefficient hyperparameter features, the determining module is specifically configured to: input the coefficient hyperparameter features into a mean hyperparameter decoding network to obtain the initial mean features corresponding to the current image block, and determine the target mean features based on the initial mean features; wherein, the mean hyperparameter decoding network includes an activation layer, and the activation layer uses a target activation function with upper and lower clamping to nonlinearly adjust the features; wherein, the target activation function corresponds to an upper limit threshold and a lower limit threshold, and the output features of the activation layer are not greater than the upper limit threshold and not less than the lower limit threshold.

[0344] Exemplarily, when determining the target mean features based on the initial mean features, the determining module is specifically configured to: use the initial mean features as the target mean features; or, input the initial mean features and the obtained reconstruction features of the current image block into a context model to obtain the target mean features; wherein, the context model includes an activation layer, and the activation layer uses a target activation function with upper and lower clamping to nonlinearly adjust the features; wherein, the target activation function corresponds to an upper limit threshold and a lower limit threshold, and the output features of the activation layer are not greater than the upper limit threshold and not less than the lower limit threshold.

[0345] Exemplarily, the determining module determines probability distribution parameters based on the coefficient hyperparameter features, and when determining the target mean features corresponding to the current image block based on the coefficient hyperparameter features, specifically: inputs the coefficient hyperparameter features into a hyperparameter decoding network to obtain reference features corresponding to the current image block; inputs the reference features and the obtained reconstruction features of the current image block into a context model to obtain the probability distribution parameters and the target mean features corresponding to the current image block; wherein, the probability hyperparameter decoding network includes an activation layer, and the activation layer performs non-linear adjustment on the features using a target activation function with upper and lower clamping; and / or, the context model includes an activation layer, and the activation layer performs non-linear adjustment on the features using a target activation function with upper and lower clamping; wherein, the target activation function corresponds to an upper threshold and a lower threshold, and the output features of the activation layer are not greater than the upper threshold and not less than the lower threshold.

[0346] Exemplarily, the upper threshold corresponding to the target activation function is a configured fixed upper threshold, and the lower threshold corresponding to the target activation function is a configured fixed lower threshold; or, the upper threshold corresponding to the target activation function is an adaptively learned upper threshold, and the lower threshold corresponding to the target activation function is an adaptively learned lower threshold.

[0347] Exemplarily, when adaptively learning the upper threshold, the same upper threshold is learned for all channels of the activation layer, or, the upper threshold is learned separately for each channel of the activation layer, and the upper thresholds corresponding to different channels are the same or different;

[0348] When adaptively learning the lower threshold, the same lower threshold is learned for all channels of the activation layer, or, the lower threshold is learned separately for each channel of the activation layer, and the lower thresholds corresponding to different channels are the same or different.

[0349] Based on the same application concept as the above method, an embodiment of this application further provides a decoding device, which can be applied to a decoding end (also referred to as a video decoder). The decoding device includes: a decoding module, configured to decode a first bitstream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block; input the coefficient hyperparameter feature into a probability hyperparameter decoding network to obtain probability distribution parameters, and decode a second bitstream corresponding to the current image block based on the probability distribution parameters to obtain a residual feature corresponding to the current image block; a determination module, configured to input the coefficient hyperparameter feature into a mean hyperparameter decoding network to obtain an initial mean feature corresponding to the current image block, and determine a target mean feature corresponding to the current image block based on the initial mean feature; determine a reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature; an acquisition module, configured to determine a reconstructed image block corresponding to the current image block based on the reconstruction feature; wherein, the probability hyperparameter decoding network includes an activation layer, and the output feature of the activation layer is not greater than an upper limit threshold, and / or, the output feature of the activation layer is not less than a lower limit threshold; and / or, the mean hyperparameter decoding network includes an activation layer, and the output feature of the activation layer is not greater than an upper limit threshold, and / or, the output feature of the activation layer is not less than a lower limit threshold.

[0350] Exemplarily, the activation layer uses a target activation function with upper and lower clamping to non-linearly adjust the feature. The target activation function corresponds to an upper limit threshold and a lower limit threshold. The output feature of the activation layer is not greater than the upper limit threshold, and the output feature of the activation layer is not less than the lower limit threshold; wherein, the upper limit threshold corresponding to the target activation function is a configured fixed upper limit threshold, and the lower limit threshold corresponding to the target activation function is a configured fixed lower limit threshold; or, the upper limit threshold corresponding to the target activation function is an adaptively learned upper limit threshold, and the lower limit threshold corresponding to the target activation function is an adaptively learned lower limit threshold; wherein, when adaptively learning the upper limit threshold, the same upper limit threshold is learned for all channels of the activation layer, or, the upper limit threshold is learned separately for each channel of the activation layer, and the upper limit thresholds corresponding to different channels are the same or different; when adaptively learning the lower limit threshold, the same lower limit threshold is learned for all channels of the activation layer, or, the lower limit threshold is learned separately for each channel of the activation layer, and the lower limit thresholds corresponding to different channels are the same or different.

[0351] Based on the same application concept as the above method, an embodiment of the present application further provides a decoding device, which can be applied to a decoding end (also referred to as a video decoder). The decoding device includes: a decoding module for decoding a first bitstream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block; a determination module for inputting the coefficient hyperparameter feature into a hyperparameter decoding network to obtain a reference feature corresponding to the current image block; inputting the reference feature and the already obtained reconstruction feature of the current image block into a context model to obtain a probability distribution parameter and a target mean feature corresponding to the current image block; the decoding module is further configured to decode a second bitstream corresponding to the current image block based on the probability distribution parameter to obtain a residual feature corresponding to the current image block; the determination module is further configured to determine the reconstruction feature based on the target mean feature and the residual feature; an acquisition module for determining a reconstructed image block corresponding to the current image block based on the reconstruction feature; wherein, the hyperparameter decoding network includes an activation layer, and the output feature of the activation layer is not greater than an upper threshold value, and / or, the output feature of the activation layer is not less than a lower threshold value; and / or, the context model includes an activation layer, and the output feature of the activation layer is not greater than an upper threshold value, and / or, the output feature of the activation layer is not less than a lower threshold value.

[0352] Exemplarily, the activation layer uses a target activation function with upper and lower clamping to non-linearly adjust the feature. The target activation function corresponds to an upper threshold value and a lower threshold value. The output feature of the activation layer is not greater than the upper threshold value, and the output feature of the activation layer is not less than the lower threshold value; wherein, the upper threshold value corresponding to the target activation function is a configured fixed upper threshold value, and the lower threshold value corresponding to the target activation function is a configured fixed lower threshold value; or, the upper threshold value corresponding to the target activation function is an adaptively learned upper threshold value, and the lower threshold value corresponding to the target activation function is an adaptively learned lower threshold value; wherein, when adaptively learning the upper threshold value, the same upper threshold value is learned for all channels of the activation layer, or, an upper threshold value is separately learned for each channel of the activation layer, and the upper threshold values corresponding to different channels are the same or different; when adaptively learning the lower threshold value, the same lower threshold value is learned for all channels of the activation layer, or, a lower threshold value is separately learned for each channel of the activation layer, and the lower threshold values corresponding to different channels are the same or different.

[0353] Based on the same application concept as the above method, an embodiment of the present application also proposes a decoding device, which can be applied to a decoding end (also known as a video decoder). The decoding device includes: a decoding module, configured to decode a first bitstream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block; determine probability distribution parameters based on the coefficient hyperparameter feature, and decode a second bitstream corresponding to the current image block based on the probability distribution parameters to obtain a residual feature corresponding to the current image block; a determination module, configured to input the coefficient hyperparameter feature into a mean hyperparameter decoding network to obtain an initial mean feature corresponding to the current image block, and determine a target mean feature corresponding to the current image block based on the initial mean feature; wherein, the mean hyperparameter decoding network sequentially includes a processing layer, an upsampling layer for performing an upsampling operation, and a cropping layer; wherein, the feature output by the upsampling layer is obtained as the initial mean feature after passing through the cropping layer; determine a reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature; an acquisition module, configured to determine a reconstructed image block corresponding to the current image block based on the reconstruction feature.

[0354] Exemplarily, the processing layer does not include an upsampling layer for performing an upsampling operation, and the processing layer includes at least one of the following: one or more ordinary convolutional layers, one or more transposed convolutional layers, one or more deformable convolutional layers, one or more depthwise separable convolutional layers, one or more grouped convolutional layers, one or more dilated convolutional layers, one or more global processing unit layers, one or more activation layers using a Relu activation function, one or more activation layers using a target activation function with upper and lower clamping, one or more activation layers using a LeakyRelu activation function; wherein, there are arbitrary skip connections between the layers of the processing layer.

[0355] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. The present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. The embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code. The above are only the embodiments of the present application and are not used to limit the present application.

[0356] For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. An image decoding method, characterized in that, The method includes: Obtaining the coefficient hyperparameter feature corresponding to the current image block; Inputting the coefficient hyperparameter feature into the mean hyperparameter decoding network to obtain the initial mean feature corresponding to the current image block, and determining the target mean feature corresponding to the current image block based on the initial mean feature; Determining the reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature corresponding to the current image block; Determining the reconstructed image block corresponding to the current image block based on the reconstruction feature; Wherein, the activation layer included in the mean hyperparameter decoding network uses a target activation function with upper and lower clamping to perform non-linear adjustment on the feature, the target activation function corresponds to a fixed upper limit threshold and a fixed lower limit threshold, the output feature of the activation layer is not greater than the fixed upper limit threshold, and the output feature of the activation layer is not less than the fixed lower limit threshold.

2. The method according to claim 1, wherein The obtaining the coefficient hyperparameter feature corresponding to the current image block includes: decoding the first bitstream corresponding to the current image block to obtain the coefficient hyperparameter feature corresponding to the current image block; Before determining the reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature corresponding to the current image block, the method further includes: Determining probability distribution parameters based on the coefficient hyperparameter feature, and decoding the second bitstream corresponding to the current image block based on the probability distribution parameters to obtain the residual feature corresponding to the current image block.

3. The method according to claim 1, wherein The fixed upper limit threshold is 6, and the fixed lower limit threshold is 0.

4. The method according to any one of claims 1-3, characterized in that The mean hyperparameter decoding network includes at least one convolutional layer, at least one activation layer, a transposed convolutional layer, and a Crop layer.

5. The method according to claim 4, wherein The mean hyperparameter decoding network sequentially includes: a convolutional layer, a transposed convolutional layer, a Crop layer, an activation layer, a convolutional layer, an activation layer, and a convolutional layer.

6. The method according to any one of claims 1-5, characterized in that The target activation function is the Relu6 activation function.

7. An image decoding device, characterized in that, The apparatus includes: A decoding module, configured to obtain the coefficient hyperparameter feature corresponding to the current image block; A determination module, configured to input the coefficient hyperparameter feature into the mean hyperparameter decoding network to obtain the initial mean feature corresponding to the current image block, and determine the target mean feature corresponding to the current image block based on the initial mean feature; determine the reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature corresponding to the current image block; An obtaining module, configured to determine the reconstructed image block corresponding to the current image block based on the reconstruction feature; Wherein, the activation layer included in the mean hyperparameter decoding network uses a target activation function with upper and lower clamping to perform non-linear adjustment on the feature, the target activation function corresponds to a fixed upper limit threshold and a fixed lower limit threshold, the output feature of the activation layer is not greater than the fixed upper limit threshold, and the output feature of the activation layer is not less than the fixed lower limit threshold.

8. A decoding end device, characterized in that, The decoding end device includes: a processor and a machine-readable storage medium, and the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; The processor is configured to execute machine-executable instructions to implement the method according to any one of claims 1-6.

9. A machine-readable storage medium, characterized in that, A number of computer instructions are stored on the machine-readable storage medium, and when the computer instructions are executed by the processor, the method according to any one of claims 1-6 is implemented.

10. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1-6 is implemented.