Image decoding method and apparatus, and device
By using the nonlinear residual network layer and mean hyperparameter decoding network optimization in the image decoding method, the problems of neural network decoding performance and complexity are solved, and efficient image decoding effect is achieved.
Patent Information
- Application Number
- PCT/CN2025/073103
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-18
- Filing Date
- 2025-01-17
- Publication Date
- 2025-07-24
AI Technical Summary
The existing neural network-based image decoding methods have problems with poor decoding performance and high complexity.
The nonlinear residual network layer is adopted, including the activation layer and the jump connection structure, and the features are nonlinearly adjusted through the target activation function of the upper and lower clamps, optimized the complexity of the activation function, and combined with the optimization of the mean hyperparameter decoding network, the overall complexity of the neural network is reduced.
While maintaining low complexity, the performance of image decoding is improved, the phenomenon of decoding is reduced, and the quality of reconstructed images is improved.
Smart Images

Figure CN2025073103_24072025_PF_FP_ABST
Abstract
Description
Image decoding method, device and equipment Technical Field
[0001] The present application relates to the field of coding and decoding technology, and in particular to an image decoding method, apparatus and device thereof. Background Art
[0002] To save space, video images are encoded before transmission. Complete video encoding involves prediction, transformation, quantization, entropy coding, and filtering. Prediction can include intra-frame prediction and inter-frame prediction. Inter-frame prediction leverages temporal correlations in the video to predict the current pixel using pixels from adjacent coded images, effectively removing temporal redundancy. Intra-frame prediction leverages spatial correlations in the video to predict the current pixel using pixels from coded blocks in the current frame, effectively removing spatial redundancy.
[0003] With the rapid development of deep learning, it has achieved success in many high-level computer vision problems, such as image classification and object detection. Deep learning is also gradually being applied to codecs, using neural networks to encode and decode images. Although neural network-based codecs demonstrate significant performance potential, they still suffer from poor decoding performance and high complexity. Summary of the Invention
[0004] In view of this, the present application provides an image decoding method, apparatus and device thereof, which can improve decoding performance and reduce complexity.
[0005] The present application provides an image decoding method, which is executed by a decoding end, and the method includes: decoding a first code stream corresponding to a current image block to obtain coefficient hyperparameter features corresponding to the current image block; determining probability distribution parameters based on the coefficient hyperparameter features, and determining target mean features corresponding to the current image block based on the coefficient hyperparameter features; decoding a second code stream corresponding to the current image block based on the probability distribution parameters to obtain residual features corresponding to the current image block; determining reconstruction features corresponding to the current image block based on the target mean features and the residual features; inputting the reconstruction features into a synthetic transformation network to obtain a reconstructed image block corresponding to the current image block; wherein the synthetic transformation network includes a nonlinear residual network layer, the nonlinear residual network layer includes at least an activation layer and a skip connection structure; the output feature of the activation layer is not greater than an upper threshold, and / or the output feature of the activation layer is not less than a lower threshold.
[0006] In one embodiment, the activation layer uses an upper and lower clamped target activation function to perform nonlinear adjustment on the features, the target activation function corresponds to an upper threshold and a lower threshold, and the output feature of the activation layer is not greater than the upper threshold and not less than the lower threshold; or, the activation layer uses a Relu activation function to perform nonlinear adjustment on the features, the Relu activation function corresponds to the lower threshold, and the output feature of the activation layer is not less than the lower threshold.
[0007] In one embodiment, the nonlinear residual network layer includes an activation layer, a processing layer and a skip connection structure in sequence, and the processing layer includes at least one of the following: one or more ordinary convolution layers, one or more transposed convolution layers, one or more deformable convolution layers, one or more depth separation convolution layers, one or more grouped convolution layers, one or more dilated convolution layers, and one or more global processing unit layers.
[0008] In one embodiment, determining the target mean feature corresponding to the current image block based on the coefficient hyperparameter feature includes: inputting the coefficient hyperparameter feature into a mean hyperparameter decoding network to obtain the initial mean feature corresponding to the current image block, and determining the target mean feature based on the initial mean feature; wherein, the mean hyperparameter decoding network includes an activation layer, and the activation layer uses an upper and lower clamped target activation function to perform nonlinear adjustment on the feature; wherein, the target activation function corresponds to an upper threshold and a lower threshold, and the output feature of the activation layer is not greater than the upper threshold and not less than the lower threshold.
[0009] In one embodiment, determining the target mean feature based on the initial mean feature includes: using the initial mean feature as the target mean feature; or inputting the initial mean feature and the obtained reconstructed feature of the current image block into a context model to obtain the target mean feature; wherein, the context model includes an activation layer, and the activation layer uses a lower-clamped target activation function to perform nonlinear adjustment on the feature; wherein, the target activation function corresponds to a lower threshold, and the output feature of the activation layer is not less than the lower threshold.
[0010] In one embodiment, the method of determining the probability distribution parameters and the target mean features corresponding to the current image block based on the coefficient hyperparameter features includes: inputting the coefficient hyperparameter features into a hyperparameter decoding network to obtain the reference features corresponding to the current image block; inputting the reference features and the obtained reconstructed features of the current image block into a context model to obtain the probability distribution parameters and the target mean features corresponding to the current image block; wherein the hyperparameter decoding network includes an activation layer, and the activation layer uses an upper and lower clamped target activation function to perform nonlinear adjustment on the features, wherein the target activation function corresponds to an upper threshold and a lower threshold, and the output features of the activation layer are not greater than the upper threshold and not less than the lower threshold; and / or the context model includes an activation layer, and the activation layer uses an upper and lower clamped target activation function to perform nonlinear adjustment on the features, wherein the target activation function corresponds to an upper threshold and a lower threshold, and the output features of the activation layer are not greater than the upper threshold and not less than the lower threshold.
[0011] In one embodiment, the upper threshold corresponding to the target activation function is a configured fixed upper threshold, and the lower threshold corresponding to the target activation function is a configured fixed lower threshold; or, the upper threshold corresponding to the target activation function is the upper threshold of adaptive learning, and the lower threshold corresponding to the target activation function is the lower threshold of adaptive learning.
[0012] In one embodiment, when adaptively learning the upper threshold, the same upper threshold is learned for all channels of the activation layer, or the upper threshold is learned separately for each channel of the activation layer, and the upper thresholds corresponding to different channels are the same or different; when adaptively learning the lower threshold, the same lower threshold is learned for all channels of the activation layer, or the lower threshold is learned separately for each channel of the activation layer, and the lower thresholds corresponding to different channels are the same or different.
[0013] The present application provides an image decoding method, which is performed by a decoding end, and the method includes: decoding a first bitstream corresponding to a current image block to obtain coefficient hyperparameter features corresponding to the current image block; inputting the coefficient hyperparameter features into a probabilistic hyperparameter decoding network to obtain probability distribution parameters, and decoding a second bitstream corresponding to the current image block based on the probability distribution parameters to obtain residual features corresponding to the current image block; inputting the coefficient hyperparameter features into a mean hyperparameter decoding network to obtain initial mean features corresponding to the current image block, and determining a target mean feature corresponding to the current image block based on the initial mean features; determining a reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature; and determining a reconstructed image block corresponding to the current image block based on the reconstruction feature; wherein the probabilistic hyperparameter decoding network includes an activation layer, the output feature of the activation layer is not greater than an upper threshold, and / or the output feature of the activation layer is not less than a lower threshold; and / or the mean hyperparameter decoding network includes an activation layer, the output feature of the activation layer is not greater than an upper threshold, and / or the output feature of the activation layer is not less than a lower threshold.
[0014] In one embodiment, the activation layer of the probabilistic hyperparameter decoding network and / or the activation layer of the mean hyperparameter decoding network uses an upper and lower clamped target activation function to perform nonlinear adjustment on the features, the target activation function corresponds to an upper threshold and a lower threshold, and the output feature of the activation layer is not greater than the upper threshold and not less than the lower threshold; wherein the upper threshold corresponding to the target activation function is a configured fixed upper threshold, and the lower threshold corresponding to the target activation function is a configured fixed lower threshold; or, the upper threshold corresponding to the target activation function is an upper threshold for adaptive learning, and the lower threshold corresponding to the target activation function is a lower threshold for adaptive learning; wherein, when adaptively learning the upper threshold, the same upper threshold is learned for all channels of the activation layer, or, the upper threshold is learned separately for each channel of the activation layer, and the upper thresholds corresponding to different channels are the same or different; when adaptively learning the lower threshold, the same lower threshold is learned for all channels of the activation layer, or, the lower threshold is learned separately for each channel of the activation layer, and the lower thresholds corresponding to different channels are the same or different.
[0015] The present application provides an image decoding method, which is performed by a decoding end, and the method includes: decoding a first bitstream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block; inputting the coefficient hyperparameter feature into a hyperparameter decoding network to obtain a reference feature corresponding to the current image block; inputting the reference feature and the obtained reconstruction feature of the current image block into a context model to obtain probability distribution parameters and a target mean feature corresponding to the current image block; decoding a second bitstream corresponding to the current image block based on the probability distribution parameters to obtain a residual feature corresponding to the current image block, determining the reconstruction feature based on the target mean feature and the residual feature; and determining a reconstructed image block corresponding to the current image block based on the reconstruction feature; wherein the hyperparameter decoding network includes an activation layer, the output feature of the activation layer is not greater than an upper threshold, and / or the output feature of the activation layer is not less than a lower threshold; and / or the context model includes an activation layer, the output feature of the activation layer is not greater than an upper threshold, and / or the output feature of the activation layer is not less than a lower threshold.
[0016] The present application provides an image decoding method, which is executed by a decoding end, and the method includes: decoding a first code stream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block; determining a probability distribution parameter based on the coefficient hyperparameter feature, and decoding a second code stream corresponding to the current image block based on the probability distribution parameter to obtain a residual feature corresponding to the current image block; inputting the coefficient hyperparameter feature into a mean hyperparameter decoding network to obtain an initial mean feature corresponding to the current image block, and determining a target mean feature corresponding to the current image block based on the initial mean feature; wherein the mean hyperparameter decoding network includes a first processing layer, an upsampling layer for performing an upsampling operation, a cropping layer, and a second processing layer in sequence; wherein the feature output by the second processing layer is used as the initial mean feature; determining a reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature; and determining a reconstructed image block corresponding to the current image block based on the reconstruction feature.
[0017] In one embodiment, the first processing layer does not include an upsampling layer for performing an upsampling operation, and the first processing layer includes at least one of the following: one or more ordinary convolution layers, one or more transposed convolution layers, one or more deformable convolution layers, one or more depth-separated convolution layers, one or more grouped convolution layers, one or more expanded convolution layers, one or more global processing unit layers, one or more activation layers using a Relu activation function, one or more activation layers using a target activation function with upper and lower clamping, and one or more activation layers using a LeakyRelu activation function; wherein, there are arbitrary jump connections between the layers of the first processing layer.
[0018] In one embodiment, the second processing layer includes at least one of the following: one or more ordinary convolution layers, one or more deformable convolution layers, one or more depth separation convolution layers, one or more grouped convolution layers, one or more dilated convolution layers, one or more global processing unit layers, one or more activation layers using Relu activation function, one or more activation layers using upper and lower clamped target activation function, and one or more activation layers using LeakyRelu activation function; wherein, there are arbitrary jump connections between the layers of the second processing layer.
[0019] The present application provides an image decoding method, which includes: obtaining coefficient hyperparameter features corresponding to a current image block; determining a target mean feature corresponding to the current image block based on the coefficient hyperparameter features; determining a reconstruction feature corresponding to the current image block based on the target mean feature and a residual feature corresponding to the current image block; inputting the reconstruction feature into a synthetic transformation network to obtain a reconstructed image block corresponding to the current image block; wherein the synthetic transformation network includes an activation layer that uses an upper and lower clamped target activation function to perform nonlinear adjustment on the features, the target activation function corresponds to a fixed upper threshold and a fixed lower threshold, the output feature of the activation layer is not greater than the fixed upper threshold, and the output feature of the activation layer is not less than the fixed lower threshold.
[0020] In one embodiment, the synthetic transformation network includes a nonlinear residual network layer, and the nonlinear residual network layer includes at least an activation layer and a skip connection structure.
[0021] In one embodiment, the nonlinear residual network layer includes an activation layer, a processing layer and a skip connection structure in sequence; wherein, when there is one activation layer and the processing layer includes multiple network layers, all network layers are located behind the activation layer.
[0022] In one embodiment, the fixed upper threshold is 6, and the fixed lower threshold is 0.
[0023] In one embodiment, the synthetic transformation network includes a residual ResAU layer, which serves as the nonlinear residual network layer; wherein the ResAU layer includes an activation layer, a grouped convolution layer and a convolution layer in sequence; wherein the input features of the ResAU layer are connected to the output features of the convolution layer; wherein the input features of the ResAU layer pass through the activation layer, the grouped convolution layer and the convolution layer in sequence to obtain post-convolution features; and feature multiplication and feature addition operations are performed on the input features of the ResAU layer and the post-convolution features in sequence to obtain the output features of the ResAU layer.
[0024] In one embodiment, if the synthesis transformation network includes a low-complexity synthesis transformation network, the low-complexity synthesis transformation network includes in sequence: a lightweight residual block LRB layer, a convolution layer, an upsampling Pixshuffle layer, a ResAU layer, a cropping Crop layer, a convolution layer, a Pixshuffle layer, a Crop layer, a ResAU layer, a convolution layer, a ResAU layer, a convolution layer, a Pixshuffle layer, and a Crop layer; if the synthesis transformation network includes a medium-complexity synthesis transformation network, the medium-complexity synthesis transformation network includes in sequence: an LRB layer, a transposed convolution layer, a ResAU layer, a Crop layer Layer, transposed convolution layer, Crop layer, ResAU layer, convolution layer, ResAU layer, convolution layer, Pixshuffle layer, Crop layer; if the synthetic transformation network includes a high-complexity synthetic transformation network, the high-complexity synthetic transformation network includes, in sequence: residual block RB layer, transposed convolution layer, ResAU layer, Crop layer, transposed convolution layer, convolution basic attention block CAB layer, Crop layer, ResAU layer, convolution layer, Pixshuffle layer, Transformer-based attention block TAM layer, Crop layer, ResAU layer, transposed convolution layer, Crop layer.
[0025] The present application provides an image decoding method, which includes: obtaining coefficient hyperparameter features corresponding to a current image block; inputting the coefficient hyperparameter features into a mean hyperparameter decoding network to obtain an initial mean feature corresponding to the current image block, and determining a target mean feature corresponding to the current image block based on the initial mean feature; determining a reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature corresponding to the current image block; and determining a reconstructed image block corresponding to the current image block based on the reconstruction feature; wherein the activation layer included in the mean hyperparameter decoding network uses an upper and lower clamped target activation function to perform nonlinear adjustment on the features, the target activation function corresponds to a fixed upper threshold and a fixed lower threshold, the output feature of the activation layer is not greater than the fixed upper threshold, and the output feature of the activation layer is not less than the fixed lower threshold.
[0026] In one embodiment, obtaining the coefficient hyperparameter features corresponding to the current image block includes: decoding the first code stream corresponding to the current image block to obtain the coefficient hyperparameter features corresponding to the current image block; before determining the reconstruction features corresponding to the current image block based on the target mean features and the residual features corresponding to the current image block, the method also includes: determining probability distribution parameters based on the coefficient hyperparameter features, and decoding the second code stream corresponding to the current image block based on the probability distribution parameters to obtain the residual features corresponding to the current image block.
[0027] In one embodiment, the fixed upper threshold is 6, and the fixed lower threshold is 0.
[0028] In one embodiment, the mean hyperparameter decoding network includes at least one convolutional layer, at least one activation layer, a transposed convolutional layer, and a cropping layer.
[0029] In one embodiment, the mean super-parameter decoding network includes, in sequence: a convolutional layer, a transposed convolutional layer, a crop layer, an activation layer, a convolutional layer, an activation layer, and a convolutional layer.
[0030] In one embodiment, the target activation function is a Relu6 activation function.
[0031] The present application provides an image decoding method, which includes: obtaining coefficient hyperparameter features corresponding to a current image block; inputting the coefficient hyperparameter features into a mean hyperparameter decoding network to obtain an initial mean feature corresponding to the current image block, and determining a target mean feature corresponding to the current image block based on the initial mean feature; determining a reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature corresponding to the current image block; inputting the reconstruction feature into a synthetic transformation network to obtain a reconstructed image block corresponding to the current image block; wherein the activation layer included in the mean hyperparameter decoding network and the synthetic transformation network uses an upper and lower clamped target activation function to perform nonlinear adjustment on the features, the target activation function corresponds to a fixed upper threshold and a fixed lower threshold, the output feature of the activation layer is not greater than the fixed upper threshold, and the output feature of the activation layer is not less than the fixed lower threshold.
[0032] In one embodiment, the synthetic transformation network includes a nonlinear residual network layer, which in turn includes an activation layer, a processing layer and a skip connection structure; wherein, when there is an activation layer and the processing layer includes multiple network layers, all network layers are located behind the activation layer.
[0033] In one embodiment, the synthetic transformation network includes a residual ResAU layer, which serves as the nonlinear residual network layer; wherein the ResAU layer includes an activation layer, a grouped convolution layer and a convolution layer in sequence; wherein the input features of the ResAU layer are connected to the output features of the convolution layer; wherein the input features of the ResAU layer pass through the activation layer, the grouped convolution layer and the convolution layer in sequence to obtain post-convolution features; and feature multiplication and feature addition operations are performed on the input features of the ResAU layer and the post-convolution features in sequence to obtain the output features of the ResAU layer.
[0034] In one embodiment, if the synthesis transformation network includes a low-complexity synthesis transformation network, the low-complexity synthesis transformation network includes in sequence: a lightweight residual block LRB layer, a convolution layer, an upsampling Pixshuffle layer, a ResAU layer, a cropping Crop layer, a convolution layer, a Pixshuffle layer, a Crop layer, a ResAU layer, a convolution layer, a ResAU layer, a convolution layer, a Pixshuffle layer, and a Crop layer; if the synthesis transformation network includes a medium-complexity synthesis transformation network, the medium-complexity synthesis transformation network includes in sequence: an LRB layer, a transposed convolution layer, a ResAU layer, a Crop layer Layer, transposed convolution layer, Crop layer, ResAU layer, convolution layer, ResAU layer, convolution layer, Pixshuffle layer, Crop layer; if the synthetic transformation network includes a high-complexity synthetic transformation network, the high-complexity synthetic transformation network includes, in sequence: residual block RB layer, transposed convolution layer, ResAU layer, Crop layer, transposed convolution layer, convolution basic attention block CAB layer, Crop layer, ResAU layer, convolution layer, Pixshuffle layer, Transformer-based attention block TAM layer, Crop layer, ResAU layer, transposed convolution layer, Crop layer.
[0035] In one embodiment, the mean hyperparameter decoding network includes at least one convolutional layer, at least one activation layer, a transposed convolutional layer, and a cropping layer.
[0036] In one embodiment, the mean super-parameter decoding network includes, in sequence: a convolutional layer, a transposed convolutional layer, a crop layer, an activation layer, a convolutional layer, an activation layer, and a convolutional layer.
[0037] The present application provides an image decoding device, which is applied to a decoding end, and the device includes: a decoding module, which is used to decode a first code stream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block; determine a probability distribution parameter based on the coefficient hyperparameter feature, and decode a second code stream corresponding to the current image block based on the probability distribution parameter to obtain a residual feature corresponding to the current image block; a determination module, which is used to determine a target mean feature corresponding to the current image block based on the coefficient hyperparameter feature; determine a reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature; and an acquisition module, which is used to input the reconstruction feature into a synthetic transformation network to obtain a reconstructed image block corresponding to the current image block; wherein the synthetic transformation network includes a nonlinear residual network layer, and the nonlinear residual network layer includes at least an activation layer and a skip connection structure; wherein the output feature of the activation layer is not greater than an upper threshold, and / or the output feature of the activation layer is not less than a lower threshold.
[0038] The present application provides an image decoding device, which is applied to a decoding end. The device includes: a decoding module, which is used to decode a first code stream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block; input the coefficient hyperparameter feature into a probability hyperparameter decoding network to obtain a probability distribution parameter, and decode the second code stream corresponding to the current image block based on the probability distribution parameter to obtain a residual feature corresponding to the current image block; a determination module, which is used to input the coefficient hyperparameter feature into a mean hyperparameter decoding network to obtain an initial mean feature corresponding to the current image block, and based on the initial mean feature Determine the target mean feature corresponding to the current image block; determine the reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature; an acquisition module is used to determine the reconstructed image block corresponding to the current image block based on the reconstruction feature; wherein the probabilistic hyperparameter decoding network includes an activation layer, the output feature of the activation layer is not greater than an upper threshold, and / or, the output feature of the activation layer is not less than a lower threshold; and / or, the mean hyperparameter decoding network includes an activation layer, the output feature of the activation layer is not greater than an upper threshold, and / or, the output feature of the activation layer is not less than a lower threshold.
[0039] The present application provides an image decoding device, applied to a decoding end, the device comprising: a decoding module, configured to decode a first bitstream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block; a determination module, configured to input the coefficient hyperparameter feature into a hyperparameter decoding network to obtain a reference feature corresponding to the current image block; input the reference feature and the obtained reconstruction feature of the current image block into a context model to obtain probability distribution parameters and a target mean feature corresponding to the current image block; the decoding module is further configured to decode a second bitstream corresponding to the current image block based on the probability distribution parameters to obtain a residual feature corresponding to the current image block; the determination module is further configured to determine the reconstruction feature based on the target mean feature and the residual feature; and an acquisition module is configured to determine a reconstructed image block corresponding to the current image block based on the reconstruction feature; wherein the hyperparameter decoding network includes an activation layer, the output feature of the activation layer is not greater than an upper threshold value, and / or the output feature of the activation layer is not less than a lower threshold value; and / or the context model includes an activation layer, the output feature of the activation layer is not greater than an upper threshold value, and / or the output feature of the activation layer is not less than a lower threshold value.
[0040] The present application provides an image decoding device, which is applied to a decoding end, and the device includes: a decoding module, which is used to decode a first code stream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block; determine a probability distribution parameter based on the coefficient hyperparameter feature, and decode a second code stream corresponding to the current image block based on the probability distribution parameter to obtain a residual feature corresponding to the current image block; a determination module, which is used to input the coefficient hyperparameter feature into a mean hyperparameter decoding network to obtain an initial mean feature corresponding to the current image block, and determine a target mean feature corresponding to the current image block based on the initial mean feature; wherein the mean hyperparameter decoding network includes a first processing layer, an upsampling layer for performing an upsampling operation, a cropping layer, and a second processing layer in sequence; wherein the feature output by the second processing layer is used as the initial mean feature; determine the reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature; and an acquisition module, which is used to determine the reconstructed image block corresponding to the current image block based on the reconstruction feature.
[0041] The present application provides an image decoding device, which includes: a decoding module for obtaining coefficient hyperparameter features corresponding to a current image block; a determination module for determining a target mean feature corresponding to the current image block based on the coefficient hyperparameter features; a reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature corresponding to the current image block; an acquisition module for inputting the reconstruction feature into a synthetic transformation network to obtain a reconstructed image block corresponding to the current image block; wherein the synthetic transformation network includes an activation layer that uses an upper and lower clamped target activation function to perform nonlinear adjustment on the feature, the target activation function corresponds to a fixed upper threshold and a fixed lower threshold, the output feature of the activation layer is not greater than the fixed upper threshold, and the output feature of the activation layer is not less than the fixed lower threshold.
[0042] The present application provides an image decoding device, which includes: a decoding module for obtaining coefficient hyperparameter features corresponding to a current image block; a determination module for inputting the coefficient hyperparameter features into a mean hyperparameter decoding network to obtain an initial mean feature corresponding to the current image block, and determining a target mean feature corresponding to the current image block based on the initial mean feature; determining a reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature corresponding to the current image block; an acquisition module for determining a reconstructed image block corresponding to the current image block based on the reconstruction feature; wherein the activation layer included in the mean hyperparameter decoding network uses an upper and lower clamped target activation function to perform nonlinear adjustment on the features, the target activation function corresponds to a fixed upper threshold and a fixed lower threshold, the output feature of the activation layer is not greater than the fixed upper threshold, and the output feature of the activation layer is not less than the fixed lower threshold.
[0043] The present application provides an image decoding device, which includes: a decoding module for obtaining coefficient hyperparameter features corresponding to a current image block; a determination module for inputting the coefficient hyperparameter features into a mean hyperparameter decoding network to obtain an initial mean feature corresponding to the current image block, and determining a target mean feature corresponding to the current image block based on the initial mean feature; determining a reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature corresponding to the current image block; an acquisition module for inputting the reconstruction feature into a synthetic transformation network to obtain a reconstructed image block corresponding to the current image block; wherein the activation layer included in the mean hyperparameter decoding network and the synthetic transformation network uses an upper and lower clamped target activation function to perform nonlinear adjustment on the features, the target activation function corresponds to a fixed upper threshold and a fixed lower threshold, the output feature of the activation layer is not greater than the fixed upper threshold, and the output feature of the activation layer is not less than the fixed lower threshold.
[0044] The present application provides a decoding end device, which includes a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the above-mentioned decoding method.
[0045] The present application provides an electronic device, which includes a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the above-mentioned decoding method.
[0046] The present application provides a machine-readable storage medium, on which a number of computer instructions are stored. When the computer instructions are executed by a processor, the above-mentioned decoding method can be implemented.
[0047] The present application provides a computer program, which implements the above-mentioned decoding method when executed by a processor.
[0048] As can be seen from the above technical solutions, in the embodiment of the present application, an end-to-end video image compression method is proposed, which can realize the decoding of video images based on a neural network. By optimizing the activation function, the target activation function with upper and lower clamping is used to perform nonlinear adjustment on the features, thereby reducing the complexity of the activation function, and then reducing the complexity of the neural network. The target activation function can improve the robustness of the neural network and reduce the phenomenon of garbled code when the neural network decodes high-frequency images. While maintaining low complexity, the neural network effectively guarantees the quality of the reconstructed image block, improves the quality of the reconstructed image, achieves the purpose of improving decoding performance, and reduces complexity. By optimizing the mean hyperparameter decoding network, the complexity of the neural network is greatly reduced, and the performance of the neural network is kept unchanged, thereby achieving the purpose of improving decoding performance and reducing complexity. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] FIG1 is a schematic diagram of a three-dimensional feature matrix in one embodiment of the present application;
[0050] FIG2 is a schematic flow chart of a decoding method in one embodiment of the present application;
[0051] FIG3 is a schematic flow chart of an encoding method in one embodiment of the present application;
[0052] 4A-4D are schematic diagrams of a coding and decoding framework in one embodiment of the present application;
[0053] FIG5A is a schematic diagram of a probabilistic super-parameter decoding network in one embodiment of the present application;
[0054] FIG5B is a schematic diagram of a ReLU activation function in one embodiment of the present application;
[0055] FIG5C is a schematic diagram of a probabilistic super-parameter decoding network in one embodiment of the present application;
[0056] FIG5D is a schematic diagram of a Relu6 activation function in one embodiment of the present application;
[0057] FIG5E is a schematic diagram of an activation function in one embodiment of the present application;
[0058] 6A and 6B are schematic diagrams of a mean super-parameter decoding network in one embodiment of the present application;
[0059] FIG6C and FIG6D are schematic diagrams of a context model in one embodiment of the present application;
[0060] 7A-7I are schematic diagrams of a synthetic transformation network in one embodiment of the present application;
[0061] FIG8A is a schematic diagram of a ResAU layer in one embodiment of the present application;
[0062] FIG8B is a schematic diagram of a LeakyRelu activation function in one embodiment of the present application;
[0063] 8C-8H are schematic diagrams of a ResAU layer in one embodiment of the present application;
[0064] 9A-9C are schematic diagrams of upsampling operations in one embodiment of the present application;
[0065] FIG9D is a schematic diagram of a mean super-parameter decoding network in one embodiment of the present application;
[0066] 10A-10F are schematic diagrams of convolution operations in one embodiment of the present application;
[0067] FIG11A is a hardware structure diagram of a decoding end device in one embodiment of the present application;
[0068] FIG11B is a hardware structure diagram of an encoding terminal device in one embodiment of the present application. DETAILED DESCRIPTION
[0069] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, rather than for limiting the present application. The singular forms of "a", "said" and "the" used in the embodiments of the present application and the claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to any or all possible combinations of one or more associated listed items. It should be understood that although the terms first, second, third, etc. may be used to describe various information in the embodiments of the present application, these information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information, depending on the context. In addition, the word "if" used can be interpreted as "at the time of...", or "when...", or "in response to determination".
[0070] The embodiments of the present application provide a decoding method, apparatus, and device thereof, which may involve the following concepts:
[0071] Entropy encoding: Entropy encoding is a method of encoding that does not lose any information during the encoding process according to the entropy principle. Information entropy is the average amount of information in the source (a measure of uncertainty). Entropy coding methods include, but are not limited to, Shannon coding, Huffman coding, and arithmetic coding.
[0072] Neural Network (NN): A neural network refers to an artificial neural network. A NN is a computational model composed of a large number of interconnected nodes (or neurons). In a NN, neurons (often called processing units) can represent different objects, such as features, letters, concepts, or some meaningful abstract patterns. Processing units in a NN can be divided into three types: input units, output units, and hidden units. Input units receive signals and data from the external world; output units output the processed results; and hidden units are located between the input and output units and cannot be observed from outside the system. The connection weights between neurons reflect the strength of the connections between units, and the representation and processing of information is reflected in the connections between processing units. A NN is a non-programmed, brain-like information processing method. Its essence is to achieve parallel and distributed information processing capabilities through the transformations and dynamics of the neural network, mimicking the information processing capabilities of the human brain to varying degrees and levels. In the field of video processing, commonly used NNs include, but are not limited to, convolutional neural networks (CNNs), recurrent neural networks (RNNs), and fully connected networks.
[0073] Convolutional Neural Network (CNN): A convolutional neural network (CNN) is a feedforward neural network and one of the most representative network structures in deep learning technology. Its artificial neurons can respond to surrounding units within a certain coverage area, making it an excellent choice for processing large images. The basic structure of a CNN consists of two layers: the feature extraction layer (also known as the convolution layer). Each neuron's input is connected to the local receptive field of the previous layer, extracting local features. Once a local feature is extracted, its positional relationship with other features is determined. The second layer is the feature mapping layer (also known as the activation layer). Each computational layer of the neural network consists of multiple feature maps. Each feature map is a plane where all neurons have equal weights. Feature mapping structures can use sigmoid functions, Rectified Linear Unit (ReLU) functions, Leaky-ReLU functions, Parametric Rectified Linear Unit (PReLU) functions, and Generalized Divisive Normalization (GDN) functions as activation functions. Furthermore, because neurons within a mapping plane share weights, the number of free parameters in the network is reduced.
[0074] For example, one of the advantages of convolutional neural networks over image processing algorithms is that they avoid complex pre-processing of images (such as extracting artificial features) and can directly input the original image for end-to-end learning. One of the advantages of convolutional neural networks over ordinary neural networks is that ordinary neural networks use a fully connected approach, that is, all neurons from the input layer to the hidden layer are connected. This will result in a huge number of parameters, making network training time-consuming or even difficult to train. Convolutional neural networks avoid this difficulty through methods such as local connections and weight sharing.
[0075] Deconvolution: Also known as a transposed convolution layer, the deconvolution layer works very similarly to the convolution layer. The main difference is that the deconvolution layer uses padding to make the output larger than the input (or the same). If the stride is 1, the output size is equal to the input size; if the stride is N, the width of the output feature is N times the width of the input feature, and the height of the output feature is N times the height of the input feature.
[0076] Generalization Ability: refers to the ability of a machine learning algorithm to adapt to new samples. The purpose of learning is to learn the patterns hidden behind the data, so that the trained network can also give appropriate output for data outside the learning set with the same patterns. This ability can be called generalization ability.
[0077] Feature: The features referred to in this application are three-dimensional feature matrices or tensors of C*W*H. Figure 1 shows a schematic diagram of a three-dimensional feature matrix. In this matrix, C represents the number of channels, H represents the feature height, and W represents the feature width. A three-dimensional feature matrix can be the input or output of a neural network.
[0078] Rate-Distortion Optimized (RDO): Coding efficiency is evaluated using two key metrics: bit rate and PSNR (Peak Signal to Noise Ratio). A smaller bitrate indicates a higher compression ratio, while a higher PSNR indicates better reconstructed image quality. When selecting a mode, the discriminant formula is essentially a comprehensive evaluation of these two metrics. For example, the cost of a mode is: J(mode) = D + λ*R, where D represents distortion. This is typically measured using the SSE (sum-square error) metric, which is the mean squared sum of the differences between the reconstructed image block and the source image. For cost considerations, the SAD (sum of absolute difference) metric can also be used. λ is the Lagrange multiplier, and R is the actual number of bits required to encode the image block in that mode, including the bits required for coding mode information, motion information, and residual information. When selecting a mode, if the rate-distortion principle is used to make comparison decisions on the coding mode, the best coding performance can usually be guaranteed.
[0079] Numerous coding tools have been proposed for various modules on the encoding side, and each tool often has multiple modes. Different coding tools often produce the best coding performance for different video sequences. Therefore, during the encoding process, Rate-Distortion Optimization (RDO) is often used to compare the coding performance of different tools or modes to select the optimal mode. After determining the optimal tool or mode, the decision-making information for the tool or mode is conveyed by encoding marker information in the bitstream. While this approach introduces higher coding complexity, it can adaptively select the optimal mode combination for different content, achieving optimal coding performance. The decoding side can obtain relevant mode information by directly parsing the marker information, with minimal complexity impact.
[0080] The general end-to-end image coding and decoding framework mainly includes the main feature information component and the super-prior side information component. The main feature information component includes an analysis network, quantization, normal entropy coding, normal entropy decoding, and a synthesis network, while the super-prior side information component includes a super-prior analysis network, quantization, factored entropy coding, factored entropy decoding, and a super-prior synthesis network. The image components are compressed, encoded, and reconstructed and restored by the analysis network and synthesis network of the main feature information component, respectively. The super-prior side information component is mainly used to model the probability of the main feature information and guide the entropy coding and decoding of the main feature information. The general end-to-end image coding and decoding framework has problems such as the high complexity of the mean hyperparameter decoding network, which first upsampling and then convolution processing, and the high hardware complexity of the LeakyRelu function of the nonlinear residual network layer of the synthesis transform network.
[0081] In response to the above findings, in an embodiment of the present application, a feature clipping and network simplification technology for video image decoding is proposed. For the nonlinear residual network layer of the synthetic transformation network, an upper and lower clamped activation function (used to replace the LeakyRelu function) can be used to perform nonlinear adjustment on the features, thereby solving the problem of high hardware complexity of the LeakyRelu function. For the mean hyperparameter decoding network, convolution processing can be performed first and then upsampling to solve the problem of high complexity.
[0082] The following describes in detail the decoding method and encoding method in the embodiments of the present application in conjunction with specific embodiments.
[0083] In an embodiment of the present application, a decoding method is proposed. FIG2 is a flowchart of the decoding method. The method can be applied to a decoding end (also called a video decoder). The method may include:
[0084] Step 201: Decode a first code stream corresponding to a current image block to obtain coefficient hyperparameter features corresponding to the current image block.
[0085] Step 202: Determine a probability distribution parameter based on the coefficient hyperparameter feature, and determine a target mean feature corresponding to the current image block based on the coefficient hyperparameter feature.
[0086] Step 203: Decode the second code stream corresponding to the current image block based on the probability distribution parameters to obtain residual features corresponding to the current image block.
[0087] Step 204: Determine a reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature.
[0088] Step 205: Input the reconstructed features into a synthetic transformation network to obtain a reconstructed image block corresponding to the current image block. Exemplarily, the synthetic transformation network includes a nonlinear residual network layer, which may include at least an activation layer and a skip connection structure. The output features of the activation layer are no greater than an upper threshold, and / or no less than a lower threshold.
[0089] Exemplarily, the nonlinear residual network layer can be a residual activation unit (ResAU) network layer, or other types of network layers, without limitation, as long as it can implement nonlinear residual processing. The following description will be made using the ResAU network layer as an example.
[0090] Exemplarily, the output features of the activation layer of the nonlinear residual network layer may be no greater than an upper threshold, and the output features of the activation layer may be no less than a lower threshold, that is, the upper and lower thresholds of the activation layer are simultaneously limited. In this case, the activation layer may use an upper and lower clamped target activation function to perform nonlinear adjustment on the features, and the target activation function corresponds to an upper and lower threshold, and the output features of the activation layer are no greater than the upper threshold, and the output features of the activation layer are no less than the lower threshold.
[0091] Alternatively, the output features of the activation layer of the nonlinear residual network layer can be no less than the lower threshold, that is, only the lower threshold of the activation layer is limited, and the upper threshold of the activation layer is not limited. In this case, the activation layer can use a Relu activation function to perform nonlinear adjustment on the features, and the Relu activation function corresponds to the lower threshold, and the output features of the activation layer are no less than the lower threshold.
[0092] Alternatively, the output features of the activation layer of the nonlinear residual network layer may be no greater than the upper threshold, that is, only the upper threshold of the activation layer is limited, and the lower threshold of the activation layer is not limited. In this case, the activation layer may use an activation function, such as that shown in FIG5E , to perform nonlinear adjustment on the features, and the activation function corresponds to the upper threshold, and the output features of the activation layer are no greater than the upper threshold.
[0093] Exemplarily, the nonlinear residual network layer may further include a processing layer. For example, the nonlinear residual network layer may include an activation layer, a processing layer, and a skip connection structure in sequence. The processing layer may include, but is not limited to, at least one of the following: one or more ordinary convolutional layers, one or more transposed convolutional layers, one or more deformable convolutional layers, one or more depth-separated convolutional layers, one or more grouped convolutional layers, one or more dilated convolutional layers, and one or more global processing unit layers. The global processing unit layer is used to implement global processing functions. It may be a transformer layer or other types of network layers, as long as it can implement global processing functions. The transformer layer will be used as an example for explanation.
[0094] For example, when there is one activation layer and the processing layer includes a network layer, the network layer can be located before or after the activation layer. When there is one activation layer and the processing layer includes multiple network layers, all network layers can be located before or after the activation layer, or some network layers can be located before or after the activation layer, while the remaining network layers can be located after the activation layer.
[0095] For another example, when there are multiple active layers (taking the first and second active layers as an example), and the processing layer includes a network layer, the network layer may be located before the first active layer, between the first and second active layers, or behind the second active layer. If the processing layer includes multiple network layers, all network layers may be located before the first active layer, between the first and second active layers, or behind the second active layer. Alternatively, some network layers may be located before the first active layer, while the remaining network layers may be located between the first and second active layers. Alternatively, some network layers may be located before the first active layer, while the remaining network layers may be located behind the second active layer. Alternatively, some network layers may be located before the first active layer, while the remaining network layers may be located between the first and second active layers, while the remaining network layers may be located behind the second active layer. Alternatively, some network layers may be located before the first active layer, while the remaining network layers may be located between the first and second active layers, while the remaining network layers may be located behind the second active layer. Alternatively, some network layers may be located between the first and second active layers, while the remaining network layers may be located behind the second active layer.
[0096] Exemplarily, determining the probability distribution parameters based on the coefficient hyperparameter feature may include, but is not limited to: inputting the coefficient hyperparameter feature into a probability hyperparameter decoding network to obtain the probability distribution parameters. Exemplarily, the probability hyperparameter decoding network may include at least an activation layer, and the activation layer uses an upper and lower clamped target activation function (such as a Relu6 activation function) to perform nonlinear adjustment on the features. The target activation function may correspond to an upper threshold and a lower threshold, and the output feature of the activation layer is not greater than the upper threshold, and the output feature of the activation layer is not less than the lower threshold.
[0097] Exemplarily, determining the target mean feature corresponding to the current image block based on the coefficient hyperparameter feature may include but is not limited to: inputting the coefficient hyperparameter feature into the mean hyperparameter decoding network to obtain the initial mean feature corresponding to the current image block, and determining the target mean feature based on the initial mean feature. Exemplarily, the mean hyperparameter decoding network includes an activation layer, and the activation layer uses an upper and lower clamped target activation function to perform nonlinear adjustment on the feature. Wherein, the target activation function corresponds to an upper threshold and a lower threshold, the output feature of the activation layer is not greater than the upper threshold, and the output feature of the activation layer is not less than the lower threshold.
[0098] Exemplarily, determining the target mean feature based on the initial mean feature may include, but is not limited to: using the initial mean feature as the target mean feature; or inputting the initial mean feature and the obtained reconstructed feature of the current image block into a context model to obtain the target mean feature. Exemplarily, the context model includes an activation layer, and the activation layer uses an upper and lower clamped target activation function to perform nonlinear adjustment on the feature; wherein the target activation function corresponds to an upper threshold and a lower threshold, and the output feature of the activation layer is not greater than the upper threshold, and the output feature of the activation layer is not less than the lower threshold.
[0099] Exemplarily, determining the target mean feature based on the initial mean feature may include, but is not limited to: using the initial mean feature as the target mean feature; or inputting the initial mean feature and the obtained reconstructed feature of the current image block into a context model to obtain the target mean feature. Exemplarily, the context model includes an activation layer, and the activation layer uses a lower-clamped target activation function (such as a ReLU activation function) to perform nonlinear adjustment on the feature; wherein the target activation function corresponds to a lower threshold, and the output feature of the activation layer is not less than the lower threshold.
[0100] Exemplarily, determining the probability distribution parameters based on the coefficient hyperparameter feature, and determining the target mean feature corresponding to the current image block based on the coefficient hyperparameter feature may include but is not limited to: inputting the coefficient hyperparameter feature into a hyperparameter decoding network to obtain a reference feature corresponding to the current image block; inputting the reference feature and the obtained reconstructed feature of the current image block into a context model to obtain the probability distribution parameters and the target mean feature corresponding to the current image block. Exemplarily, the probabilistic hyperparameter decoding network includes an activation layer, and the activation layer uses a target activation function with upper and lower clamping to perform nonlinear adjustment on the feature, wherein the target activation function corresponds to an upper threshold and a lower threshold, and the output feature of the activation layer is not greater than the upper threshold and not less than the lower threshold; and / or, the context model includes an activation layer, and the activation layer uses a target activation function with upper and lower clamping to perform nonlinear adjustment on the feature, wherein the target activation function corresponds to an upper threshold and a lower threshold, and the output feature of the activation layer is not greater than the upper threshold and not less than the lower threshold.
[0101] Exemplarily, the upper threshold corresponding to the target activation function is a configured fixed upper threshold, and the lower threshold corresponding to the target activation function is a configured fixed lower threshold; or, the upper threshold corresponding to the target activation function is the upper threshold of adaptive learning, and the lower threshold corresponding to the target activation function is the lower threshold of adaptive learning.
[0102] Exemplarily, when adaptively learning the upper limit threshold, the same upper limit threshold is learned for all channels of the activation layer, or the upper limit threshold is learned separately for each channel of the activation layer, and the upper limit thresholds corresponding to different channels are the same or different.
[0103] Exemplarily, when adaptively learning the lower threshold, the same lower threshold is learned for all channels of the activation layer, or the lower threshold is learned separately for each channel of the activation layer, and the lower thresholds corresponding to different channels are the same or different.
[0104] In an embodiment of the present application, a decoding method is proposed, which can be applied to a decoding end. The method may include:
[0105] Step S11: Decode the first code stream corresponding to the current image block to obtain coefficient hyperparameter features corresponding to the current image block.
[0106] Step S12: Input the coefficient hyperparameter feature into the probability hyperparameter decoding network to obtain probability distribution parameters, and decode the second code stream corresponding to the current image block based on the probability distribution parameters to obtain the residual feature corresponding to the current image block.
[0107] Step S13: Input the coefficient hyperparameter feature into the mean hyperparameter decoding network to obtain the initial mean feature corresponding to the current image block, and determine the target mean feature corresponding to the current image block based on the initial mean feature.
[0108] Step S14: determining the reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature.
[0109] Step S15: Determine a reconstructed image block corresponding to the current image block based on the reconstruction feature.
[0110] In one possible embodiment, the probabilistic super-parameter decoding network may include an activation layer, and the output feature of the activation layer is not greater than an upper threshold, and / or the output feature of the activation layer is not less than a lower threshold. For example, the output feature of the activation layer may be not greater than the upper threshold, and the output feature of the activation layer is not less than the lower threshold, that is, the upper and lower thresholds of the activation layer are simultaneously limited. In this case, the activation layer may use an upper and lower clamped target activation function to perform nonlinear adjustment on the feature, and the target activation function corresponds to an upper threshold and a lower threshold, the output feature of the activation layer is not greater than the upper threshold, and the output feature of the activation layer is not less than the lower threshold. Alternatively, the output feature of the activation layer may be not less than the lower threshold, that is, only the lower threshold of the activation layer is limited, and the upper threshold of the activation layer is not limited. In this case, the activation layer may use a Relu activation function to perform nonlinear adjustment on the feature, and the Relu activation function corresponds to the lower threshold, and the output feature of the activation layer is not less than the lower threshold. Alternatively, the output feature of the activation layer may be not greater than the upper threshold, that is, only the upper threshold of the activation layer is limited, and the lower threshold of the activation layer is not limited. In this case, the activation layer may use an activation function such as that shown in FIG5E to perform nonlinear adjustment on the features, and the activation function corresponds to an upper threshold, and the output features of the activation layer are not greater than the upper threshold.
[0111] In one possible implementation, the mean hyperparameter decoding network may include an activation layer, and the output feature of the activation layer is not greater than an upper threshold, and / or the output feature of the activation layer is not less than a lower threshold. For example, the output feature of the activation layer may be not greater than the upper threshold, and the output feature of the activation layer is not less than the lower threshold, that is, the upper and lower thresholds of the activation layer are simultaneously limited. In this case, the activation layer may use an upper and lower clamped target activation function to perform nonlinear adjustment on the feature, and the target activation function corresponds to an upper threshold and a lower threshold, the output feature of the activation layer is not greater than the upper threshold, and the output feature of the activation layer is not less than the lower threshold. Alternatively, the output feature of the activation layer may be not less than the lower threshold, that is, only the lower threshold of the activation layer is limited, and the upper threshold of the activation layer is not limited. In this case, the activation layer may use a Relu activation function to perform nonlinear adjustment on the feature, and the Relu activation function corresponds to the lower threshold, and the output feature of the activation layer is not less than the lower threshold. Alternatively, the output feature of the activation layer may be not greater than the upper threshold, that is, only the upper threshold of the activation layer is limited, and the lower threshold of the activation layer is not limited. In this case, the activation layer may use an activation function such as that shown in FIG5E to perform nonlinear adjustment on the features, and the activation function corresponds to an upper threshold, and the output features of the activation layer are not greater than the upper threshold.
[0112] Exemplarily, the upper threshold corresponding to the target activation function is a configured fixed upper threshold, and the lower threshold corresponding to the target activation function is a configured fixed lower threshold; or, the upper threshold corresponding to the target activation function is the upper threshold of adaptive learning, and the lower threshold corresponding to the target activation function is the lower threshold of adaptive learning.
[0113] Exemplarily, when adaptively learning the upper limit threshold, the same upper limit threshold is learned for all channels of the activation layer, or the upper limit threshold is learned separately for each channel of the activation layer, and the upper limit thresholds corresponding to different channels are the same or different.
[0114] Exemplarily, when adaptively learning the lower threshold, the same lower threshold is learned for all channels of the activation layer, or the lower threshold is learned separately for each channel of the activation layer, and the lower thresholds corresponding to different channels are the same or different.
[0115] In an embodiment of the present application, a decoding method is proposed, which can be applied to a decoding end. The method may include:
[0116] Step S21: Decode the first code stream corresponding to the current image block to obtain coefficient hyperparameter features corresponding to the current image block.
[0117] Step S22: Input the coefficient hyperparameter feature into the hyperparameter decoding network to obtain the reference feature corresponding to the current image block.
[0118] Step S23: Input the reference feature and the obtained reconstructed feature of the current image block into the context model to obtain the probability distribution parameter and the target mean feature corresponding to the current image block.
[0119] Step S24: decoding the second code stream corresponding to the current image block based on the probability distribution parameters to obtain a residual feature corresponding to the current image block, and determining a reconstruction feature based on the target mean feature and the residual feature.
[0120] Step S25: Determine a reconstructed image block corresponding to the current image block based on the reconstruction feature.
[0121] In one possible implementation, the super-parameter decoding network may include an activation layer, and the output feature of the activation layer is not greater than an upper threshold, and / or the output feature of the activation layer is not less than a lower threshold. For example, the output feature of the activation layer may be not greater than the upper threshold, and the output feature of the activation layer is not less than the lower threshold, that is, the upper and lower thresholds of the activation layer are simultaneously limited. In this case, the activation layer may use an upper and lower clamped target activation function to perform nonlinear adjustment on the feature, and the target activation function corresponds to an upper threshold and a lower threshold, the output feature of the activation layer is not greater than the upper threshold, and the output feature of the activation layer is not less than the lower threshold. Alternatively, the output feature of the activation layer may be not less than the lower threshold, that is, only the lower threshold of the activation layer is limited, and the upper threshold of the activation layer is not limited. In this case, the activation layer may use a Relu activation function to perform nonlinear adjustment on the feature, and the Relu activation function corresponds to the lower threshold, and the output feature of the activation layer is not less than the lower threshold. Alternatively, the output feature of the activation layer may be not greater than the upper threshold, that is, only the upper threshold of the activation layer is limited, and the lower threshold of the activation layer is not limited. In this case, the activation layer can use a certain activation function to perform nonlinear adjustment on the features, and the activation function corresponds to an upper threshold, and the output features of the activation layer are not greater than the upper threshold.
[0122] In one possible implementation, the context model may include an activation layer, and the output feature of the activation layer is not greater than an upper threshold, and / or the output feature of the activation layer is not less than a lower threshold. For example, the output feature of the activation layer may be not greater than the upper threshold, and the output feature of the activation layer is not less than the lower threshold, that is, the upper and lower thresholds of the activation layer are simultaneously limited. In this case, the activation layer may use an upper and lower clamped target activation function to perform nonlinear adjustment on the feature, and the target activation function corresponds to an upper threshold and a lower threshold, the output feature of the activation layer is not greater than the upper threshold, and the output feature of the activation layer is not less than the lower threshold. Alternatively, the output feature of the activation layer may be not less than the lower threshold, that is, only the lower threshold of the activation layer is limited, and the upper threshold of the activation layer is not limited. In this case, the activation layer may use a Relu activation function to perform nonlinear adjustment on the feature, and the Relu activation function corresponds to the lower threshold, and the output feature of the activation layer is not less than the lower threshold. Alternatively, the output feature of the activation layer may be not greater than the upper threshold, that is, only the upper threshold of the activation layer is limited, and the lower threshold of the activation layer is not limited. In this case, the activation layer can use a certain activation function to perform nonlinear adjustment on the features, and the activation function corresponds to an upper threshold, and the output features of the activation layer are not greater than the upper threshold.
[0123] Exemplarily, the upper threshold corresponding to the target activation function is a configured fixed upper threshold, and the lower threshold corresponding to the target activation function is a configured fixed lower threshold; or, the upper threshold corresponding to the target activation function is the upper threshold of adaptive learning, and the lower threshold corresponding to the target activation function is the lower threshold of adaptive learning.
[0124] Exemplarily, when adaptively learning the upper limit threshold, the same upper limit threshold is learned for all channels of the activation layer, or the upper limit threshold is learned separately for each channel of the activation layer, and the upper limit thresholds corresponding to different channels are the same or different.
[0125] Exemplarily, when adaptively learning the lower threshold, the same lower threshold is learned for all channels of the activation layer, or the lower threshold is learned separately for each channel of the activation layer, and the lower thresholds corresponding to different channels are the same or different.
[0126] In an embodiment of the present application, a decoding method is proposed, which can be applied to a decoding end. The method may include:
[0127] Step S31: Decode the first code stream corresponding to the current image block to obtain coefficient hyperparameter features corresponding to the current image block.
[0128] Step S32: Determine a probability distribution parameter based on the coefficient hyperparameter feature, and decode the second code stream corresponding to the current image block based on the probability distribution parameter to obtain a residual feature corresponding to the current image block.
[0129] Step S33: Input the coefficient hyperparameter feature into the mean hyperparameter decoding network to obtain the initial mean feature corresponding to the current image block, and determine the target mean feature corresponding to the current image block based on the initial mean feature. Exemplarily, the mean hyperparameter decoding network may include a first processing layer, an upsampling layer for performing upsampling operations, a cropping layer, and a second processing layer in sequence. The feature output by the second processing layer is used as the initial mean feature, that is, the second processing layer is located at the end of the mean hyperparameter decoding network, and the feature output by the upsampling layer passes through the cropping layer and the second processing layer to obtain the output feature of the mean hyperparameter decoding network.
[0130] Exemplarily, the cropping layer is used to crop and align the pixel positions after upsampling. The cropping layer can be a Crop layer or other network layer as long as it can realize the cropping and alignment function. The following description will be made using the Crop layer as an example.
[0131] Step S34: Determine the reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature.
[0132] Step S35: Determine a reconstructed image block corresponding to the current image block based on the reconstruction feature.
[0133] Exemplarily, the first processing layer of the mean superparameter decoding network does not include an upsampling layer for performing an upsampling operation, that is, the upsampling layer for performing an upsampling operation is deployed after the first processing layer, so that the first processing layer does not perform an upsampling operation. The upsampling layer can be located after the first processing layer, and the features output by the upsampling layer are passed through the Crop layer and the second processing layer to obtain the output features of the mean superparameter decoding network. Alternatively, an additional convolution layer (such as a 1*1 convolution layer or a 3*3 convolution layer, which only performs a simple convolution operation) can be deployed between the upsampling layer and the Crop layer. After the features output by the upsampling layer pass through a convolution layer, a Crop layer, and a second processing layer, the output features (initial mean features) of the mean superparameter decoding network are obtained. For the convenience of description, in the subsequent embodiments, the mean superparameter decoding network can sequentially include a first processing layer, an upsampling layer for performing an upsampling operation, a Crop layer, and a second processing layer as an example.
[0134] Exemplarily, based on the mean hyperparameter decoding network, the coefficient hyperparameter feature can be processed by the first processing layer to obtain the processed feature; the processed feature can be upsampled by the upsampling layer to obtain the upsampled feature; the upsampled feature can be cropped and aligned by the Crop layer to obtain the cropped feature; the cropped feature can be processed by the second processing layer to obtain the initial mean feature.
[0135] Exemplarily, upsampling the processed features through an upsampling layer to obtain upsampled features may include but is not limited to: upsampling the processed features by pixel rearrangement to obtain upsampled features; or, upsampling the processed features by transposed convolution to obtain upsampled features; or, upsampling the processed features by nearest neighbor interpolation to obtain upsampled features.
[0136] Exemplarily, the first processing layer may include but is not limited to at least one of the following: one or more ordinary convolution layers, one or more deformable convolution layers, one or more depth separation convolution layers, one or more grouped convolution layers, one or more dilated convolution layers, one or more global processing unit layers, one or more activation layers using Relu activation function, one or more activation layers using upper and lower clamped target activation functions, and one or more activation layers using LeakyRelu activation function; there are arbitrary jump connections between the layers of the first processing layer.
[0137] Exemplarily, the second processing layer may include but is not limited to at least one of the following: one or more ordinary convolution layers, one or more deformable convolution layers, one or more depth separation convolution layers, one or more grouped convolution layers, one or more dilated convolution layers, one or more global processing unit layers, one or more activation layers using Relu activation function, one or more activation layers using upper and lower clamped target activation functions, and one or more activation layers using LeakyRelu activation function; there are arbitrary jump connections between the layers of the second processing layer.
[0138] For example, the above execution order is only for the convenience of describing the examples given. In actual applications, the execution order between the steps can also be changed, and this execution order is not limited. Moreover, in other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in this specification, and the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; multiple steps described in this specification may also be combined into a single step for description in other embodiments.
[0139] As can be seen from the above technical solutions, in the embodiment of the present application, an end-to-end video image compression method is proposed, which can realize the decoding of video images based on a neural network. By optimizing the activation function, the target activation function with upper and lower clamping is used to perform nonlinear adjustment on the features, thereby reducing the complexity of the activation function, and then reducing the complexity of the neural network. The target activation function can improve the robustness of the neural network and reduce the phenomenon of garbled code when the neural network decodes high-frequency images. While maintaining low complexity, the neural network effectively guarantees the quality of the reconstructed image block, improves the quality of the reconstructed image, achieves the purpose of improving decoding performance, and reduces complexity. By optimizing the mean hyperparameter decoding network, the complexity of the neural network is greatly reduced, and the performance of the neural network is kept unchanged, thereby achieving the purpose of improving decoding performance and reducing complexity.
[0140] In an embodiment of the present application, an encoding method is proposed. FIG3 is a flow chart of the encoding method. The method can be applied to an encoding end (also called a video encoder). The method may include:
[0141] Step 301: The encoder decodes the first bitstream corresponding to the current image block to obtain coefficient hyperparameter features corresponding to the current image block.
[0142] Step 302: The encoder determines a probability distribution parameter based on the coefficient hyperparameter feature, and determines a target mean feature corresponding to the current image block based on the coefficient hyperparameter feature.
[0143] Step 303: The encoder decodes the second code stream corresponding to the current image block based on the probability distribution parameters to obtain residual features corresponding to the current image block.
[0144] Step 304: The encoder determines a reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature.
[0145] In step 305, the encoder inputs the reconstructed features into a synthetic transformation network to obtain a reconstructed image block corresponding to the current image block. The synthetic transformation network includes a nonlinear residual network layer, which may include at least an activation layer and a skip connection structure. The output features of the activation layer are not greater than an upper threshold and / or the output features of the activation layer are not less than a lower threshold.
[0146] Exemplarily, the output features of the activation layer of the nonlinear residual network layer may be no greater than an upper threshold, and the output features of the activation layer may be no less than a lower threshold, that is, the upper and lower thresholds of the activation layer are simultaneously limited. In this case, the activation layer may use an upper and lower clamped target activation function to perform nonlinear adjustment on the features, and the target activation function corresponds to an upper and lower threshold, and the output features of the activation layer are no greater than the upper threshold, and the output features of the activation layer are no less than the lower threshold.
[0147] Alternatively, the output features of the activation layer of the nonlinear residual network layer can be no less than the lower threshold, that is, only the lower threshold of the activation layer is limited, and the upper threshold of the activation layer is not limited. In this case, the activation layer can use a Relu activation function to perform nonlinear adjustment on the features, and the Relu activation function corresponds to the lower threshold, and the output features of the activation layer are no less than the lower threshold.
[0148] Alternatively, the output features of the activation layer of the nonlinear residual network layer may be no greater than the upper threshold, that is, only the upper threshold of the activation layer is limited, and the lower threshold of the activation layer is not limited. In this case, the activation layer may use an activation function, such as that shown in FIG5E , to perform nonlinear adjustment on the features, and the activation function corresponds to the upper threshold, and the output features of the activation layer are no greater than the upper threshold.
[0149] In an embodiment of the present application, a coding method is proposed, which is applied to the coding end, including: decoding the first code stream corresponding to the current image block to obtain the coefficient hyperparameter feature corresponding to the current image block. Inputting the coefficient hyperparameter feature into the probability hyperparameter decoding network to obtain the probability distribution parameter, decoding the second code stream corresponding to the current image block based on the probability distribution parameter, and obtaining the residual feature corresponding to the current image block. Inputting the coefficient hyperparameter feature into the mean hyperparameter decoding network to obtain the initial mean feature corresponding to the current image block, and determining the target mean feature corresponding to the current image block based on the initial mean feature. Determining the reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature. Determine the reconstructed image block corresponding to the current image block based on the reconstruction feature.
[0150] In one possible implementation, the probabilistic hyperparameter decoding network may include an activation layer, and the output feature of the activation layer is not greater than an upper threshold, and / or the output feature of the activation layer is not less than a lower threshold. And / or, the mean hyperparameter decoding network may include an activation layer, and the output feature of the activation layer is not greater than an upper threshold, and / or the output feature of the activation layer is not less than a lower threshold.
[0151] In an embodiment of the present application, a coding method is proposed, which can be applied to the coding end, and the method includes: decoding the first code stream corresponding to the current image block to obtain the coefficient hyperparameter feature corresponding to the current image block. The coefficient hyperparameter feature is input to the hyperparameter decoding network to obtain the reference feature corresponding to the current image block. The reference feature and the obtained reconstruction feature of the current image block are input to the context model to obtain the probability distribution parameter and the target mean feature corresponding to the current image block. Based on the probability distribution parameter, the second code stream corresponding to the current image block is decoded to obtain the residual feature corresponding to the current image block, and the reconstruction feature is determined based on the target mean feature and the residual feature. The reconstructed image block corresponding to the current image block is determined based on the reconstruction feature.
[0152] In one possible implementation, the super-parameter decoding network may include an activation layer, and the output feature of the activation layer is not greater than an upper threshold, and / or the output feature of the activation layer is not less than a lower threshold. And / or, the context model may include an activation layer, and the output feature of the activation layer is not greater than an upper threshold, and / or the output feature of the activation layer is not less than a lower threshold.
[0153] In an embodiment of the present application, a coding method is proposed. The method can be applied to an encoding end. The method may include: decoding a first bitstream corresponding to a current image block to obtain coefficient hyperparameter features corresponding to the current image block. Determining probability distribution parameters based on the coefficient hyperparameter features, and decoding a second bitstream corresponding to the current image block based on the probability distribution parameters to obtain residual features corresponding to the current image block. Inputting the coefficient hyperparameter features into a mean hyperparameter decoding network to obtain initial mean features corresponding to the current image block, and determining a target mean feature corresponding to the current image block based on the initial mean features. The mean hyperparameter decoding network may sequentially include a processing layer, an upsampling layer for performing upsampling operations, and a cropping layer. Determining a reconstruction feature corresponding to the current image block based on the target mean features and the residual features. Determining a reconstructed image block corresponding to the current image block based on the reconstruction features. For the mean hyperparameter decoding network, the features output by the upsampling layer are passed through the cropping layer to obtain the initial mean features. That is, the upsampling layer is located at the end of the mean hyperparameter decoding network. Thus, after the features output by the upsampling layer pass through the cropping layer, the output features of the mean hyperparameter decoding network (i.e., the initial mean features) are obtained.
[0154] Exemplarily, the processing process of the encoding end is similar to that of the decoding end, and the similarities are not repeated here. The processing process of the decoding end can be applied to the encoding end, that is, the encoding end adopts the same processing method as the decoding end.
[0155] For example, the above execution order is only for the convenience of describing the examples given. In actual applications, the execution order between the steps can also be changed, and this execution order is not limited. Moreover, in other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in this specification, and the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; multiple steps described in this specification may also be combined into a single step for description in other embodiments.
[0156] As can be seen from the above technical solutions, in the embodiment of the present application, an end-to-end video image compression method is proposed, which can realize the encoding of video images based on neural networks. By optimizing the activation function, the target activation function with upper and lower clamping is used to perform nonlinear adjustment on the features, thereby reducing the complexity of the activation function, and then reducing the complexity of the neural network. The target activation function can improve the robustness of the neural network and reduce the phenomenon of garbled encoding of the neural network on high-frequency images. While maintaining low complexity, the neural network effectively guarantees the quality of the reconstructed image block, improves the quality of the reconstructed image, achieves the purpose of improving encoding performance, and reduces complexity. By optimizing the mean hyperparameter coding network, the complexity of the neural network is greatly reduced, and the performance of the neural network is kept unchanged, thereby achieving the purpose of improving encoding performance and reducing complexity.
[0157] In the method described in the above embodiment, the processing process of the encoding end can be referred to as shown in FIG4A . Of course, FIG4A is only an example of the processing process of the encoding end and does not limit the processing process of the encoding end.
[0158] As shown in Figure 4A , after obtaining the current image block x (which may be the original image block x, i.e., the input image block), the encoder can perform an analysis transformation on the current image block x using an analysis transformation network (i.e., a neural network) to obtain image features y corresponding to the current image block x. Performing feature transformation on the current image block x using the analysis transformation network refers to transforming the current image block x into image features y in the latent domain, thereby facilitating all subsequent processes to operate in the latent domain.
[0159] Exemplarily, an image can be divided into one image block or multiple image blocks. If the image is divided into one image block, the current image block x can also be an image, that is, the encoding and decoding process for the image block can also be directly used for the image.
[0160] After obtaining the image feature y, the encoding end performs a coefficient hyperparameter feature transformation on the image feature y to obtain a coefficient hyperparameter feature z. For example, the image feature y can be input into a hyperparameter encoding network (i.e., a neural network), and the hyperparameter encoding network performs a coefficient hyperparameter feature transformation on the image feature y to obtain a coefficient hyperparameter feature z. The hyperparameter encoding network can be a trained neural network, and there is no restriction on the training process of this hyperparameter encoding network. It only needs to be able to perform a coefficient hyperparameter feature transformation on the image feature y. The image feature y in the latent domain obtains the super-prior latent information z after passing through the hyperparameter encoding network.
[0161] After obtaining the coefficient hyperparameter feature z, the encoding end can quantize the coefficient hyperparameter feature z to obtain the superparameter quantization feature corresponding to the coefficient hyperparameter feature z, that is, the Q operation in Figure 4A is the quantization process. After obtaining the superparameter quantization feature corresponding to the coefficient hyperparameter feature z, the superparameter quantization feature is encoded to obtain Bitstream#1 (that is, the first code stream) corresponding to the current image block, that is, the AE operation in Figure 4A represents the encoding process, such as the entropy encoding process. Alternatively, the encoding end can also directly encode the coefficient hyperparameter feature z to obtain Bitstream#1 corresponding to the current image block. Among them, the superparameter quantization feature or coefficient hyperparameter feature z carried in Bitstream#1 is mainly used to obtain the parameters of the mean and probability distribution model. That is, the first code stream is a code stream encoded with the coefficient hyperparameter feature z corresponding to the current image block x.
[0162] After obtaining the Bitstream#1 corresponding to the current image block, the encoder can send the Bitstream#1 corresponding to the current image block to the decoder. For the processing process of the Bitstream#1 corresponding to the current image block by the decoder, please refer to the subsequent embodiments.
[0163] After obtaining Bitstream#1 corresponding to the current image block, the encoder can also decode Bitstream#1 to obtain super-parameter quantization features. That is, AD in Figure 4A represents the decoding process. Then, the super-parameter quantization features are dequantized to obtain the coefficient super-parameter feature z_hat. The coefficient super-parameter feature z_hat and the coefficient super-parameter feature z can be the same or different. The IQ operation in Figure 4A is the dequantization process. Alternatively, after obtaining Bitstream#1 corresponding to the current image block, the encoder can decode Bitstream#1 to obtain the coefficient super-parameter feature z_hat without dequantizing the super-parameter quantization features.
[0164] For the encoding process of Bitstream#1, an encoding method of a fixed probability density model can be used. For the decoding process of Bitstream#1, a decoding method of a fixed probability density model can be used. There is no restriction on this encoding and decoding process.
[0165] After obtaining the coefficient hyperparameter feature z_hat, the encoding end can input the coefficient hyperparameter feature z_hat into the mean hyperparameter decoding network, and the mean hyperparameter decoding network processes it based on the coefficient hyperparameter feature z_hat to obtain the initial mean feature m (the initial mean feature m is an intermediate parameter for the mean feature). There is no restriction on the processing process of this mean hyperparameter decoding network.
[0166] After obtaining the initial mean feature m, the encoding end can use the initial mean feature m as the target mean feature mu (i.e., the predicted value mu). Alternatively, the encoding end can input the initial mean feature m and the decoded reconstruction feature y_hat (i.e., the obtained reconstruction feature of the current image block, and the determination process of the reconstruction feature y_hat can be found in the subsequent embodiments) into the context model, and the context model performs a context-based prediction process to obtain the target mean feature mu (i.e., mean mu) corresponding to the current image block. For example, for the prediction process of the context model, the input data of the context model includes the initial mean feature m and the decoded reconstruction feature y_hat, and the two are jointly input to obtain a more accurate target mean feature mu, and the target mean feature mu is used to obtain the residual r_hat by subtracting from the original feature and adding it to the decoded residual to obtain the reconstructed feature y_hat.
[0167] The mean hyperparameter decoding network and context model are optional neural networks, that is, there may be no mean hyperparameter decoding network and context model, that is, there is no need to determine the target mean feature mu through the mean hyperparameter decoding network and context model.
[0168] After obtaining the image feature y, the encoding end can determine the residual feature r based on the image feature y and the target mean feature mu, such as taking the difference between the image feature y and the target mean feature mu as the residual feature r. Then, the residual feature r is feature processed to obtain the image feature s. There is no restriction on this feature processing process, and it can be any feature processing method. In this case, it is necessary to deploy a mean hyperparameter decoding network and a context model to provide the target mean feature mu. Alternatively, after obtaining the image feature y, the encoding end can perform feature processing on the image feature y to obtain the image feature s. There is no restriction on this feature processing process, and it can be any feature processing method. In this case, there is no need to deploy a mean hyperparameter decoding network and a context model.
[0169] After obtaining the image feature s, the encoder can quantize the image feature s to obtain the image quantization feature corresponding to the image feature s. That is, the Q operation in Figure 4A is the quantization process. After obtaining the image quantization feature corresponding to the image feature s, the encoder can encode the image quantization feature to obtain Bitstream#2 (i.e., the second code stream) corresponding to the current image block. That is, the AE operation in Figure 4A represents the encoding process, such as the entropy encoding process. Alternatively, the encoder can directly encode the image feature s to obtain Bitstream#2 corresponding to the current image block without involving the quantization process of the image feature s.
[0170] Alternatively, after obtaining the residual feature r, the encoder may not perform feature processing on the residual feature r, but may directly quantize the residual feature r to obtain the image quantization feature corresponding to the residual feature r. The encoder then encodes the image quantization feature to obtain Bitstream#2 (i.e., the second bitstream) corresponding to the current image block. Alternatively, after obtaining the residual feature r, the encoder directly encodes the residual feature r to obtain Bitstream#2 corresponding to the current image block without involving a quantization process. That is, the second bitstream is a bitstream encoded with the residual feature r corresponding to the current image block x.
[0171] After obtaining Bitstream#2 corresponding to the current image block, the encoder can send Bitstream#2 corresponding to the current image block to the decoder. For details about the processing of Bitstream#2 corresponding to the current image block by the decoder, see the subsequent embodiments.
[0172] After obtaining the Bitstream#2 corresponding to the current image block, the encoding end can also decode Bitstream#2 to obtain the image quantization feature, that is, AD in Figure 4A represents the decoding process. Then, the encoding end can dequantize the image quantization feature to obtain the image feature s'. The image feature s' can be the same as or different from the image feature s. The IQ operation in Figure 4A is the dequantization process. Alternatively, after obtaining the Bitstream#2 corresponding to the current image block, the encoding end can also decode Bitstream#2 to obtain the image feature s' without involving the dequantization process of the image quantization feature. After obtaining the image feature s', the encoding end can perform feature recovery (that is, the inverse process of feature processing) on the image feature s'. There is no restriction on this feature recovery, and it can be any feature recovery method to obtain the residual feature r_hat. The residual feature r_hat is the same as or different from the residual feature r.
[0173] Alternatively, after obtaining Bitstream#2 corresponding to the current image block, the encoder can also decode Bitstream#2 to obtain image quantization features, and then the encoder can dequantize the image quantization features to obtain the residual features r_hat. Alternatively, after obtaining Bitstream#2 corresponding to the current image block, the encoder can also decode Bitstream#2 to obtain the residual features r_hat without involving the dequantization process of the image quantization features.
[0174] After obtaining the residual feature r_hat, the encoder determines the image feature y_hat (i.e., the reconstructed feature) based on the residual feature r_hat and the target mean feature mu. The image feature y_hat is the same as or different from the image feature y. For example, the sum of the residual feature r_hat and the target mean feature mu can be used as the reconstructed feature y_hat. In this case, it is necessary to deploy a mean hyperparameter decoding network and a context model, and the target mean feature mu is provided by the mean hyperparameter decoding network and the context model. Alternatively, after the encoder obtains Bitstream#2 corresponding to the current image block, it can directly obtain the reconstructed feature y_hat after decoding and dequantization operations. In this case, there is no need to deploy a mean hyperparameter decoding network and a context model.
[0175] After obtaining the reconstructed feature y_hat, the encoding end can perform a synthetic transformation on the reconstructed feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the reconstructed feature y_hat is input into the synthetic transformation network, and the synthetic transformation network performs a synthetic transformation on the reconstructed feature y_hat to obtain the reconstructed image block x_hat. At this point, the image reconstruction process is completed.
[0176] For example, when encoding features to obtain Bitstream #2 corresponding to the current image block, the encoder must first determine a probability distribution model and then encode the features based on the probability distribution model. Furthermore, when decoding Bitstream #2, the encoder must also first determine a probability distribution model and then decode Bitstream #2 based on the probability distribution model.
[0177] In order to obtain a probability distribution model, as shown in Figure 4A, after obtaining the coefficient superparameter feature z_hat, the encoding end can perform a coefficient superparameter feature inverse transformation on the coefficient superparameter feature z_hat to obtain the probability distribution parameters. For example, the coefficient superparameter feature z_hat is input into the probability superparameter decoding network, and the probability superparameter decoding network performs a coefficient superparameter feature inverse transformation on the coefficient superparameter feature z_hat to obtain the probability distribution parameter p. After obtaining the probability distribution parameters, a probability distribution model can be generated based on the probability distribution parameters. Among them, the probability superparameter decoding network can be a trained neural network. There is no restriction on the training process of this probability superparameter decoding network. It is sufficient to be able to perform a coefficient superparameter feature inverse transformation on the coefficient superparameter feature z_hat.
[0178] In one possible implementation, the processing at the encoding end may be performed by a deep learning model or a neural network model, thereby achieving an end-to-end image compression and encoding process, without any restriction on the encoding process.
[0179] In the method described in the above embodiment, the processing process of the decoding end can be referred to as shown in FIG4B . Of course, FIG4B is only an example of the processing process of the decoding end and does not limit the processing process of the decoding end.
[0180] After obtaining Bitstream#1 corresponding to the current image block, the decoder can decode Bitstream#1 to obtain the super-parameter quantization features. That is, AD in Figure 4B represents the decoding process. The super-parameter quantization features are then dequantized to obtain the coefficient super-parameter features z_hat. The IQ operation in Figure 4B represents the dequantization process. Alternatively, after obtaining Bitstream#1 corresponding to the current image block, the decoder can decode Bitstream#1 to obtain the coefficient super-parameter features z_hat without involving the dequantization process.
[0181] For the decoding process of Bitstream#1, a decoding method using a fixed probability density model may be used, and no restriction is imposed on this.
[0182] The image can be divided into one image block or multiple image blocks. If the image is divided into one image block, the current image block x can also be an image, that is, the decoding process for the image block can also be directly used for the image.
[0183] After obtaining the coefficient hyperparameter feature z_hat, the decoding end can input the coefficient hyperparameter feature z_hat into the mean hyperparameter decoding network, and the mean hyperparameter decoding network processes it based on the coefficient hyperparameter feature z_hat to obtain the initial mean feature m (the initial mean feature m is an intermediate parameter for the mean feature). There is no restriction on the processing process of this mean hyperparameter decoding network.
[0184] After obtaining the initial mean feature m, the decoding end can use the initial mean feature m as the target mean feature mu (i.e., the predicted value mu). Alternatively, the decoding end can input the initial mean feature m and the decoded reconstruction feature y_hat (i.e., the obtained reconstruction feature of the current image block, and the determination process of the reconstruction feature y_hat can be found in the subsequent embodiments) into the context model, and the context model performs a context-based prediction process to obtain the target mean feature mu (i.e., mean mu) corresponding to the current image block. For example, for the prediction process of the context model, the input data of the context model includes the initial mean feature m and the decoded reconstruction feature y_hat, and the two are jointly input to obtain a more accurate target mean feature mu, and the target mean feature mu is used to obtain the residual r_hat by subtracting from the original feature and adding it to the decoded residual to obtain the reconstructed feature y_hat.
[0185] The mean hyperparameter decoding network and context model are optional neural networks, that is, there may be no mean hyperparameter decoding network and context model, that is, there is no need to determine the target mean feature mu through the mean hyperparameter decoding network and context model.
[0186] After obtaining the Bitstream#2 corresponding to the current image block, the decoding end can also decode Bitstream#2 to obtain the image quantization feature, that is, AD in Figure 4B represents the decoding process. Then, the decoding end can dequantize the image quantization feature to obtain the image feature s'. The IQ operation in Figure 4B is the dequantization process. Alternatively, after obtaining the Bitstream#2 corresponding to the current image block, the decoding end can also decode Bitstream#2 to obtain the image feature s' without involving the dequantization process of the image quantization feature. After obtaining the image feature s', the decoding end can perform feature recovery on the image feature s'. There is no restriction on this feature recovery, and it can be any feature recovery method to obtain the residual feature r_hat.
[0187] Alternatively, after obtaining Bitstream#2 corresponding to the current image block, the decoding end may further decode Bitstream#2 to obtain image quantization features, and then, the decoding end may dequantize the image quantization features to obtain residual features r_hat. Alternatively, after obtaining Bitstream#2 corresponding to the current image block, the decoding end may further decode Bitstream#2 to obtain residual features r_hat without involving a dequantization process of the image quantization features.
[0188] After obtaining the residual feature r_hat, the decoding end determines the image feature y_hat (i.e., the reconstructed feature) based on the residual feature r_hat and the target mean feature mu. The image feature y_hat is the same as or different from the image feature y. For example, the sum of the residual feature r_hat and the target mean feature mu can be used as the reconstructed feature y_hat. In this case, it is necessary to deploy a mean hyperparameter decoding network and a context model, and the target mean feature mu is provided by the mean hyperparameter decoding network and the context model. Alternatively, after the decoding end obtains Bitstream#2 corresponding to the current image block, it can directly obtain the reconstructed feature y_hat after performing operations such as decoding and dequantization. In this case, there is no need to deploy a mean hyperparameter decoding network and a context model.
[0189] After obtaining the reconstructed feature y_hat, the decoding end can perform a synthetic transformation on the reconstructed feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the reconstructed feature y_hat is input into the synthetic transformation network, and the synthetic transformation network performs a synthetic transformation on the reconstructed feature y_hat to obtain the reconstructed image block x_hat. At this point, the image reconstruction process is completed.
[0190] Exemplarily, when decoding Bitstream#2, the decoding end needs to first determine the probability distribution model, and then decode Bitstream#2 based on the probability distribution model. To obtain the probability distribution model, referring to FIG4B , after obtaining the coefficient superparameter feature z_hat, the decoding end can also perform a coefficient superparameter feature inverse transformation on the coefficient superparameter feature z_hat to obtain the probability distribution parameters. For example, the decoding end inputs the coefficient superparameter feature z_hat into the probability superparameter decoding network, which performs a coefficient superparameter feature inverse transformation on the coefficient superparameter feature z_hat to obtain the probability distribution parameter p. After obtaining the probability distribution parameters, the decoding end can generate the probability distribution model based on the probability distribution parameters.
[0191] Among them, the probabilistic hyperparameter decoding network can be a trained neural network. There is no restriction on the training process of this probabilistic hyperparameter decoding network. It only needs to be able to perform the coefficient hyperparameter feature inverse transformation on the coefficient hyperparameter feature z_hat.
[0192] In one possible implementation, the processing at the decoding end may be performed by a deep learning model or a neural network model, thereby achieving an end-to-end image compression and decoding process, without any restriction on the decoding process.
[0193] In the method described in the above embodiment, the processing process of the encoding end can be referred to as shown in FIG4C . Of course, FIG4C is only an example of the processing process of the encoding end and does not limit the processing process of the encoding end.
[0194] After obtaining the current image block x, the encoder can perform an analysis transform on the current image block x through the analysis transform network to obtain the image features y corresponding to the current image block x. Performing feature transformation on the current image block x through the analysis transform network means transforming the current image block x into the image features y in the latent domain, facilitating all subsequent processes to operate in the latent domain.
[0195] Exemplarily, an image can be divided into one image block or multiple image blocks. If the image is divided into one image block, the current image block x can also be an image, that is, the encoding and decoding process for the image block can also be directly used for the image.
[0196] After obtaining image feature y, the encoder performs a coefficient hyperparameter feature transformation on it to obtain coefficient hyperparameter feature z. For example, image feature y is input into a hyperparameter encoding network, which then performs a coefficient hyperparameter feature transformation on it to obtain coefficient hyperparameter feature z. The hyperparameter encoding network is a trained neural network that can perform coefficient hyperparameter feature transformation on image feature y. After passing through the hyperparameter encoding network, the latent domain image feature y obtains super-prior latent information z.
[0197] After obtaining the coefficient hyperparameter feature z, the encoder can quantize the coefficient hyperparameter feature z to obtain the corresponding hyperparameter quantization feature z, encode the hyperparameter quantization feature, and obtain Bitstream #1 (i.e., the first bitstream) corresponding to the current image block. Alternatively, the coefficient hyperparameter feature z can be directly encoded to obtain Bitstream #1 corresponding to the current image block. In other words, the first bitstream is the bitstream encoded with the coefficient hyperparameter feature z corresponding to the current image block x.
[0198] After obtaining the Bitstream#1 corresponding to the current image block, the encoder can send the Bitstream#1 corresponding to the current image block to the decoder. For the processing process of the Bitstream#1 corresponding to the current image block by the decoder, please refer to the subsequent embodiments.
[0199] After obtaining Bitstream#1 corresponding to the current image block, the encoder can also decode Bitstream#1 to obtain the hyperparameter quantization feature, dequantize the hyperparameter quantization feature, and obtain the coefficient hyperparameter feature z_hat. Alternatively, after obtaining Bitstream#1 corresponding to the current image block, the encoder can decode Bitstream#1 to obtain the coefficient hyperparameter feature z_hat.
[0200] For the encoding process of Bitstream#1, an encoding method of a fixed probability density model can be used. For the decoding process of Bitstream#1, a decoding method of a fixed probability density model can be used. There is no restriction on this encoding and decoding process.
[0201] After obtaining the coefficient hyperparameter feature z_hat, the encoding end can input the coefficient hyperparameter feature z_hat into the hyperparameter decoding network, and the superparameter decoding network processes it based on the coefficient hyperparameter feature z_hat to obtain the reference feature m (the reference feature m is an intermediate parameter for the mean feature and the probability distribution parameter). There is no restriction on the processing process of this superparameter decoding network.
[0202] After obtaining the reference feature m, the encoding end can input the reference feature m and the decoded reconstructed feature y_hat (i.e., the obtained reconstructed feature of the current image block, and the determination process of the reconstructed feature y_hat can be found in the subsequent embodiments) into the context model, and the context model performs a context-based prediction process to obtain the target mean feature mu (i.e., the predicted value mu, i.e., the mean mu) and the probability distribution parameter p corresponding to the current image block. For example, for the prediction process of the context model, the input data of the context model includes the reference feature m and the decoded reconstructed feature y_hat. The two are jointly input to obtain a more accurate target mean feature mu and probability distribution parameter p. The target mean feature mu is used to obtain the residual r_hat by subtracting it from the original feature and adding it to the decoded residual to obtain the reconstructed feature y_hat. The probability distribution parameter p is used to encode Bitstream#2.
[0203] After obtaining the image feature y, the encoding end can determine the residual feature r based on the image feature y and the target mean feature mu, such as taking the difference between the image feature y and the target mean feature mu as the residual feature r. Then, the residual feature r is feature-processed to obtain the image feature s. There is no restriction on the feature processing process, and it can be any feature processing method. Alternatively, after obtaining the image feature y, the encoding end can perform feature processing on the image feature y to obtain the image feature s. There is no restriction on the feature processing process, and it can be any feature processing method. After obtaining the image feature s, the encoding end can quantize the image feature s to obtain the image quantization feature corresponding to the image feature s, encode the image quantization feature, and obtain the Bitstream#2 (i.e., the second code stream) corresponding to the current image block. Alternatively, after obtaining the image feature s, the encoding end can also directly encode the image feature s to obtain the Bitstream#2 corresponding to the current image block without involving the quantization process of the image feature s.
[0204] Alternatively, after obtaining the residual feature r, the encoder may not perform feature processing on the residual feature r, but may directly quantize the residual feature r to obtain the image quantization feature corresponding to the residual feature r. The encoder then encodes the image quantization feature to obtain Bitstream#2 (i.e., the second bitstream) corresponding to the current image block. Alternatively, after obtaining the residual feature r, the encoder directly encodes the residual feature r to obtain Bitstream#2 corresponding to the current image block without involving a quantization process. That is, the second bitstream is a bitstream encoded with the residual feature r corresponding to the current image block x.
[0205] For example, when encoding the features to obtain Bitstream#2 corresponding to the current image block, the encoder may generate a probability distribution model based on the probability distribution parameter p, and then encode the features based on the probability distribution model.
[0206] After obtaining Bitstream#2 corresponding to the current image block, the encoder can send Bitstream#2 corresponding to the current image block to the decoder. For details about the processing of Bitstream#2 corresponding to the current image block by the decoder, see the subsequent embodiments.
[0207] After obtaining Bitstream#2 corresponding to the current image block, the encoder can also decode Bitstream#2 to obtain image quantization features, and dequantize the image quantization features to obtain image features s'. Alternatively, after obtaining Bitstream#2 corresponding to the current image block, the encoder can also decode Bitstream#2 to obtain image features s' without involving the dequantization process of the image quantization features. After obtaining image features s', the encoder can perform feature recovery on image features s'. There is no restriction on this feature recovery method, and any feature recovery method can be used to obtain residual features r_hat.
[0208] Alternatively, after obtaining Bitstream#2 corresponding to the current image block, the encoder can also decode Bitstream#2 to obtain image quantization features, and then the encoder can dequantize the image quantization features to obtain the residual features r_hat. Alternatively, after obtaining Bitstream#2 corresponding to the current image block, the encoder can also decode Bitstream#2 to obtain the residual features r_hat without involving the dequantization process of the image quantization features.
[0209] For example, when the encoder decodes Bitstream#2, it may generate a probability distribution model based on the probability distribution parameter p, and then decode Bitstream#2 based on the probability distribution model. There is no restriction on this decoding process.
[0210] After obtaining the residual feature r_hat, the encoder determines the image feature y_hat (i.e., the reconstruction feature) based on the residual feature r_hat and the target mean feature mu, such as taking the sum of the residual feature r_hat and the target mean feature mu as the reconstruction feature y_hat.
[0211] After obtaining the reconstructed feature y_hat, the encoding end can perform a synthetic transformation on the reconstructed feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the reconstructed feature y_hat is input into the synthetic transformation network, and the synthetic transformation network performs a synthetic transformation on the reconstructed feature y_hat to obtain the reconstructed image block x_hat. At this point, the image reconstruction process is completed.
[0212] In one possible implementation, the processing at the encoding end may be performed by a deep learning model or a neural network model, thereby achieving an end-to-end image compression and encoding process, without any restriction on the encoding process.
[0213] In the method described in the above embodiment, the processing process of the decoding end can be referred to as shown in FIG4D . Of course, FIG4D is only an example of the processing process of the decoding end and does not limit the processing process of the decoding end.
[0214] After obtaining Bitstream#1 corresponding to the current image block, the decoder can also decode Bitstream#1 to obtain the hyperparameter quantization feature, dequantize the hyperparameter quantization feature, and obtain the coefficient hyperparameter feature z_hat. Alternatively, after obtaining Bitstream#1 corresponding to the current image block, the decoder can decode Bitstream#1 to obtain the coefficient hyperparameter feature z_hat.
[0215] For the decoding process of Bitstream#1, a decoding method using a fixed probability density model may be used, and no restriction is imposed on this.
[0216] The image can be divided into one image block or multiple image blocks. If the image is divided into one image block, the current image block x can also be an image, that is, the decoding process for the image block can also be directly used for the image.
[0217] After obtaining the coefficient hyperparameter feature z_hat, the decoding end can input the coefficient hyperparameter feature z_hat into the hyperparameter decoding network, and the hyperparameter decoding network processes it based on the coefficient hyperparameter feature z_hat to obtain the reference feature m (the reference feature m is an intermediate parameter for the mean feature and the probability distribution parameter). There is no restriction on the processing process of this hyperparameter decoding network.
[0218] After obtaining the reference feature m, the decoding end can input the reference feature m and the decoded reconstructed feature y_hat (i.e., the obtained reconstructed feature of the current image block, and the determination process of the reconstructed feature y_hat can be found in the subsequent embodiments) into the context model, and the context model performs a context-based prediction process to obtain the target mean feature mu (i.e., the predicted value mu, i.e., the mean mu) and the probability distribution parameter p corresponding to the current image block. For example, for the prediction process of the context model, the input data of the context model includes the reference feature m and the decoded reconstructed feature y_hat. The two are jointly input to obtain a more accurate target mean feature mu and probability distribution parameter p. The target mean feature mu is used to obtain the residual r_hat by subtracting it from the original feature and adding it to the decoded residual to obtain the reconstructed feature y_hat. The probability distribution parameter p is used to decode Bitstream#2.
[0219] After obtaining Bitstream#2 corresponding to the current image block, the decoder can also decode Bitstream#2 to obtain image quantization features, and dequantize the image quantization features to obtain image features s'. Alternatively, after obtaining Bitstream#2 corresponding to the current image block, the decoder can also decode Bitstream#2 to obtain image features s' without involving the dequantization process of the image quantization features. After obtaining image features s', the decoder can perform feature recovery on image features s'. There is no restriction on this feature recovery method, and any feature recovery method can be used to obtain residual features r_hat.
[0220] Alternatively, after obtaining Bitstream#2 corresponding to the current image block, the decoding end may further decode Bitstream#2 to obtain image quantization features, and then, the decoding end may dequantize the image quantization features to obtain residual features r_hat. Alternatively, after obtaining Bitstream#2 corresponding to the current image block, the decoding end may further decode Bitstream#2 to obtain residual features r_hat without involving a dequantization process of the image quantization features.
[0221] For example, when decoding Bitstream#2, the decoding end may generate a probability distribution model based on the probability distribution parameter p, and then decode Bitstream#2 based on the probability distribution model. There is no restriction on this decoding process.
[0222] After obtaining the residual feature r_hat, the decoding end determines the image feature y_hat (i.e., the reconstructed feature) based on the residual feature r_hat and the target mean feature mu, such as taking the sum of the residual feature r_hat and the target mean feature mu as the reconstructed feature y_hat.
[0223] After obtaining the reconstructed feature y_hat, the decoding end can perform a synthetic transformation on the reconstructed feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the reconstructed feature y_hat is input into the synthetic transformation network, and the synthetic transformation network performs a synthetic transformation on the reconstructed feature y_hat to obtain the reconstructed image block x_hat. At this point, the image reconstruction process is completed.
[0224] In one possible implementation, the processing at the decoding end may be performed by a deep learning model or a neural network model, thereby achieving an end-to-end image compression and decoding process, without any restriction on the decoding process.
[0225] In the above embodiment, for the encoding end and the decoding end, after obtaining the first bitstream (Bitstream#1) corresponding to the current image block, the first bitstream corresponding to the current image block can be decoded to obtain the coefficient hyperparameter feature z_hat corresponding to the current image block. That is, the first bitstream corresponding to the current image block is a bitstream encoded with the coefficient hyperparameter feature z corresponding to the current image block x. The probability distribution parameter p can be determined based on the coefficient hyperparameter feature z_hat. For example, the coefficient hyperparameter feature z_hat is input into the probability hyperparameter decoding network to obtain the probability distribution parameter p. The second bitstream (Bitstream#2) corresponding to the current image block can be decoded based on the probability distribution parameter p to obtain the residual feature r_hat corresponding to the current image block. That is, the second bitstream corresponding to the current image block is a bitstream encoded with the residual feature r corresponding to the current image block x.
[0226] The target mean feature mu corresponding to the current image block can be determined based on the coefficient hyperparameter feature z_hat. For example, the coefficient hyperparameter feature z_hat can be input into the mean hyperparameter decoding network to obtain the initial mean feature m corresponding to the current image block, and the initial mean feature m can be used as the target mean feature mu. Alternatively, the coefficient hyperparameter feature z_hat can be input into the mean hyperparameter decoding network to obtain the initial mean feature m corresponding to the current image block, and the initial mean feature m and the obtained reconstructed feature y_hat of the current image block can be input into the context model to obtain the target mean feature mu.
[0227] Based on the target mean feature mu and the residual feature r_hat, the reconstructed feature y_hat corresponding to the current image block is determined, and based on the reconstructed feature y_hat, the reconstructed image block x_hat corresponding to the current image block is determined. For example, the reconstructed feature y_hat can be input into the synthesis transformation network to obtain the reconstructed image block x_hat corresponding to the current image block.
[0228] In the above embodiment, for the encoding end and the decoding end, after obtaining the first bitstream (Bitstream#1) corresponding to the current image block, the first bitstream corresponding to the current image block can be decoded to obtain the coefficient hyperparameter feature z_hat corresponding to the current image block. That is, the first bitstream corresponding to the current image block is a bitstream encoded with the coefficient hyperparameter feature z corresponding to the current image block x. The probability distribution parameter p and the target mean feature mu corresponding to the current image block can be determined based on the coefficient hyperparameter feature z_hat. For example, the coefficient hyperparameter feature z_hat can be input into the hyperparameter decoding network to obtain the reference feature m corresponding to the current image block. Then, the reference feature m and the obtained reconstructed feature y_hat of the current image block can be input into the context model to obtain the probability distribution parameter p and the target mean feature mu.
[0229] The second bitstream (Bitstream#2) corresponding to the current image block can be decoded based on the probability distribution parameter p to obtain the residual feature r_hat corresponding to the current image block. That is, the second bitstream corresponding to the current image block is a bitstream encoded with the residual feature r corresponding to the current image block x. The reconstructed feature y_hat corresponding to the current image block is determined based on the target mean feature mu and the residual feature r_hat, and the reconstructed image block x_hat corresponding to the current image block is determined based on the reconstructed feature y_hat. For example, the reconstructed feature y_hat can be input into a synthesis transformation network to obtain the reconstructed image block x_hat corresponding to the current image block.
[0230] In the above embodiments, for the encoding end and the decoding end, the coefficient super-parameter feature z_hat can be input into the probabilistic super-parameter decoding network to obtain the probability distribution parameter p. The probabilistic super-parameter decoding network can be a trained neural network for performing a coefficient super-parameter feature inverse transformation on the coefficient super-parameter feature z_hat. For example, the structure of the probabilistic super-parameter decoding network can be shown in Figure 5A. However, Figure 5A is only an example of a probabilistic super-parameter decoding network. This application does not limit the structure of this probabilistic super-parameter decoding network. The probabilistic super-parameter decoding network of Figure 5A will be used as an example later.
[0231] Exemplarily, the probabilistic hyperparameter decoding network may include in sequence: a convolutional layer (such as a convolutional layer with a size of 1*1, the number of input channels is c, and the number of output channels is c), a Relu activation layer, a convolutional layer (such as a convolutional layer with a size of 3*3, the number of input channels is c, and the number of output channels is c), a Relu activation layer, a convolutional layer (such as a convolutional layer with a size of 1*1, the number of input channels is c, and the number of output channels is c), a Pixshuffle (upsampling) layer, and a Crop layer. The Crop layer is used to crop and align the pixel positions after upsampling. The Crop layer can be located after the Pixshuffle layer to ensure that the feature map sizes are aligned.
[0232] As shown in Figure 5A, the input data of the probabilistic hyperparameter decoding network is the coefficient hyperparameter feature z_hat. After being processed by the probabilistic hyperparameter decoding network, the probability distribution parameter p can be obtained and the probability distribution parameter p is output.
[0233] As shown in FIG5A , the probabilistic hyperparameter decoding network may include an activation layer, and the activation layer may use a Relu activation function. For example, the Relu activation function may be shown in FIG5B , where the Relu activation function corresponds to a lower threshold. As the input feature decreases, the lower limit of the output feature is limited so that the activation layer does not output an output feature with a very small value.
[0234] In this embodiment, the activation layer of the probabilistic superparameter decoding network can also be optimized so that the activation layer adopts a target activation function that is clamped up and down, that is, the activation layer adopts a target activation function that is clamped up and down to perform nonlinear adjustment on the features. For example, the target activation function can be a Relu6 activation function, or other activation function, as long as the upper and lower clamping is achieved at the same time. The Relu6 activation function will be used as an example for explanation below. Referring to Figure 5C, which is a structural diagram of the probabilistic superparameter decoding network, the Relu activation layer is replaced by the Relu6 activation layer, that is, the Relu6 activation function is used for processing. The activation layer of the probabilistic superparameter decoding network can also be optimized so that the activation layer adopts a certain activation function, and the activation function corresponds to an upper limit threshold, that is, the upper limit of the activation layer is limited, but the lower limit of the activation layer is not limited.
[0235] The Relu6 activation function (i.e., the target activation function for upper and lower clamping) can be seen in Figure 5D. The Relu6 activation function corresponds to an upper threshold and a lower threshold. As the input feature decreases, the lower limit of the output feature will be restricted so that the output feature of the activation layer is not less than the lower threshold (0 is used as an example in Figure 5D). As the input feature increases, the upper limit of the output feature will be restricted so that the output feature of the activation layer is not greater than the upper threshold (6 is used as an example in Figure 5D). The Relu6 activation function limits the lower and upper limits so that the output feature is within the specified interval (between the lower threshold and the upper threshold), so that the activation layer will not output output features with very large or very small values.
[0236] Exemplarily, the upper threshold corresponding to the Relu6 activation function may be a configured fixed upper threshold, that is, a fixed upper threshold is used as the upper threshold corresponding to the Relu6 activation function by default, such as the fixed upper threshold in 5D is 6. The lower threshold corresponding to the Relu6 activation function may be a configured fixed lower threshold, that is, a fixed lower threshold is used as the lower threshold corresponding to the Relu6 activation function by default, such as the fixed lower threshold in 5D is 0. Of course, both the fixed upper threshold and the fixed lower threshold can be configured based on experience, and in this embodiment, there is no restriction on the fixed upper threshold and the fixed lower threshold.
[0237] Exemplarily, the upper threshold corresponding to the Relu6 activation function may be an upper threshold for adaptive learning, i.e., the upper threshold corresponding to the Relu6 activation function is determined through a training process, rather than a fixed upper threshold. In this embodiment, the learning process for this upper threshold is not restricted. Furthermore, the lower threshold corresponding to the Relu6 activation function may be an lower threshold for adaptive learning, i.e., the lower threshold corresponding to the Relu6 activation function is determined through a training process, rather than a fixed lower threshold.
[0238] Alternatively, the upper threshold corresponding to the Relu6 activation function may be an upper threshold for adaptive learning rather than a fixed upper threshold, and the lower threshold corresponding to the Relu6 activation function may be a fixed lower threshold.
[0239] Alternatively, the lower threshold corresponding to the Relu6 activation function may be a lower threshold for adaptive learning rather than a fixed lower threshold, and the upper threshold corresponding to the Relu6 activation function may be a fixed upper threshold.
[0240] For example, when adaptively learning the upper threshold corresponding to the Relu6 activation function, the same upper threshold can be learned for all channels of the activation layer. That is, only one upper threshold needs to be adaptively learned, and this upper threshold serves as the upper threshold for all channels of the activation layer. Alternatively, when adaptively learning the upper threshold corresponding to the Relu6 activation function, an upper threshold can be learned separately for each channel of the activation layer. That is, assuming that the activation layer corresponds to K channels (i.e., the number of channels of the input feature, K is a positive integer), K upper thresholds need to be adaptively learned. The K upper thresholds correspond one-to-one to the K channels, and each channel corresponds to an upper threshold separately, and the upper thresholds corresponding to different channels can be the same or different.
[0241] For example, when adaptively learning the lower threshold corresponding to the Relu6 activation function, the same lower threshold can be learned for all channels of the activation layer. That is, only one lower threshold needs to be adaptively learned, and this lower threshold serves as the lower threshold for all channels of the activation layer. Alternatively, when adaptively learning the lower threshold corresponding to the Relu6 activation function, a lower threshold can be learned separately for each channel of the activation layer. That is, assuming that the activation layer corresponds to K channels (i.e., the number of channels of the input feature, K is a positive integer), K lower thresholds need to be adaptively learned. The K lower thresholds correspond one-to-one to the K channels, and each channel corresponds to a lower threshold separately, and the lower thresholds corresponding to different channels can be the same or different.
[0242] In the above embodiment, for both the encoding and decoding ends, the coefficient hyperparameter feature z_hat can be input into the mean hyperparameter decoding network to obtain the initial mean feature m corresponding to the current image block. For example, the structure of the mean hyperparameter decoding network can be seen in Figure 6A. However, Figure 6A is only an example of a mean hyperparameter decoding network, and this application does not limit the structure of this mean hyperparameter decoding network. The mean hyperparameter decoding network of Figure 6A will be used as an example in the following.
[0243] Exemplarily, the mean hyperparameter decoding network may include, in sequence: a convolution layer (such as a convolution layer of size 1*1, with c input channels and c output channels), a transposed convolution layer TConv (such as a convolution layer of size 4*4, with c input channels and c output channels, and a convolution step of s2), a Crop layer, a Relu activation layer, a convolution layer (such as a convolution layer of size 3*3, with c input channels and c output channels), a Relu activation layer, and a convolution layer (such as a convolution layer of size 3*3, with c input channels and c output channels). Referring to FIG6A , the input data of the mean hyperparameter decoding network is the coefficient hyperparameter feature z_hat. After processing by the mean hyperparameter decoding network, the initial mean feature m can be obtained and output.
[0244] As shown in Figure 6A, the mean hyperparameter decoding network may include an activation layer, and the activation layer may adopt a Relu activation function. For example, the Relu activation function can be shown in Figure 5B. On this basis, the activation layer of the mean hyperparameter decoding network can also be optimized so that the activation layer adopts a target activation function that is clamped up and down, that is, the target activation function that is clamped up and down is used to perform nonlinear adjustment on the features. For example, the target activation function can be a Relu6 activation function, or other activation functions, as long as the upper and lower clamping is achieved at the same time. As shown in Figure 6B, it is a structural diagram of the mean hyperparameter decoding network, in which the Relu activation layer is replaced by a Relu6 activation layer, that is, the Relu6 activation function is used for processing. Of course, the activation layer of the mean hyperparameter decoding network can also be optimized so that the activation layer adopts a certain activation function, and the activation function corresponds to an upper limit threshold, that is, the upper limit of the activation layer is limited, but the lower limit of the activation layer is not limited.
[0245] The Relu6 activation function can be shown in FIG5D , where the Relu6 activation function corresponds to an upper threshold and a lower threshold. The upper threshold corresponding to the Relu6 activation function can be a configured fixed upper threshold, and the lower threshold corresponding to the Relu6 activation function can be a configured fixed lower threshold. Alternatively, the upper threshold corresponding to the Relu6 activation function can be an upper threshold for adaptive learning, and / or the lower threshold corresponding to the Relu6 activation function can be a lower threshold for adaptive learning.
[0246] Exemplarily, when adaptively learning the upper threshold corresponding to the Relu6 activation function, the same upper threshold can be learned for all channels of the activation layer, or an upper threshold can be learned separately for each channel of the activation layer.
[0247] Exemplarily, when adaptively learning the lower threshold corresponding to the Relu6 activation function, the same lower threshold may be learned for all channels of the activation layer, or a lower threshold may be learned separately for each channel of the activation layer.
[0248] In the above embodiment, at the encoding and decoding ends, the initial mean feature m and the obtained reconstructed feature y_hat of the current image block can be input into the context model to obtain the target mean feature mu. For example, the structure of the context model can be seen in Figure 6C. However, Figure 6C is only an example of a context model, and this application does not limit the structure of this context model. The context model of Figure 6C will be used as an example for the following description.
[0249] Exemplarily, the context model may include multiple parameter fusion networks. For each parameter fusion network, the parameter fusion network may include, in sequence: a convolutional layer (e.g., a convolutional layer of size 1*1), a Relu activation layer, a convolutional layer (e.g., a convolutional layer of size 1*1), a Relu activation layer, and a convolutional layer (e.g., a convolutional layer of size 1*1). Referring to FIG6C , the input data of the context model is the initial mean feature m and the reconstructed feature y_hat. After processing by the context model, the target mean feature mu can be obtained and output. There is no restriction on the processing process of this context model.
[0250] As shown in FIG6C , the parameter fusion network of the context model may include an activation layer, and the activation layer may adopt a Relu activation function. The Relu activation function can be shown in FIG5B . On this basis, the activation layer of the parameter fusion network can be optimized so that the activation layer adopts a target activation function that is clamped up and down, that is, the target activation function that is clamped up and down is used to perform nonlinear adjustment on the features. For example, the target activation function can be a Relu6 activation function, or other activation functions, as long as the upper and lower clamping is achieved at the same time. FIG6D shows the structure of the parameter fusion network of the context model. As shown in FIG6D , the Relu activation layer of the parameter fusion network is replaced by a Relu6 activation layer. Of course, the activation layer of the context model can also be optimized so that the activation layer adopts a certain activation function, and the activation function corresponds to an upper limit threshold, that is, the upper limit of the activation layer is limited, but the lower limit of the activation layer is not limited.
[0251] The Relu6 activation function can be shown in FIG5D , where the Relu6 activation function corresponds to an upper threshold and a lower threshold. The upper threshold corresponding to the Relu6 activation function can be a configured fixed upper threshold, and the lower threshold corresponding to the Relu6 activation function can be a configured fixed lower threshold. Alternatively, the upper threshold corresponding to the Relu6 activation function can be an upper threshold for adaptive learning, and / or the lower threshold corresponding to the Relu6 activation function can be a lower threshold for adaptive learning.
[0252] Exemplarily, when adaptively learning the upper threshold corresponding to the Relu6 activation function, the same upper threshold can be learned for all channels of the activation layer, or an upper threshold can be learned separately for each channel of the activation layer.
[0253] Exemplarily, when adaptively learning the lower threshold corresponding to the Relu6 activation function, the same lower threshold may be learned for all channels of the activation layer, or a lower threshold may be learned separately for each channel of the activation layer.
[0254] In the above embodiment, the Relu activation function in the probabilistic hyperparameter decoding network, the mean hyperparameter decoding network, and the context model can be replaced with the Relu6 activation function. If the probabilistic hyperparameter decoding network has an activation layer that uses the LeakyRelu activation function, the LeakyRelu activation function can also be replaced with the Relu6 activation function, and there is no restriction on this process. And / or, if the mean hyperparameter decoding network has an activation layer that uses the LeakyRelu activation function, the LeakyRelu activation function can also be replaced with the Relu6 activation function, and there is no restriction on this process. And / or, if the context model has an activation layer that uses the LeakyRelu activation function, the LeakyRelu activation function can also be replaced with the Relu6 activation function, and there is no restriction on this process.
[0255] In the above embodiment, for the encoding end and the decoding end, the coefficient hyperparameter feature z_hat can be input into the hyperparameter decoding network to obtain the reference feature m, and the reference feature m and the obtained reconstructed feature y_hat of the current image block can be input into the context model to obtain the probability distribution parameter p and the target mean feature mu.
[0256] The structure of the context model can be shown in Figure 6C. The context model can include multiple parameter fusion networks, which can include: convolutional layer, Relu activation layer, convolutional layer, Relu activation layer, and convolutional layer. The structure of the super-parameter decoding network is not limited in this embodiment, and the super-parameter decoding network can include at least one Relu activation layer.
[0257] For the context model and the hyperparameter decoding network, the activation layer can adopt the Relu activation function, and the Relu activation function can be shown in Figure 5B. On this basis, the activation layer of the context model and / or the hyperparameter decoding network can be optimized so that the activation layer adopts the upper and lower clamped target activation function, that is, the upper and lower clamped target activation function is used to perform nonlinear adjustment on the features. For example, the target activation function can be the Relu6 activation function, that is, the Relu activation layer is replaced by the Relu6 activation layer. The activation layer of the context model and / or the hyperparameter decoding network can also be optimized so that the activation layer adopts a certain activation function, and the activation function corresponds to an upper limit threshold, that is, the upper limit of the activation layer is limited, but the lower limit of the activation layer is not limited.
[0258] The Relu6 activation function can be shown in FIG5D , where the Relu6 activation function corresponds to an upper threshold and a lower threshold. The upper threshold corresponding to the Relu6 activation function can be a configured fixed upper threshold, and the lower threshold corresponding to the Relu6 activation function can be a configured fixed lower threshold. Alternatively, the upper threshold corresponding to the Relu6 activation function can be an upper threshold for adaptive learning, and / or the lower threshold corresponding to the Relu6 activation function can be a lower threshold for adaptive learning.
[0259] Exemplarily, when adaptively learning the upper threshold corresponding to the Relu6 activation function, the same upper threshold can be learned for all channels of the activation layer, or an upper threshold can be learned separately for each channel of the activation layer.
[0260] Exemplarily, when adaptively learning the lower threshold corresponding to the Relu6 activation function, the same lower threshold may be learned for all channels of the activation layer, or a lower threshold may be learned separately for each channel of the activation layer.
[0261] In the above embodiment, for the encoding end and the decoding end, the reconstructed feature y_hat can be input into the synthesis transformation network to obtain the reconstructed image block x_hat corresponding to the current image block. The synthesis transformation network can be a trained neural network for performing a synthesis transformation on the reconstructed feature y_hat. For example, the synthesis transformation network can correspond to three branches, namely a high-complexity synthesis transformation network, a medium-complexity synthesis transformation network, and a low-complexity synthesis transformation network, and the image quality output by the three branches is different. At least one of the high-complexity synthesis transformation network, the medium-complexity synthesis transformation network, and the low-complexity synthesis transformation network can be deployed according to computing power. For example, for a decoding end with higher computing power, a high-complexity synthesis transformation network, a medium-complexity synthesis transformation network, and a low-complexity synthesis transformation network can be deployed simultaneously. If a reconstructed image block x_hat with higher image quality needs to be output, the reconstructed feature y_hat can be input into the high-complexity synthesis transformation network. If a reconstructed image block x_hat with lower image quality needs to be output, the reconstructed feature y_hat can be input into the low-complexity synthesis transformation network. If a reconstructed image block x_hat with moderate image quality needs to be output, the reconstructed feature y_hat can be input into the medium-complexity synthesis transformation network. For another example, for a decoding end with lower computing power, only a low-complexity synthesis transformation network can be deployed, and the reconstructed feature y_hat can be input into the low-complexity synthesis transformation network. Taking the deployment of three branches as an example, see Figure 7A, which is a structural diagram of the synthesis transformation network. Figure 7A shows the three branches of the synthesis transformation network.
[0262] For example, the structure of the low-complexity synthesis transformation network can be shown in Figure 7B. However, Figure 7B is only an example of the low-complexity synthesis transformation network, and this application does not limit this. Figure 7B is used as an example in the following. The low-complexity synthetic transformation network may include in sequence: LRB layer, convolution layer (Conv, such as a convolution layer with a size of 2*2, the number of input channels is c, the number of output channels is c, and the convolution step is 2), Pixshuffle layer (such as a 2x upsampling layer), ResAU layer, Crop layer, convolution layer (Conv, such as a convolution layer with a size of 2*2, the number of input channels is c1, the number of output channels is c2, and the convolution step is s2), Pixshuffle layer (such as a 2x upsampling layer), Crop layer, ResAU layer, convolution layer (such as a convolution layer with a size of 3*3, the number of input channels is c, the number of output channels is c), ResAU layer, convolution layer (such as a convolution layer with a size of 1*1, the number of input channels is c, the number of output channels is 16*cout), Pixshuffle layer (such as a 4x upsampling layer), and Crop layer. The Crop layer is used to crop and align the pixel positions after upsampling. The Crop layer can be placed after the Pixshuffle layer to ensure that the feature map sizes are aligned.
[0263] As shown in FIG7B , the input data of the low-complexity synthesis transformation network is the reconstructed feature y_hat. After being processed by the low-complexity synthesis transformation network, a reconstructed image block x_hat can be obtained, and the reconstructed image block x_hat is output.
[0264] For example, the structure of the medium complexity synthesis transformation network can be shown in Figure 7C. However, Figure 7C is only an example of the medium complexity synthesis transformation network, and this application does not limit this. Figure 7C will be taken as an example later. The medium complexity synthesis transformation network may include in sequence: LRB layer, transposed convolution layer (TConv, such as a convolution layer with a size of 4*4, the number of input channels is c1, the number of output channels is c2, and the convolution step is S2), ResAU layer, Crop layer, transposed convolution layer (TConv, such as a convolution layer with a size of 4*4, the number of input channels is c2, the number of output channels is c3, and the convolution step is S2), Crop layer, ResAU layer, convolution layer (Conv, such as a convolution layer with a size of 3*3, the number of input channels is c, the number of output channels is c), ResAU layer, convolution layer (Conv, such as a convolution layer with a size of 3*3, the number of input channels is c, the number of output channels is c), Pixshuffle layer (such as a 4x upsampling layer), and Crop layer. The input data of the medium complexity synthetic transformation network is the reconstructed feature y_hat. After being processed by the medium complexity synthetic transformation network, the reconstructed image block x_hat can be obtained, and the reconstructed image block x_hat is output.
[0265] For example, the structure of the high-complexity synthesis transformation network can be shown in Figure 7D. However, Figure 7D is only an example of a high-complexity synthesis transformation network, and this application does not limit this. Figure 7D is used as an example in the following. The high-complexity synthetic transformation network may include in sequence: RB layer, transposed convolution layer (TConv, such as a convolution layer with a size of 3*3, the number of input channels is c, the number of output channels is c, and the convolution step is S2), ResAU layer, Crop layer, transposed convolution layer (TConv, such as a convolution layer with a size of 3*3, the number of input channels is c, the number of output channels is c, and the convolution step is S2), CAB layer, Crop layer, ResAU layer, convolution layer (Conv, such as a convolution layer with a size of 1*1, the number of input channels is c, the number of output channels is c), Pixshuffle layer (such as a 2x upsampling layer), TAM layer, Crop layer, ResAU layer, transposed convolution layer (TConv, such as a convolution layer with a size of 3*3, the number of input channels is c2, the number of output channels is c3, and the convolution step is S2), and Crop layer.
[0266] As shown in FIG7D , the input data of the high-complexity synthesis transformation network is the reconstructed feature y_hat. After being processed by the high-complexity synthesis transformation network, a reconstructed image block x_hat can be obtained, and the reconstructed image block x_hat is output.
[0267] In FIG7D , the CAB layer is a convolution-base attention block, and the structure of the CAB layer can be seen in FIG7E . However, FIG7E is only an example and this application does not limit this.
[0268] In Figure 7D, the TAM layer is a Transformer-based attention block. The structure of the TAM layer can be shown in Figure 7F. However, Figure 7F is only an example and this application does not limit this. Alternatively, the structure of the TAM layer can be shown in Figure 7G. However, Figure 7G is only an example and this application does not limit this.
[0269] In Figures 7B and 7C , the LRB layer is a lightweight residual block (LRB), and the structure of the LRB layer can be seen in Figure 7H . However, Figure 7H is only an example and this application does not limit this. In Figure 7D , the RB layer is a residual block (Residual Block), and the structure of the RB layer can be seen in Figure 7I . However, Figure 7I is only an example and this application does not limit this.
[0270] In the above-mentioned synthetic transformation network, an activation layer using a Relu activation function (such as the Relu activation function in the LRB layer and the RB layer) may be involved. On this basis, the Relu activation function in the synthetic transformation network can be optimized, and the Relu activation function can be replaced with a target activation function using upper and lower clamping (such as the Relu6 activation function), that is, the target activation function using upper and lower clamping is used to perform nonlinear adjustment on the features. Among them, the Relu6 activation function corresponds to an upper threshold and a lower threshold. The activation layer of the synthetic transformation network can also be optimized so that the activation layer uses a certain activation function, and the activation function corresponds to an upper threshold, that is, the upper limit of the activation layer is limited, but the lower limit of the activation layer is not limited.
[0271] The upper threshold corresponding to the Relu6 activation function may be a configured fixed upper threshold, and the lower threshold corresponding to the Relu6 activation function may be a configured fixed lower threshold. Alternatively, the upper threshold corresponding to the Relu6 activation function may be an upper threshold for adaptive learning, and / or the lower threshold corresponding to the Relu6 activation function may be a lower threshold for adaptive learning.
[0272] In the above embodiment, for high-complexity synthetic transformation networks, medium-complexity synthetic transformation networks, and low-complexity synthetic transformation networks, a ResAU layer can be included. The structure of the ResAU layer can be shown in Figure 8A. The ResAU layer can sequentially include a LeakyRelu activation layer (i.e., an activation layer using the LeakyRelu activation function), a grouped convolution layer (Group-Conv, such as a convolution layer with a size of 3*3, the number of input channels is c, the number of output channels is c, the convolution step is s1, and the number of group groups g is 4), and a convolution layer (Conv, such as a convolution layer with a size of 1*1, the number of input channels is c, the number of output channels is c, and the convolution step is s1). In addition, the ResAU layer can also include a skip connection structure, that is, the input features of the ResAU layer are connected to the output features of the convolution layer, and operations such as feature addition, feature multiplication, and feature splicing can be performed. For example, assuming the input feature of the ResAU layer is feature X1, the LeakyRelu activation layer processes feature X1 to obtain feature X2, the group convolution layer Group-Conv processes feature X2 to obtain feature X3, and the convolution layer Conv processes feature X3 to obtain feature X4. Through the skip connection structure, feature X1 and feature X4 can be added together to obtain the output feature of the ResAU layer.
[0273] As can be seen from Figure 8A, the input of the ResAU layer has a residual connection (i.e., a skip connection structure). The input features of the ResAU layer pass through the LeakyRelu function (LeakyRelu activation layer), group convolution (group convolution layer Group-Conv), and 1x1 regular convolution (convolution layer Conv) in sequence. The output is added to the residual connection to obtain the output features of the ResAU layer.
[0274] The LeakyRelu activation layer of the ResAU layer adopts the LeakyRelu activation function. The LeakyRelu activation function can be seen in Figure 8B, that is, the LeakyRelu activation function does not correspond to the lower threshold value, nor does it correspond to the upper threshold value. Obviously, as the input feature increases, the upper limit of the output feature will not be restricted, resulting in the activation layer possibly outputting an output feature with a large value. As the input feature decreases, the lower limit of the output feature will not be restricted, resulting in the activation layer possibly outputting an output feature with a very small value. The above-mentioned LeakyRelu activation function has a high complexity relative to the hardware, resulting in a high complexity on the decoding end.
[0275] On this basis, in one implementation of this embodiment, the activation layer of the ResAU layer can be optimized so that the activation layer adopts a target activation function that is clamped up and down, that is, the activation layer adopts a target activation function that is clamped up and down to perform nonlinear adjustment on the features. For example, the target activation function can be a Relu6 activation function or other activation function, as long as the upper and lower clamping are achieved at the same time. The Relu6 activation function will be used as an example for explanation. Referring to Figure 8C, which is a schematic diagram of the structure of the ResAU layer, the LeakyRelu activation layer is replaced by the Relu6 activation layer, that is, the Relu6 activation function is used for processing.
[0276] The Relu6 activation function (i.e., the target activation function with upper and lower clamping) can be seen in Figure 5D. The Relu6 activation function corresponds to an upper threshold and a lower threshold. As the input feature decreases, the lower limit of the output feature will be restricted so that the output feature of the activation layer is not less than the lower threshold. As the input feature increases, the upper limit of the output feature will be restricted so that the output feature of the activation layer is not greater than the upper threshold. The Relu6 activation function restricts the lower and upper limits so that the output feature is between the lower and upper thresholds, so that the activation layer will not output output features with very large or very small values.
[0277] The upper threshold corresponding to the Relu6 activation function can be a configured fixed upper threshold, that is, a fixed upper threshold is used as the upper threshold corresponding to the Relu6 activation function by default. The lower threshold corresponding to the Relu6 activation function can be a configured fixed lower threshold, that is, a fixed lower threshold is used as the lower threshold corresponding to the Relu6 activation function by default.
[0278] The upper threshold corresponding to the Relu6 activation function can be an upper threshold for adaptive learning, that is, the upper threshold corresponding to the Relu6 activation function is determined through the training process, rather than a fixed upper threshold. The lower threshold corresponding to the Relu6 activation function can be a lower threshold for adaptive learning, that is, the lower threshold corresponding to the Relu6 activation function is determined through the training process, rather than a fixed lower threshold. Alternatively, the upper threshold corresponding to the Relu6 activation function can be an upper threshold for adaptive learning, and the lower threshold corresponding to the Relu6 activation function can be a fixed lower threshold. Alternatively, the lower threshold corresponding to the Relu6 activation function can be a lower threshold for adaptive learning, and the upper threshold corresponding to the Relu6 activation function can be a fixed upper threshold.
[0279] When adaptively learning the upper threshold corresponding to the Relu6 activation function, the same upper threshold can be learned for all channels of the activation layer. Only one upper threshold needs to be adaptively learned, and this upper threshold serves as the upper threshold for all channels of the activation layer. Alternatively, when adaptively learning the upper threshold corresponding to the Relu6 activation function, a separate upper threshold can be learned for each channel of the activation layer. Assuming that the activation layer corresponds to K channels, K upper thresholds need to be adaptively learned. The K upper thresholds correspond one-to-one to the K channels, and each channel corresponds to a separate upper threshold. The upper thresholds corresponding to different channels can be the same or different.
[0280] When adaptively learning the lower threshold corresponding to the Relu6 activation function, the same lower threshold can be learned for all channels of the activation layer. Only one lower threshold needs to be adaptively learned, and this lower threshold serves as the lower threshold for all channels of the activation layer. Alternatively, when adaptively learning the lower threshold corresponding to the Relu6 activation function, a separate lower threshold can be learned for each channel of the activation layer. Assuming that the activation layer corresponds to K channels, K lower thresholds need to be adaptively learned. The K lower thresholds correspond one-to-one to the K channels, and each channel corresponds to a separate lower threshold. The lower thresholds corresponding to different channels can be the same or different.
[0281] When adaptively learning the upper and lower thresholds corresponding to the Relu6 activation function, the Relu6 activation function can also be called the ClipRelu6 activation function. When the ClipRelu6 activation function is extended to the channel level, that is, when the upper and lower thresholds are learned separately for each channel, the Relu6 activation function can also be called the chs-wise ClipRelu6 activation function.
[0282] In another implementation of this embodiment, the activation layer of the ResAU layer is optimized so that the activation layer uses the Relu activation function. That is, the activation layer uses the Relu activation function to perform nonlinear adjustment of features. Referring to FIG8D , which is a schematic diagram of the structure of the ResAU layer, the LeakyRelu activation layer is replaced by a Relu activation layer, i.e., the Relu activation function is used for processing.
[0283] The Relu activation function can be shown in Figure 5B. The Relu activation function can correspond to a lower threshold. As the input feature decreases, the lower limit of the output feature is restricted so that the output feature of the activation layer is not less than the lower threshold. By restricting the lower limit, the Relu activation function prevents the activation layer from outputting output features with very small values.
[0284] The lower threshold corresponding to the Relu activation function can be a configured fixed lower threshold, that is, a fixed lower threshold is used as the lower threshold corresponding to the Relu activation function by default. Alternatively, the lower threshold corresponding to the Relu activation function can be an adaptively learned lower threshold, that is, the lower threshold corresponding to the Relu activation function is determined through the training process rather than a fixed lower threshold.
[0285] When adaptively learning the lower threshold corresponding to the Relu activation function, the same lower threshold can be learned for all channels of the activation layer. Only one lower threshold needs to be adaptively learned, and this lower threshold serves as the lower threshold for all channels of the activation layer. Alternatively, when adaptively learning the lower threshold corresponding to the Relu activation function, a separate lower threshold can be learned for each channel of the activation layer. Assuming that the activation layer corresponds to K channels, K lower thresholds need to be adaptively learned. The K lower thresholds correspond one-to-one to the K channels, and each channel corresponds to a separate lower threshold. The lower thresholds corresponding to different channels can be the same or different.
[0286] In another implementation of this embodiment, the activation layer of the ResAU layer is optimized so that the activation layer adopts an activation function as shown in Figure 5E, that is, the activation layer adopts the activation function to perform nonlinear adjustment on the features. The activation function may correspond to an upper threshold. As the input features increase, the upper limit of the output features will be limited so that the output features of the activation layer are not greater than the upper threshold. The upper threshold corresponding to the activation function may be a configured fixed upper threshold, that is, a fixed upper threshold is used as the upper threshold corresponding to the activation function by default. Alternatively, the upper threshold corresponding to the activation function may be an upper threshold for adaptive learning, that is, the upper threshold corresponding to the activation function is determined through a training process.
[0287] When adaptively learning the upper threshold corresponding to the activation function, the same upper threshold can be learned for all channels of the activation layer. Only one upper threshold needs to be adaptively learned, and this upper threshold serves as the upper threshold for all channels of the activation layer. Alternatively, when adaptively learning the upper threshold corresponding to the activation function, a separate upper threshold can be learned for each channel of the activation layer. Assuming that the activation layer corresponds to K channels, K upper thresholds need to be adaptively learned. The K upper thresholds correspond one-to-one to the K channels, and each channel corresponds to a separate upper threshold. The upper thresholds corresponding to different channels may be the same or different.
[0288] As shown in Figures 8C and 8D , the ResAU layer (i.e., the ResAU network layer) can include an activation layer (using a Relu6 activation function or a Relu activation function), a processing layer, and a skip connection structure. The skip connection structure (i.e., residual connection) is used to connect the input features of the ResAU layer with the output features of the processing layer, and can perform operations such as feature addition, feature multiplication, and feature concatenation. Of course, it can also be a skip connection structure at other locations, such as a skip connection structure used to connect the output features of the activation layer with the output features of the processing layer. There is no restriction on the location of this skip connection structure, and it can be a skip connection structure at any location.
[0289] Exemplarily, the ResAU layer includes at least an activation layer and a skip connection structure. In addition to the activation layer and the skip connection structure, the ResAU layer may or may not include a processing layer. When a processing layer is included, the ResAU layer may sequentially include an activation layer, a processing layer, and a skip connection structure. In Figures 8C and 8D, the processing layer is included as an example. In addition to the network structures of Figures 8C and 8D, the network structure of the ResAU layer can also be seen in Figures 8E, 8F, 8G, and 8H.
[0290] In Figure 8E, the ResAU layer may include a convolutional layer (such as a convolutional layer of size 3*3, with c input channels, c output channels, and a convolution step size of s1), an activation layer (using a Relu6 activation function or a Relu activation function), and a skip connection structure. In Figure 8F, the ResAU layer may include a convolutional layer (such as a convolutional layer of size 1*1, with c input channels, c output channels, and a convolution step size of s1), an activation layer (using a Relu6 activation function or a Relu activation function), a convolutional layer (such as a convolutional layer of size 3*3, with c input channels, c output channels, and a convolution step size of s1), and a skip connection structure. In Figure 8G, the ResAU layer may include a convolutional layer (such as a convolutional layer with a size of 3*3, the number of input channels is c, the number of output channels is c, and the convolution step is s1), an activation layer (using Relu6 activation function or Relu activation function), a convolutional layer (such as a convolutional layer with a size of 1*1, the number of input channels is c, the number of output channels is c, and the convolution step is s1), a convolutional layer (such as a convolutional layer with a size of 3*3, the number of input channels is c, the number of output channels is c, and the convolution step is s1) and a skip connection structure. In Figure 8H, the ResAU layer may include a convolutional layer (e.g., a convolutional layer of size 3*3, with c input channels, c output channels, and a convolution step size of s1), an activation layer (using a Relu6 activation function or a Relu activation function), a convolutional layer (e.g., a convolutional layer of size 1*1, with c input channels, c output channels, and a convolution step size of s1), an activation layer (using a Relu6 activation function or a Relu activation function), a convolutional layer (e.g., a convolutional layer of size 3*3, with c input channels, c output channels, and a convolution step size of s1), and a skip connection structure. Of course, the above are just a few examples and are not limiting.
[0291] In a possible embodiment, the processing layer (also called a processing unit) for the ResAU layer may include, but is not limited to, at least one of the following: one or more ordinary convolutional layers, one or more transposed convolutional layers, one or more deformable convolutional layers, one or more depthwise separable convolutional layers, one or more grouped convolutional layers, one or more dilated convolutional layers, and one or more transformer layers. For each convolutional layer in the processing layer (such as an ordinary convolutional layer, a transposed convolutional layer, a deformable convolutional layer, a depthwise separable convolutional layer, a grouped convolutional layer, a dilated convolutional layer, a transformer layer, etc.), a quantized convolution may also be used, that is, the parameters in the convolutional layer are quantized, and the input features are quantized, and then a convolution operation is performed. In other words, each convolutional layer in the processing layer is a quantized convolutional layer (or an integerized convolutional layer, the integerized bit width can be 2, 4, 8, 16, etc.), that is, the parameters in the convolutional layer are integerized parameters.
[0292] Each convolutional layer in the processing layer can be stacked repeatedly, with no restrictions on the stacking structure. For example, you can stack two normal convolutional layers, then two grouped convolutional layers, then deploy a dilated convolutional layer, then deploy a normal convolutional layer, and so on. Of course, this is just an example, and there is no restriction on the structure of this processing layer.
[0293] In one possible implementation, for the encoding framework and decoding framework shown in FIG4A and FIG4B , a Relu6 activation function, or a ClipRelu6 activation function, or a chs-wise ClipRelu6 activation function may be connected to any layer of the probabilistic hyperparameter decoding network. A Relu6 activation function, or a ClipRelu6 activation function, or a chs-wise ClipRelu6 activation function may be connected to any layer of the mean hyperparameter decoding network. A Relu6 activation function, or a ClipRelu6 activation function, or a chs-wise ClipRelu6 activation function may be connected to any layer of the context model. A Relu6 activation function, or a ClipRelu6 activation function, or a chs-wise ClipRelu6 activation function may be connected to any layer of a synthesis transformation network (such as a high-complexity synthesis transformation network, a medium-complexity synthesis transformation network, or a low-complexity synthesis transformation network). For the encoding and decoding frameworks shown in Figures 4C and 4D, a Relu6 activation function, ClipRelu6 activation function, or chs-wise ClipRelu6 activation function can be connected after any layer of the super-parameter decoding network. A Relu6 activation function, ClipRelu6 activation function, or chs-wise ClipRelu6 activation function can be connected after any layer of the context model. A Relu6 activation function, ClipRelu6 activation function, or chs-wise ClipRelu6 activation function can be connected after any layer of the synthesis transformation network (such as the high-complexity synthesis transformation network, the medium-complexity synthesis transformation network, and the low-complexity synthesis transformation network).
[0294] In a possible implementation, for each of the above embodiments, the complexity of the model can be reduced by adopting the Relu6 activation function, the ClipRelu6 activation function, or the chs-wise ClipRelu6 activation function.
[0295] In the above embodiment, for the encoding end and the decoding end, the coefficient hyperparameter feature z_hat can be input into the mean hyperparameter decoding network to obtain the initial mean feature m corresponding to the current image block, that is, the input of the mean hyperparameter decoding network is the decoding information of the first code stream (such as the coefficient hyperparameter feature z_hat), and the output of the mean hyperparameter decoding network is the mean mu (that is, the initial mean feature m is used as the target mean feature mu), or, the output of the mean hyperparameter decoding network is the potential feature m of the mean (that is, the initial mean feature m needs to be input into the context model, and the context model outputs the target mean feature mu).
[0296] In one possible implementation, the mean hyperparameter decoding network can be upsampled first and then convolution processed. The complexity of this form of mean hyperparameter decoding network is relatively high, resulting in a higher complexity at the decoding end. In view of the above findings, in this embodiment, the mean hyperparameter decoding network can be simplified (optimized). The mean hyperparameter decoding network may include a processing layer, an upsampling layer, and a Crop layer in sequence. The upsampling layer is located behind the processing layer, and the upsampling layer is located in front of the Crop layer. Since the upsampling layer is deployed behind the processing layer, the mean hyperparameter decoding network can be processed (such as convolution processing, etc.) before upsampling. The complexity of this form of mean hyperparameter decoding network is relatively low, which reduces the complexity of the decoding end.
[0297] Based on the mean hyperparameter decoding network, the coefficient hyperparameter feature z_hat can be input into the processing layer of the mean hyperparameter decoding network. The coefficient hyperparameter feature z_hat is processed by the processing layer (such as convolution processing and / or activation processing, etc.) to obtain the processed feature, and the processed feature is input into the upsampling layer of the mean hyperparameter decoding network. The processed feature is upsampled by the upsampling layer to obtain the upsampled feature, and the upsampled feature is input into the Crop layer of the mean hyperparameter decoding network. The upsampled feature is cropped and aligned by the Crop layer to obtain the initial mean feature m, which is used as the output feature of the mean hyperparameter decoding network. Exemplarily, the Crop layer is used to crop and align the pixel positions after upsampling, and can be located after the upsampling layer to ensure that the feature map size is aligned. There is no restriction on the processing process of this Crop layer.
[0298] For example, the upsampling layer (also called the sampling layer or Pixshuffle layer) can be located after the processing layer (all network layers except the upsampling layer and the crop layer are divided into the processing layer), that is, the upsampling layer is located at the end of the mean hyperparameter decoding network. The form of the upsampling layer can include but is not limited to pixel rearrangement, pixshuffle, upshuffle, transposed convolution, and nearest neighbor interpolation. There is no restriction on the implementation method of this upsampling layer, as long as it can upsample the input features.
[0299] For example, pixel rearrangement can be used to upsample the processed features (i.e., the output features of the processing layer) to obtain upsampled features. Pixel rearrangement is used to rearrange channel domain information into the spatial domain to achieve an upsampling effect.
[0300] For example, the processed features (i.e., the output features of the processing layer) can be upsampled using the pixshuffle method to obtain upsampled features. Pixshuffle is a spatial operation that rearranges feature points, transforming the three-dimensional representation [4C, H, W] to [C, 2H, 2W], thereby increasing spatial resolution. Figure 9A shows a schematic diagram of upsampling the processed features using the pixshuffle method to obtain upsampled features.
[0301] For example, an upshuffle can be used to upsample the processed features (i.e., the output features of the processing layer) to obtain upsampled features. Upshuffle reorders the feature points, transforming the three-dimensional representation [4C, H, W] to [C, 2H, 2W], thereby increasing spatial resolution. Compared to pixshuffle, upshuffle uses a different permutation criterion. Figure 9B shows an example of upsampling the processed features using the upshuffle method.
[0302] For example, the processed features can be upsampled using a transposed convolution method to obtain upsampled features.
[0303] For example, the processed features can be upsampled using nearest neighbor interpolation (also known as nearest neighbor upsampling) to obtain upsampled features. Nearest neighbor interpolation plays the role of amplifying the spatial resolution of the features. Nearest neighbor interpolation rearranges the three-dimensional representation [C, H, W] to obtain [C, 2H, 2W]. The extra pixels are copied to the nearest point in the spatial domain as the current interpolation point. See Figure 9C for a schematic diagram of the nearest neighbor interpolation method.
[0304] Of course, the above are just a few examples of how to process the upsampling layer, and there is no limitation to this.
[0305] In one possible embodiment, the processing layer (also called a processing unit or processing link) of the mean hyperparameter decoding network includes but is not limited to at least one of the following: one or more ordinary convolution layers, one or more transposed convolution layers, one or more deformable convolution layers, one or more depth separation convolution layers, one or more grouped convolution layers, one or more dilated convolution layers, one or more transformer layers, one or more activation layers using a Relu activation function, one or more activation layers using a target activation function with upper and lower clamping, and one or more activation layers using a LeakyRelu activation function.
[0306] Exemplarily, for each convolution layer in the processing layer (such as a normal convolution layer, a transposed convolution layer, a deformable convolution layer, a depth-separated convolution layer, a grouped convolution layer, an expanded convolution layer, a transformer layer, etc.), a quantized convolution can also be used, that is, the parameters in the convolution layer are quantized, and the input features are quantized, and then the convolution operation is performed. When the parameters in the convolution layer are quantized, the bit width can be 2, 4, 8, 16, etc., and there is no restriction on this. For example, each convolution layer in the processing layer is a quantized convolution layer (or an integerized convolution layer, the integerized bit width can be 2, 4, 8, 16, etc.), that is, the parameters in the convolution layer are integerized parameters. For example, the parameters in the normal convolution layer are quantized parameters, the parameters in the transposed convolution layer are quantized parameters, the parameters in the grouped convolution layer are quantized parameters, and so on.
[0307] Each convolutional layer in the processing layer can be stacked repeatedly, with no restrictions on the stacking structure. For example, you can stack two normal convolutional layers, then two grouped convolutional layers, then deploy a dilated convolutional layer, then deploy a normal convolutional layer, and so on. Of course, this is just an example, and there is no restriction on the structure of this processing layer.
[0308] For each layer in the processing layer, there is an arbitrary jump connection (i.e., jump connection structure, also called residual connection) between each layer of the processing layer. The jump connection structure is used to connect the input features of a certain network layer (such as any network layer) with the output features of another network layer (such as any network layer), and can perform operations such as feature addition, feature multiplication, and feature splicing.
[0309] If the processing layer includes convolutional layers and activation layers, the convolutional layers and activation layers can be interleaved, and the convolutional layers or activation layers can be repeatedly stacked. Any input node and output node (that is, any network layer) from the front to the back can be added through residual connections (that is, there are any skip connections between the layers of the processing layer, and they can be added through skip connections). For example, the input features of the first network layer are added to the output features of the fifth network layer, the input features of the second network layer are added to the output features of the fifth network layer, the input features of the third network layer are added to the output features of the sixth network layer, and so on.
[0310] In one possible implementation, the structure of the mean hyperparameter decoding network (i.e., the optimized mean hyperparameter decoding network) can be shown in FIG9D . However, FIG9D is only an example of a mean hyperparameter decoding network, and this application does not limit the structure of this mean hyperparameter decoding network. The mean hyperparameter decoding network of FIG9D will be used as an example in the following. The mean hyperparameter decoding network may include a processing layer, an upsampling layer, and a Crop layer in sequence, and the processing layer may include a convolutional layer (Conv, such as a convolutional layer with a size of 3*3, the number of input channels is c, and the number of output channels is c), a Relu activation layer (using a Relu activation function), a convolutional layer (Conv, such as a convolutional layer with a size of 3*3, the number of input channels is c, and the number of output channels is c), a Relu activation layer (using a Relu activation function), a grouped convolutional layer (gConv, such as a convolutional layer with a size of 3*3, the number of input channels is c, the number of output channels is 4c, and the number of grouping groups is 4), and a convolutional layer (such as a convolutional layer with a size of 1*1, the number of input channels is 4c, and the number of output channels is c).
[0311] The Relu activation layer in Figure 9D can also be replaced with a Relu6 activation layer (using the Relu6 activation function), a ClipRelu6 activation layer (using the ClipRelu6 activation function), or a chs-wise ClipRelu6 activation layer (using the chs-wise ClipRelu6 activation function). For example, any Relu activation layer can be replaced with a Relu6 activation layer, or both Relu activation layers can be replaced with a Relu6 activation layer.
[0312] In the above embodiments, ordinary convolution layers, transposed convolution layers, deformable convolution layers, depthwise separation convolution layers, grouped convolution layers, expanded convolution layers, transformer layers, etc. may be involved.
[0313] For ordinary convolution layers, ordinary convolution layers can also be called conventional convolution layers - Convolution (Conv). Ordinary convolution layers start with a small weight matrix, that is, the convolution kernel (kernel), and let it gradually "scan" on the two-dimensional input data. While the convolution kernel "slides", the product of the weight matrix and the scanned data matrix is calculated, and then the result is summarized into an output pixel. Referring to Figure 10A, it is a schematic diagram of an ordinary convolution layer. Figure 10A shows the convolution kernel matrix, input data, and output data, and the dotted white part is the filling value of the input data. For ordinary convolution layers, the convolution kernel can be operated from left to right and from top to bottom in the form of a sliding window. The stride parameter determines the step size of each slide.
[0314] For the transposed convolution layer - Transposed Convolution (Tconv), for some tasks, the data needs to be upsampled. Conventional convolution operations can only keep the output feature map and the input feature map resolution equal or similar, and cannot achieve the purpose of upsampling. Therefore, transposed convolution can be used. Unlike conventional convolution, the stride parameter in the transposed convolution does not indicate the sliding step size, but the number of 0s filled between the input data. For example, when stride = 1, the transposed convolution is consistent with the conventional convolution. When stride = 2, see Figure 10B, which is a schematic diagram of the transposed convolution layer. The input feature map can be inserted, and (stride-1) 0s will be inserted between each feature point in the spatial domain. After expanding the input spatial domain resolution, it is calculated with the convolution weight.
[0315] Regarding the dilated convolution layer (Dconv), the dilated convolution layer can also be called a dilated convolution layer or an expanded convolution layer. This refers to inserting holes between the points of the normal convolution kernel, which is relative to the normal discrete convolution. For example, for a normal convolution with a stride of 2 and a padding of 1, see Figure 10C, which is a schematic diagram of the dilated convolution layer, which is a dilation insertion of the convolution.
[0316] For the group convolution layer - Group-Convolution (gconv), the size of the input x is [C, H, W]. The group convolution operation will split x into channels and divide it into groups of the same number as group. Each group is convolved separately to obtain the output of each group, and then the output of each group is channel-joined. For example, see Figure 10D, which is a schematic diagram of the group convolution layer.
[0317] For the depth separation convolution layer, the depth separation convolution layer can be divided into two stages. The first stage is spatial separability and the second stage is depth separability. Spatial separability is expressed as: the input [C, H, W] is divided into C groups, and the data of each group is [1, H, W]. Each group is convolved separately to obtain the output and then spliced. For example, see Figure 10E, which shows a schematic diagram of the spatial separability of the depth separation convolution layer. The second stage is the depth separable stage, which is used to operate the output of the first stage with a convolution kernel with a kernel size of 1x1. For example, see Figure 10E, which shows a schematic diagram of the depth separability of the depth separation convolution layer.
[0318] Regarding the transformer layer, the transformer layer is a transformer-based attention block (also known as a TAM layer). The structure of the transformer layer can be seen in Figure 7F or Figure 7G, and will not be repeated here.
[0319] For each convolution layer in the processing layer (such as ordinary convolution layer, transposed convolution layer, deformable convolution layer, depthwise separation convolution layer, grouped convolution layer, dilated convolution layer, transformer layer, etc.), quantized convolution (quantized convolution can also be called integerized convolution) can also be used, that is, the parameters in the convolution layer are quantized. This process can also be called quantized convolution. All of the above-mentioned operation forms can be quantized. Quantized convolution is not a specific form of convolution. It means that the weights inside the convolution are quantized to a certain bit width. For example, the original weight data of the convolution kernel is represented by float for floating point. The weight data is scaled, offset and integerized to ensure that the data can be represented by 4 bits, 8 bits, or 16 bits.
[0320] Illustratively, the above embodiments may be implemented individually or in combination. For example, each of the above embodiments may be implemented individually, or at least two of them may be implemented in combination.
[0321] Illustratively, in the above embodiments, the content of the encoding end can also be applied to the decoding end, that is, the decoding end can be processed in the same way, and the content of the decoding end can also be applied to the encoding end, that is, the encoding end can be processed in the same way.
[0322] Based on the same concept as the above method, an embodiment of the present application also proposes a decoding device, which is applied to the decoding end, and the device includes: a memory, which is configured to store video data; a decoder, which is configured to implement the decoding method in the above embodiments, that is, the processing flow of the decoding end.
[0323] Based on the same concept as the above method, an encoding device is also proposed in an embodiment of the present application. The device is applied to the encoding end, and the device includes: a memory, which is configured to store video data; an encoder, which is configured to implement the encoding method in the above embodiments, that is, the processing flow of the encoding end.
[0324] Based on the same concept as the above-mentioned method, the decoding end device (also referred to as a video decoder) provided in the embodiments of the present application, from a hardware perspective, has a hardware architecture diagram as shown in FIG11A . As shown in FIG11A , the decoding end device 10 includes a processor 111 and a machine-readable storage medium 112 . The machine-readable storage medium 112 stores machine-executable instructions that can be executed by the processor 111 . The processor 111 is configured to execute the machine-executable instructions to implement the decoding methods of the above-mentioned embodiments of the present application.
[0325] Based on the same concept as the above-mentioned method, the encoding end device (also referred to as a video encoder) provided in the embodiments of the present application, from a hardware perspective, has a hardware architecture diagram specifically shown in FIG11B . As shown in FIG11B , the encoding end device 20 includes a processor 113 and a machine-readable storage medium 114 . The machine-readable storage medium 114 stores machine-executable instructions that can be executed by the processor 113 . The processor 113 is configured to execute the machine-executable instructions to implement the encoding methods of the above-mentioned embodiments of the present application.
[0326] Based on the same concept as the above-mentioned method, an embodiment of the present application provides an electronic device. The electronic device may include a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is configured to execute the machine-executable instructions to implement the decoding method or encoding method of each of the above-mentioned embodiments of the present application.
[0327] The aforementioned machine-readable storage medium, such as the machine-readable storage medium 112 in the decoding end device 10, the machine-readable storage medium 114 in the encoding end device 20, and the machine-readable storage medium in the electronic device, can be any electronic, magnetic, optical, or other physical storage device that can contain or store information, such as executable instructions, data, and the like. For example, the machine-readable storage medium can be a random access memory (RAM), a volatile memory, a non-volatile memory, a flash memory, a storage drive (such as a hard disk drive), a solid-state drive, any type of storage disk (such as a CD, DVD, etc.), or similar storage media, or a combination thereof.
[0328] Based on the same concept as the above method, an embodiment of the present application also provides a machine-readable storage medium, on which a number of computer instructions are stored. When the computer instructions are executed by a processor, the method disclosed in the above examples of the present application can be implemented, such as the decoding method or encoding method in the above embodiments.
[0329] Based on the same concept as the above method, an embodiment of the present application also provides a computer program, which, when executed by a processor, can implement the decoding method or encoding method disclosed in the above example of the present application.
[0330] Based on the same concept as the above method, an embodiment of the present application also proposes a decoding device, which can be applied to a decoding end (also called a video decoder), and the decoding device includes: a decoding module, which is used to decode the first code stream corresponding to the current image block to obtain the coefficient hyperparameter feature corresponding to the current image block; determine the probability distribution parameter based on the coefficient hyperparameter feature, and decode the second code stream corresponding to the current image block based on the probability distribution parameter to obtain the residual feature corresponding to the current image block; a determination module, which is used to determine the target mean feature corresponding to the current image block based on the coefficient hyperparameter feature; determine the reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature; an acquisition module, which is used to input the reconstruction feature into a synthetic transformation network to obtain the reconstructed image block corresponding to the current image block; wherein the synthetic transformation network includes a nonlinear residual network layer, and the nonlinear residual network layer includes at least an activation layer and a skip connection structure; wherein the output feature of the activation layer is not greater than an upper threshold, and / or the output feature of the activation layer is not less than a lower threshold.
[0331] Exemplarily, the activation layer uses an upper and lower clamped target activation function to perform nonlinear adjustment on the features, the target activation function corresponds to an upper threshold and a lower threshold, and the output feature of the activation layer is not greater than the upper threshold and not less than the lower threshold; or, the activation layer uses a Relu activation function to perform nonlinear adjustment on the features, the Relu activation function corresponds to the lower threshold, and the output feature of the activation layer is not less than the lower threshold.
[0332] Exemplarily, the nonlinear residual network layer includes an activation layer, a processing layer and a skip connection structure in sequence, and the processing layer includes at least one of the following: one or more ordinary convolution layers, one or more transposed convolution layers, one or more deformable convolution layers, one or more depth separation convolution layers, one or more grouped convolution layers, one or more expanded convolution layers, and one or more global processing unit layers.
[0333] Exemplarily, when the decoding module determines the probability distribution parameters based on the coefficient hyperparameter features, it is specifically used to: input the coefficient hyperparameter features into the probability hyperparameter decoding network to obtain the probability distribution parameters; wherein, the probability hyperparameter decoding network includes an activation layer, and the activation layer uses an upper and lower clamped target activation function to perform nonlinear adjustment on the features; wherein, the target activation function corresponds to an upper threshold and a lower threshold, the output feature of the activation layer is not greater than the upper threshold, and the output feature of the activation layer is not less than the lower threshold.
[0334] Exemplarily, when the determination module determines the target mean feature corresponding to the current image block based on the coefficient hyperparameter feature, it is specifically used to: input the coefficient hyperparameter feature into the mean hyperparameter decoding network to obtain the initial mean feature corresponding to the current image block, and determine the target mean feature based on the initial mean feature; wherein, the mean hyperparameter decoding network includes an activation layer, and the activation layer uses an upper and lower clamped target activation function to perform nonlinear adjustment on the feature; wherein, the target activation function corresponds to an upper threshold and a lower threshold, and the output feature of the activation layer is not greater than the upper threshold and not less than the lower threshold.
[0335] Exemplarily, when the determination module determines the target mean feature based on the initial mean feature, it is specifically used to: use the initial mean feature as the target mean feature; or input the initial mean feature and the obtained reconstructed feature of the current image block into a context model to obtain the target mean feature; wherein, the context model includes an activation layer, and the activation layer uses a lower-clamped target activation function (such as a Relu activation function) to perform nonlinear adjustment on the feature; wherein, the target activation function corresponds to a lower limit threshold, and the output feature of the activation layer is not less than the lower limit threshold.
[0336] Exemplarily, the determination module determines the probability distribution parameters based on the coefficient hyperparameter features, and determines the target mean features corresponding to the current image block based on the coefficient hyperparameter features, and is specifically used to: input the coefficient hyperparameter features into the hyperparameter decoding network to obtain the reference features corresponding to the current image block; input the reference features and the obtained reconstructed features of the current image block into the context model to obtain the probability distribution parameters and the target mean features corresponding to the current image block; wherein the hyperparameter decoding network includes an activation layer, and the activation layer uses an upper and lower clamped target activation function to perform nonlinear adjustment on the features, wherein the target activation function corresponds to an upper threshold and a lower threshold, and the output features of the activation layer are not greater than the upper threshold and not less than the lower threshold; and / or, the context model includes an activation layer, and the activation layer uses an upper and lower clamped target activation function to perform nonlinear adjustment on the features, wherein the target activation function corresponds to an upper threshold and a lower threshold, and the output features of the activation layer are not greater than the upper threshold and not less than the lower threshold.
[0337] Exemplarily, the upper threshold corresponding to the target activation function is a configured fixed upper threshold, and the lower threshold corresponding to the target activation function is a configured fixed lower threshold; or, the upper threshold corresponding to the target activation function is the upper threshold of adaptive learning, and the lower threshold corresponding to the target activation function is the lower threshold of adaptive learning.
[0338] Exemplarily, when adaptively learning the upper threshold, the same upper threshold is learned for all channels of the activation layer, or an upper threshold is learned separately for each channel of the activation layer, and the upper thresholds corresponding to different channels are the same or different; when adaptively learning the lower threshold, the same lower threshold is learned for all channels of the activation layer, or a lower threshold is learned separately for each channel of the activation layer, and the lower thresholds corresponding to different channels are the same or different.
[0339] Based on the same concept as the above method, the embodiment of the present application also proposes a decoding device, which can be applied to a decoding end (also called a video decoder). The decoding device includes: a decoding module for decoding a first code stream corresponding to a current image block to obtain coefficient hyperparameter features corresponding to the current image block; inputting the coefficient hyperparameter features into a probability hyperparameter decoding network to obtain probability distribution parameters, and decoding the second code stream corresponding to the current image block based on the probability distribution parameters to obtain residual features corresponding to the current image block; a determination module for inputting the coefficient hyperparameter features into a mean hyperparameter decoding network to obtain an initial mean feature corresponding to the current image block. Features, and based on the initial mean features, determine the target mean features corresponding to the current image block; based on the target mean features and the residual features, determine the reconstruction features corresponding to the current image block; an acquisition module is used to determine the reconstructed image block corresponding to the current image block based on the reconstruction features; wherein the probabilistic hyperparameter decoding network includes an activation layer, the output features of the activation layer are not greater than the upper threshold, and / or, the output features of the activation layer are not less than the lower threshold; and / or, the mean hyperparameter decoding network includes an activation layer, the output features of the activation layer are not greater than the upper threshold, and / or, the output features of the activation layer are not less than the lower threshold.
[0340] Exemplarily, the activation layer of the probabilistic hyperparameter decoding network and / or the activation layer of the mean hyperparameter decoding network uses an upper and lower clamped target activation function to perform nonlinear adjustment on the features, the target activation function corresponds to an upper threshold and a lower threshold, and the output feature of the activation layer is not greater than the upper threshold and not less than the lower threshold; wherein the upper threshold corresponding to the target activation function is a configured fixed upper threshold, and the lower threshold corresponding to the target activation function is a configured fixed lower threshold; or, the upper threshold corresponding to the target activation function is an upper threshold for adaptive learning, and the lower threshold corresponding to the target activation function is a lower threshold for adaptive learning; wherein, when adaptively learning the upper threshold, the same upper threshold is learned for all channels of the activation layer, or, the upper threshold is learned separately for each channel of the activation layer, and the upper thresholds corresponding to different channels are the same or different; when adaptively learning the lower threshold, the same lower threshold is learned for all channels of the activation layer, or, the lower threshold is learned separately for each channel of the activation layer, and the lower thresholds corresponding to different channels are the same or different.
[0341] Based on the same application concept as the above method, the embodiment of the present application also proposes a decoding device, which can be applied to a decoding end (also called a video decoder), and the decoding device includes: a decoding module for decoding a first code stream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block; a determination module for inputting the coefficient hyperparameter feature into a hyperparameter decoding network to obtain a reference feature corresponding to the current image block; inputting the reference feature and the obtained reconstructed feature of the current image block into a context model to obtain probability distribution parameters and a target mean feature corresponding to the current image block; the decoding module is also used to obtain a target mean feature corresponding to the current image block based on the probability distribution parameter. The second code stream corresponding to the current image block is decoded to obtain the residual feature corresponding to the current image block; the determination module is further used to determine the reconstruction feature based on the target mean feature and the residual feature; the acquisition module is used to determine the reconstructed image block corresponding to the current image block based on the reconstruction feature; wherein the hyperparameter decoding network includes an activation layer, the output feature of the activation layer is not greater than the upper threshold, and / or, the output feature of the activation layer is not less than the lower threshold; and / or, the context model includes an activation layer, the output feature of the activation layer is not greater than the upper threshold, and / or, the output feature of the activation layer is not less than the lower threshold.
[0342] Exemplarily, the activation layer uses an upper and lower clamped target activation function to perform nonlinear adjustment on the features, the target activation function corresponds to an upper threshold and a lower threshold, the output feature of the activation layer is not greater than the upper threshold, and the output feature of the activation layer is not less than the lower threshold; wherein, the upper threshold corresponding to the target activation function is a configured fixed upper threshold, and the lower threshold corresponding to the target activation function is a configured fixed lower threshold; or, the upper threshold corresponding to the target activation function is an upper threshold of adaptive learning, and the lower threshold corresponding to the target activation function is a lower threshold of adaptive learning; wherein, when adaptively learning the upper threshold, the same upper threshold is learned for all channels of the activation layer, or, the upper threshold is learned separately for each channel of the activation layer, and the upper thresholds corresponding to different channels are the same or different; when adaptively learning the lower threshold, the same lower threshold is learned for all channels of the activation layer, or, the lower threshold is learned separately for each channel of the activation layer, and the lower thresholds corresponding to different channels are the same or different.
[0343] Based on the same concept as the above method, an embodiment of the present application also proposes a decoding device, which can be applied to a decoding end (also called a video decoder), and the decoding device includes: a decoding module, which is used to decode a first code stream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block; determine a probability distribution parameter based on the coefficient hyperparameter feature, and decode a second code stream corresponding to the current image block based on the probability distribution parameter to obtain a residual feature corresponding to the current image block; a determination module, which is used to input the coefficient hyperparameter feature into a mean hyperparameter decoding network to obtain an initial mean feature corresponding to the current image block, and determine a target mean feature corresponding to the current image block based on the initial mean feature; wherein the mean hyperparameter decoding network includes a first processing layer, an upsampling layer for performing an upsampling operation, a cropping layer, and a second processing layer in sequence; wherein the feature output by the second processing layer is used as the initial mean feature; determine the reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature; and an acquisition module, which is used to determine the reconstructed image block corresponding to the current image block based on the reconstruction feature.
[0344] Exemplarily, the first processing layer does not include an upsampling layer for performing an upsampling operation, and the first processing layer includes at least one of the following: one or more ordinary convolution layers, one or more transposed convolution layers, one or more deformable convolution layers, one or more depth-separated convolution layers, one or more grouped convolution layers, one or more expanded convolution layers, one or more global processing unit layers, one or more activation layers using a Relu activation function, one or more activation layers using a target activation function with upper and lower clamping, and one or more activation layers using a LeakyRelu activation function; wherein, there are arbitrary jump connections between the layers of the first processing layer.
[0345] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. The present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. The embodiments of the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The foregoing is merely an embodiment of the present application and is not intended to limit the present application.
[0346] For those skilled in the art, various modifications and variations are possible in this application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this application should be included in the scope of the claims of this application.
Claims
1. An image decoding method, characterized in that, Performed at the decoding end, the method includes: Decoding a first bitstream corresponding to a current image block to obtain coefficient hyperparameter features corresponding to the current image block; Determining probability distribution parameters and target mean features corresponding to the current image block based on the coefficient hyperparameter features; Decoding a second bitstream corresponding to the current image block based on the probability distribution parameters to obtain residual features corresponding to the current image block; Determining reconstructed features corresponding to the current image block based on the target mean features and the residual features; Inputting the reconstructed features into a synthesis transformation network to obtain a reconstructed image block corresponding to the current image block; wherein, the synthesis transformation network includes a non-linear residual network layer, and the non-linear residual network layer at least includes an activation layer and a skip connection structure; the output features of the activation layer are not greater than an upper limit threshold, and / or, the output features of the activation layer are not less than a lower limit threshold.
2. The method according to claim 1, wherein The activation layer uses a target activation function with upper and lower clamping to non-linearly adjust the features, the target activation function corresponding to the upper limit threshold and the lower limit threshold, and the output features of the activation layer are not greater than the upper limit threshold and not less than the lower limit threshold; or, The activation layer uses a Relu activation function to non-linearly adjust the features, the Relu activation function corresponding to the lower limit threshold, and the output features of the activation layer are not less than the lower limit threshold.
3. The method according to claim 1, wherein The non-linear residual network layer sequentially includes an activation layer, a processing layer, and a skip connection structure, and the processing layer includes at least one of the following: one or more ordinary convolutional layers, one or more transposed convolutional layers, one or more deformable convolutional layers, one or more depthwise separable convolutional layers, one or more grouped convolutional layers, one or more dilated convolutional layers, one or more global processing unit layers.
4. The method according to claim 1, characterized in that, The determining the target mean features corresponding to the current image block based on the coefficient hyperparameter features includes: inputting the coefficient hyperparameter features into a mean hyperparameter decoding network to obtain initial mean features corresponding to the current image block, and determining the target mean features based on the initial mean features; Wherein, the mean hyperparameter decoding network includes an activation layer, and the activation layer uses a target activation function with upper and lower clamping to non-linearly adjust the features; wherein, the target activation function corresponds to the upper limit threshold and the lower limit threshold, and the output features of the activation layer are not greater than the upper limit threshold and not less than the lower limit threshold.
5. The method according to claim 4, characterized in that The determining the target mean features based on the initial mean features includes: using the initial mean features as the target mean features; or, inputting the initial mean features and the already obtained reconstructed features of the current image block into a context model to obtain the target mean features; Wherein, the context model includes an activation layer, and the activation layer uses a target activation function with lower clamping to non-linearly adjust the features; wherein, the target activation function corresponds to the lower limit threshold, and the output features of the activation layer are not less than the lower limit threshold.
6. The method according to claim 1, characterized in that The determining the probability distribution parameters and the target mean features corresponding to the current image block based on the coefficient hyperparameter features includes: Input the coefficient hyperparameter feature into the hyperparameter decoding network to obtain the reference feature corresponding to the current image block; Input the reference feature and the obtained reconstruction feature of the current image block into the context model to obtain the probability distribution parameter and the target mean feature corresponding to the current image block; Wherein, the hyperparameter decoding network includes an activation layer, and the activation layer uses a target activation function with upper and lower clamping to non-linearly adjust the feature. Wherein, the target activation function corresponds to an upper threshold and a lower threshold, and the output feature of the activation layer is not greater than the upper threshold and not less than the lower threshold; and / or, the context model includes an activation layer, and the activation layer uses a target activation function with upper and lower clamping to non-linearly adjust the feature. Wherein, the target activation function corresponds to an upper threshold and a lower threshold, and the output feature of the activation layer is not greater than the upper threshold and not less than the lower threshold.
7. The method according to any one of claims 2 or 4 - 6, wherein The upper threshold corresponding to the target activation function is a configured fixed upper threshold, and the lower threshold corresponding to the target activation function is a configured fixed lower threshold; or, the upper threshold corresponding to the target activation function is an adaptively learned upper threshold, and the lower threshold corresponding to the target activation function is an adaptively learned lower threshold.
8. The method according to claim 7, wherein When adaptively learning the upper threshold, learn the same upper threshold for all channels of the activation layer, or learn the upper threshold separately for each channel of the activation layer, and the upper thresholds corresponding to different channels are the same or different; When adaptively learning the lower threshold, learn the same lower threshold for all channels of the activation layer, or learn the lower threshold separately for each channel of the activation layer, and the lower thresholds corresponding to different channels are the same or different.
9. An image decoding method, characterized in that, Executed by the decoding end, the method includes: Decode the first bitstream corresponding to the current image block to obtain the coefficient hyperparameter feature corresponding to the current image block; Input the coefficient hyperparameter feature into the probability hyperparameter decoding network to obtain the probability distribution parameter, and decode the second bitstream corresponding to the current image block based on the probability distribution parameter to obtain the residual feature corresponding to the current image block; Input the coefficient hyperparameter feature into the mean hyperparameter decoding network to obtain the initial mean feature corresponding to the current image block, and determine the target mean feature corresponding to the current image block based on the initial mean feature; Determine the reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature; Determine the reconstructed image block corresponding to the current image block based on the reconstruction feature; Wherein, the probability hyperparameter decoding network includes an activation layer, and the output feature of the activation layer is not greater than the upper threshold, and / or, the output feature of the activation layer is not less than the lower threshold; and / or, the mean hyperparameter decoding network includes an activation layer, and the output feature of the activation layer is not greater than the upper threshold, and / or, the output feature of the activation layer is not less than the lower threshold.
10. The method according to claim 9, characterized in that, The activation layer of the probability hyperparameter decoding network and / or the activation layer of the mean hyperparameter decoding network non-linearly adjusts the features using a target activation function with upper and lower clamping. The target activation function corresponds to an upper threshold and a lower threshold, and the output features of the activation layer are not greater than the upper threshold and not less than the lower threshold; Among them, the upper threshold corresponding to the target activation function is a configured fixed upper threshold, and the lower threshold corresponding to the target activation function is a configured fixed lower threshold; or, the upper threshold corresponding to the target activation function is an adaptively learned upper threshold, and the lower threshold corresponding to the target activation function is an adaptively learned lower threshold; Among them, when adaptively learning the upper threshold, the same upper threshold is learned for all channels of the activation layer, or, the upper threshold is learned separately for each channel of the activation layer, and the upper thresholds corresponding to different channels are the same or different; When adaptively learning the lower threshold, the same lower threshold is learned for all channels of the activation layer, or, the lower threshold is learned separately for each channel of the activation layer, and the lower thresholds corresponding to different channels are the same or different.
11. An image decoding method, characterized in that, Executed by the decoding end, the method includes: Decoding the first bitstream corresponding to the current image block to obtain the coefficient hyperparameter features corresponding to the current image block; Determining probability distribution parameters based on the coefficient hyperparameter features, and decoding the second bitstream corresponding to the current image block based on the probability distribution parameters to obtain the residual features corresponding to the current image block; Inputting the coefficient hyperparameter features into the mean hyperparameter decoding network to obtain the initial mean features corresponding to the current image block, and determining the target mean features corresponding to the current image block based on the initial mean features; among them, the mean hyperparameter decoding network sequentially includes a first processing layer, an upsampling layer for performing upsampling operations, a cropping layer, and a second processing layer; among them, the features output by the second processing layer are used as the initial mean features; Determining the reconstructed features corresponding to the current image block based on the target mean features and the residual features; Determining the reconstructed image block corresponding to the current image block based on the reconstructed features.
12. The method according to claim 11, wherein The first processing layer does not include an upsampling layer for performing upsampling operations, and the first processing layer includes at least one of the following: One or more ordinary convolutional layers, one or more transposed convolutional layers, one or more deformable convolutional layers, one or more depthwise separable convolutional layers, one or more grouped convolutional layers, one or more dilated convolutional layers, one or more global processing unit layers, one or more activation layers using the Relu activation function, one or more activation layers using the target activation function with upper and lower clamping, one or more activation layers using the LeakyRelu activation function; Among them, there are arbitrary skip connections between the layers of the first processing layer.
13. The method according to claim 11, wherein The second processing layer includes at least one of the following: one or more ordinary convolutional layers, one or more deformable convolutional layers, one or more depthwise separable convolutional layers, one or more grouped convolutional layers, one or more dilated convolutional layers, one or more global processing unit layers, one or more activation layers using the Relu activation function, one or more activation layers using a target activation function with upper and lower clamping, one or more activation layers using the LeakyRelu activation function; Among them, there are arbitrary skip connections between the layers of the second processing layer.
14. An image decoding method, characterized in that, The method includes: Obtaining a coefficient hyperparameter feature corresponding to the current image patch; Determining a target mean feature corresponding to the current image patch based on the coefficient hyperparameter feature; Determining a reconstructed feature corresponding to the current image patch based on the target mean feature and the residual feature corresponding to the current image patch; Inputting the reconstructed feature into a synthesis transformation network to obtain a reconstructed image patch corresponding to the current image patch; Among them, the activation layer included in the synthesis transformation network uses a target activation function with upper and lower clamping to perform non-linear adjustment on the feature. The target activation function corresponds to a fixed upper limit threshold and a fixed lower limit threshold. The output feature of the activation layer is not greater than the fixed upper limit threshold, and the output feature of the activation layer is not less than the fixed lower limit threshold.
15. The method according to claim 14, wherein The synthesis transformation network includes a non-linear residual network layer, and the non-linear residual network layer at least includes an activation layer and a skip connection structure.
16. The method according to claim 15, wherein The non-linear residual network layer sequentially includes an activation layer, a processing layer, and a skip connection structure; among them, when there is one activation layer and the processing layer includes multiple network layers, all network layers are located behind the activation layer.
17. The method according to claim 14, wherein The fixed upper limit threshold is 6, and the fixed lower limit threshold is 0.
18. The method according to claim 15 or 16, wherein The synthesis transformation network includes a Residual Attention Unit (ResAU) layer, and the ResAU layer serves as the non-linear residual network layer; Among them, the ResAU layer sequentially includes an activation layer, a grouped convolutional layer, and a convolutional layer; Among them, the input feature of the ResAU layer is connected to the output feature of the convolutional layer; Among them, the input feature of the ResAU layer passes through the activation layer, the grouped convolutional layer, and the convolutional layer in sequence to obtain a post-convolution feature; the input feature of the ResAU layer and the post-convolution feature are subjected to a feature multiplication operation and a feature addition operation in sequence to obtain the output feature of the ResAU layer.
19. The method according to claim 18, wherein If the synthesis transformation network includes a low-complexity synthesis transformation network, the low-complexity synthesis transformation network sequentially includes: a lightweight residual block (LRB) layer, a convolutional layer, an upsampling Pixshuffle layer, a ResAU layer, a cropping (Crop) layer, a convolutional layer, a Pixshuffle layer, a Crop layer, a ResAU layer, a convolutional layer, a ResAU layer, a convolutional layer, a Pixshuffle layer, a Crop layer; If the synthesis transformation network includes a medium-complexity synthesis transformation network, the medium-complexity synthesis transformation network sequentially includes: an LRB layer, a transposed convolution layer, a ResAU layer, a Crop layer, a transposed convolution layer, a Crop layer, a ResAU layer, a convolution layer, a ResAU layer, a convolution layer, a Pixshuffle layer, and a Crop layer; If the synthesis transformation network includes a high-complexity synthesis transformation network, the high-complexity synthesis transformation network sequentially includes: a residual block RB layer, a transposed convolution layer, a ResAU layer, a Crop layer, a transposed convolution layer, a convolutional basic attention block CAB layer, a Crop layer, a ResAU layer, a convolution layer, a Pixshuffle layer, a Transformer-based attention block TAM layer, a Crop layer, a ResAU layer, a transposed convolution layer, and a Crop layer.
20. An image decoding method, characterized in that, The method includes: Obtaining coefficient hyperparameter features corresponding to a current image block; Inputting the coefficient hyperparameter features into a mean hyperparameter decoding network to obtain initial mean features corresponding to the current image block, and determining target mean features corresponding to the current image block based on the initial mean features; Determining reconstruction features corresponding to the current image block based on the target mean features and residual features corresponding to the current image block; Determining a reconstructed image block corresponding to the current image block based on the reconstruction features; Wherein, the activation layer included in the mean hyperparameter decoding network uses a target activation function with upper and lower clamping to perform non-linear adjustment on the features. The target activation function corresponds to a fixed upper limit threshold and a fixed lower limit threshold. The output features of the activation layer are not greater than the fixed upper limit threshold, and the output features of the activation layer are not less than the fixed lower limit threshold.
21. The method according to claim 20, wherein The obtaining of the coefficient hyperparameter features corresponding to the current image block includes: decoding a first bitstream corresponding to the current image block to obtain the coefficient hyperparameter features corresponding to the current image block; Before determining the reconstruction features corresponding to the current image block based on the target mean features and the residual features corresponding to the current image block, the method further includes: Determining probability distribution parameters based on the coefficient hyperparameter features, and decoding a second bitstream corresponding to the current image block based on the probability distribution parameters to obtain the residual features corresponding to the current image block.
22. The method according to claim 20, wherein The fixed upper limit threshold is 6, and the fixed lower limit threshold is 0.
23. The method according to any one of claims 20-22, characterized in that, The mean hyperparameter decoding network includes at least one convolution layer, at least one activation layer, a transposed convolution layer, and a Crop layer.
24. The method according to claim 23, characterized in that, The mean hyperparameter decoding network sequentially includes: a convolution layer, a transposed convolution layer, a Crop layer, an activation layer, a convolution layer, an activation layer, and a convolution layer.
25. The method according to any one of claims 20-24, characterized in that, The target activation function is a Relu6 activation function.
26. An image decoding method, characterized in that, The method includes: Obtaining coefficient hyperparameter features corresponding to a current image block; Input the coefficient hyperparameter feature into the mean hyperparameter decoding network to obtain the initial mean feature corresponding to the current image block, and determine the target mean feature corresponding to the current image block based on the initial mean feature; Determine the reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature corresponding to the current image block; Input the reconstruction feature into the synthesis transformation network to obtain the reconstructed image block corresponding to the current image block; Among them, the activation layers included in the mean hyperparameter decoding network and the synthesis transformation network use a target activation function with upper and lower clamping to perform non-linear adjustment on the feature. The target activation function corresponds to a fixed upper limit threshold and a fixed lower limit threshold. The output feature of the activation layer is not greater than the fixed upper limit threshold, and the output feature of the activation layer is not less than the fixed lower limit threshold.
27. The method according to claim 26, wherein The synthesis transformation network includes a non-linear residual network layer, and the non-linear residual network layer sequentially includes an activation layer, a processing layer, and a skip connection structure; among them, when there is one activation layer and the processing layer includes multiple network layers, all network layers are located behind the activation layer.
28. The method according to claim 27, wherein The synthesis transformation network includes a Residual Attention Unit (ResAU) layer, and the ResAU layer serves as the non-linear residual network layer; Among them, the ResAU layer sequentially includes an activation layer, a grouped convolution layer, and a convolution layer; Among them, the input feature of the ResAU layer is connected to the output feature of the convolution layer; Among them, the input feature of the ResAU layer passes through the activation layer, the grouped convolution layer, and the convolution layer in sequence to obtain a convolution feature; perform a feature multiplication operation and a feature addition operation on the input feature of the ResAU layer and the convolution feature in sequence to obtain the output feature of the ResAU layer.
29. The method according to claim 28, wherein If the synthesis transformation network includes a low-complexity synthesis transformation network, the low-complexity synthesis transformation network sequentially includes: a lightweight residual block (LRB) layer, a convolution layer, an upsampling Pixshuffle layer, a ResAU layer, a cropping (Crop) layer, a convolution layer, a Pixshuffle layer, a Crop layer, a ResAU layer, a convolution layer, a ResAU layer, a convolution layer, a Pixshuffle layer, a Crop layer; If the synthesis transformation network includes a medium-complexity synthesis transformation network, the medium-complexity synthesis transformation network sequentially includes: an LRB layer, a transposed convolution layer, a ResAU layer, a Crop layer, a transposed convolution layer, a Crop layer, a ResAU layer, a convolution layer, a ResAU layer, a convolution layer, a Pixshuffle layer, a Crop layer; If the synthesis transformation network includes a high-complexity synthesis transformation network, the high-complexity synthesis transformation network sequentially includes: a residual block RB layer, a transposed convolution layer, a ResAU layer, a Crop layer, a transposed convolution layer, a convolutional basic attention block CAB layer, a Crop layer, a ResAU layer, a convolutional layer, a Pixshuffle layer, a Transformer-based attention block TAM layer, a Crop layer, a ResAU layer, a transposed convolution layer, and a Crop layer.
30. The method according to claim 26, wherein The mean hyperparameter decoding network includes at least one convolutional layer, at least one activation layer, a transposed convolution layer, and a Crop layer.
31. The method according to claim 30, wherein The mean hyperparameter decoding network sequentially includes: a convolutional layer, a transposed convolution layer, a Crop layer, an activation layer, a convolutional layer, an activation layer, and a convolutional layer.
32. A decoding device, characterized in that, Applied to the decoding end, the device includes: A decoding module, configured to decode a first bitstream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block; determine a probability distribution parameter based on the coefficient hyperparameter feature, and decode a second bitstream corresponding to the current image block based on the probability distribution parameter to obtain a residual feature corresponding to the current image block; A determination module, configured to determine a target mean feature corresponding to the current image block based on the coefficient hyperparameter feature; determine a reconstructed feature corresponding to the current image block based on the target mean feature and the residual feature; An acquisition module, configured to input the reconstructed feature into a synthesis transformation network to obtain a reconstructed image block corresponding to the current image block; wherein, the synthesis transformation network includes a non-linear residual network layer, and the non-linear residual network layer at least includes an activation layer and a skip connection structure; wherein, the output feature of the activation layer is not greater than an upper threshold, and / or, the output feature of the activation layer is not less than a lower threshold.
33. A decoding device, characterized in that, Applied to the decoding end, the device includes: A decoding module, configured to decode a first bitstream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block; input the coefficient hyperparameter feature into a probability hyperparameter decoding network to obtain a probability distribution parameter, and decode a second bitstream corresponding to the current image block based on the probability distribution parameter to obtain a residual feature corresponding to the current image block; A determination module, configured to input the coefficient hyperparameter feature into a mean hyperparameter decoding network to obtain an initial mean feature corresponding to the current image block, and determine a target mean feature corresponding to the current image block based on the initial mean feature; determine a reconstructed feature corresponding to the current image block based on the target mean feature and the residual feature; An acquisition module, configured to determine a reconstructed image block corresponding to the current image block based on the reconstructed feature; wherein, the probability hyperparameter decoding network includes an activation layer, and the output feature of the activation layer is not greater than an upper threshold, and / or, the output feature of the activation layer is not less than a lower threshold; and / or, the mean hyperparameter decoding network includes an activation layer, and the output feature of the activation layer is not greater than an upper threshold, and / or, the output feature of the activation layer is not less than a lower threshold.
34. An image decoding device, characterized in that, Applied to the decoding end, the device includes: A decoding module, configured to decode a first bitstream corresponding to a current image block to obtain coefficient hyperparameter features corresponding to the current image block; determine probability distribution parameters based on the coefficient hyperparameter features, and decode a second bitstream corresponding to the current image block based on the probability distribution parameters to obtain residual features corresponding to the current image block; A determination module, configured to input the coefficient hyperparameter features into a mean hyperparameter decoding network to obtain initial mean features corresponding to the current image block, and determine target mean features corresponding to the current image block based on the initial mean features; wherein, the mean hyperparameter decoding network sequentially includes a first processing layer, an upsampling layer for performing an upsampling operation, a cropping layer, and a second processing layer; wherein, features output by the second processing layer are used as the initial mean features; determine reconstruction features corresponding to the current image block based on the target mean features and the residual features; An acquisition module, configured to determine a reconstructed image block corresponding to the current image block based on the reconstruction features.
35. An image decoding device, characterized in that, The apparatus includes: A decoding module, configured to obtain coefficient hyperparameter features corresponding to a current image block; A determination module, configured to determine target mean features corresponding to the current image block based on the coefficient hyperparameter features; determine reconstruction features corresponding to the current image block based on the target mean features and residual features corresponding to the current image block; An acquisition module, configured to input the reconstruction features into a synthesis transformation network to obtain a reconstructed image block corresponding to the current image block; wherein, an activation layer included in the synthesis transformation network uses a target activation function with upper and lower clamping to perform non-linear adjustment on features, the target activation function corresponds to a fixed upper limit threshold and a fixed lower limit threshold, output features of the activation layer are not greater than the fixed upper limit threshold, and output features of the activation layer are not less than the fixed lower limit threshold.
36. An image decoding device, characterized in that, The apparatus includes: A decoding module, configured to obtain coefficient hyperparameter features corresponding to a current image block; A determination module, configured to input the coefficient hyperparameter features into a mean hyperparameter decoding network to obtain initial mean features corresponding to the current image block, and determine target mean features corresponding to the current image block based on the initial mean features; determine reconstruction features corresponding to the current image block based on the target mean features and residual features corresponding to the current image block; An acquisition module, configured to determine a reconstructed image block corresponding to the current image block based on the reconstruction features; wherein, an activation layer included in the mean hyperparameter decoding network uses a target activation function with upper and lower clamping to perform non-linear adjustment on features, the target activation function corresponds to a fixed upper limit threshold and a fixed lower limit threshold, output features of the activation layer are not greater than the fixed upper limit threshold, and output features of the activation layer are not less than the fixed lower limit threshold.
37. An image decoding device, characterized in that, The apparatus includes: A decoding module, configured to obtain coefficient hyperparameter features corresponding to a current image block; A determination module, configured to input the coefficient hyperparameter feature to a mean hyperparameter decoding network to obtain an initial mean feature corresponding to the current image block, and determine a target mean feature corresponding to the current image block based on the initial mean feature; determine a reconstruction feature corresponding to the current image block based on the target mean feature and the residual feature corresponding to the current image block; An acquisition module, configured to input the reconstruction feature to a synthesis transformation network to obtain a reconstructed image block corresponding to the current image block; Wherein, the activation layers included in the mean hyperparameter decoding network and the synthesis transformation network perform non-linear adjustment on features by using a target activation function with upper and lower clamping, the target activation function corresponds to a fixed upper limit threshold and a fixed lower limit threshold, the output feature of the activation layer is not greater than the fixed upper limit threshold, and the output feature of the activation layer is not less than the fixed lower limit threshold.
38. A decoding end device, characterized in that, The decoding end device includes a processor and a machine-readable storage medium, and the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; The processor is configured to execute the machine-executable instructions to implement the method according to any one of claims 1-31.
39. A machine-readable storage medium, characterized in that, A number of computer instructions are stored on the machine-readable storage medium, and when the computer instructions are executed by the processor, the method according to any one of claims 1-31 is implemented.
40. A computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1-31 is implemented.
Citation Information
Patent Citations
Video image coding and decoding method and related equipment
CN115118972A
Coding and decoding method and device of regional enhancement layer
CN116939218A
Deep neural network training device and method for executing statistical regularization
WO2023085852A1
Cited By
Single-pixel imaging method based on non-training residual decoder and denoising prior
CN122176111A