Decoding, encoding methods, apparatus, devices and media
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
- Filing Date
- 2024-04-12
- Publication Date
- 2026-05-26
Smart Images

Figure 0007866154000007 
Figure 0007866154000008 
Figure 0007866154000009
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of encoding and decoding, and particularly to decoding, encoding methods, apparatuses, devices, and media thereof.
Background Art
[0002] To save space, all video images are encoded before transmission, and complete video encoding may include processes such as prediction, transformation, quantization, entropy encoding, filtering, etc. Regarding the prediction process, the prediction process may include intra-frame prediction and inter-frame prediction. Inter-frame prediction utilizes the temporal correlation of the video to predict the current pixel using the pixels of adjacent encoded images, thereby achieving the purpose of effectively removing the temporal redundancy of the video. Intra-frame prediction utilizes the spatial correlation of the video to predict the current pixel using the pixels of the encoded blocks in the current frame image, thereby achieving the purpose of removing the spatial redundancy of the video.
[0003] With the rapid development of deep learning, deep learning has achieved results in many high-level computer vision problems such as image classification and object detection. Deep learning has also gradually begun to be applied in the field of encoding and decoding, that is, it has become possible to encode and decode images using neural networks. Although the encoding and decoding methods based on neural networks show great performance potential, the encoding and decoding methods based on neural networks still have problems such as low encoding performance, low decoding performance, and high complexity.
Summary of the Invention
Means for Solving the Problems
[0004] In view of this, the present invention provides a decoding, encoding method, apparatus, device, and medium to improve encoding performance and decoding performance.
[0005] The present invention relates to a decoding method applied to the decoding side, The steps include decoding the first bitstream of the current image block to obtain coefficient hyperparameter features of each stage subblock of the current image block, For each stage subblock, the steps include determining the probability distribution parameters based on the coefficient hyperparameter features of the stage subblock, decoding the second bitstream of the current image block based on the probability distribution parameters to obtain the residual features of the stage subblock, A step of determining the reconstruction features of the stage subblock based on the residual features and the mean features of the stage subblock, A decoding method is provided, which includes the step of determining a reconstructed image block corresponding to the current image block based on the reconstruction features of each stage subblock.
[0006] The present invention relates to an encoding method applied to the encoding side, The steps include inputting the current image block into an analysis and transformation network to obtain a feature block corresponding to the current image block, The steps include dividing the feature block into features to be encoded in multiple sub-blocks, For each stage subblock corresponding to the current image block, the steps include obtaining the coefficient hyperparameter features of the stage subblock and encoding the coefficient hyperparameter features of the stage subblock into the first bitstream of the current image block. A step of determining the residual features of the stage subblock based on the features to be encoded of the stage subblock and the mean value features of the stage subblock, The present invention provides an encoding method comprising the steps of: determining probability distribution parameters based on the coefficient hyperparameter features of the stage subblock; and encoding the residual features of the stage subblock into a second bitstream of the current image block based on the probability distribution parameters.
[0007] The present invention relates to a decoding method applied to the decoding side, The steps include decoding the first bitstream of the current image block to obtain the coefficient hyperparameter features of the current image block, The steps include determining probability distribution parameters based on the coefficient hyperparameter features, decoding the second bitstream of the current image block based on the probability distribution parameters to obtain residual features of the current image block, and determining reconstruction features of the current image block based on the residual features. The present invention provides a decoding method that includes the steps of decoding an auxiliary bitstream corresponding to the current image block to obtain bitrate control parameters corresponding to the current image block, and inputting the reconstruction features and the bitrate control parameters into a composite transformation network to obtain a reconstructed image block corresponding to the current image block.
[0008] The present invention relates to a decoding device applied to the decoding side, A decoding module for decoding the first bitstream of the current image block to obtain coefficient hyperparameter features of each stage subblock of the current image block, determining probability distribution parameters for each stage subblock based on the coefficient hyperparameter features of the stage subblock, decoding the second bitstream of the current image block based on the probability distribution parameters to obtain residual features of the stage subblock, A decoding device is provided, which includes a determination module for determining the reconstruction features of a stage subblock based on the residual features and mean features of the stage subblock, and for determining the reconstructed image block corresponding to the current image block based on the reconstruction features of each stage subblock.
[0009] The present invention relates to an encoding device applied to the encoding side, An acquisition module for inputting the current image block into an analysis and transformation network to obtain a feature block corresponding to the current image block, and for dividing the feature block into features to be encoded in multiple stage subblocks, For each stage sub-block corresponding to the current image block, obtain the coefficient hyperparameter feature of the stage sub-block, and a coding module for coding the coefficient hyperparameter feature of the stage sub-block into the first bit stream of the current image block, and a determination module for determining the residual feature of the stage sub-block based on the feature to be coded and the average value feature of the stage sub-block. The coding module further provides a coding device that determines probability distribution parameters based on the coefficient hyperparameter feature of the stage sub-block and is used to code the residual feature of the stage sub-block into the second bit stream of the current image block based on the probability distribution parameters. The present invention relates to a decoding method applied to the decoding side, The steps include: decoding the bitstream of the current image block to obtain coefficient hyperparameter features of each subblock of the current image block; For any stage subblock, the steps include determining a probability distribution parameter based on the coefficient hyperparameter features of the stage subblock, decoding the bitstream of the current image block based on the probability distribution parameter to obtain the residual features of the stage subblock, For the first stage subblock, obtain the mean feature of the stage subblock via an mean prediction network based on the coefficient hyperparameter features of the stage subblock, and / or, for the i-th stage subblock, obtain the reference feature of the i-th stage subblock based on the reconstruction features of the i-1 stage subblocks preceding the i-th stage subblock, and obtain the mean feature of the stage subblock via an mean prediction network based on the coefficient hyperparameter features of the stage subblock and the reference feature, wherein i is greater than 1. A decoding method is provided, which includes the steps of: determining the reconstruction features of any stage subblock based on the residual features and mean features of the stage subblock; and determining the reconstructed image block corresponding to the current image block based on the reconstruction features of a plurality of stage subblocks. In one possible implementation, the reference feature includes all reconstruction features of the previous i-1 stage subblocks, or some reconstruction features of the previous i-1 stage subblocks, or reconstruction features of the i-1th stage subblock. In one possible implementation, the mean prediction network includes a second prediction network, and the step of obtaining the mean feature of the stage subblock via the mean prediction network based on the coefficient hyperparameter features of the stage subblock and the reference feature is: The steps include obtaining a second prediction feature corresponding to the reference feature via a second prediction network, The process includes the step of obtaining the mean value feature of the stage subblock based on the second prediction feature and the coefficient hyperparameter feature of the stage subblock. In one possible implementation, the step of obtaining a second predictive feature corresponding to the reference feature via the second predictive network is: If the reference feature includes at least two reconstruction features, the process includes combining all reconstruction features in the reference feature according to the channel dimension, and performing a convolution operation on the combined features to obtain the second prediction feature, or if the reference feature includes one reconstruction feature, the process includes performing a convolution operation on that reconstruction feature to obtain the second prediction feature. In one possible implementation, the mean prediction network includes a first prediction network, and the step of obtaining the mean feature of the stage subblock based on the second prediction feature and the coefficient hyperparameter feature of the stage subblock is: The steps include: combining the second prediction feature and the coefficient hyperparameter feature of the stage subblock to obtain a combined feature, and inputting the combined feature into the first prediction network; A step of obtaining a first predicted feature corresponding to the combined feature via a first prediction network, [[ID=3B]] The process includes the step of determining the mean value feature of the stage subblock based on the first predictive feature. In one possible implementation, the step of obtaining a first predictive feature corresponding to the coefficient hyperparameter feature or the combined feature via a first predictive network is: The process includes the step of obtaining the first predicted feature by performing a feature enhancement operation on the coefficient hyperparameter feature or the combined feature via a first prediction network, wherein the feature enhancement operation includes a convolution operation and an activation operation, and the activation operation includes a ReLU activation operation. In one possible implementation, the step of obtaining a first predictive feature corresponding to the coefficient hyperparameter feature or the combined feature via a first predictive network is: The process includes the step of obtaining the first predicted feature by sequentially performing convolution, activation, convolution, activation, and convolution operations on the coefficient hyperparameter feature or the combined feature via the first prediction network, wherein the activation operation includes a ReLU activation operation. In one possible implementation, after determining the reconstruction features of the stage subblock based on the residual features and the mean features of the stage subblock, The process includes a step of performing feature enhancement on the reconstruction features of the stage subblock to obtain enhanced reconstruction features, wherein the enhanced reconstruction features are used to determine the mean feature of the stage subblock, and the enhanced reconstruction features are used to determine the reconstruction image block. The present invention relates to an encoding method applied to the encoding side, The steps include inputting the current image block into an analysis and transformation network to obtain a feature block corresponding to the current image block, <00Q0107> The steps include dividing the feature block into features to be encoded in multiple sub-blocks, For any step subblock corresponding to the current image block, the steps include obtaining the coefficient hyperparameter features of the step subblock and encoding the coefficient hyperparameter features of the step subblock into the bitstream of the current image block. For the first stage subblock, obtain the mean value feature of the stage subblock via the mean value prediction network based on the coefficient hyperparameter feature of the stage subblock, or obtain a set default reference feature, obtain the mean value feature of the stage subblock via the mean value prediction network based on the coefficient hyperparameter feature and the default reference feature, and / or, for the i-th stage subblock, obtain the reference feature of the i-th stage subblock based on the reconstruction features of the i-1 stage subblocks preceding the i-th stage subblock, obtain the mean value feature of the stage subblock via the mean value prediction network based on the coefficient hyperparameter feature and the reference feature, wherein i is greater than 1. The present invention provides an encoding method comprising the steps of: determining the residual features of any stage subblock based on the features to be encoded of the stage subblock and the mean features of the stage subblock; determining probability distribution parameters based on the coefficient hyperparameter features of the stage subblock; and encoding the residual features of the stage subblock into the bitstream of the current image block based on the probability distribution parameters. The present invention relates to a decoding method applied to the decoding side, The steps include: decoding the bitstream of the current image block to obtain coefficient hyperparameter features of each subblock of the current image block; For any stage subblock, the steps include determining a probability distribution parameter based on the coefficient hyperparameter features of the stage subblock, decoding the bitstream of the current image block based on the probability distribution parameter to obtain the residual features of the stage subblock, A step of determining the reconstruction features of the stage subblock based on the residual features and the mean features of the stage subblock, The present invention provides a decoding method that includes the steps of: performing feature aggregation on the reconstruction features of multiple stage subblocks to obtain aggregated features; determining the reconstructed image block corresponding to the current image block based on the aggregated features; or inputting the reconstruction features of multiple stage subblocks into a composite transformation network to obtain a block-partitioned reconstructed image block corresponding to the stage subblock; and merging the block-partitioned reconstructed image blocks corresponding to all stage subblocks to obtain a reconstructed image block corresponding to the current image block. In one possible implementation, the step of determining the reconstructed image block corresponding to the current image block based on the aggregated features is: The aggregated features are input to the synthesis transformation network to obtain a reconstructed image block corresponding to the current image block. Alternatively, the process includes the steps of performing block division on the aggregated features to obtain a plurality of block division features, inputting the plurality of block division features into a composite transformation network to obtain a block division reconstructed image block corresponding to the block division feature, and merging the block division reconstructed image blocks corresponding to the plurality of block division features to obtain a reconstructed image block corresponding to the current image block. In one possible implementation, the step of performing feature aggregation on the reconstruction features of multiple stage subblocks to obtain aggregated features is: The reconstruction features of multiple stage subblocks are aggregated according to their phase to obtain the aggregated features, or, The process includes the steps of aggregating the reconstruction features of multiple stage subblocks according to their phases to obtain multiple phase-aggregated features, and then combining the multiple phase-aggregated features according to their channels to obtain the final aggregated feature. In one possible implementation, the step of performing block partitioning on the aggregated feature to obtain multiple block partitioning features is: Determine the target size of the block division feature, and based on the target size, divide the aggregated feature evenly into multiple block division features in the order from top to bottom and left to right, such that the size of any one of the block division features becomes the target size, or The process includes the steps of determining the actual block division size and overlap size of the block division feature, and dividing the aggregated feature into a plurality of block division features based on the actual block division size and overlap size such that the size of any one of the block division features becomes the target size, The actual block division size is the block division size obtained by removing the overlap portion from the block division features, and the overlap size is the size of the overlap portion of adjacent blocks. In one possible implementation, the step of merging the block-partitioned reconstructed image blocks corresponding to the plurality of block-partitioning features to obtain a reconstructed image block corresponding to the current image block is: The multiple block division reconstruction image blocks corresponding to the multiple block division features are sorted from top to bottom and from left to right, and the sorted multiple block division reconstruction image blocks are sequentially combined to obtain the reconstruction image block corresponding to the current image block, or, The process includes the steps of determining the actual block division size and overlap size of multiple block division reconstruction image blocks, sorting the multiple block division reconstruction image blocks corresponding to the multiple block division features from top to bottom and from left to right, and sequentially combining the sorted multiple block division reconstruction image blocks based on the actual block division size and overlap size to obtain a reconstruction image block corresponding to the current image block, The size of the non-overlapping portion of a block-partitioned reconstructed image block is the actual block division size of that block-partitioned reconstructed image block, and the size of the overlapping portion between a block-partitioned reconstructed image block and an adjacent block-partitioned reconstructed image block is the overlap size of that block-partitioned reconstructed image block. For the overlapping portion of the two block-reconstructed image blocks (left and right), the value of the overlapping portion is the value of the left block-reconstructed image block, or the value of the overlapping portion is the average value of the two block-reconstructed image blocks. For the overlapping portion of the upper and lower block-divided reconstructed image blocks, the value of the overlapping portion is either the value of the upper block-divided reconstructed image block or the average value of the two block-divided reconstructed image blocks. The present invention relates to an encoding method applied to the encoding side, The steps include inputting the current image block into an analysis and transformation network to obtain a feature block corresponding to the current image block, The steps include dividing the feature block into features to be encoded in multiple sub-blocks, For any step subblock corresponding to the current image block, the steps include obtaining the coefficient hyperparameter features of the step subblock and encoding the coefficient hyperparameter features of the step subblock into the bitstream of the current image block. For any stage subblock, the residual features of the stage subblock are determined based on the features to be encoded and the mean features of the stage subblock; probability distribution parameters are determined based on the coefficient hyperparameter features of the stage subblock; and the residual features of the stage subblock are encoded into the bitstream of the current image block based on the probability distribution parameters. A step of determining the reconstruction features of the stage subblock based on the residual features and the mean features of the stage subblock, The present invention provides an encoding method that includes the steps of: performing feature aggregation on the reconstruction features of multiple stage subblocks to obtain aggregated features; determining the reconstructed image block corresponding to the current image block based on the aggregated features; or inputting the reconstruction features of multiple stage subblocks into a composite transformation network to obtain a block-partitioned reconstructed image block corresponding to the stage subblock; and merging the block-partitioned reconstructed image blocks corresponding to all stage subblocks to obtain a reconstructed image block corresponding to the current image block. The present invention relates to a decoding method applied to the decoding side, The steps include: decoding the bitstream of the current image block to obtain coefficient hyperparameter features of each subblock of the current image block; For any stage subblock, the steps include determining a probability distribution parameter based on the coefficient hyperparameter features of the stage subblock, decoding the bitstream of the current image block based on the probability distribution parameter to obtain the residual features of the stage subblock, A step of determining the reconstruction features of the stage subblock based on the residual features and the mean features of the stage subblock, A decoding method is provided, which includes the steps of decoding a bitstream corresponding to the current image block to obtain a bitrate control parameter for the current image block, and determining a reconstructed image block corresponding to the current image block based on the reconstruction features of a plurality of stage subblocks and the bitrate control parameter. In one possible implementation, the step of determining the reconstructed image block corresponding to the current image block based on the reconstruction features of a plurality of stage subblocks and the bitrate control parameters is: A step of obtaining aggregated features by performing feature aggregation on the reconstruction features of multiple stage subblocks, The steps include: inputting the bitrate control parameters and the aggregated features into a synthesis transformation network, processing the aggregated features via the synthesis transformation network to obtain a first feature; processing the bitrate control parameters via the synthesis transformation network to obtain a second feature; and generating a third feature based on the first and second features. The process includes the step of determining a reconstructed image block corresponding to the current image block based on the third feature described above. In one possible implementation, the step of determining the reconstructed image block corresponding to the current image block based on the reconstruction features of a plurality of stage subblocks and the bitrate control parameters is: A step of performing feature aggregation on the reconstruction features of multiple stage subblocks to obtain aggregated features, and performing block partitioning on the aggregated features to obtain multiple block partitioning features, wherein the bitrate control parameter of the current image block includes the bitrate control parameters of the multiple block partitioning features. The process involves: inputting the block partitioning feature and its bitrate control parameters into a composite transformation network for any block partitioning feature, processing the block partitioning feature through the composite transformation network to obtain a first feature; processing the bitrate control parameters of the block partitioning feature through the composite transformation network to obtain a second feature; generating a third feature based on the first and second features, and determining the block partitioning reconstruction image block corresponding to the block partitioning feature based on the third feature; The process includes the step of merging block-partitioned reconstructed image blocks corresponding to multiple block-partitioning features to obtain a reconstructed image block corresponding to the current image block. In one possible implementation, the step of determining the reconstructed image block corresponding to the current image block based on the reconstruction features of a plurality of stage subblocks and the bitrate control parameters is: The bitrate control parameters of the current image block include the bitrate control parameters of multiple stage subblocks, and for any stage subblock, the reconstruction features and bitrate control parameters of the stage subblock are input to a composite transformation network, the reconstruction features of the stage subblock are processed via the composite transformation network to obtain a first feature, the bitrate control parameters of the stage subblock are processed via the composite transformation network to obtain a second feature, a third feature is generated based on the first and second features, and a block-partitioned reconstructed image block corresponding to the stage subblock is determined based on the third feature. The process includes the step of merging the block-partitioned reconstructed image blocks corresponding to all stage subblocks to obtain a reconstructed image block corresponding to the current image block. In one possible implementation, the step of determining the reconstruction features of the stage subblock based on the residual features and the mean features of the stage subblock is: For the first stage subblock, the process includes obtaining the mean value feature of the stage subblock via an mean value prediction network based on the coefficient hyperparameter feature of the stage subblock, or obtaining a set default reference feature, and obtaining the mean value feature of the stage subblock via the mean value prediction network based on the coefficient hyperparameter feature and the default reference feature of the stage subblock. In one possible implementation, the step of determining the reconstruction features of the stage subblock based on the residual features and the mean features of the stage subblock is: The steps include: obtaining a reference feature of the i-th stage subblock based on the reconstruction features of the i-1 stage subblocks preceding the i-th stage subblock; obtaining a mean feature of the stage subblock via a mean prediction network based on the coefficient hyperparameter features of the stage subblock and the reference feature; and determining the reconstruction feature of any stage subblock based on the residual features of the stage subblock and the mean feature of the stage subblock, wherein i is greater than 1. In one possible implementation, the reference feature includes all reconstruction features of the previous i-1 stage subblocks, or some reconstruction features of the previous i-1 stage subblocks, or reconstruction features of the i-1th stage subblock. In one possible implementation, the mean prediction network includes a second prediction network, and the step of obtaining the mean feature of the stage subblock via the mean prediction network based on the coefficient hyperparameter features of the stage subblock and the reference feature is: The steps include obtaining a second prediction feature corresponding to the reference feature via a second prediction network, The process includes the step of obtaining the mean value feature of the stage subblock based on the second prediction feature and the coefficient hyperparameter feature of the stage subblock. The present invention relates to an encoding method applied to the encoding side, The steps include inputting the current image block into an analysis and transformation network to obtain a feature block corresponding to the current image block, The steps include dividing the feature block into features to be encoded in multiple sub-blocks, For any step subblock corresponding to the current image block, the steps include obtaining the coefficient hyperparameter features of the step subblock and encoding the coefficient hyperparameter features of the step subblock into the bitstream of the current image block. For any stage subblock, the residual features of the stage subblock are determined based on the features to be encoded and the mean features of the stage subblock; probability distribution parameters are determined based on the coefficient hyperparameter features of the stage subblock; and the residual features of the stage subblock are encoded into the bitstream of the current image block based on the probability distribution parameters. The present invention provides an encoding method comprising the steps of: determining the reconstruction features of a stage subblock based on the residual features and mean features of the stage subblock; obtaining the bitrate control parameters of the current image block; determining a reconstructed image block corresponding to the current image block based on the reconstruction features of a plurality of stage subblocks and the bitrate control parameters; and encoding the bitrate control parameters into a bitstream corresponding to the current image block.
[0010] The present invention relates to a decoding device comprising a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor. The processor provides a decoding device used to execute the above decoding method by executing machine-executable instructions.
[0011] The present invention relates to an encoding device comprising a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor. The processor provides an encoding device used to execute machine-executable instructions and implement the above-described encoding method.
[0012] The present invention provides an electronic device comprising a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor, and the processor is used to execute the machine-executable instructions to carry out the above-described decoding method, or the processor is used to execute the machine-executable instructions to carry out the above-described encoding method.
[0013] The present invention provides a machine-readable storage medium in which a plurality of computer instructions are stored, wherein when the computer instructions are executed by a processor, the above-described decoding method or the above-described encoding method is performed.
[0014] The present invention provides a computer application in which, when executed by a processor, the above decoding method is performed, or when executed by a processor, the above encoding method is performed.
[0015] As can be seen from the above technical proposals, the embodiments of the present invention propose a variable and adjustable bitrate encoding and decoding method for neural network-based encoding and decoding technology, improving parallelism, effectively saving cache for storing features, achieving higher bitrate control accuracy, reducing the loss of encoding performance, and providing better encoding performance and bitrate control accuracy. By adopting block partitioning encoding and decoding, peak memory occupation is reduced, the decoding time for a single block is shortened, and high-speed parallel decoding capability is achieved, maintaining a low complexity of the neural network, effectively guaranteeing the quality of reconstructed image blocks, improving encoding and decoding performance, and reducing complexity. [Brief explanation of the drawing]
[0016] [Figure 1] This is a schematic diagram of a 3D feature matrix according to one embodiment of the present invention. [Figure 2] This is a flowchart of a decoding method according to one embodiment of the present invention. [Figure 3]This is a flowchart of the encoding method according to one embodiment of the present invention. [Figure 4] This is a schematic diagram of the encoding process according to one embodiment of the present invention. [Figure 5] This is a schematic diagram of the decoding process according to one embodiment of the present invention. [Figure 6] This is a schematic diagram of the encoding process according to one embodiment of the present invention. [Figure 7] This is a schematic diagram of the decoding process according to one embodiment of the present invention. [Figure 8A] This is a schematic diagram of the mean prediction network according to one embodiment of the present invention. [Figure 8B] This is a schematic diagram of the mean prediction network according to one embodiment of the present invention. [Figure 8C] This is a schematic diagram of the mean prediction network according to one embodiment of the present invention. [Figure 8D] This is a schematic diagram of the mean prediction network according to one embodiment of the present invention. [Figure 8E] This is a schematic diagram of the mean prediction network according to one embodiment of the present invention. [Figure 8F] This is a schematic diagram of the mean prediction network according to one embodiment of the present invention. [Figure 8G] This is a schematic diagram of the mean prediction network according to one embodiment of the present invention. [Figure 8H] This is a schematic diagram of the mean prediction network according to one embodiment of the present invention. [Figure 9A] This is a schematic diagram of aggregation and block division according to one embodiment of the present invention. [Figure 9B] This is a schematic diagram of aggregation and block division according to one embodiment of the present invention. [Figure 9C] This is a schematic diagram of aggregation and block division according to one embodiment of the present invention. [Figure 9D] This is a schematic diagram of aggregation and block division according to one embodiment of the present invention. [Figure 9E]This is a schematic diagram of aggregation and block division according to one embodiment of the present invention. [Figure 9F] This is a schematic diagram of aggregation and block division according to one embodiment of the present invention. [Figure 10A] This is a schematic diagram of a synthesis-transformation network according to one embodiment of the present invention. [Figure 10B] This is a schematic diagram of a synthesis-transformation network according to one embodiment of the present invention. [Figure 10C] This is a schematic diagram of a synthesis-transformation network according to one embodiment of the present invention. [Figure 11A] This is a schematic diagram of the processing of a convolutional layer according to one embodiment of the present invention. [Figure 11B] This is a schematic diagram of the processing of a convolutional layer according to one embodiment of the present invention. [Figure 12] This is a schematic diagram of the determination of the λ parameter according to one embodiment of the present invention. [Figure 13A] This is a hardware structure diagram of a decoding device according to one embodiment of the present invention. [Figure 13B] This is a hardware structure diagram of an encoding device according to one embodiment of the present invention. [Modes for carrying out the invention]
[0017] The terms used in the embodiments of this invention are merely for the purpose of describing specific embodiments and are not intended to limit the invention. The singular forms “one kind,” “the said,” and “the” used in the embodiments and claims of this invention are also intended to include the plural form unless the context clearly indicates otherwise. Furthermore, it should be understood that the term “and / or” used in this invention means including any or all possible combinations of one or more related enumerated items. The embodiments of this invention may use terms such as first, second, third, etc. to describe various types of information, but it should be understood that this information is not limited to these terms. These terms are used only to distinguish the same type of information. For example, as long as it does not deviate from the scope of the embodiments of this invention, depending on the context, first information may be called second information, and similarly, second information may be called first information. Furthermore, the word “…case” used herein may be interpreted as “…and,” “…when,” or “in response to a decision.”
[0018] Embodiments of the present invention provide decoding methods and encoding methods and may relate to the following concepts.
[0019] Entropy coding: Entropy coding is a method of coding in which no information is lost during the coding process, according to the principle of entropy. Information entropy is the average amount of information (degree of uncertainty) of the information source. Entropy coding schemes may include, but are not limited to, Shannon coding, Huffman coding, and arithmetic coding.
[0020] Neural Network (NN): A neural network refers to an artificial neural network, which is a computational model composed of a large number of nodes (called neurons) connected to one another. In a neural network, neuron processing units can represent different objects, such as features, alphabets, concepts, or several meaningful abstract modes. There are three types of processing units in a neural network: input units, output units, and hidden units. Input units receive external signals and data, output units realize the output of the processing results, and hidden units are units that are between the input and output units and cannot be observed from outside the system. The connection weights between neurons reflect the strength of the connections between units, and the representation and processing of information are reflected in the connection relationships of the processing units. A neural network is an unprogrammed, brain-like information processing method, and the essence of a neural network is to acquire parallel and distributed information processing capabilities through the transformation and dynamic actions of the neural network, mimicking the information processing capabilities of the human brain and nervous system to different degrees and levels. In the field of video processing, commonly used neural networks may include, but are not limited to, convolutional neural networks (CNNs), recurrent neural networks (RNNs), and fully connected networks.
[0021] Convolutional Neural Networks (CNNs): Convolutional neural networks are feedforward neural networks and one of the representative network structures in deep learning techniques. The artificial neurons in a convolutional neural network can respond to peripheral units within a certain coverage area, exhibiting excellent performance in large-scale image processing. The basic structure of a convolutional neural network consists of two layers: one is a feature extraction layer (also called a convolutional layer), where the input of each neuron is connected to the local receptive field of the previous layer, and the local features are extracted. Once the local features are extracted, their positional relationship with other features is also determined. The other is a feature mapping layer (also called an activation layer), where each computational layer of the neural network consists of multiple feature mappings, each feature mapping is a plane, and the weights of all neurons in the plane are equal. The feature mapping structure may use functions such as the Sigmoid function, ReLU function (normalized linear unit), Leaky-ReLU (Leaky Rectified Linear Unit), PReLU (Parametric Rectified Linear Unit), or GDN (Generalized Difference Network) as activation functions for the convolutional network. Furthermore, because neurons in a single mapping plane share weights, the number of free parameters in the network is reduced.
[0022] For example, one advantage of convolutional neural networks compared to image processing algorithms is that they can avoid complex pre-processing processes for images (such as extracting artificial features), directly inputting original images and performing end-to-end learning. Another advantage of convolutional neural networks compared to general neural networks is that while general neural networks employ a fully connected architecture, meaning all neurons from the input layer to the hidden layer are connected, resulting in a huge number of parameters and making network training time-consuming and difficult, convolutional neural networks mitigate this difficulty through methods such as local connectivity and weight sharing.
[0023] Deconvolution: Also known as transposed convolution, deconvolution and convolution layers operate similarly. The main difference is that deconvolution uses padding to make the output larger than the input (though they can also be the same). A stride of 1 indicates that the output size is equal to the input size. A stride of N indicates that the width of the output features is N times the width of the input features, and the height of the output features is N times the height of the input features.
[0024] Generalization Ability: Generalization ability may also refer to a machine learning algorithm's ability to adapt to fresh samples. The goal of learning is to learn the underlying rules of the data, and the trained network can produce appropriate outputs even for data outside the training set that has the same rules. This ability may be called generalization ability.
[0025] Feature: The feature according to the present invention is a 3D feature matrix or tensor of type C*W*H. Figure 1 is a schematic diagram of the 3D feature matrix, where C represents the number of channels, H represents the feature height, and W represents the feature width. The 3D feature matrix may be the input to a neural network or the output of a neural network.
[0026] The Rate-Distortion Optimization principle: Encoding efficiency is evaluated using two metrics: bitrate and PSNR (Peak Signal to Noise Ratio). A smaller bitstream results in greater compression, and a higher PSNR results in better reconstructed image quality. When selecting a mode, the discriminant is essentially a comprehensive evaluation of both. For example, the cost corresponding to a mode is J(mode) = D + λ*R, where D represents distortion, which can usually be evaluated using the SSE (Sum of the Squared Errors) metric, where SSE is the mean square sum of the differences between the reconstructed image block and the source image. To consider the cost, the SAD metric may also be used, where SAD is the sum of the absolute differences between the reconstructed image block and the source image, λ is the Lagrangian multiplier, and R is the actual number of bits required to encode the image block in that mode, including the total number of bits required to encode mode information, motion information, residuals, etc. When selecting a mode, comparing and determining the encoding mode using the rate distortion principle usually guarantees optimal encoding performance.
[0027] For each module on the encoding side, a great many encoding tools have been proposed, and each tool usually has many modes. The encoding tool that yields optimal encoding performance often differs for different video sequences. Therefore, in the encoding process, Rate-Distortion Optimize (RDO) is typically used to compare the encoding performance of different tools or modes and select the optimal tool or mode. After determining the optimal tool or mode, the decision information is transmitted by encoding mark information into the bitstream. While this method introduces encoding complexity, it allows for the adaptive selection of the optimal tool or mode combination for different content, thereby achieving optimal encoding performance. On the decoding side, the relevant tool or mode information can be obtained by directly analyzing the mark information, minimizing the impact of complexity.
[0028] The decoding method and encoding method according to the embodiment of the present invention will be described in detail below with reference to several specific examples.
[0029] Example 1: An embodiment of the present invention provides a decoding method, Figure 2 is a schematic flowchart of the decoding method, which may be applied to the decoding side (also called a video decoder), and which may include the following steps.
[0030] In step 201, the first bitstream of the current image block is decoded to obtain the coefficient hyperparameter features of each stage sub-block of the current image block. The current image block is divided into multiple stage sub-blocks, i.e., the current image block contains multiple stage sub-blocks.
[0031] In step 202, for each stage subblock, probability distribution parameters are determined based on the coefficient hyperparameter features of the stage subblock, and the second bitstream of the current image block is decoded based on the probability distribution parameters to obtain the residual features of the stage subblock.
[0032] In step 203, the reconstructed features of the stage subblock are determined based on the residual features and the mean features of the stage subblock. The mean features of the stage subblock are obtained based on the coefficient hyperparameter features and / or the reference features of the stage subblock, for example, based on the coefficient hyperparameter features, or based on the coefficient hyperparameter features and the reference features.
[0033] In step 204, the reconstructed image block corresponding to the current image block is determined based on the reconstruction features of each stage subblock.
[0034] For example, for the first stage subblock, the mean value feature of the stage subblock is obtained via an mean value prediction network based on the coefficient hyperparameter feature of the stage subblock, or a set default reference feature is obtained, and the mean value feature of the stage subblock is obtained via an mean value prediction network based on the coefficient hyperparameter feature and default reference feature of the stage subblock.
[0035] For the i-th stage subblock, a reference feature of the i-th stage subblock is obtained based on the reconstruction features of the i-1 stage subblocks preceding it, and the mean feature of the stage subblock is obtained via a mean prediction network based on the coefficient hyperparameter features and reference feature of the stage subblock. i is greater than 1, and the reference feature includes all the reconstruction features of the i-1 preceding stage subblocks, or some of the reconstruction features of the i-1 preceding stage subblocks, or the reconstruction features of the i-1th stage subblock.
[0036] Exemplary, the mean prediction network may include a first prediction network, a second prediction network, and a prediction fusion network, and the step of obtaining the mean feature of the stage subblock via the mean prediction network based on the coefficient hyperparameter features and reference features of the stage subblock may include the step of obtaining a first prediction feature corresponding to the coefficient hyperparameter features via the first prediction network, the step of obtaining a second prediction feature corresponding to the reference features via the second prediction network, the step of feature-combining the first prediction feature and the second prediction feature to obtain a combined feature, inputting the combined feature into the prediction fusion network, processing the combined feature via the prediction fusion network to obtain the mean feature of the stage subblock.
[0037] For example, the step of obtaining a first predicted feature corresponding to a coefficient hyperparameter feature via a first predicted network may include the step of obtaining the first predicted feature by performing a feature enhancement operation and an upsampling operation on the coefficient hyperparameter feature via the first predicted network. The feature enhancement operation includes a convolution operation, or a convolution operation and an activation operation, and the upsampling operation may include a deconvolution operation, a crop operation and an activation operation, or a deconvolution operation, a crop operation, an activation operation and a convolution operation. For example, the activation operation is a ReLU operation.
[0038] Exemplary, the step of obtaining a second predictive feature corresponding to a reference feature via a second predictive network may include the steps of: combining all reconstructed features in the reference feature according to the channel dimension, and performing a convolution operation on the combined features to obtain a second predictive feature; or adding features to all reconstructed features in the reference feature, and performing a convolution operation on the added features to obtain a second predictive feature. The reference feature includes all reconstructed features of the previous i-1 stage subblocks, or some reconstructed features of the previous i-1 stage subblocks, or the reconstructed features of the i-1th stage subblock.
[0039] Exemplary, the step of processing post-combined features via a predictive fusion network to obtain the mean feature of the step subblock may include, but is not limited to, a step of performing a convolution operation and at least one fusion operation on the post-combined features via a predictive fusion network to obtain the mean feature of the step subblock. The fusion operation includes an activation operation and a convolution operation. Exemplary, the activation operation is a ReLU operation.
[0040] For example, after determining the reconstruction features of the stage subblock based on the residual features and mean features of the stage subblock, feature enhancement may be performed on the reconstruction features of the stage subblock to obtain enhanced reconstruction features, which are used to determine the mean features of the stage subblock, and which are used to determine the reconstructed image block.
[0041] For example, the step of determining the reconstructed image block corresponding to the current image block based on the reconstruction features of each stage subblock may include, but is not limited to, the step of performing feature aggregation on the reconstruction features of each stage subblock to obtain aggregated features, inputting the aggregated features into a composite transformation network to obtain the reconstructed image block corresponding to the current image block. Alternatively, feature aggregation may be performed on the reconstruction features of each stage subblock to obtain aggregated features, block partitioning may be performed on the aggregated features to obtain a plurality of block partitioning features, input each block partitioning feature into a composite transformation network to obtain a block partitioned reconstructed image block corresponding to the block partitioning feature, and the block partitioned reconstructed image blocks corresponding to the plurality of block partitioning features may be merged to obtain the reconstructed image block corresponding to the current image block. Alternatively, the reconstruction features of each stage subblock may be input into a composite transformation network to obtain a block partitioned reconstructed image block corresponding to the stage subblock, and the block partitioned reconstructed image blocks corresponding to all stage subblocks may be merged to obtain the reconstructed image block corresponding to the current image block.
[0042] For example, the step of performing feature aggregation on the reconstructed features of each stage subblock to obtain an aggregated feature may include, but is not limited to, a step of aggregating the reconstructed features of each stage subblock according to phase to obtain an aggregated feature, or a step of aggregating the reconstructed features of each stage subblock according to phase to obtain multiple phase-aggregated features, and combining the multiple phase-aggregated features according to channel to obtain an aggregated feature.
[0043] For example, the step of performing block partitioning on the aggregated feature to obtain multiple block partitioning features may include determining the target size of the block partitioning features, and based on the target size, evenly dividing the aggregated feature into multiple block partitioning features in the order from top to bottom and left to right so that the size of each block partitioning feature becomes the target size, or determining the actual block partitioning size and overlap size of the block partitioning features, where the actual block partitioning size is the block partitioning size obtained by removing the overlap portion from the block partitioning feature, and the overlap size is the size of the overlap portion of adjacent blocks, and based on the actual block partitioning size and overlap size, dividing the aggregated feature into multiple block partitioning features so that the size of each block partitioning feature becomes the target size. If the actual block partitioning size is tile_c*tile_w*tile_h and the overlap size is tile_c*padding_w*padding_h, then the target size is tile_c*(tile_w+2*padding_w)*(tile_h+2*padding_h).
[0044] For example, the step of merging block-partitioned reconstructed image blocks corresponding to multiple block-partitioning features to obtain a reconstructed image block corresponding to the current image block may include, but is not limited to, the step of sorting the block-partitioned reconstructed image blocks corresponding to multiple block-partitioning features from top to bottom and from left to right, and sequentially combining the sorted block-partitioned reconstructed image blocks corresponding to multiple block-partitioning features to obtain a reconstructed image block corresponding to the current image block. Alternatively, the actual block-partitioning size and overlap size of each block-partitioned reconstructed image block may be determined, the block-partitioned reconstructed image blocks corresponding to multiple block-partitioning features may be sorted from top to bottom and from left to right, and sequentially combining the sorted block-partitioned reconstructed image blocks corresponding to multiple block-partitioning features based on the actual block-partitioning size and overlap size to obtain a reconstructed image block corresponding to the current image block. The reconstructed image block includes multiple block-partitioned reconstructed image blocks, the size of the non-overlapping portion of a block-partitioned reconstructed image block is the actual block-partitioning size of the block-partitioned reconstructed image block, and the size of the overlap portion between a block-partitioned reconstructed image block and an adjacent block-partitioned reconstructed image block is the overlap size of the block-partitioned reconstructed image block. For the overlapping portion of two left and right block-reconstructed image blocks, the value of the overlapping portion is the value of the left block-reconstructed image block, or the value of the overlapping portion is the average value of the two block-reconstructed image blocks. Similarly, for the overlapping portion of two upper and lower block-reconstructed image blocks, the value of the overlapping portion is the value of the upper block-reconstructed image block, or the value of the overlapping portion is the average value of the two block-reconstructed image blocks.
[0045] For example, the step of inputting aggregated features into a composite transformation network to obtain a reconstructed image block corresponding to the current image block may include the step of decoding an auxiliary bitstream corresponding to the current image block to obtain the bitrate control parameters of the current image block, and inputting the bitrate control parameters and aggregated features into a composite transformation network to obtain a reconstructed image block corresponding to the current image block.
[0046] The step of inputting each block division feature into a composite transformation network to obtain a block division reconstructed image block corresponding to the block division feature may include the step of decoding an auxiliary bitstream corresponding to the current image block to obtain a bitrate control parameter for each block division feature, and inputting the block division feature and its bitrate control parameter into a composite transformation network to obtain a block division reconstructed image block corresponding to the block division feature.
[0047] The step of inputting the reconstruction features of each stage subblock into a composite transformation network to obtain a block-partitioned reconstructed image block corresponding to the stage subblock may include the step of decoding the auxiliary bitstream corresponding to the current image block to obtain the bitrate control parameters of each stage subblock, inputting the reconstruction features of the stage subblock and the bitrate control parameters of the stage subblock into a composite transformation network to obtain a block-partitioned reconstructed image block corresponding to the stage subblock.
[0048] For example, the step of inputting bitrate control parameters and aggregated features into a composite transformation network to obtain a reconstructed image block corresponding to the current image block may include the steps of processing the aggregated features via the composite transformation network to obtain a first feature, processing the bitrate control parameters via the composite transformation network to obtain a second feature, and generating a third feature based on the first and second features, and determining a reconstructed image block corresponding to the current image block based on the third feature.
[0049] The step of inputting a block division feature and the bitrate control parameter of the block division feature into a composite transformation network to obtain a block division reconstructed image block corresponding to the block division feature may include the steps of processing the block division feature via the composite transformation network to obtain a first feature, processing the bitrate control parameter of the block division feature via the composite transformation network to obtain a second feature, and generating a third feature based on the first and second features, and determining a block division reconstructed image block corresponding to the block division feature based on the third feature.
[0050] The step of inputting the reconstruction features of a stage subblock and the bitrate control parameters of the stage subblock into a composite transformation network to obtain a block-partitioned reconstructed image block corresponding to the stage subblock may include the steps of: processing the reconstruction features of the stage subblock via the composite transformation network to obtain a first feature; processing the bitrate control parameters of the stage subblock via the composite transformation network to obtain a second feature; and generating a third feature based on the first and second features, and determining a block-partitioned reconstructed image block corresponding to the stage subblock based on the third feature.
[0051] For illustrative purposes, the above execution order is merely illustrative for the sake of clarity, and in actual applications, the order of execution between steps may be changed and is not limiting. Furthermore, in other embodiments, the steps of the corresponding method may not necessarily be performed in the order shown and described herein, and the method may include more or fewer steps than those described herein. Also, a single step described herein may be broken down into multiple steps in other embodiments, and multiple steps described herein may be combined into a single step in other embodiments.
[0052] As can be seen from the above technical proposals, the embodiments of the present invention propose a variable and adjustable bitrate encoding and decoding method for neural network-based encoding and decoding technology, improving parallelism, effectively saving cache for storing features, achieving higher bitrate control accuracy, reducing the loss of encoding performance, and providing better encoding performance and bitrate control accuracy. By adopting block partitioning encoding and decoding, peak memory occupation is reduced, the decoding time for a single block is shortened, and high-speed parallel decoding capability is achieved, maintaining a low complexity of the neural network, effectively guaranteeing the quality of reconstructed image blocks, improving encoding and decoding performance, and reducing complexity.
[0053] Example 2: An embodiment of the present invention provides an encoding method, Figure 3 is a schematic flowchart of the encoding method, which may be applied to the encoding side (also called a video encoder), and which may include the following steps.
[0054] In step 301, the current image block is input to the analysis and transformation network to obtain the feature block corresponding to the current image block.
[0055] In step 302, the feature block is divided into features to be encoded in multiple stage subblocks.
[0056] In step 303, for each stage subblock corresponding to the current image block, the coefficient hyperparameter features of the stage subblock are obtained, and the coefficient hyperparameter features of the stage subblock are encoded into the first bitstream of the current image block.
[0057] In step 304, the residual features of the stage subblock are determined based on the features to be encoded of the stage subblock and the mean features of the stage subblock. The mean features of the stage subblock are obtained based on the coefficient hyperparameter features and / or the reference features of the stage subblock, for example, based on the coefficient hyperparameter features, or based on the coefficient hyperparameter features and the reference features.
[0058] In step 305, probability distribution parameters are determined based on the coefficient hyperparameter features of the stage subblock, and the residual features of the stage subblock are encoded into the second bitstream of the current image block based on the probability distribution parameters.
[0059] For example, for the first stage subblock, the mean value feature of the stage subblock is obtained via an mean value prediction network based on the coefficient hyperparameter feature of the stage subblock, or a set default reference feature is obtained, and the mean value feature of the stage subblock is obtained via an mean value prediction network based on the coefficient hyperparameter feature and default reference feature of the stage subblock.
[0060] For the i-th stage subblock, a reference feature of the i-th stage subblock is obtained based on the reconstruction features of the i-1 stage subblocks preceding it, and the mean feature of the stage subblock is obtained via a mean prediction network based on the coefficient hyperparameter features and reference feature of the stage subblock. i is greater than 1, and the reference feature includes all the reconstruction features of the i-1 preceding stage subblocks, or some of the reconstruction features of the i-1 preceding stage subblocks, or the reconstruction features of the i-1th stage subblock.
[0061] For example, the encoding process is similar to the decoding process, but the same parts are not repeated.
[0062] For illustrative purposes, the above execution order is merely illustrative for the sake of clarity, and in actual applications, the order of execution between steps may be changed and is not limiting. Furthermore, in other embodiments, the steps of the corresponding method may not necessarily be performed in the order shown and described herein, and the method may include more or fewer steps than those described herein. Also, a single step described herein may be broken down into multiple steps in other embodiments, and multiple steps described herein may be combined into a single step in other embodiments.
[0063] As can be seen from the above technical proposals, the embodiments of the present invention propose a variable and adjustable bitrate encoding and decoding method for neural network-based encoding and decoding technology, improving parallelism, effectively saving cache for storing features, achieving higher bitrate control accuracy, reducing the loss of encoding performance, and providing better encoding performance and bitrate control accuracy. By adopting block partitioning encoding and decoding, peak memory occupation is reduced, the decoding time for a single block is shortened, and high-speed parallel decoding capability is achieved, maintaining a low complexity of the neural network, effectively guaranteeing the quality of reconstructed image blocks, improving encoding and decoding performance, and reducing complexity.
[0064] Example 3: Compared to Examples 1 and 2, the encoding-side processing process is shown in Figure 4. Of course, Figure 4 is merely one example of the encoding-side processing process and is not limited to this encoding-side processing process.
[0065] The encoding side may, after obtaining the current image block x (which may be the original image block x, i.e., the input image block), perform an analytical transformation on the current image block x via an analytical transformation network (i.e., a neural network) to obtain the image features y corresponding to the current image block x. Here, performing a feature transformation on the current image block x via an analytical transformation network means transforming the current image block x into image features y in the latent domain, so that all subsequent processes can operate in the latent domain.
[0066] An image may be divided into one image block or into multiple image blocks. If an image is divided into one image block, the current image block x may be an image itself, that is, the encoding and decoding processes for the image block may be applied directly to the image.
[0067] The encoding side obtains image features y, then performs a coefficient hyperparameter feature transformation on image features y to obtain coefficient hyperparameter features z. For example, image features y are input to a hyperparameter coding network (i.e., a neural network), and the hyperparameter coding network performs a coefficient hyperparameter feature transformation on image features y to obtain coefficient hyperparameter features z. The hyperparameter coding network may be a trained neural network, and the training process of this hyperparameter coding network is not limited; it is sufficient that a coefficient hyperparameter feature transformation can be performed on image features y. Here, after the image features y of the latent domain pass through the hyperparameter coding network, hyperprior latent information z (i.e., coefficient hyperparameter features z) is obtained.
[0068] After obtaining the coefficient hyperparameter feature z, the encoding side may quantize the coefficient hyperparameter feature z to obtain the corresponding hyperparameter quantization feature; that is, the Q operation in Figure 4 is a quantization process. After obtaining the hyperparameter quantization feature corresponding to the coefficient hyperparameter feature z, the hyperparameter quantization feature is encoded to obtain Bitstream#1 (i.e., the first bitstream) corresponding to the current image block; that is, the AE operation in Figure 4 represents an encoding process such as an entropy encoding process. Alternatively, the encoding side may directly encode the coefficient hyperparameter feature z to obtain Bitstream#1 corresponding to the current image block. The hyperparameter quantization feature or coefficient hyperparameter feature z contained in Bitstream#1 is mainly used to obtain the mean value and the parameters of the probability distribution model (i.e., probability distribution parameters).
[0069] After obtaining Bitstream#1 corresponding to the current image block, the encoding side may send Bitstream#1 corresponding to the current image block to the decoding side. The decoding side's processing process for Bitstream#1 corresponding to the current image block will be described in subsequent embodiments.
[0070] After obtaining Bitstream#1 corresponding to the current image block, the encoding side may decode Bitstream#1 to obtain the hyperparameter quantization feature, i.e., AD in Figure 4 represents the decoding process, and then the hyperparameter quantization feature is inversely quantized to obtain the coefficient hyperparameter feature z_hat, which may be the same as or different from the coefficient hyperparameter feature z, and the IQ operation in Figure 4 is the inverse quantization process. Alternatively, after obtaining Bitstream#1 corresponding to the current image block, the encoding side may decode Bitstream#1 to obtain the coefficient hyperparameter feature z_hat without performing the inverse quantization process of the coefficient hyperparameter feature z_hat.
[0071] The encoding process for Bitstream#1 may employ a fixed probability density model encoding method, and the decoding process for Bitstream#1 may employ a fixed probability density model decoding method; however, the encoding and decoding processes are not limited to these methods.
[0072] After obtaining the coefficient hyperparameter feature z_hat, the encoding side may perform a context-based prediction based on the coefficient hyperparameter feature z_hat of the current image block and the image feature y_hat of the previous image block (the process for determining the image feature y_hat will be described in subsequent examples) to obtain a predicted value mu (i.e., mean mu) corresponding to the current image block. For example, the coefficient hyperparameter feature z_hat and the image feature y_hat may be input to a mean prediction network, and the mean prediction network may determine the predicted value mu based on the coefficient hyperparameter feature z_hat and the image feature y_hat. This prediction process is not limited. Here, for the context-based prediction process, the input includes the coefficient hyperparameter feature z_hat and the decoded image feature y_hat. These two are combined and input to obtain a more accurate predicted value mu. The predicted value mu is used to obtain the residual by calculating the difference from the original feature, and is added to the decoded residual to obtain the reconstructed feature y_hat.
[0073] Note that the mean prediction network is a selectable neural network; that is, it is not necessary to have a mean prediction network, meaning that it is not necessary to determine the predicted value mu via the mean prediction network. The dashed box in Figure 4 indicates that the mean prediction network is selectable.
[0074] The encoding side can obtain image features y and then determine residual features r based on image features y and predicted values mu. For example, residual features r are defined as the difference between image features y and predicted values mu. Subsequently, feature processing is performed on residual features r to obtain image features s. This feature processing process is not limited to any specific method. In this case, an average value prediction network is required, and the predicted values mu are provided by this network. Alternatively, the encoding side may obtain image features s after obtaining image features y, and this feature processing process is not limited to any specific method. In this case, an average value prediction network is not required, and the dashed box indicates that the residual process is a selectable process.
[0075] After obtaining the image feature s, the encoding side may quantize the image feature s to obtain the image quantized feature corresponding to the image feature s; that is, the Q operation in Figure 4 is a quantization process. After obtaining the image quantized feature corresponding to the image feature s, the encoding side may encode the image quantized feature to obtain Bitstream #2 (i.e., the second bitstream) corresponding to the current image block; that is, the AE operation in Figure 4 represents an encoding process such as an entropy encoding process. Alternatively, the encoding side may directly encode the image feature s to obtain Bitstream #2 corresponding to the current image block without performing a quantization process of the image feature s.
[0076] After obtaining Bitstream#2 corresponding to the current image block, the encoding side may send Bitstream#2 corresponding to the current image block to the decoding side. The decoding side's processing process for Bitstream#2 corresponding to the current image block will be described in subsequent embodiments.
[0077] After obtaining Bitstream#2 corresponding to the current image block, the encoding side may decode Bitstream#2 to obtain image quantization features, i.e., AD in Figure 4 represents the decoding process, and then the encoding side may dequantize the image quantization features to obtain image features s', image features s' may be the same as or different from image features s, and the IQ operation in Figure 4 is the dequantization process. Alternatively, after obtaining Bitstream#2 corresponding to the current image block, the encoding side may decode Bitstream#2 to obtain image features s' without performing the dequantization process of image quantization features.
[0078] The encoding side may, after obtaining image features s', perform feature reconstruction (i.e., the reverse process of feature processing) on image features s'. This feature reconstruction process is not limited to any specific method; any feature reconstruction method may be used, resulting in a residual feature r_hat, which may be the same as or different from the residual feature r. After obtaining the residual feature r_hat, the encoding side determines the image feature y_hat based on the residual feature r_hat and the predicted value mu. The image feature y_hat may be the same as or different from the image feature y. For example, the sum of the residual feature r_hat and the predicted value mu is taken as the image feature y_hat. In this case, it is necessary to implement an mean prediction network, which provides the predicted value mu. Alternatively, the encoding side may, after obtaining image features s', perform feature reconstruction (i.e., the reverse process of feature processing) on image features s' to obtain the image feature y_hat. The image feature y_hat may be the same as or different from the image feature y. In this case, it is not necessary to implement an mean prediction network, and the dashed box indicates that the residual process is a selectable process.
[0079] The encoding side may, after obtaining the image feature y_hat, perform a composite transformation on the image feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the image feature y_hat can be input to a composite transformation network, and the composite transformation network can perform a composite transformation on the image feature y_hat to obtain the reconstructed image block x_hat. At this point, the image reconstruction process is complete.
[0080] In one possible embodiment, when the encoding side encodes image quantization features or image features s to obtain Bitstream#2 corresponding to the current image block, the encoding side first needs to determine a probability distribution model, and then encodes the image quantization features or image features s based on the probability distribution model. Similarly, when decoding Bitstream#2, the encoding side first needs to determine a probability distribution model, and then decodes Bitstream#2 based on the probability distribution model.
[0081] To obtain a probability distribution model, as shown in Figure 4, the encoding side may, after obtaining the coefficient hyperparameter feature z_hat, perform an inverse coefficient hyperparameter feature transformation on the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter p. Alternatively, the coefficient hyperparameter feature z_hat may be input to a probabilistic hyperparameter decoding network, and the probabilistic hyperparameter decoding network may perform an inverse coefficient hyperparameter feature transformation on the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter p. Or, after obtaining the probability distribution parameter p, a probability distribution model may be generated based on the probability distribution parameter p. Here, the probabilistic hyperparameter decoding network may be a trained neural network, and the training process of this probabilistic hyperparameter decoding network is not limited; it is sufficient that an inverse coefficient hyperparameter feature transformation can be performed on the coefficient hyperparameter feature z_hat.
[0082] In one possible embodiment, the encoding process may be performed by a deep learning model or a neural network model to realize an end-to-end image compression and encoding process, but is not limited to this encoding process.
[0083] Example 4: Compared to Examples 1 and 2, the decoding process is shown in Figure 5. Of course, Figure 5 is merely one example of the decoding process and is not limited to this decoding process.
[0084] After obtaining Bitstream#1 corresponding to the current image block, the decoding side may decode Bitstream#1 to obtain the hyperparameter quantization feature, i.e., AD in Figure 5 represents the decoding process, and then the hyperparameter quantization feature is inversely quantized to obtain the coefficient hyperparameter feature z_hat, which may be the same as or different from the coefficient hyperparameter feature z, and the IQ operation in Figure 5 is the inverse quantization process. Alternatively, after obtaining Bitstream#1 corresponding to the current image block, the decoding side may decode Bitstream#1 to obtain the coefficient hyperparameter feature z_hat without performing the inverse quantization process of the coefficient hyperparameter feature z_hat.
[0085] For the decoding process of Bitstream#1, a fixed probability density model decoding method may be used, but is not limited to this method.
[0086] The image may be divided into one image block or into multiple image blocks. If the image is divided into one image block, the current image block x may be an image itself, that is, the decoding process for the image block may be applied directly to the image.
[0087] After obtaining the coefficient hyperparameter feature z_hat, the decoding side may perform a context-based prediction based on the coefficient hyperparameter feature z_hat of the current image block and the image feature y_hat of the previous image block (the process for determining the image feature y_hat will be described in subsequent examples) to obtain a predicted value mu (i.e., mean mu) corresponding to the current image block. For example, the coefficient hyperparameter feature z_hat and the image feature y_hat may be input to a mean prediction network, and the mean prediction network may determine the predicted value mu based on the coefficient hyperparameter feature z_hat and the image feature y_hat. This prediction process is not limited. Here, for the context-based prediction process, the input includes the coefficient hyperparameter feature z_hat and the decoded image feature y_hat, and both are combined and input to obtain a more accurate predicted value mu.
[0088] Note that the mean prediction network is a selectable neural network; that is, it is not necessary to have a mean prediction network, meaning that it is not necessary to determine the predicted value mu via the mean prediction network. The dashed box in Figure 5 indicates that the mean prediction network is selectable.
[0089] After obtaining Bitstream#2 corresponding to the current image block, the decoding side may decode Bitstream#2 to obtain image quantization features, i.e., AD in Figure 5 represents the decoding process, and then the decoding side may dequantize the image quantization features to obtain image features s', image features s' may be the same as or different from image features s, and the IQ operation in Figure 5 is the dequantization process. Alternatively, after obtaining Bitstream#2 corresponding to the current image block, the decoding side may decode Bitstream#2 to obtain image features s' without performing the dequantization process of image quantization features.
[0090] The decoding side may, after obtaining image features s', perform feature reconstruction (i.e., the reverse process of feature processing) on image features s' to obtain residual features r_hat, where residual features r_hat may be the same as or different from residual features r. After obtaining residual features r_hat, the decoding side determines image features y_hat based on residual features r_hat and predicted values mu, where image features y_hat may be the same as or different from image features y, for example, the sum of residual features r_hat and predicted values mu is taken as image features y_hat. In this case, it is necessary to set up an mean prediction network, which provides the predicted values mu. Alternatively, the decoding side may, after obtaining image features s', perform feature reconstruction on image features s' to obtain image features y_hat, where image features y_hat may be the same as or different from image features y. In this case, it is not necessary to set up an mean prediction network, and the dashed box indicates that the residual process is a selectable process.
[0091] The decoding side may, after obtaining the image feature y_hat, perform a composite transformation on the image feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the image feature y_hat may be input to a composite transformation network, and the composite transformation network may perform a composite transformation on the image feature y_hat to obtain the reconstructed image block x_hat. At this point, the image reconstruction process is complete.
[0092] In one possible embodiment, the decoding side must first determine a probability distribution model when decoding Bitstream#2, and then decode Bitstream#2 based on this probability distribution model. To obtain the probability distribution model, as shown in Figure 5, the decoding side may first obtain the coefficient hyperparameter feature z_hat, then perform an inverse coefficient hyperparameter feature transformation on the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter p. For example, the coefficient hyperparameter feature z_hat may be input to a probabilistic hyperparameter decoding network, and the probabilistic hyperparameter decoding network may perform an inverse coefficient hyperparameter feature transformation on the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter p. Alternatively, after obtaining the probability distribution parameter p, a probability distribution model may be generated based on the probability distribution parameter p. Here, the probabilistic hyperparameter decoding network may be a trained neural network, and the training process of this probabilistic hyperparameter decoding network is not limited; it is sufficient that the probability distribution parameter p can be obtained by performing an inverse coefficient hyperparameter feature transformation on the coefficient hyperparameter feature z_hat.
[0093] In one possible embodiment, the decoding process may be performed by a deep learning model or a neural network model to realize an end-to-end image compression and encoding process, and is not limited to this decoding process.
[0094] Example 5: For Examples 1, 2, 3, and 4, an average prediction network may or may not be deployed. An example is given of deploying an average prediction network to improve the performance of feature coding. When deploying an average prediction network, it is necessary to use the reconstructed values of adjacent features to obtain accurate predicted values of image features (i.e., context-based prediction is performed based on the coefficient hyperparameter feature z_hat of the current image block and the image feature y_hat of the previous image block to obtain the predicted value mu corresponding to the current image block). Therefore, the feature value on the right (using the reconstructed value of the left or upper-left feature as a reference) can only undergo the coding process of the current feature after the coding and reconstruction of the left or upper-left feature value is complete. Due to this dependency, coding between features can only be performed serially and parallel execution is impossible, thus increasing complexity.
[0095] In response to the above findings, this embodiment employs block partitioning coding and decoding to improve parallelism, effectively conserve the cache for storing features, achieve higher bitrate control accuracy, reduce coding performance loss, have better coding performance and bitrate control accuracy, lower peak memory occupancy, shorter single-block decoding time, and possess high-speed parallel decoding capability.
[0096] The encoding process is shown in Figure 6, and this is not an exhaustive description of the encoding process.
[0097] 1. After obtaining the original image x, the block division unit decides whether or not to divide the original image x into blocks based on the image resolution. For example, if the image resolution is greater than a threshold, the original image x is divided into blocks; otherwise, the original image x is not divided into blocks. If block division is performed, the original image x is divided into multiple image blocks x_p, with overlapping portions between adjacent image blocks. If block division is not performed, image blocks x_p are the original image x, meaning that the original image x contains only one image block, and each image block x_p can be recorded as the current image block. To facilitate explanation, the processing process of one current image block x_p will be described below as an example.
[0098] 2. After obtaining the current image block x_p, an analytical transformation is performed on the current image block x_p via an analytical transformation network to obtain the image feature y_p (feature block y_p) corresponding to the current image block x_p. For example, by transforming the current image block x_p into an image feature y_p in the latent domain via the analytical transformation network, all subsequent processes can be operated in the latent domain.
[0099] 3. The combined classification unit performs combined classification on the image feature y_p to obtain features y_i (i=1...N) to be encoded by N stage subblocks, that is, it divides the image feature y_p into features y_i to be encoded by N stage subblocks.
[0100] 4. For each stage subblock corresponding to the current image block, obtain the coefficient hyperparameter feature corresponding to that stage subblock. For example, perform a coefficient hyperparameter feature transformation on the feature y_i to be encoded in stage subblock i to obtain the coefficient hyperparameter feature z_i corresponding to stage subblock i. Alternatively, the feature y_i to be encoded may be input to a hyperparameter coding network (i.e., a neural network), and the hyperparameter coding network may perform a coefficient hyperparameter feature transformation on the feature y_i to be encoded to obtain the coefficient hyperparameter feature z_i.
[0101] 5. For each stage subblock corresponding to the current image block, the coefficient hyperparameter features of the stage subblock are encoded into the first bitstream (Bitstream #1) of the current image block. For example, the coefficient hyperparameter feature z_i corresponding to stage subblock i is quantized to obtain a hyperparameter quantization feature corresponding to the coefficient hyperparameter feature z_i, and this hyperparameter quantization feature is encoded to obtain the first bitstream corresponding to the current image block. Alternatively, the coefficient hyperparameter feature z_i corresponding to stage subblock i is directly encoded to obtain the first bitstream corresponding to the current image block. After obtaining the first bitstream corresponding to the current image block, the first bitstream corresponding to the current image block is transmitted to the decoding side, and the decoding side's processing process for the first bitstream corresponding to the current image block will be described in subsequent embodiments.
[0102] 6. After obtaining the first bitstream corresponding to the current image block, the first bitstream may be decoded to obtain coefficient hyperparameter features for each stage subblock of the current image block, for example, the coefficient hyperparameter feature z_hat_i corresponding to stage subblock i (i=1...N).
[0103] For example, the first bitstream may be decoded to obtain the hyperparameter quantization feature of stage subblock i, and then the hyperparameter quantization feature may be inversely quantized to obtain the coefficient hyperparameter feature z_hat_i corresponding to stage subblock i. Alternatively, the first bitstream may be decoded without performing the inverse quantization process to obtain the coefficient hyperparameter feature z_hat_i corresponding to stage subblock i.
[0104] The encoding process for the first bitstream may employ a fixed probability density model encoding method, and the decoding process for the first bitstream may employ a fixed probability density model decoding method; however, the encoding and decoding processes are not limited to these methods.
[0105] 7. For each stage subblock corresponding to the current image block, the probability distribution parameter is determined based on the coefficient hyperparameter features of the stage subblock. For example, the coefficient hyperparameter feature z_hat_i corresponding to stage subblock i may be subjected to an inverse transformation of the coefficient hyperparameter features to obtain the probability distribution parameter p_i corresponding to stage subblock i. Alternatively, the coefficient hyperparameter feature z_hat_i can be input into a probabilistic hyperparameter decoding network, the probabilistic hyperparameter decoding network can perform an inverse transformation of the coefficient hyperparameter feature z_hat_i to obtain the probability distribution parameter p_i, and after obtaining the probability distribution parameter p_i, a probability distribution model can be generated based on the probability distribution parameter p_i.
[0106] 8. After obtaining the features to be encoded for N stage subblocks, for the feature y_1 of the first stage subblock to be encoded, input the coefficient hyperparameter feature z_hat_1 of the stage subblock into the mean prediction network to obtain the mean feature of the stage subblock (i.e., predicted value mu_1, mean value mu_1). Alternatively, obtain the set default reference feature, input the coefficient hyperparameter feature z_hat_1 of the stage subblock and the default reference feature into the mean prediction network to obtain the mean feature of the stage subblock.
[0107] Based on the feature y_1 to be encoded and the mean feature mu_1 of the first stage subblock, the residual feature r_1 is determined. For example, the difference between the feature y_1 to be encoded and the mean feature mu_1 is taken as the residual feature r_1. Then, feature processing is performed on the residual feature r_1 to obtain the residual feature r_1 after feature processing. This feature processing process is not limited to this method, and any feature processing method may be used. Of course, feature processing is an optional step, and feature processing is not required for the residual feature r_1.
[0108] The residual feature r_1 (or the residual feature r_1 after feature processing) is quantized to obtain an image quantization feature, and this image quantization feature is encoded to obtain a second bitstream (Bitstream #2) corresponding to the current image block. Alternatively, the residual feature r_1 (or the residual feature r_1 after feature processing) may be directly encoded to obtain a second bitstream corresponding to the current image block. After obtaining the second bitstream corresponding to the current image block, the second bitstream corresponding to the current image block is transmitted to the decoding side, and the decoding side's processing process for the second bitstream corresponding to the current image block will be described in subsequent embodiments. Here, when encoding the image quantization feature or residual feature r_1, the image quantization feature or residual feature r_1 may be encoded using a probability distribution model corresponding to the probability distribution parameter p_1 of the first stage subblock to obtain the second bitstream.
[0109] After obtaining the second bitstream, the second bitstream is decoded using a probability distribution model corresponding to the probability distribution parameter p_1 of the first stage subblock to obtain image quantization features, the image quantization features are dequantized, feature reconstruction is performed on the dequantized features to obtain residual features r_hat_1, or, after dequantizing the image quantization features, residual features r_hat_1 may be obtained directly without performing the feature reconstruction process. Alternatively, after decoding the second bitstream, feature reconstruction is performed on the decoded features to obtain residual features r_hat_1, or, after decoding the second bitstream, residual features r_hat_1 may be obtained directly.
[0110] After obtaining the residual feature r_hat_1, the reconstruction feature y_hat_1 of the first stage subblock is determined based on the residual feature r_hat_1 and the mean feature mu_1. For example, the sum of the residual feature r_hat_1 and the mean feature mu_1 is taken as the reconstruction feature y_hat_1.
[0111] 9. After obtaining the features to be encoded for N stage subblocks, for the feature y_i to be encoded for the i-th (i>1 and i is less than or equal to N)-th stage subblock, the coefficient hyperparameter feature z_hat_i of stage subblock i and the reference feature of stage subblock i are input into the mean prediction network to obtain the mean feature of stage subblock i (i.e., predicted value mu_i, mean value mu_i).
[0112] The reference feature of stage subblock i (i.e., the i-th stage subblock) may be obtained based on the reconstructed features of the i-1 stage subblocks preceding stage subblock i, i.e., based on the reconstructed features y_hat_1 to y_hat_i-1. For example, the reference feature of stage subblock i may be obtained based on all the reconstructed features in y_hat_1 to y_hat_i-1, i.e., the reference feature includes y_hat_1 to y_hat_i-1. Or, the reference feature of stage subblock i may be obtained based on some of the reconstructed features in y_hat_1 to y_hat_i-1, i.e., the reference feature includes some of the reconstructed features in y_hat_1 to y_hat_i-1. Or, the reference feature of stage subblock i may be obtained based on the reconstructed features of the i-1th stage subblock, i.e., the reference feature of stage subblock i is y_hat_i-1.
[0113] Based on the feature y_i to be encoded and the mean feature mu_i of the i-th stage subblock, the residual feature r_i is determined. For example, the difference between the feature y_i to be encoded and the mean feature mu_i is taken as the residual feature r_i. Feature processing is performed on the residual feature r_i to obtain the residual feature r_i after feature processing. Of course, feature processing is not required for the residual feature r_i.
[0114] The residual feature r_i (or the residual feature r_i after feature processing) is quantized to obtain the image quantization feature, and the image quantization feature is encoded to obtain the second bitstream corresponding to the current image block. Alternatively, the residual feature r_i (or the residual feature r_i after feature processing) is directly encoded to obtain the second bitstream corresponding to the current image block. The second bitstream corresponding to the current image block is transmitted to the decoding side, and the decoding side's processing process for the second bitstream corresponding to the current image block will be described in subsequent embodiments. Here, when encoding the image quantization feature or residual feature r_i, the image quantization feature or residual feature r_i may be encoded using a probability distribution model corresponding to the probability distribution parameter p_i of the i-th stage subblock to obtain the second bitstream.
[0115] After obtaining the second bitstream, the second bitstream is decoded using a probability distribution model corresponding to the probability distribution parameter p_i of the i-th stage subblock to obtain image quantization features, the image quantization features are dequantized, feature reconstruction is performed on the dequantized features to obtain residual features r_hat_i, or the residual features r_hat_i may be obtained directly after dequantizing the image quantization features. Alternatively, the residual features r_hat_i may be obtained by decoding the second bitstream and then performing feature reconstruction on the decoded features, or the residual features r_hat_i may be obtained directly after decoding the second bitstream.
[0116] After obtaining the residual feature r_hat_i, the reconstruction feature y_hat_i of the i-th stage subblock is determined based on the residual feature r_hat_i and the mean feature mu_i. For example, the sum of the residual feature r_hat_i and the mean feature mu_i is taken as the reconstruction feature y_hat_i.
[0117] 10. The clustering block partitioning unit may obtain reconstruction features y_hat_i (i=1...N) of N stage subblocks, perform feature aggregation on the reconstruction features y_hat_i of the N stage subblocks, and obtain aggregated features. Then, the clustering block partitioning unit partitions the aggregated features into blocks to obtain at least one block partitioning feature y_hat_p, i.e., a feature block y_hat_p.
[0118] 11. For each block partitioning feature y_hat_p, the clustering block partitioning unit inputs the block partitioning feature y_hat_p into the composite transformation network, and the composite transformation network outputs the block partitioning reconstructed image block x_hat_p corresponding to the block partitioning feature y_hat_p.
[0119] 12. The merging unit may obtain block-partitioned reconstructed image blocks x_hat_p corresponding to multiple block-partitioning features y_hat_p, and merge these block-partitioned reconstructed image blocks x_hat_p to obtain a reconstructed image block x_hat corresponding to the current image block. Clearly, if the original image x is divided into multiple current image blocks x_p, then the reconstructed image blocks x_hat corresponding to all current image blocks x_p can constitute a reconstructed image. If the original image x corresponds to one current image block x_p, then the reconstructed image block x_hat is a reconstructed image.
[0120] In the above process, the encoding side may assign different bitrate control parameters to different stage subblocks of different image blocks based on region of interest or bitrate information, so that different stage subblocks of different image blocks consume different bitrates.
[0121] For example, if it is necessary to perform special processing on a region of interest (e.g., non-uniform processing to improve the quality of the region of interest) and to avoid uniform quantization distortion in all image blocks, the bitrate control parameter may also be a quantization parameter (the bitrate control parameter needs to be transmitted to the decoding side via the bitstream), and as shown in Figure 6, the bitrate control parameter affects the quantization process (Q) and inverse quantization process (IQ) of the first bitstream, and affects the quantization process (Q) and inverse quantization process (IQ) of the second bitstream.
[0122] For example, when quantizing the coefficient hyperparameter feature z_i corresponding to stage subblock i, if stage subblock i is located within the region of interest, the bitrate control parameter is used to quantize the coefficient hyperparameter feature z_i corresponding to stage subblock i, and this bitrate control parameter is used to ensure that the image quality corresponding to stage subblock i is higher. Similarly, when dequantizing the hyperparameter quantization feature corresponding to stage subblock i, it is necessary to dequantize the hyperparameter quantization feature corresponding to stage subblock i using the same bitrate control parameter.
[0123] When quantizing the residual feature r_i (or the residual feature r_i after feature processing) corresponding to stage subblock i, if stage subblock i is located within the region of interest, the bitrate control parameter is used to quantize the residual feature r_i corresponding to stage subblock i, and this bitrate control parameter is used to ensure that the image quality corresponding to stage subblock i is higher. Similarly, when dequantizing the image quantization features corresponding to stage subblock i, it is necessary to dequantize the image quantization features corresponding to stage subblock i using the same bitrate control parameter.
[0124] For example, to achieve precise bitrate control, the bitrate control parameter may also be a control parameter of the composite transformation network (the bitrate control parameter needs to be transmitted to the decoding side via the bitstream), and as shown in Figure 6, the bitrate control parameter affects both the analytical transformation network and the composite transformation network. For example, when analyzing and transforming the current image block x_p via the analytical transformation network, the bitrate control parameter and the current image block x_p are input to the analytical transformation network, and the analytical transformation network outputs the image feature y_p corresponding to the current image block x_p. When inputting the block partitioning feature y_hat_p to the composite transformation network, the block partitioning feature y_hat_p and the bitrate control parameter are input to the composite transformation network, and the composite transformation network outputs the block partitioned reconstructed image block x_hat_p corresponding to the block partitioning feature y_hat_p.
[0125] Based on the set target bitrate (whether each image block corresponds to one target bitrate, or each image corresponds to one target bitrate, or multiple images correspond to one target bitrate), the encoding side attempts to adopt different bitrate control parameters. After adopting a bitrate control parameter, if the bitrate corresponding to the bitstream for the current image block (e.g., the first bitstream and the second bitstream) is the target bitrate or close to the target bitrate (i.e., the difference between the two is less than the threshold), then this bitrate control parameter is the bitrate control parameter that is ultimately adopted, and the encoding side can transmit this bitrate control parameter to the decoding side.
[0126] Example 6: The decoding process is shown in Figure 7, and this decoding process is not the only one we can use.
[0127] 1. After obtaining the first bitstream (Bitstream#1) corresponding to the current image block, the first bitstream is decoded to obtain the coefficient hyperparameter features of each stage subblock of the current image block, for example, the coefficient hyperparameter feature z_hat_i corresponding to stage subblock i (i=1...N).
[0128] For example, the first bitstream may be decoded to obtain the hyperparameter quantization feature of stage subblock i, and then the hyperparameter quantization feature may be inversely quantized to obtain the coefficient hyperparameter feature z_hat_i corresponding to stage subblock i. Alternatively, the first bitstream may be decoded without performing the inverse quantization process to obtain the coefficient hyperparameter feature z_hat_i corresponding to stage subblock i. The decoding process for the first bitstream may employ a decoding method of a fixed probability density model, and is not limited to this decoding process.
[0129] 2. For each stage subblock corresponding to the current image block, the probability distribution parameter is determined based on the coefficient hyperparameter features of the stage subblock. For example, the coefficient hyperparameter feature z_hat_i corresponding to stage subblock i may be subjected to an inverse transformation of the coefficient hyperparameter features to obtain the probability distribution parameter p_i corresponding to stage subblock i. Alternatively, the coefficient hyperparameter feature z_hat_i can be input into a probabilistic hyperparameter decoding network, the probabilistic hyperparameter decoding network can perform an inverse transformation of the coefficient hyperparameter feature z_hat_i to obtain the probability distribution parameter p_i, and after obtaining the probability distribution parameter p_i, a probability distribution model can be generated based on the probability distribution parameter p_i.
[0130] 3. For each stage subblock corresponding to the current image block, the second bitstream (Bitstream#2) of the current image block is decoded based on the probability distribution parameters (probability distribution model) of the stage subblock, and the residual features of the stage subblock are obtained.
[0131] For example, the second bitstream may be decoded using a probability distribution model corresponding to the probability distribution parameter p_i of the i-th (i=1...N)-th stage subblock to obtain image quantization features, the image quantization features may be dequantized, feature reconstruction may be performed on the dequantized features to obtain residual features r_hat_i, or the residual features r_hat_i may be obtained directly after dequantizing the image quantization features. Alternatively, the second bitstream may be decoded and then feature reconstruction may be performed on the decoded features to obtain residual features r_hat_i, or the residual features r_hat_i may be obtained directly after decoding the second bitstream.
[0132] 4. After obtaining the residual features of N stage subblocks, input the coefficient hyperparameter feature z_hat_1 of the first stage subblock into the mean prediction network to obtain the mean feature (i.e., predicted value mu_1, mean value mu_1) of the stage subblock. Alternatively, obtain the set default reference feature, input the coefficient hyperparameter feature z_hat_1 of the stage subblock and the default reference feature into the mean prediction network to obtain the mean feature of the stage subblock.
[0133] Subsequently, the reconstruction feature y_hat_1 of the first stage subblock is determined based on the residual feature r_hat_1 and the mean feature mu_1. For example, the sum of the residual feature r_hat_1 and the mean feature mu_1 may be used as the reconstruction feature y_hat_1.
[0134] 5. After obtaining the residual features of N stage subblocks, for the residual feature r_hat_i of the i-th (i>1 and i is less than or equal to N) stage subblock, the coefficient hyperparameter feature z_hat_i of stage subblock i and the reference feature of stage subblock i are input into the mean prediction network to obtain the mean feature of stage subblock i (i.e., predicted value mu_i, mean value mu_i).
[0135] The reference feature of stage subblock i (i.e., the i-th stage subblock) may be obtained based on the reconstructed features of the i-1 stage subblocks preceding stage subblock i, i.e., based on the reconstructed features y_hat_1 to y_hat_i-1. For example, the reference feature of stage subblock i may be obtained based on all the reconstructed features in y_hat_1 to y_hat_i-1, i.e., the reference feature includes y_hat_1 to y_hat_i-1. Or, the reference feature of stage subblock i may be obtained based on some of the reconstructed features in y_hat_1 to y_hat_i-1, i.e., the reference feature includes some of the reconstructed features in y_hat_1 to y_hat_i-1. Or, the reference feature of stage subblock i may be obtained based on the reconstructed features of the i-1th stage subblock, i.e., the reference feature of stage subblock i is y_hat_i-1.
[0136] Here, to control the cache size, the number of reconstructed features included in the reference feature cannot exceed a predetermined number, where M is the predetermined number. If the number of reconstructed features is less than or equal to M, the reference feature includes all reconstructed features in y_hat_1 to y_hat_i-1. If the number of reconstructed features is greater than M, the reference feature includes M reconstructed features (i.e., some reconstructed features) in y_hat_1 to y_hat_i-1. To select M reconstructed features from y_hat_1 to y_hat_i-1, the M reconstructed features can be selected based on the default policy. For example, y_hat_i-1, y_hat_i-2, y_hat_i-3, etc., are selected until the M reconstructed features whose decoding order is closest to the current stage subblock are selected. Of course, the M reconstructed features can be selected in other ways, and this is not limited to that. Alternatively, the bitstream can be decoded to obtain the index values of the M reconstructed features, and the M reconstructed features can be selected based on the obtained index values.
[0137] Subsequently, the reconstruction feature y_hat_i of the i-th stage subblock is determined based on the residual feature r_hat_i and the mean feature mu_i. For example, the sum of the residual feature r_hat_i and the mean feature mu_i may be used as the reconstruction feature y_hat_i.
[0138] 6. The clustering block partitioning unit may obtain reconstruction features y_hat_i (i=1...N) of N stage subblocks, perform feature aggregation on the reconstruction features y_hat_i of the N stage subblocks, and obtain aggregated features. Then, the clustering block partitioning unit partitions the aggregated features into blocks to obtain at least one block partitioning feature y_hat_p, i.e., a feature block y_hat_p.
[0139] 7. For each block partitioning feature y_hat_p, the clustering block partitioning unit inputs the block partitioning feature y_hat_p into the composite transformation network, and the composite transformation network outputs the block partitioning reconstructed image block x_hat_p corresponding to the block partitioning feature y_hat_p.
[0140] 8. The merging unit may obtain block-partitioned reconstructed image blocks x_hat_p corresponding to multiple block-partitioning features y_hat_p, and merge these block-partitioned reconstructed image blocks x_hat_p to obtain a reconstructed image block x_hat corresponding to the current image block. Clearly, if the original image x is divided into multiple current image blocks x_p, then the reconstructed image blocks x_hat corresponding to all current image blocks x_p can constitute a reconstructed image. If the original image x corresponds to one current image block x_p, then the reconstructed image block x_hat is a reconstructed image.
[0141] In the above process, different bitrate control parameters may be assigned to different stage subblocks of different image blocks based on region of interest or bitrate information, thereby causing different stage subblocks of different image blocks to consume different bitrates.
[0142] For example, if a special processing is required for a region of interest (e.g., non-uniform processing to improve the quality of the region of interest), the bitrate control parameter may also be a quantization parameter (the bitrate control parameter needs to be transmitted to the decoding side via the bitstream), and as shown in Figure 7, the bitrate control parameter affects the inverse quantization process (IQ) of the first bitstream and the inverse quantization process (IQ) of the second bitstream.
[0143] For example, when dequantizing the hyperparameter quantization features corresponding to stage subblock i, if bitrate control parameters corresponding to stage subblock i are decoded from the bitstream, and these bitrate control parameters are intended to ensure higher image quality for the stage subblock i, then the hyperparameter quantization features corresponding to stage subblock i can be dequantized using these bitrate control parameters.
[0144] Furthermore, for example, when dequantizing the image quantization features corresponding to stage subblock i, if bitrate control parameters corresponding to stage subblock i are decoded from the bitstream, and these bitrate control parameters are intended to ensure higher image quality for the image corresponding to stage subblock i, then these bitrate control parameters can be used to dequantize the image quantization features corresponding to stage subblock i.
[0145] For example, to perform precise bitrate control, the bitrate control parameter may also be a control parameter of the composite transformation network (the bitrate control parameter needs to be transmitted to the decoding side via the bitstream, and the control parameter of the composite transformation network means that the composite transformation network can be determined based on the bitrate control parameter), and as shown in Figure 7, the bitrate control parameter affects the composite transformation network. For example, when inputting a block partitioning feature y_hat_p into the composite transformation network, the block partitioning feature y_hat_p and the bitrate control parameter are input into the composite transformation network, and the composite transformation network outputs a block partitioning reconstructed image block x_hat_p corresponding to the block partitioning feature y_hat_p.
[0146] Example 7: In Examples 5 and 6, for the i-th (i=1...N)-th stage subblock, the reconstruction feature y_hat_i (i.e., the sum of the residual feature r_hat_i and the mean feature mu_i) of the i-th stage subblock is obtained, and the reconstruction feature y_hat_i is used to determine the mean feature of the subsequent stage subblock, and / or the reconstruction feature y_hat_i is used to determine the reconstructed image block. Here, the use of the reconstruction feature y_hat_i to determine the mean feature of the subsequent stage subblock means that the reconstruction feature y_hat_i is used as a reference feature of the subsequent stage subblock to determine the mean feature of the subsequent stage subblock. The use of the reconstruction feature y_hat_i to determine the reconstructed image block means that the reconstruction feature y_hat_i is input to the clustering block partitioning unit to determine the reconstructed image block corresponding to the current image block.
[0147] For the i-th (i=1...N)-th stage subblock, after obtaining the reconstructed feature y_hat_i of the i-th stage subblock, feature enhancement may be performed on the reconstructed feature y_hat_i to obtain the enhanced reconstructed feature y_hat_i'. The enhanced reconstructed feature y_hat_i' is used to determine the mean feature of the subsequent stage subblock, and / or the enhanced reconstructed feature y_hat_i' is used to determine the reconstructed image block. Here, the use of the enhanced reconstructed feature y_hat_i' to determine the mean feature of the subsequent stage subblock means that the enhanced reconstructed feature y_hat_i' is used as a reference feature of the subsequent stage subblock to determine the mean feature of the subsequent stage subblock. The use of the enhanced reconstructed feature y_hat_i' to determine the reconstructed image block means that the enhanced reconstructed feature y_hat_i' is input into the clustering block division unit to determine the reconstructed image block corresponding to the current image block. For example, the bitstream is decoded to obtain emphasis information, the reconstruction feature y_hat_i is enhanced based on the emphasis information to obtain a new reconstruction feature y_hat_i', and the new reconstruction feature y_hat_i' may be used to predict subsequent y_hat_i+1~y_hat_N, or it may be used in a subsequent clustering block partitioning process.
[0148] In one possible embodiment, the following method may be employed for the enhancement process of the reconstructed feature y_hat.
[0149] Method 1: The input is the unenhanced reconstructed feature y_hat, Threshold[idx], GreaterFlag[idx], Scale1[idx], and Scale2[idx], and the output is the enhanced reconstructed feature y_hat_en. Threshold represents the threshold value, and different idx may correspond to the same or different threshold values. idx represents an index such as 0, 1, 2, 3, etc. GreaterFlag represents the flag bit, and different idx may correspond to the same or different flag bit. Scale1 represents the scaling factor, and different idx may correspond to the same or different Scale1. Scale2 represents the scaling factor, and different idx may correspond to the same or different Scale2.
[0150] Step 1: Determine the mask[c,i,j] in the filter (index is idx) for each set. For example, mask[c,i,j] can be determined using the following formula, and of course, the following formula is just one example and is not limiting.
number
[0151] σ[c,i,j] represents the probability distribution parameter value of [c,i,j], where c represents the channel, i represents the x-coordinate, and j represents the y-coordinate. If the probability distribution parameter value is greater than Threshold[idx], then GreaterFlag[idx] is TRUE and mask[idx,c,i,j] is 1. If the probability distribution parameter value is less than Threshold[idx], then GreaterFlag[idx] is False and mask[idx,c,i,j] is 1. Otherwise, mask[idx,c,i,j] is 0.
[0152] Step 2: The filter for each set (index is idx) is executed sequentially, and if mask[idx,c,i,j] is 1, then y_hat_en[c,i,j]=y_hat[c,i,j]+mean_hat[c,i,j]*Scale2[idx]+residual_hat[c,i,j]*Scale1[idx]; y_hat[c,i,j]=y_hat_en[c,i,j]. If there are multiple sets of filters, y_hat is compensated multiple times. Here, mean_hat represents the mean feature, i.e., mu_i, and residual_hat represents the residual feature, i.e., r_hat.
[0153] In Method 1, the scaling factor of each channel c is the same, while in Method 2, different channels c are allowed to have different scaling factors, i.e., the scaling factors are Scale1[c,idx] and Scale2[c,idx], and a control switch may be used to control whether or not to scale the current feature, or the scaling factor of a channel may be changed to 0 to turn off scaling. Method 2 differs from Method 1 in that the emphasis method is y_hat_en[c,i,j]=y_hat[c,i,j]+mean_hat[c,i,j]*Scale2[c,idx]+residual_hat[c,i,j]*Scale1[c,idx].
[0154] Comparing Method 3 with Method 1, the implementation methods of both are similar, and the difference is that the emphasis method is y_hat_en[c,i,j]=y_hat[c,i,j]*Scale1[idx]+mean_hat[c,i,j]*Scale2[idx].
[0155] Comparing Method 4 with Method 1, the implementation methods of both are similar, the difference being that the emphasis method is y_hat_en[c,i,j]=y_hat[c,i,j]*Scale1[idx]+residual_hat[c,i,j]*Scale2[idx]. In Method 4, Scale1[idx] may be 1 by default, and only the change in the value of Scale2[idx] affects the adjustment method.
[0156] Example 8: In Examples 5 and 6, the coefficient hyperparameter feature z_hat_1 of the first stage subblock may be input to the mean prediction network, and the mean prediction network will output the mean feature of the first stage subblock, without limiting the prediction process of this mean prediction network. For the i-th (i=1...N)th stage subblock, the coefficient hyperparameter feature z_hat_i of the i-th stage subblock and the reference feature of the i-th stage subblock (which may be the default reference feature) may be input to the mean prediction network, and the mean prediction network will output the mean feature of the i-th stage subblock, and the prediction process of the mean prediction network will be described below.
[0157] Figure 8A shows a schematic diagram of the mean prediction network, which may include a first prediction network, a second prediction network, and a prediction fusion network. Of course, this is merely an example, and the structure of this mean prediction network is not limited.
[0158] The coefficient hyperparameter feature z_hat_i of the i-th stage subblock may be input to the first prediction network, and the first prediction feature corresponding to the coefficient hyperparameter feature z_hat_i may be obtained via the first prediction network. Alternatively, the reference feature of the i-th stage subblock (e.g., reconstructed feature y_hat_1~y_hat_i-1) may be input to the second prediction network, and the second prediction feature corresponding to the reference feature may be obtained via the second prediction network. Then, the first prediction feature and the second prediction feature are feature-joined to obtain a combined feature, the combined feature is input to the prediction fusion network, and the combined feature is processed via the prediction fusion network to obtain the mean feature mu_i of the i-th stage subblock.
[0159] For the first prediction network, feature enhancement and upsampling operations are performed on the coefficient hyperparameter feature z_hat_i via the first prediction network to obtain the first predicted feature. The feature enhancement operation may include, but is not limited to, a convolution operation, or a convolution operation and an activation operation. Of course, the above is merely an example of a feature enhancement operation and is not limited to this feature enhancement operation. The upsampling operation may include, but is not limited to, a deconvolution operation, a crop operation and an activation operation, or a deconvolution operation, a crop operation, an activation operation and a convolution operation. Of course, the above is merely an example of an upsampling operation and is not limited to this upsampling operation. Exemplarily, the activation operation is a ReLU operation.
[0160] For example, Figure 8B is a schematic diagram of the first prediction network, which consists of three enhanced networks and two 2x upsampling networks. Of course, the number of enhanced networks may be more or less, the number of 2x upsampling networks may be more or less, and the 2x upsampling networks may be replaced with upsampling networks of other magnifications.
[0161] After the coefficient hyperparameter feature z_hat_i passes through enhancement network 1 and 2x upsampling network 1, enhanced large-size feature 1 is obtained, and the spatial size (width, height) of enhanced large-size feature 1 is twice that of the coefficient hyperparameter feature z_hat_i. After enhanced large-size feature 1 passes through enhancement network 2 and 2x upsampling network 2, enhanced large-size feature 2 is obtained, and the spatial size (width, height) of enhanced large-size feature 2 is twice that of enhanced large-size feature 1. After enhanced large-size feature 2 passes through enhancement network 3, the first prediction feature is obtained.
[0162] For example, collaborative network 1 and collaborative network 2 may be the same or different, collaborative network 1 and collaborative network 3 may be the same or different, and collaborative network 2 and collaborative network 3 may be the same or different.
[0163] For example, one embodiment of a collaborative network (e.g., collaborative network 1, and / or collaborative network 2, and / or collaborative network 3) is shown in Figure 8C, namely, the collaborative network consists of one convolutional layer, which may be a 1×1 convolutional layer, a 3×3 convolutional layer, or a 5×5 convolutional layer, and is not limited thereto; for example, a 3×3 convolutional layer may be selected.
[0164] Furthermore, for example, one embodiment of a collaborative network (e.g., collaborative network 1, and / or collaborative network 2, and / or collaborative network 3) is shown in Figure 8D, namely, the collaborative network may consist of a convolutional layer 1, an activation layer 1, a convolutional layer 2, an activation layer 2, and a convolutional layer 3. The convolutional layers (e.g., convolutional layer 1, and / or convolutional layer 2, and / or convolutional layer 3) may be 1×1 convolutional layers, 3×3 convolutional layers, or 5×5 convolutional layers, and are not limited thereto; for example, a 3×3 convolutional layer may be selected.
[0165] The activation layer (for example, activation layer 1 and / or activation layer 2) may be a relu layer, a leaky relu layer, a sigmoid layer, a tanh (hyperbolic tangent) layer, or a Gelu (Gaussian Error Linear Units) layer, and the type of activation layer is not limited; for example, a relu layer may be selected.
[0166] Of course, the above are just two examples of collaborative networks, and we are not limited to these; any network that can achieve collaborative functionality is acceptable.
[0167] For example, 2x upsampling network 1 and 2x upsampling network 2 may be the same or different.
[0168] For example, one embodiment of a 2x upsampling network (e.g., 2x upsampling network 1 and / or 2x upsampling network 2) is shown in Figure 8E, where the 2x upsampling network may consist of a 2x deconvolutional layer, a crop layer, and an activation layer. The deconvolutional layer may be a 2x2 deconvolutional layer, a 3x3 deconvolutional layer, a 4x4 deconvolutional layer, or a 5x5 deconvolutional layer, and is not limited thereto; for example, a 4x4 deconvolutional layer may be selected. The activation layer may be a relu layer, a leaky relu layer, a sigmoid layer, a tanh layer, or a gelu layer, and is not limited thereto; for example, a relu layer may be selected. The crop layer is used to crop features so that the spatial resolution of the features is reduced without changing the number of channels.
[0169] Furthermore, for example, one embodiment of a 2x upsampling network (e.g., 2x upsampling network 1 and / or 2x upsampling network 2) is shown in Figure 8F, where the 2x upsampling network may consist of a 2x deconvolutional layer, a crop layer, an activation layer, and a convolutional layer. The deconvolutional layer may be a 2x2 deconvolutional layer, a 3x3 deconvolutional layer, a 4x4 deconvolutional layer, or a 5x5 deconvolutional layer. The activation layer may be a relu layer, or a leaky relu layer, or a sigmoid layer, or a tanh layer, or a gelu layer. The crop layer is used to crop features so that the spatial resolution of the features is reduced without changing the number of channels. The convolutional layer may be a 1x1 convolutional layer, a 3x3 convolutional layer, or a 5x5 convolutional layer, and is not limited thereto; for example, a 3x3 convolutional layer is selected.
[0170] Of course, the above are just two examples of 2x upsampling networks, and we are not limited to these; any method that can achieve upsampling functionality is acceptable.
[0171] For example, the convolutional layer in the first prediction network may be a normal 2D convolution, or it may be a group convolution, such as a double group convolution. Embodiments of group convolution will be described in subsequent examples and will not be described here.
[0172] For the second prediction network, feature concatenation is performed on all reconstructed features in the reference feature of the i-th stage subblock according to the channel dimension via the second prediction network, and a convolution operation is performed on the concatenated features to obtain the second prediction feature. Alternatively, feature addition is performed on all reconstructed features in the reference feature of the i-th stage subblock via the second prediction network, and a convolution operation is performed on the added features to obtain the second prediction feature. The reference feature includes all reconstructed features of the previous i-1 stage subblocks, or some reconstructed features of the previous i-1 stage subblocks, or the reconstructed features of the i-1th stage subblock.
[0173] For example, Figure 8G is a schematic diagram of the second prediction network, which may consist of a feature concatenation layer and a convolutional layer, and is not limited to this structure. If the reference feature contains all the reconstructed features of the previous i-1 stage subblocks (e.g., y_hat_1 to y_hat_i-1), the feature concatenation layer can concatenate these reconstructed features according to the channel dimension to obtain the concatenated feature. Alternatively, if the reference feature contains some of the reconstructed features of the previous i-1 stage subblocks (e.g., some of y_hat_1 to y_hat_i-1), the feature concatenation layer can concatenate these reconstructed features according to the channel dimension to obtain the concatenated feature. Alternatively, if the reference feature contains the reconstructed features of the i-1th stage subblock, the feature concatenation layer can directly use the reconstructed features of the i-1th stage subblock as the concatenated feature.
[0174] The feature coupling layer inputs the combined features to a convolutional layer, which then performs a convolution operation on the combined features to obtain a second predictive feature. The convolutional layer may be a 1x1, 3x3, or 5x5 convolutional layer, and is not limited to these; for example, a 3x3 convolutional layer may be selected. The convolutional layer in the second predictive network may be a normal 2D convolution, or a group convolution, such as a double group convolution, and embodiments of group convolution will be described in subsequent examples.
[0175] Furthermore, for example, the second prediction network may consist of a feature summing layer and a convolutional layer. If the reference feature includes all the reconstructed features of the previous i-1 stage subblocks, the feature summing layer performs feature summing on these reconstructed features to obtain the summed features. Alternatively, if the reference feature includes some of the reconstructed features of the previous i-1 stage subblocks, the feature summing layer performs feature summing on these reconstructed features to obtain the summed features. Alternatively, if the reference feature includes the reconstructed features of the i-1th stage subblock, the feature summing layer can directly use the reconstructed features of the i-1th stage subblock as the summed features. The feature summing layer inputs the summed features into the convolutional layer, which performs a convolution operation on the summed features to obtain the second prediction features.
[0176] Here, in order to control the cache size, the number of reconstructed features included in the reference feature cannot exceed a predetermined number, and if the predetermined number is M, then if the number of reconstructed features is less than or equal to M, the reference feature includes all reconstructed features in y_hat_1 to y_hat_i-1. Alternatively, if the number of reconstructed features is greater than M, the reference feature includes M reconstructed features in y_hat_1 to y_hat_i-1, and for example, M reconstructed features can be selected based on the default policy, for example, the M reconstructed features whose decoding order is closest to the current stage subblock may be selected. Alternatively, the bitstream may be decoded to obtain the index values of the M reconstructed features, and M reconstructed features may be selected based on the obtained index values.
[0177] For a predictive fusion network, a convolution operation and at least one fusion operation are performed on the combined feature (i.e., the combined feature of the first and second predictive features) via the predictive fusion network to obtain the mean feature mu_i of the i-th stage subblock, and each fusion operation may include an activation operation and a convolution operation. Exemplarily, the activation operation is a ReLU operation.
[0178] For example, Figure 8H is a schematic diagram of a predictive fusion network, which may consist of a convolutional layer and N fusion layers, where N is a positive integer, for example, N is 1 or 2, and each fusion layer may consist of an activation layer and a convolutional layer, and the structure of this predictive fusion network is not limited. Clearly, after the input features (i.e., the combined features) pass through the convolutional layer and N "activation layer + convolutional layer", the output features, i.e., the average feature mu_i of the i-th stage subblock, are obtained.
[0179] The convolutional layer may be a 1x1 convolutional layer, a 3x3 convolutional layer, or a 5x5 convolutional layer, and is not limited thereto; for example, a 1x1 convolutional layer may be selected. The convolutional layer in the predictive fusion network may be a normal 2D convolution, or a group convolution, for example, a double group convolution, and embodiments of group convolution will be described in subsequent examples. The activation layer may be a ReLU layer, or a leaky ReLU layer, or a sigmoid layer, or a Tanh layer, or a Gelu layer; for example, a ReLU layer may be selected.
[0180] Example 9: For both the encoding and decoding sides, it is necessary to determine the reconstructed image block corresponding to the current image block based on the reconstruction features of each stage subblock. For example, in Examples 5 and 6, the clustering block partitioning unit performs feature aggregation on the reconstruction features (y_hat_1 to y_hat_N) of N stage subblocks to obtain aggregated features (which may be called full-size aggregated features y_hat_all), and performs block partitioning on the aggregated features y_hat_all to obtain at least one block partitioning feature y_hat_p, for example, block partitioning features y_hat_p_1 to block partitioning features y_hat_p_K. For each block partitioning feature y_hat_p, the clustering block partitioning unit inputs the block partitioning feature y_hat_p into a composite transformation network, and the composite transformation network outputs a block partitioned reconstructed image block x_hat_p corresponding to the block partitioning feature y_hat_p. The merging unit obtains block-partitioned reconstructed image blocks x_hat_p corresponding to multiple block-partitioning features y_hat_p, and merges these block-partitioned reconstructed image blocks x_hat_p to obtain a reconstructed image block x_hat corresponding to the current image block.
[0181] First, the clustering block partitioning unit performs feature aggregation on the reconstructed features of N stage subblocks to obtain aggregated features.
[0182] As an example, as shown in Figure 9A, the reconstructed features (y_hat_1~y_hat_N) of N step subblocks can be aggregated according to the phase to obtain the aggregated feature y_hat_all. Figure 9A uses N=4 as an example. For example, one example of aggregation according to the phase is shown below, and after y_hat_group1 is obtained, y_hat_group1 may be used as the aggregated feature y_hat_all.
number
[0183] For example, as shown in Figure 9B, the reconstructed features (y_hat_1 to y_hat_N) of N step subblocks can be aggregated according to their phase to obtain multiple phase-aggregated features, and these multiple phase-aggregated features can be combined according to their channels to obtain the aggregated feature y_hat_all. For example, in the case of aggregation according to phase, refer to the above formula and aggregate according to phase to obtain y_hat_group1 and y_hat_group2. If there are many reconstructed features, after obtaining y_hat_group1 and y_hat_group2, the aggregated feature y_hat_all can be obtained by combining the multiple aggregated y_hat_groups according to their channels.
[0184] For example, before aggregating the reconstructed features of N stage subblocks, redundant features may be removed. For instance, if the encoding side divides a full-size feature into four parts because its height or width is odd, some of the reconstructed features of the stage subblocks may need to be padded to ensure that the size of the reconstructed features of each stage subblock is the same. As a padding method, one column can be added to the right of the full-size feature, or one row can be added to the bottom. Based on this, redundant features may need to be removed after aggregating the reconstructed features of the N stage subblocks. Figure 9C shows an example where the width is odd, and if the height is odd, the bottommost row is removed. Of course, the above is for the case of four divisions; in the case of sixteen divisions, the features to be removed are not limited to one row or one column.
[0185] Secondly, the clustering block partitioning unit partitions the aggregated feature y_hat_all into blocks to obtain K block partitioning features y_hat_p.
[0186] For example, the aggregated feature y_hat_all may be block-partitioned to reduce subsequent memory overhead (the synthesis transformation operation requires upsampling the feature multiple times so that the size of the feature data exceeds the acceptable size of memory). Of course, this feature block-partitioning process is not mandatory, and if the size of the current feature is small, or if it is determined that the current feature does not need to be block-partitioned based on the bitstream information, the following block-partitioning process can be skipped directly, and y_hat_p will be equal to y_hat_all, i.e., there will be only one block-partitioned feature.
[0187] To partition the aggregated feature y_hat_all into blocks, a feature block partitioning method without overlapping parts can be employed. Figure 9D shows an example in which the aggregated feature y_hat_all is partitioned into four block-partitioned features y_hat_p.
[0188] In a feature block partitioning method where there are no overlapping parts, the target size of the block partitioning features is determined, and based on the target size, the aggregated feature y_hat_all is evenly divided into multiple block partitioning features in the order from top to bottom and from left to right, and the size of each block partitioning feature is the target size, for example, it can be divided into block partitioning features y_hat_p_1, y_hat_p_2, y_hat_p_3, and y_hat_p_4.
[0189] For example, the target size is the size of each block partitioning feature, and the target size may be the size read from the bitstream, or it may be determined based on the size of the feature to be currently partitioned, without limitation. If the number of channels of each block partitioning feature is tile_c, the width is tile_w, and the height is tile_h, then the aggregated feature y_hat_all is evenly divided into multiple tile_c*tile_w*tile_h block partitioning features, for example, y_hat_p_1, y_hat_p_2, y_hat_p_3, and y_hat_p_4, in the order from top to bottom and from left to right.
[0190] To partition the aggregated feature y_hat_all into blocks, a feature block partitioning method with overlapping portions can be employed. Figure 9E shows an example of partitioning the aggregated feature y_hat_all into four block partitioned features y_hat_p.
[0191] In a feature block partitioning method where overlapping portions exist, the actual block partitioning size and overlap size of the block partitioning features may be determined. The actual block partitioning size may be the block partitioning size obtained by removing the overlapping portion from the block partitioning feature, and the overlap size may be the size of the overlapping portion of adjacent blocks. Based on the actual block partitioning size and overlap size, the aggregated features are divided into multiple block partitioning features such that the size of each block partitioning feature becomes the target size, for example, divided into block partitioning features y_hat_p_1, y_hat_p_2, y_hat_p_3, and y_hat_p_4.
[0192] For example, the actual block partition size and overlap size are determined, where the actual block partition size is the block partition size after removing the overlap portion, and the overlap size is the size of the overlap portion of adjacent blocks. This size information may be read from the bitstream or determined based on the size of the feature to be partitioned, and is not limited to this. As shown in Figure 9E, if the actual block partition size is tile_c*tile_w*tile_h and the overlap size is tile_c*padding_w*padding_h, then if there is overlap in the top, bottom, left, and right directions, the size of each block partition feature (i.e., the target size) is tile_c*(tile_w+2*padding_w)*(tile_h+2*padding_h). Based on this, for the aggregated feature y_hat_all, the features from the first to (tile_w+2*padding_w) columns and the first to (tile_h+2*padding_h) rows are divided into the first block partition feature. The feature of the tile_h+padding_h+1 to (2*tile_h+3*padding_h) row in the 2*tile_w+3*padding_w column from tile_w+padding_w+1 is divided into the second block partitioning feature, and so on.
[0193] As shown in Figure 9E, the aggregated 4×4×4 feature y_hat_all can be divided into four 4×3×3 block partitioning features, where tile_c is 4, padding_w and padding_h are 1, and tile_w and tile_h are 1.
[0194] For example, for the last row or column of blocks, if there are not enough blocks, the overlap area may be increased so that the final block division size meets the demand. Figure 9F shows an example where the size of y_hat_all is 11×13, padding_w and padding_h are 1, and tile_w and tile_h are 3. The size of the first and second blocks in the first row is 5×5, and for the third block, an overlap area of two columns was extended to the left to ensure that the size of that block is also 5×5.
[0195] Third, the clustering block partitioning unit inputs each block partitioning feature y_hat_p (for example, K block partitioning features y_hat_p) into the composite transformation network, and the composite transformation network outputs a block partitioning reconstructed image block x_hat_p corresponding to each block partitioning feature y_hat_p.
[0196] For each block partitioning feature y_hat_p, after inputting the block partitioning feature y_hat_p into the composite transformation network, a composite transformation is performed on the block partitioning feature y_hat_p via the composite transformation network to obtain a reconstructed block partitioned image block x_hat_p, and this is not limited to this block.
[0197] Fourth, the merging unit merges multiple block-divided reconstructed image blocks x_hat_p to obtain a reconstructed image block x_hat.
[0198] For example, the merging unit obtains K block-partitioned reconstructed image blocks x_hat_p corresponding to K block-partitioned features y_hat_p, and merges the K block-partitioned reconstructed image blocks x_hat_p to obtain a reconstructed image block x_hat.
[0199] For example, when the aggregated feature y_hat_all is partitioned into blocks using a feature block partitioning method that does not involve overlap, the merger unit merges K block partitioned reconstructed image blocks x_hat_p using a feature merger method that does not involve overlap. For instance, the block partitioned reconstructed image blocks x_hat_p corresponding to the K block partitioning features are sorted from top to bottom and from left to right, and the sorted block partitioned reconstructed image blocks x_hat_p corresponding to the K block partitioning features are joined in order to obtain the reconstructed image block x_hat corresponding to the current image block. The joining process may also be the reverse process of Figure 9D, which will not be explained here.
[0200] For example, when the aggregated feature y_hat_all is partitioned into blocks using a feature partitioning scheme that includes overlapping portions, the merger unit merges K partitioned reconstructed image blocks x_hat_p using a feature merger scheme that includes overlapping portions. For example, the merger unit may determine the actual partitioning size and overlap size of each partitioned reconstructed image block x_hat_p. The merger unit sorts the partitioned reconstructed image blocks x_hat_p corresponding to the K partitioned features y_hat_p from top to bottom and left to right, and then sequentially combines the sorted K partitioned reconstructed image blocks x_hat_p corresponding to the y_hat_p to obtain the reconstructed image block x_hat corresponding to the current image block. A reconstructed image block may contain K block-partitioned reconstructed image blocks, the size of the non-overlapping portion of a block-partitioned reconstructed image block may be the actual block-partition size of that block-partitioned reconstructed image block, and the size of the overlap portion between a block-partitioned reconstructed image block and an adjacent block-partitioned reconstructed image block may be the overlap size of that block-partitioned reconstructed image block. The joining process may be the reverse process of Figure 9E, which will not be explained here.
[0201] For example, the actual block division size and overlap size are determined, where the actual block division size is the block division size after removing the overlap portion, and the overlap size is the size of the overlap portion of adjacent blocks. This size information may be read from the bitstream or determined based on the size of the feature to be currently divided into blocks, and is not limited to this. If the actual block division size is x_c*xtile_w*xtile_h, the overlap width in the overlap size is xpadding_w, and the overlap height in the overlap size is xpadding_h, then the multiple block division reconstructed image blocks x_hat_p are joined in order according to the order in which they are received.
[0202] For the overlapping portion of two block-reconstructed image blocks (left and right), the value of the overlapping portion is the value of the left block-reconstructed image block, and the value of the right block-reconstructed image block is discarded. In other words, the value of the earliest (leftmost) block-reconstructed image block is adopted, and for the overlapping portion of subsequent subblocks, the value of the overlapping portion is directly discarded. Alternatively, the value of the overlapping portion is the average value of the two block-reconstructed image blocks.
[0203] For the overlapping portion of two block-reconstructed image blocks, the value of the overlapping portion is the value of the upper block-reconstructed image block, and the value of the lower block-reconstructed image block is discarded. In other words, the value of the earliest (upper) block-reconstructed image block is adopted, and for the overlapping portion of subsequent subblocks, the value of the overlapping portion is directly discarded. Alternatively, the value of the overlapping portion is the average value of the two block-reconstructed image blocks.
[0204] Example 10: On both the encoding and decoding sides, it is necessary to determine the reconstructed image block corresponding to the current image block based on the reconstruction features of each stage subblock. For example, the clustering block partitioning unit performs feature aggregation on the reconstruction features (y_hat_1 to y_hat_N) of N stage subblocks to obtain aggregated features (which may be called full-size aggregated features y_hat_all). The clustering block partitioning unit inputs the aggregated features y_hat_all into a composite transformation network, and the composite transformation network outputs a reconstructed image block corresponding to the aggregated features y_hat_all. This reconstructed image block is the reconstructed image block x_hat corresponding to the current image block. Compared to Example 9, block partitioning is not performed on the aggregated features y_hat_all, and no merging is performed on multiple block-partitioned reconstructed image blocks x_hat_p.
[0205] First, the clustering block partitioning unit performs feature aggregation on the reconstructed features of N stage subblocks to obtain aggregated features.
[0206] For example, the reconstructed features (y_hat_1~y_hat_N) of N step subblocks can be aggregated according to their phase to obtain the aggregated feature y_hat_all. Alternatively, the reconstructed features (y_hat_1~y_hat_N) of N step subblocks can be aggregated according to their phase to obtain multiple phase-aggregated features, and these multiple phase-aggregated features can be combined according to their channels to obtain the aggregated feature y_hat_all. Feature aggregation is explained in Example 9, and will not be explained here.
[0207] Secondly, the clustering block partitioning unit inputs the aggregated feature y_hat_all into the composite transformation network, performs a composite transformation on the aggregated feature y_hat_all via the composite transformation network, and without being limited to this composite transformation method, outputs a reconstructed image block x_hat corresponding to the aggregated feature y_hat_all, i.e., a reconstructed image block x_hat corresponding to the current image block, via the composite transformation network.
[0208] Example 11: For the encoding side and the decoding side, it is necessary to determine the reconstructed image block corresponding to the current image block based on the reconstruction characteristics of each stage sub-block. For example, after the clustering block splitting unit obtains the reconstruction characteristics (y_hat_1~y_hat_N) of N stage sub-blocks, it does not perform feature aggregation on the reconstruction characteristics, does not perform block splitting on the aggregated features, directly inputs the reconstruction characteristics of each stage sub-block into the synthesis conversion network, and the synthesis conversion network outputs the block splitting reconstructed image block x_hat_p corresponding to each stage sub-block. The merging unit obtains the block splitting reconstructed image blocks x_hat_p corresponding to a plurality of stage sub-blocks, and can merge these block splitting reconstructed image blocks x_hat_p to obtain the reconstructed image block x_hat corresponding to the current image block.
[0209] First, after the clustering block splitting unit obtains the reconstruction characteristics of N stage sub-blocks, it inputs the reconstruction characteristics of each stage sub-block into the synthesis conversion network, performs synthesis conversion on the reconstruction characteristics through the synthesis conversion network, and does not limit this synthesis conversion method. The synthesis conversion network outputs the block splitting reconstructed image block x_hat_p corresponding to each stage sub-block.
[0210] Second, the merging unit merges a plurality of block splitting reconstructed image blocks x_hat_p to obtain the reconstructed image block x_hat.
[0211] Exemplarily, the merging unit can obtain N block splitting reconstructed image blocks x_hat_p corresponding to N stage sub-blocks, and merge the N block splitting reconstructed image blocks x_hat_p to obtain the reconstructed image block x_hat.
[0212] For example, the merging unit merges N block-partitioned reconstructed image blocks x_hat_p using a feature merging method that does not involve overlap. For instance, the N block-partitioned reconstructed image blocks x_hat_p are sorted from top to bottom and left to right, and the sorted N block-partitioned reconstructed image blocks x_hat_p are sequentially joined to obtain the reconstructed image block x_hat corresponding to the current image block. Alternatively, the merging unit merges N block-partitioned reconstructed image blocks x_hat_p using a feature merging method that involves overlap. For instance, the actual block division size and overlap size of each block-partitioned reconstructed image block x_hat_p are determined, the N block-partitioned reconstructed image blocks x_hat_p are sorted from top to bottom and left to right, and the sorted N block-partitioned reconstructed image blocks x_hat_p are sequentially joined based on the actual block division size and overlap size to obtain the reconstructed image block x_hat corresponding to the current image block. A reconstructed image block contains N block-partitioned reconstructed image blocks. The size of the non-overlapping portion of a block-partitioned reconstructed image block is the actual block-partition size of that block-partitioned reconstructed image block, and the size of the overlapping portion between a block-partitioned reconstructed image block and an adjacent block-partitioned reconstructed image block is the overlap size of that block-partitioned reconstructed image block. For the overlapping portion of two block-partitioned reconstructed image blocks (left and right), the value of the overlapping portion is the value of the left block-partitioned reconstructed image block, and the value of the right block-partitioned reconstructed image block is discarded. Alternatively, the value of the overlapping portion is the average value of the two block-partitioned reconstructed image blocks. For the overlapping portion of two block-partitioned reconstructed image blocks (top and bottom), the value of the overlapping portion is the value of the top block-partitioned reconstructed image block, and the value of the bottom block-partitioned reconstructed image block is discarded. Alternatively, the value of the overlapping portion is the average value of the two block-partitioned reconstructed image blocks.
[0213] Example 12: In Examples 1 to 11, with respect to the composite transformation network, in order to perform accurate bitrate control, the encoding side may transmit bitrate control parameters for the composite transformation network to the decoding side. For example, the encoding side may encode the bitrate control parameters in an auxiliary bitstream corresponding to the current image block, and the decoding side may decode the auxiliary bitstream corresponding to the current image block to obtain the bitrate control parameters and determine the composite transformation network based on the bitrate control parameters. That is, the composite transformation network is determined based on bitstream information, particularly bitrate-related information. The synthetic transformation network may consist of at least one 2x deconvolutional layer (e.g., a 2x2, 3x3, 4x4, or 5x5 2x deconvolutional layer, e.g., a 4x4 2x deconvolutional layer may be selected, and of course, other 2x deconvolutional layers may also be selected), at least one activation layer (e.g., a ReLU layer, or a leaky ReLU layer, or a sigmoid layer, or a Tanh layer, or a Gelu layer, e.g., a ReLU layer may be selected), and at least one convolutional layer (e.g., a 1x1, 3x3, or 5x5 convolutional layer, e.g., a 3x3 convolutional layer may be selected, and the convolutional layer in the synthetic transformation network may be a normal 2D convolution, or a group convolution, e.g., a 2x group convolution), the output feature spatial resolution of the synthetic transformation network is a positive integer multiple of the input feature spatial resolution, e.g., 16 times (length and width are both 16 times), and the number of output feature channels is 1 or 3. Based on the composite transformation network, the composite transformation network can be controlled based on bitrate control parameters; that is, the bitrate control parameters are used as inputs to the composite transformation network.
[0214] In one possible embodiment, in Example 9, the clustering block partitioning unit inputs K block partitioning features y_hat_p to the composite transformation network. The encoding side encodes K bitrate control parameters corresponding to each of the K block partitioning features y_hat_p into the auxiliary bitstream corresponding to the current image block. Without limiting the source of the K bitrate control parameters, the encoding side uses these K bitrate control parameters to control the analytic transformation network and the composite transformation network (i.e., the K bitrate control parameters are input to the analytic transformation network and the composite transformation network, for example, the block partitioning feature y_hat_p and the bitrate control parameter corresponding to the block partitioning feature y_hat_p are both input features to the composite transformation network). This allows the bitrate corresponding to the bitstream of the current image block to be the target bitrate or close to the target bitrate (i.e., the difference between the two is less than a threshold). The decoding side decodes the auxiliary bitstream corresponding to the current image block to obtain the K bitrate control parameters corresponding to the K block partitioning features y_hat_p.
[0215] Based on this, the clustering block partitioning unit may obtain K block partitioning features y_hat_p and K bitrate control parameters corresponding to the K block partitioning features y_hat_p, input both the block partitioning feature y_hat_p and the corresponding bitrate control parameter to the composite transformation network for each block partitioning feature y_hat_p, and generate a reconstructed block partitioning image block x_hat_p corresponding to the block partitioning feature y_hat_p based on the block partitioning feature y_hat_p and the corresponding bitrate control parameter by the composite transformation network.
[0216] Exemplary, generating a reconstructed block x_hat_p corresponding to a block partitioning feature y_hat_p based on a composite transformation network and a bitrate control parameter corresponding to the block partitioning feature y_hat_p may include, but is not limited to, processing the block partitioning feature y_hat_p via the composite transformation network to obtain a first feature, processing the bitrate control parameter corresponding to the block partitioning feature y_hat_p via the composite transformation network to obtain a second feature, generating a third feature based on the first and second features, and determining the reconstructed block x_hat_p corresponding to the block partitioning feature y_hat_p based on the third feature.
[0217] In one possible embodiment, for Example 10, the clustering block division unit inputs the aggregated feature y_hat_all to the composite transform network. The encoding side encodes a bitrate control parameter corresponding to the aggregated feature y_hat_all, i.e., the bitrate control parameter corresponding to the current image block, into the auxiliary bitstream corresponding to the current image block. Without limiting the source of the bitrate control parameter, the encoding side uses the bitrate control parameter to control the analytic transform network and the composite transform network (i.e., the bitrate control parameter is used as input to the analytic transform network and the composite transform network, for example, both the aggregated feature y_hat_all and the bitrate control parameter are input features to the analytic transform network). This makes it possible to make the bitrate corresponding to the bitstream of the current image block equal to or close to the target bitrate. The decoding side decodes the auxiliary bitstream corresponding to the current image block to obtain the bitrate control parameter corresponding to the current image block.
[0218] Based on this, the clustering block division unit obtains the aggregated feature y_hat_all and bitrate control parameters, inputs both the aggregated feature y_hat_all and bitrate control parameters into the composite transformation network, and the composite transformation network can generate a reconstructed image block x_hat corresponding to the aggregated feature y_hat_all based on the aggregated feature y_hat_all and bitrate control parameters.
[0219] Generating a reconstructed image block x_hat corresponding to the aggregated feature y_hat_all based on the aggregated feature y_hat_all and bitrate control parameters using a composite transformation network may include processing the aggregated feature y_hat_all via the composite transformation network to obtain a first feature, processing the bitrate control parameters via the composite transformation network to obtain a second feature, generating a third feature based on the first and second features, and determining the reconstructed image block x_hat corresponding to the aggregated feature y_hat_all based on the third feature.
[0220] In one possible embodiment, in Example 11, the clustering block partitioning unit inputs the reconstruction features of N stage subblocks to the composite transform network. The encoding side encodes N bitrate control parameters corresponding to each of the N stage subblocks into the auxiliary bitstream corresponding to the current image block, and without limiting the source of the N bitrate control parameters, the encoding side uses these N bitrate control parameters to control the analytic transform network and the composite transform network (i.e., the N bitrate control parameters are input to the analytic transform network and the composite transform network, for example, inputting both the reconstruction features of the stage subblocks and the bitrate control parameters corresponding to the stage subblocks into the composite transform network), thereby making the bitrate corresponding to the bitstream of the current image block the target bitrate or close to the target bitrate. The decoding side decodes the auxiliary bitstream corresponding to the current image block to obtain N bitrate control parameters corresponding to the N stage subblocks.
[0221] Based on this, the clustering block division unit may obtain reconstruction features for N stage subblocks and N bitrate control parameters corresponding to the N stage subblocks, input both the reconstruction features and the corresponding bitrate control parameters for each stage subblock into the composite transformation network, and generate a block division reconstructed image block x_hat_p corresponding to the stage subblock based on the reconstruction features and the corresponding bitrate control parameters of the stage subblock using the composite transformation network.
[0222] Exemplary, generating a block-partitioned reconstructed image block x_hat_p corresponding to a stage subblock based on the reconstruction features of the stage subblock and the bitrate control parameters corresponding to the stage subblock by a composite transformation network may include, but is not limited to, processing the reconstruction features of the stage subblock via the composite transformation network to obtain a first feature, processing the bitrate control parameters corresponding to the stage subblock via the composite transformation network to obtain a second feature, generating a third feature based on the first and second features, and determining the block-partitioned reconstructed image block x_hat_p corresponding to the stage subblock based on the third feature.
[0223] Example 13: Based on Example 12, the composite transformation network may also be called a composite transformation network controlled by a bitrate control parameter, the bitrate control parameter can be denoted as a λ parameter, the λ parameter is obtained by decoding from an auxiliary bitstream corresponding to the current image block, Figure 10A is a schematic diagram of the composite transformation network controlled by the λ parameter, the composite transformation network is a bitrate variable composite transformation network, and output features with 1 or 3 channels can be obtained based on the composite transformation network. Of course, Figure 10A is merely one example of a composite transformation network and does not limit the structure of this composite transformation network.
[0224] As shown in Figure 10A, the composite transform network controlled by the λ parameter may include at least one Conv (convolutional layer) and at least one λ-RSTB (Residual Swin Transformer Block). In Figure 10A, four convolutional layers and four λ-RSTBs are used as an example, but the number of convolutional layers and λ-RSTBs may be more or less than this. The convolutional layers may be, for example, 1×1, 3×3, or 5×5 convolutional layers, or a 3×3 convolutional layer may be selected. The structure of the convolutional layer in the composite transform network may be a normal 2D convolution, or a group convolution, for example, a double group convolution, and is not limited to this structure. λ-RSTB represents an RSTB network layer controlled by the λ parameter, and the structure of the RSTB network layer controlled by the λ parameter is shown in Figure 10B.
[0225] As shown in Figure 10B, λ-RSTB may include, but is not limited to, network layers such as FE (Feature Embedding), λ-STB (Swin Transformer Block), and FU (Feature Un-embedding). Of course, this is merely an example and does not limit the structure of this λ-RSTB. FE is a feature embedding layer for implementing feature embedding, and FU is a feature de-embedding layer for implementing feature de-embedding.
[0226] λ-STB represents an STB network layer controlled by a λ parameter. The structure of the STB network layer controlled by the λ parameter is shown in Figure 10B. λ-STB may include, but is not limited to, LN (linear), λ-WA (weight averaging), λ-SWA (stochastic weight averaging), MLP (multilayer perceptron), etc. Of course, this is merely an example and does not limit the structure of this λ-STB. LN is a linear layer, and λ-STB may include at least one LN, without limiting the number of LNs; four LNs are used as an example here. MLP is a multilayer perceptron, and λ-STB may include at least one MLP, without limiting the number of MLPs; two MLPs are used as an example here.
[0227] The structure of λ-SWA is the same as that of λ-WA. Here, λ-WA is used as an example, representing a WA network layer controlled by the λ parameter. The structure of the WA network layer controlled by the λ parameter is shown in Figure 10C. λ-WA may include, but is not limited to, FC (Full Connection) layers, MatMul (matrix multiplication), Scale (scaling), SoftMax (logistic regression, also called normalization), MatMul, etc. Of course, this is merely an example and does not limit the structure of this λ-WA.
[0228] The FC layer may map the parameter λ to a single scaling coefficient; that is, the FC layer of the composite transformation network processes the bitrate control parameter to obtain a second feature which may be a scaling coefficient.
[0229] The input features (e.g., block split features y_hat_p, or aggregated features y_hat_all, or reconstruction features of stage sub-blocks) may be processed through a transformer network to obtain Q features, K features, and V features. The Q features, K features, and V features all originate from the input features themselves and are feature vectors generated by multiplying the corresponding weights based on the input features. This process is not limited, that is, the synthesis transformation network may include a transformer network, and the input features (e.g., block split features, or aggregated features, or reconstruction features) are processed through the transformer network of the synthesis transformation network to obtain the first feature.
[0230] After obtaining the Q features, K features, and V features, matrix multiplication operations, scaling operations, and logistic regression operations may be sequentially performed based on the Q features and K features to obtain intermediate feature 1. An intermediate feature 2 may be obtained by multiplying the scaling coefficient (obtained by mapping the parameter λ by the FC layer) and the V features. Matrix multiplication is performed on intermediate feature 1 and intermediate feature 2 to obtain the output feature of λ-WA. In summary, the scaling coefficient is the second feature, and the Q features, K features, and V features are the first features. Therefore, the third feature can be generated based on the first feature and the second feature, and the third feature is the output feature of λ-WA.
[0231] After obtaining the output feature of λ-WA, referring to FIGS. 10A and 10B, based on the output feature of λ-WA (i.e., the third feature), the output feature of the synthesis transformation network can be determined, for example, the block split reconstruction image block x_hat_p corresponding to the block split feature y_hat_p, the reconstruction image block x_hat corresponding to the aggregated feature y_hat_all, and the block split reconstruction image block x_hat_p corresponding to the stage sub-block.
[0232] Referring to λ-WA shown in FIG. 10C, the calculation method of λ-WA is shown below. By such a method, by using the λ parameter to control the outputs of the image encoder and decoder, bitstreams of different sizes and reconstructed images of different qualities can be obtained.
number
[0233] Example 14: In Examples 1 to 13, the convolutional layer may also be a regular 2D convolution, or a group convolution, for example, a double group convolution. For example, the steps for a regular 2D convolution are shown in Figure 11A, where the input feature is (H × W × C), and then C' filters (each filter having a size of (h × w × C), where C in the filter size is the same as C in the input feature) are applied, and the input layer is transformed into an output feature of size (H' × W' × C').
[0234] The steps for group convolution are shown in Figure 11B, where each group of filters may contain C' / g (for example, g is 2) filters, and the number of channels in each filter is half that of a filter in a normal 2D convolution. Each group of filters acts on 1 / g of the number of channels corresponding to the original W×H×C, i.e., W×H×C / g. Thus, each group of filters outputs features corresponding to C' / g channels. Finally, the channels can be stacked to obtain a final C' channels, achieving the same effect as the normal 2D convolution described above. Clearly, group convolution can effectively reduce the number of parameters and computational complexity.
[0235] Example 15: In Examples 1 to 14, in order to perform accurate bitrate control, the encoding side needs to send bitrate control parameters to the decoding side. For example, the bitrate control parameters are sent appended to the auxiliary bitstream corresponding to the current image block, and on the encoding side, the bitrate control parameters affect the analysis transformation network and the synthesis transformation network. To ensure that the bitrate corresponding to the bitstream of the current image block is the target bitrate or close to the target bitrate (a pre-set bitrate), the encoding side may try different bitrate control parameters and find the target bitrate control parameter. In the target bitrate control parameter, the bitrate corresponding to the bitstream of the current image block is the target bitrate or close to the target bitrate, and the encoding side encodes the target bitrate control parameter into the auxiliary bitstream corresponding to the current image block.
[0236] If the current image block corresponds to only one bitrate control parameter, for example, if the aggregated feature y_hat_all corresponds to one bitrate control parameter, the target bitrate control parameter may be obtained by sequentially trying each bitrate control parameter from multiple bitrate control parameters until the bitrate corresponding to the bitstream of the current image block is the target bitrate or close to the target bitrate.
[0237] If the current image block corresponds to multiple bitrate control parameters, for example, if K block partitioning features correspond to K bitrate control parameters and N stage subblocks correspond to N bitrate control parameters, then first, initial values are set for the multiple bitrate control parameters. If the bitrate corresponding to the bitstream of the current image block does not meet the requirements, one candidate bitrate control parameter is selected from the multiple bitrate control parameters, and the value of the candidate bitrate control parameter is reduced. If the bitrate corresponding to the bitstream of the current image block still does not meet the requirements, another candidate bitrate control parameter (which may be the same as or different from the previous candidate bitrate control parameter) is selected from the multiple bitrate control parameters, and the value of the candidate bitrate control parameter is reduced. This process is repeated until the bitrate corresponding to the bitstream of the current image block meets the requirements (i.e., the bitrate corresponding to the bitstream of the current image block is the target bitrate or close to the target bitrate), thereby obtaining the target values for the multiple bitrate control parameters.
[0238] When selecting one candidate bitrate control parameter from multiple bitrate control parameters, one bitrate control parameter may be randomly selected as the candidate bitrate control parameter, the maximum bitrate control parameter may be selected as the candidate bitrate control parameter, or the candidate bitrate control parameter may be selected based on the bitrate-distortion relationship. For example, if bitrate control parameter A is reduced from the current level to the next level, the distortion is A1, and if bitrate control parameter B is reduced from the current level to the next level, the distortion is B1, then if distortion A1 is less than distortion B1, i.e., distortion A1 is the minimum distortion, then bitrate control parameter A is used as the candidate bitrate control parameter.
[0239] In one possible embodiment, the encoding side has λ for each (feature or image) block. iOne embodiment of how to set the parameters is as follows: Each block x i For example, for a block partitioning feature or a step subblock, different λ i Encoding is performed using [this method]. Before encoding, several discrete λ values are taken for each block, and these discrete λ values are defined as quality levels. The bitrate and distortion of each block and each quality level are then statistically calculated before encoding. For each frame of the image, this process only needs to be performed once before the first encoding.
[0240] In the encoding process, a quality level λ is selected for each block and encoded. In the encoding process, first, the same quality level is adopted for all blocks and encoding is performed. This quality level is the smallest quality level that can make the total bitrate greater than the target bitrate, that is, it satisfies the following equation, R t This is the target bitrate,
number
number
[0241] Subsequently, the quality level of each block is gradually reduced until the total bitrate is less than the target bitrate. For each block, the gradient between the current quality level and the previous quality level in the rate distortion curve may be calculated, i.e., the gradient is calculated using the following formula:
number
[0242] D Bi (q) is the reconstruction loss at quality level q of block i, and when using a loss function that decreases with the bitrate (e.g., mean squared error) as an indicator, the minimum C BiSelect block i (i.e., the candidate bitrate control parameter with the least distortion) and reduce it by one quality level. Continue this process until the total bitrate is less than the target bitrate.
[0243] Figure 12 is a schematic diagram of the λ parameter determination, in which seven quality levels (i.e., seven λ parameter values) are designed in the rate-distortion curve, and at quality level 3, C B1 (3) <C B2 (3) is calculated, meaning that reducing the quality level of block 2 at quality level 3 would result in a greater loss of coding performance, therefore the quality level of block 1 is reduced.
[0244] As can be seen from the technical proposals of each embodiment described above, this embodiment proposes a variable and adjustable bitrate encoding and decoding method for neural network-based encoding and decoding technology, improving parallelism, effectively saving cache for storing features, achieving higher bitrate control accuracy, and resulting in less loss of encoding performance and better encoding performance and bitrate control accuracy. By adopting block partitioning encoding and decoding, peak memory occupation is reduced, the decoding time for a single block is shortened, high-speed parallel decoding capability is achieved, the complexity of the neural network is kept low, the quality of reconstructed image blocks is effectively guaranteed, encoding and decoding performance is improved and complexity is reduced. For example, the bitrate control algorithm of this technical proposal can achieve a bitrate control accuracy of about 99%, and the loss of encoding performance compared to no bitrate control is only about 1%. Taking the decoding of a single 2K resolution image with a block size of 256*256 as an example, by employing block partitioning coding and decoding, the peak memory usage is reduced to 1 / 10 of that of decoding the entire image, while simultaneously achieving high-speed parallel decoding capabilities with a single block decoding time of less than 0.1 seconds. Better coding and decoding performance is achieved with a smaller increase in the number of parameters.
[0245] Example 16: An embodiment of the present invention provides a decoding method which may be applied to a decoding side (also called a video decoder) and which may include the steps of: decoding a first bitstream of the current image block to obtain coefficient hyperparameter features of the current image block; determining probability distribution parameters based on the coefficient hyperparameter features; decoding a second bitstream of the current image block based on the probability distribution parameters to obtain residual features of the current image block; determining reconstruction features of the current image block based on the residual features; decoding an auxiliary bitstream corresponding to the current image block to obtain bitrate control parameters corresponding to the current image block; inputting the reconstruction features and bitrate control parameters into a composite transformation network to obtain a reconstructed image block corresponding to the current image block.
[0246] Exemplary, the step of inputting reconstruction features and bitrate control parameters into a composite transformation network to obtain a reconstructed image block corresponding to the current image block may include, but is not limited to, the steps of: processing the reconstruction features via the composite transformation network to obtain a first feature; processing the bitrate control parameters via the composite transformation network to obtain a second feature; generating a third feature based on the first and second features; and determining a reconstructed image block corresponding to the current image block based on the third feature.
[0247] Exemplary, the synthesis-transformation network can be adjusted based on Examples 3 and 4, and in Example 16, only the process of adjusting the synthesis-transformation network will be described, while other processes can refer to Examples 3 and 4.
[0248] For example, bitrate control can be performed precisely using bitrate control parameters, which affect the analytic transformation network and the synthetic transformation network. For instance, when the encoding side performs an analytic transformation on the current image block via the analytic transformation network, the bitrate control parameters and the current image block are input to the analytic transformation network, and the analytic transformation network outputs the image features corresponding to the current image block. When the encoding or decoding side inputs the image feature y_hat to the synthetic transformation network, the image feature y_hat and the bitrate control parameters are input to the synthetic transformation network, and the synthetic transformation network outputs the reconstructed image block corresponding to the current image block.
[0249] Based on the set target bitrate, the encoding side attempts to adopt different bitrate control parameters. After adopting a bitrate control parameter, if the bitrate corresponding to the bitstream for the current image block is the target bitrate or close to it, this bitrate control parameter is the one ultimately adopted, and the encoding side can transmit this bitrate control parameter to the decoding side.
[0250] For example, the encoding side may send bitrate control parameters for the composite transformation network to the decoding side. For instance, the encoding side may encode the bitrate control parameters corresponding to the current image block into an auxiliary bitstream corresponding to the current image block, and the decoding side may decode the auxiliary bitstream corresponding to the current image block to obtain the bitrate control parameters corresponding to the current image block and determine the composite transformation network based on the bitrate control parameters. The synthetic transformation network may consist of at least one 2x deconvolutional layer (e.g., a 2x2, 3x3, 4x4, or 5x5 2x deconvolutional layer, e.g., a 4x4 2x deconvolutional layer may be selected, and of course, other 2x deconvolutional layers may also be selected), at least one activation layer (e.g., a ReLU layer, or a leaky ReLU layer, or a sigmoid layer, or a Tanh layer, or a Gelu layer, e.g., a ReLU layer may be selected), and at least one convolutional layer (e.g., a 1x1, 3x3, or 5x5 convolutional layer, e.g., a 3x3 convolutional layer may be selected, and the convolutional layer in the synthetic transformation network may be a normal 2D convolution, or a group convolution, e.g., a 2x group convolution), the output feature spatial resolution of the synthetic transformation network is a positive integer multiple of the input feature spatial resolution, e.g., 16 times (length and width are both 16 times), and the number of output feature channels is 1 or 3. Based on the composite transformation network, the composite transformation network can be controlled based on bitrate control parameters; that is, the bitrate control parameters are used as inputs to the composite transformation network.
[0251] For example, the encoding side encodes a bitrate control parameter corresponding to the current image block into an auxiliary bitstream corresponding to the current image block, and without limiting the source of the bitrate control parameter, the encoding side uses the bitrate control parameter to control the analytic transformation network and the synthetic transformation network (i.e., the bitrate control parameter is the input to the analytic transformation network and the synthetic transformation network), thereby making the bitrate corresponding to the bitstream of the current image block the target bitrate or close to the target bitrate. The decoding side decodes the auxiliary bitstream corresponding to the current image block to obtain the bitrate control parameter corresponding to the current image block. Based on this, the decoding side may input both the image feature y_hat and the bitrate control parameter into the synthetic transformation network, which generates a reconstructed image block corresponding to the current image block based on the image feature y_hat and the bitrate control parameter, and outputs the reconstructed image block.
[0252] Here, the synthesis transformation network generating a reconstructed image block corresponding to the current image block based on the image feature y_hat and bitrate control parameters may include processing the image feature y_hat (i.e., the reconstructed feature y_hat) via the synthesis transformation network to obtain a first feature, processing the bitrate control parameters via the synthesis transformation network to obtain a second feature, generating a third feature based on the first and second features, and determining the reconstructed image block corresponding to the current image block based on the third feature.
[0253] For example, as shown in Figures 10A, 10B, and 10C, the composite transformation network may include at least one Conv (convolutional layer) and at least one λ-RSTB, where the λ-RSTB may include, but is not limited to, network layers such as FE, λ-STB, and FU; the λ-STB may include, but is not limited to, LN, λ-WA, λ-SWA, and MLP; the structure of the λ-SWA is the same as the structure of the λ-WA; and the λ-WA may include, but is not limited to, an FC layer, MatMul, Scale, SoftMax, and MatMul; based on this, the FC layer may map the parameter λ (i.e., the bitrate control parameter) to a single scaling coefficient, that is, the bitrate control parameter is processed by the FC layer of the composite transformation network to obtain the second feature. The image feature y_hat may be processed via the transformer network to obtain the Q feature, K feature, and V feature, that is, the image feature y_hat is processed via the transformer network of the composite transformation network to obtain the first feature. After obtaining the Q, K, and V features, matrix multiplication, scaling, and logistic regression operations may be performed sequentially based on the Q and K features to obtain intermediate feature 1. Intermediate feature 2 may be obtained by multiplying the scaling coefficient by the V feature. Intermediate feature 1 and intermediate feature 2 are matrix multiplied to obtain the output feature of λ-WA, that is, a third feature is generated based on the first and second features, and the third feature is the output feature of λ-WA. After obtaining the output feature of λ-WA, referring to Figures 10A and 10B, the output feature of the composite transformation network can be determined based on the output feature of λ-WA (i.e., the third feature), for example, the output feature may be the reconstructed image block corresponding to the current image block.
[0254] Exemplary examples, each of the above embodiments may be implemented individually or in combination. For example, each of the embodiments from Embodiments 1 to 16 may be implemented individually, or at least two of the embodiments from Embodiments 1 to 16 may be implemented in combination.
[0255] For example, in each of the above embodiments, the contents of the encoding side may be applied to the decoding side, that is, processed by the decoding side in the same manner, and the contents of the decoding side may be applied to the encoding side, that is, processed by the encoding side in the same manner.
[0256] Exemplary examples, in each of the above embodiments, the contents of different embodiments are mutually referential; for example, the contents of Embodiment 3 can be applied to Embodiment 16, the contents of Embodiment 4 can be applied to Embodiment 16, the contents of Embodiments 5 to 15 can be applied to Embodiment 16, and the contents of Embodiments 5 to 15 are mutually referential, and are not limited thereto.
[0257] Based on the same concept as described above, embodiments of the present invention further provide a decoding device which is applied to the decoding side and includes a memory configured to store video data and a decoder configured to carry out the decoding method in embodiments 1 to 16, i.e., the decoding side processing process.
[0258] For example, in one possible embodiment, the decoder is The steps include decoding the first bitstream of the current image block to obtain coefficient hyperparameter features of each stage subblock of the current image block, For each stage subblock, the steps include determining the probability distribution parameters based on the coefficient hyperparameter features of the stage subblock, decoding the second bitstream of the current image block based on the probability distribution parameters to obtain the residual features of the stage subblock, A step of determining the reconstruction features of the stage subblock based on the residual features and the mean features of the stage subblock, The system is configured to perform the steps of determining the reconstructed image block corresponding to the current image block based on the reconstruction characteristics of each stage subblock.
[0259] Based on the same concept as described above, embodiments of the present invention further provide an encoding device which is applied to the encoding side and includes a memory configured to store video data and an encoder configured to perform the encoding method of embodiments 1 to 16, i.e., the encoding side processing process.
[0260] For example, in one possible embodiment, the encoder is The steps include inputting the current image block into an analysis and transformation network to obtain a feature block corresponding to the current image block, The steps include dividing the feature block into features to be encoded in multiple sub-blocks, For each stage subblock corresponding to the current image block, the steps include obtaining the coefficient hyperparameter features of the stage subblock and encoding the coefficient hyperparameter features of the stage subblock into the first bitstream of the current image block. A step of determining the residual features of the stage subblock based on the features to be encoded of the stage subblock and the mean value features of the stage subblock, The system is configured to perform the steps of: determining probability distribution parameters based on the coefficient hyperparameter features of the stage subblock; and encoding the residual features of the stage subblock into the second bitstream of the current image block based on the probability distribution parameters.
[0261] Based on the same concept as described above, a schematic hardware architecture diagram of a decoding device (also called a video decoder) provided by an embodiment of the present invention may be shown in detail in Figure 13A. The decoding device includes a processor 1301 and a machine-readable storage medium 1302, the machine-readable storage medium 1302 storing machine-executable instructions that can be executed by the processor 1301, and the processor 1301 is used to execute the machine-executable instructions and carry out the decoding methods of embodiments 1 to 16 of the present invention. For example, in one possible embodiment, when the processor 1301 executes a machine-executable instruction, The steps include decoding the first bitstream of the current image block to obtain coefficient hyperparameter features of each stage subblock of the current image block, For each stage subblock, the steps include determining the probability distribution parameters based on the coefficient hyperparameter features of the stage subblock, decoding the second bitstream of the current image block based on the probability distribution parameters to obtain the residual features of the stage subblock, A step of determining the reconstruction features of the stage subblock based on the residual features and the mean features of the stage subblock, The following steps are performed: determining the reconstructed image block corresponding to the current image block based on the reconstruction characteristics of each stage subblock.
[0262] Based on the same concept as described above, a schematic hardware architecture diagram of an encoding device (also called a video encoder) provided by an embodiment of the present invention may be shown in detail in Figure 13B from a hardware perspective. The encoding device includes a processor 1311 and a machine-readable storage medium 1312, the machine-readable storage medium 1312 storing machine-executable instructions that can be executed by the processor 1311, and the processor 1311 is used to execute the machine-executable instructions and implement the encoding methods of embodiments 1 to 16 of the present invention. For example, in one possible embodiment, when the processor 1311 executes a machine-executable instruction, The steps include inputting the current image block into an analysis and transformation network to obtain a feature block corresponding to the current image block, The steps include dividing the feature block into features to be encoded in multiple sub-blocks, For each stage subblock corresponding to the current image block, the steps include obtaining the coefficient hyperparameter features of the stage subblock and encoding the coefficient hyperparameter features of the stage subblock into the first bitstream of the current image block. A step of determining the residual features of the stage subblock based on the features to be encoded of the stage subblock and the mean value features of the stage subblock, The process involves determining probability distribution parameters based on the coefficient hyperparameter features of the stage subblock, and encoding the residual features of the stage subblock into the second bitstream of the current image block based on the probability distribution parameters.
[0263] Based on the same concept as described above, embodiments of the present invention provide an electronic device. The electronic device includes a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor, and the processor is used to execute the machine-executable instructions and carry out the decoding or encoding methods of embodiments 1 to 16 of the present invention.
[0264] Based on the same concept as described above, embodiments of the present invention further provide a machine-readable storage medium in which several computer instructions are stored, and when the computer instructions are executed by a processor, the methods disclosed in the above examples of the present invention, such as the decoding method or encoding method in each of the above embodiments, can be performed.
[0265] Based on the same principles as described above, embodiments of the present invention further provide a computer application, and when the computer application is executed by a processor, the decoding or encoding method disclosed in the above examples of the present invention can be performed.
[0266] Based on the same concept as described above, embodiments of the present invention further provide a decoding device applicable to the decoding side, the decoding device including a decoding module for decoding a first bitstream of the current image block to obtain coefficient hyperparameter features of each stage subblock of the current image block, determining probability distribution parameters for each stage subblock based on the coefficient hyperparameter features of the stage subblock, decoding a second bitstream of the current image block based on the probability distribution parameters to obtain residual features of the stage subblock, and a determination module for determining reconstruction features of the stage subblock based on the residual features and mean features of the stage subblock, and determining a reconstructed image block corresponding to the current image block based on the reconstructed features of each stage subblock.
[0267] Exemplary, the decision module is used to obtain, for the first stage subblock, the mean feature of the stage subblock via an mean prediction network based on the coefficient hyperparameter feature of the stage subblock, or to obtain a set default reference feature and obtain the mean feature of the stage subblock via the mean prediction network based on the coefficient hyperparameter feature and the default reference feature, or, for the i-th stage subblock, to obtain a reference feature of the i-th stage subblock based on the reconstruction features of the i-1 stage subblocks preceding the i-th stage subblock, and to obtain the mean feature of the stage subblock via an mean prediction network based on the coefficient hyperparameter feature and the reference feature, where i is greater than 1, and the reference feature includes all the reconstruction features of the i-1 preceding stage subblocks, or some of the reconstruction features of the i-1 preceding stage subblocks, or the reconstruction feature of the i-1th stage subblock.
[0268] Exemplary, the mean prediction network includes a first prediction network, a second prediction network, and a prediction fusion network, and the decision module is used to obtain the mean feature of the stage subblock via the mean prediction network based on the coefficient hyperparameter features and the reference features of the stage subblock, specifically by obtaining a first prediction feature corresponding to the coefficient hyperparameter features via the first prediction network, obtaining a second prediction feature corresponding to the reference features via the second prediction network, combining the first and second prediction features to obtain a combined feature, inputting the combined feature into the prediction fusion network, and processing the combined feature via the prediction fusion network to obtain the mean feature of the stage subblock.
[0269] Exemplary, when the decision module obtains a first prediction feature corresponding to the coefficient hyperparameter feature via the first prediction network, it is used to obtain the first prediction feature by performing a feature enhancement operation and an upsampling operation on the coefficient hyperparameter feature via the first prediction network, wherein the feature enhancement operation includes a convolution operation, or a convolution operation and an activation operation, and the upsampling operation includes a deconvolution operation, a crop operation and an activation operation, or a deconvolution operation, a crop operation, an activation operation and a convolution operation. When the decision module obtains a second prediction feature corresponding to the reference feature via the second prediction network, it is used to obtain the second prediction feature by performing feature concatenation on all reconstructed features in the reference feature according to the channel dimension and performing a convolution operation on the concatenated features, or by performing feature addition on all reconstructed features in the reference feature and performing a convolution operation on the added features. The decision module is used to process the post-combined features via the predictive fusion network to obtain the mean features of the step subblock, specifically by performing a convolution operation and at least one fusion operation on the post-combined features via the predictive fusion network to obtain the mean features of the step subblock, wherein the fusion operation includes an activation operation and a convolution operation.
[0270] For example, the decision module is further used to perform feature enhancement on the reconstruction features of the stage subblock to obtain enhanced reconstruction features, the enhanced reconstruction features are used to determine the mean features of the stage subblock, and the enhanced reconstruction features are used to determine the reconstruction image block.
[0271] Illustratively, when the decision module determines a reconstructed image block corresponding to the current image block based on the reconstruction features of each stage subblock, it is used to obtain aggregated features by performing feature aggregation on the reconstruction features of each stage subblock, inputting the aggregated features into a composite transformation network to obtain a reconstructed image block corresponding to the current image block, or to obtain aggregated features by performing feature aggregation on the reconstruction features of each stage subblock, perform block partitioning on the aggregated features to obtain a plurality of block partitioning features, input each block partitioning feature into the composite transformation network to obtain a block partitioned reconstructed image block, merge the block partitioned reconstructed image blocks corresponding to the plurality of block partitioning features to obtain a reconstructed image block corresponding to the current image block, or to input the reconstruction features of each stage subblock into the composite transformation network to obtain a block partitioned reconstructed image block, merge the block partitioned reconstructed image blocks corresponding to all stage subblocks to obtain a reconstructed image block corresponding to the current image block.
[0272] Exemplary, the decision module is used to obtain an aggregated feature by performing feature aggregation on the reconstructed features of each stage subblock. Specifically, it is used to obtain the aggregated feature by aggregating the reconstructed features of each stage subblock according to phase, or by aggregating the reconstructed features of each stage subblock according to phase to obtain a plurality of phase-aggregated features, and then combining the plurality of phase-aggregated features according to channel to obtain the aggregated feature.
[0273] For example, when the decision module performs block division on the aggregated feature to obtain a plurality of block division features, it specifically determines the target size of the block division features, and based on the target size, evenly divides the aggregated feature into a plurality of block division features in the order from top to bottom and left to right, so that the size of each block division feature becomes the target size, or determines the actual block division size and overlap size of the block division features, the actual block division size being the block division size obtained by removing the overlap portion from the block division feature, and the overlap size being the size of the overlap portion of adjacent blocks, and is used to divide the aggregated feature into a plurality of block division features based on the actual block division size and the overlap size so that the size of each block division feature becomes the target size.
[0274] For example, when the decision module merges block-partitioned reconstructed image blocks corresponding to the plurality of block-partitioning features to obtain a reconstructed image block corresponding to the current image block, specifically, it sorts the plurality of block-partitioned reconstructed image blocks corresponding to the plurality of block-partitioning features from top to bottom and from left to right, and then sequentially combines the sorted plurality of block-partitioned reconstructed image blocks to obtain a reconstructed image block corresponding to the current image block, or it determines the actual block-partitioning size and overlap size of each block-partitioned reconstructed image block, sorts the plurality of block-partitioned reconstructed image blocks corresponding to the plurality of block-partitioning features from top to bottom and from left to right, and then sequentially combines the sorted plurality of block-partitioned reconstructed image blocks based on the actual block-partitioning size and overlap size to obtain a reconstructed image block corresponding to the current image block. Used to obtain image blocks, the size of the non-overlapping portion of a block-partitioned reconstructed image block is the actual block division size of the block-partitioned reconstructed image block, the size of the overlap portion between a block-partitioned reconstructed image block and an adjacent block-partitioned reconstructed image block is the overlap size of the block-partitioned reconstructed image block, for the overlap portion of two block-partitioned reconstructed image blocks (left and right), the value of the overlap portion is the value of the left block-partitioned reconstructed image block, or the value of the overlap portion is the average value of the two block-partitioned reconstructed image blocks, and for the overlap portion of two block-partitioned reconstructed image blocks (upper and lower), the value of the overlap portion is the value of the upper block-partitioned reconstructed image block, or the value of the overlap portion is the average value of the two block-partitioned reconstructed image blocks.
[0275] For example, when the decision module inputs the aggregated features into the composite transformation network to obtain a reconstructed image block corresponding to the current image block, it specifically decodes the auxiliary bitstream corresponding to the current image block to obtain the bitrate control parameter of the current image block, and inputs the bitrate control parameter and the aggregated features into the composite transformation network to obtain a reconstructed image block corresponding to the current image block, or when the decision module inputs each block division feature into the composite transformation network to obtain a block division reconstructed image block, it specifically decodes the auxiliary bitstream corresponding to the current image block to obtain the bitrate control parameter of each block division feature The decision module is used to obtain bitrate control parameters, input block division features and the bitrate control parameters of the block division features into the composite transformation network to obtain block division reconstructed image blocks corresponding to the block division features, or, when the decision module inputs the reconstruction features of each stage subblock into the composite transformation network to obtain block division reconstructed image blocks, it is used to decode the auxiliary bitstream corresponding to the current image block to obtain bitrate control parameters of each stage subblock, input the reconstruction features of the stage subblock and the bitrate control parameters of the stage subblock into the composite transformation network to obtain the block division reconstructed image blocks.
[0276] Exemplary, the decision module inputs the bitrate control parameter and the aggregated feature into the composite transformation network to obtain a reconstructed image block corresponding to the current image block, specifically by processing the aggregated feature via the composite transformation network to obtain a first feature, processing the bitrate control parameter via the composite transformation network to obtain a second feature, generating a third feature based on the first and second features, and determining the reconstructed image block corresponding to the current image block based on the third feature, or the decision module inputs the block division feature and the bitrate control parameter of the block division feature into the composite transformation network to obtain a block division reconstructed image block, specifically by processing the block division feature via the composite transformation network to obtain a first feature, via the composite transformation network The bitrate control parameters of the block division feature are processed to obtain a second feature, a third feature is generated based on the first and second features, and the block division reconstruction image block corresponding to the block division feature is determined based on the third feature. Alternatively, when the determination module inputs the reconstruction features of the stage subblock and the bitrate control parameters of the stage subblock to the composite transformation network to obtain the block division reconstruction image block, it is used to process the reconstruction features of the stage subblock via the composite transformation network to obtain a first feature, process the bitrate control parameters of the stage subblock via the composite transformation network to obtain a second feature, a third feature is generated based on the first and second features, and the block division reconstruction image block corresponding to the stage subblock is determined based on the third feature.
[0277] Based on the same concept as described above, embodiments of the present invention further provide an encoding device applied to the encoding side, the device comprising: an acquisition module for inputting the current image block into an analysis-transformation network to obtain a feature block corresponding to the current image block and dividing the feature block into features to be encoded for a plurality of stage subblocks; an encoding module for acquiring the coefficient hyperparameter features of each stage subblock corresponding to the current image block and encoding the coefficient hyperparameter features of the stage subblock into a first bitstream of the current image block; and a determination module for determining the residual features of the stage subblock based on the features to be encoded and the mean value features of the stage subblock, wherein the encoding module is further used to determine probability distribution parameters based on the coefficient hyperparameter features of the stage subblock and to encode the residual features of the stage subblock into a second bitstream of the current image block based on the probability distribution parameters.
[0278] Exemplary, the decision module is used to obtain, for the first stage subblock, the mean feature of the stage subblock via an mean prediction network based on the coefficient hyperparameter feature of the stage subblock, or to obtain a set default reference feature and obtain the mean feature of the stage subblock via the mean prediction network based on the coefficient hyperparameter feature and the default reference feature, or, for the i-th stage subblock, to obtain a reference feature of the i-th stage subblock based on the reconstruction features of the i-1 stage subblocks preceding the i-th stage subblock, and to obtain the mean feature of the stage subblock via an mean prediction network based on the coefficient hyperparameter feature and the reference feature, where i is greater than 1, and the reference feature includes all the reconstruction features of the i-1 preceding stage subblocks, or some of the reconstruction features of the i-1 preceding stage subblocks, or the reconstruction feature of the i-1th stage subblock.
[0279] Those skilled in the art will understand that embodiments of the present invention may be provided as methods, systems, or computer program products. The present invention may take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Embodiments of the present invention may take the form of computer program products implemented on one or more computer-compatible storage media (including, but not limited to, magnetic disk memory, CD-ROM, optical memory, etc.) containing computer-compatible program code. The above are merely embodiments of the present invention and do not limit the present invention.
[0280] To those skilled in the art, the present invention is subject to various modifications and changes. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. The steps include decoding the first bitstream of the current image block to obtain coefficient hyperparameter features of each stage subblock of the current image block, For each stage subblock, the steps include determining the probability distribution parameters based on the coefficient hyperparameter features of the stage subblock, decoding the second bitstream of the current image block based on the probability distribution parameters to obtain the residual features of the stage subblock, A step of determining the reconstruction features of the stage subblock based on the residual features and the mean features of the stage subblock, The process includes the step of determining the reconstructed image block corresponding to the current image block based on the reconstruction characteristics of each stage subblock, A decoding method characterized by the following:
2. Before determining the reconstruction features of the stage subblock based on the residual features and mean features of the stage subblock, For the first stage subblock, the process includes obtaining the mean value feature of the stage subblock via an mean value prediction network based on the coefficient hyperparameter feature of the stage subblock, or obtaining a set default reference feature, and obtaining the mean value feature of the stage subblock via the mean value prediction network based on the coefficient hyperparameter feature and the default reference feature of the stage subblock. The method according to feature 1.
3. Before determining the reconstruction features of the stage subblock based on the residual features and mean features of the stage subblock, The process includes the steps of obtaining a reference feature for the i-th stage subblock based on the reconstruction features of the i-1 stage subblocks preceding the i-th stage subblock, and obtaining an average feature for the stage subblock via an average prediction network based on the coefficient hyperparameter features of the stage subblock and the reference feature, i is greater than 1, and the reference feature includes all reconstruction features of the previous i-1 step subblocks, or some reconstruction features of the previous i-1 step subblocks, or reconstruction features of the i-1 step subblock. The method according to feature 1.
4. The mean prediction network includes a first prediction network, a second prediction network, and a prediction fusion network, and the step of obtaining the mean feature of the step subblock via the mean prediction network based on the coefficient hyperparameter features of the step subblock and the reference feature is: The steps include obtaining a first prediction feature corresponding to the coefficient hyperparameter feature via the first prediction network, The steps include obtaining a second predictive feature corresponding to the reference feature via the two predictive networks described above, The steps include: combining the first prediction feature and the second prediction feature to obtain a combined feature; inputting the combined feature into the prediction fusion network; and processing the combined feature through the prediction fusion network to obtain the average feature of the stage subblock. The method according to feature 3.
5. The step of obtaining a first prediction feature corresponding to the coefficient hyperparameter feature via the first prediction network is: The step of obtaining the first predicted feature is performed on the coefficient hyperparameter feature via the first prediction network by performing feature enhancement and upsampling operations. The feature enhancement operation includes a convolution operation, or a convolution operation and an activation operation, and the upsampling operation includes a deconvolution operation, a crop operation and an activation operation, or a deconvolution operation, a crop operation, an activation operation and a convolution operation. The method according to feature 4.
6. The activation operation is a normalized linear unit Relu operation. The method according to specification 5.
7. The step of obtaining a second predictive feature corresponding to the reference feature via the second predictive network is: The process includes the steps of: performing feature concatenation on all reconstructed features in the reference feature according to the channel dimension, and then performing a convolution operation on the concatenated features to obtain the second predicted feature; or performing feature addition on all reconstructed features in the reference feature, and then performing a convolution operation on the added features to obtain the second predicted feature. The method according to feature 4.
8. The step of processing the post-combined features via the predictive fusion network to obtain the mean features of the stage subblock is: The step includes performing a convolution operation and at least one fusion operation on the combined features via the predictive fusion network to obtain the mean feature of the step subblock, The aforementioned fusion operation includes an activation operation and a convolution operation. The method according to feature 4.
9. The activation operation is a normalized linear unit Relu operation. The method according to feature 8.
10. After determining the reconstruction features of the stage subblock based on the residual features and mean features of the stage subblock, The step includes performing feature enhancement on the reconstructed features of the stage subblock to obtain the enhanced reconstructed features, The enhanced reconstruction features are used to determine the mean features of the step subblocks, and the enhanced reconstruction features are used to determine the reconstructed image blocks. The method according to feature 1.
11. The step of determining the reconstructed image block corresponding to the current image block based on the reconstruction characteristics of each of the aforementioned stage subblocks is: Feature aggregation is performed on the reconstruction features of each stage subblock to obtain aggregated features, and these aggregated features are input into the synthesis transformation network to obtain the reconstructed image block corresponding to the current image block. Alternatively, feature aggregation is performed on the reconstruction features of each stage subblock to obtain aggregated features, block partitioning is performed on the aggregated features to obtain multiple block partitioning features, each block partitioning feature is input to the composite transformation network to obtain a block partitioned reconstructed image block, and the block partitioned reconstructed image blocks corresponding to the multiple block partitioning features are merged to obtain a reconstructed image block corresponding to the current image block. Alternatively, the process includes the steps of inputting the reconstruction features of each stage subblock into the composite transformation network to obtain a block-partitioned reconstructed image block, and merging the block-partitioned reconstructed image blocks corresponding to all stage subblocks to obtain a reconstructed image block corresponding to the current image block. The method according to feature 1.
12. The step of performing feature aggregation on the reconstructed features of each of the aforementioned sub-block stages to obtain aggregated features is: The reconstruction features of each stage subblock are aggregated according to their phase to obtain the aggregated features, or The reconstruction features of each stage subblock are aggregated according to their phase to obtain a plurality of phase-aggregated features, and the plurality of phase-aggregated features are combined according to the channel to obtain the aggregated feature, which includes the step of aggregating the reconstruction features of each stage subblock according to their phase to obtain a plurality of phase-aggregated features. The method according to 11, characterized by the features described above.
13. The step of performing block partitioning on the aggregated features to obtain multiple block partitioning features is: Determine the target size of the block division feature, and based on the target size, divide the aggregated feature evenly into multiple block division features in the order of top to bottom and left to right, so that the size of each block division feature becomes the target size, or The process includes the steps of determining the actual block division size and overlap size of the block division feature, and dividing the aggregated feature into a plurality of block division features based on the actual block division size and overlap size, such that the size of each block division feature becomes the target size. The actual block division size is the block division size obtained by removing the overlap portion from the block division features, and the overlap size is the size of the overlap portion of adjacent blocks. The method according to 11, characterized by the features described above.
14. The step of merging the block-divided reconstructed image blocks corresponding to the plurality of block-dividing features to obtain a reconstructed image block corresponding to the current image block is: The plurality of reconstructed image blocks corresponding to the plurality of block division features are sorted from top to bottom and from left to right, and the sorted plurality of reconstructed image blocks are sequentially combined to obtain a reconstructed image block corresponding to the current image block, or, The steps include determining the actual block division size and overlap size of each block division reconstruction image block, sorting the multiple block division reconstruction image blocks corresponding to the multiple block division features from top to bottom and from left to right, and sequentially combining the sorted multiple block division reconstruction image blocks based on the actual block division size and the overlap size to obtain a reconstruction image block corresponding to the current image block, The size of the non-overlapping portion of a block-partitioned reconstructed image block is the actual block division size of that block-partitioned reconstructed image block, and the size of the overlapping portion between a block-partitioned reconstructed image block and an adjacent block-partitioned reconstructed image block is the overlap size of that block-partitioned reconstructed image block. For the overlapping portion of the two block-reconstructed image blocks (left and right), the value of the overlapping portion is the value of the left block-reconstructed image block, or the value of the overlapping portion is the average value of the two block-reconstructed image blocks. For the overlapping portion of two block-reconstructed image blocks, the value of the overlapping portion is the value of the upper block-reconstructed image block, or the value of the overlapping portion is the average value of the two block-reconstructed image blocks. The method according to the present invention, characterized by the present invention.
15. The step of inputting the aggregated features into a composite transformation network to obtain a reconstructed image block corresponding to the current image block includes the step of decoding the auxiliary bitstream corresponding to the current image block to obtain the bitrate control parameter of the current image block, inputting the bitrate control parameter and the aggregated features into the composite transformation network to obtain a reconstructed image block corresponding to the current image block, or The step of inputting each of the aforementioned block division features into the composite transformation network to obtain a block division reconstructed image block includes the step of decoding the auxiliary bitstream corresponding to the current image block to obtain the bitrate control parameter for each block division feature, inputting the block division feature and the bitrate control parameter for the block division feature into the composite transformation network to obtain a block division reconstructed image block corresponding to the block division feature, or The step of inputting the reconstruction features of each stage subblock into the composite transformation network to obtain a block-partitioned reconstructed image block includes decoding the auxiliary bitstream corresponding to the current image block to obtain the bitrate control parameters of each stage subblock, and inputting the reconstruction features of the stage subblock and the bitrate control parameters of the stage subblock into the composite transformation network to obtain the block-partitioned reconstructed image block. The method according to 11, characterized by the features described above.
16. The step of inputting the bitrate control parameter and the aggregated feature into the composite transformation network to obtain a reconstructed image block corresponding to the current image block includes, or, processing the aggregated feature via the composite transformation network to obtain a first feature, processing the bitrate control parameter via the composite transformation network to obtain a second feature, and generating a third feature based on the first and second features, and determining a reconstructed image block corresponding to the current image block based on the third feature. The step of inputting the block division feature and the bitrate control parameter of the block division feature into the composite transformation network to obtain a block division reconstructed image block corresponding to the block division feature includes, or, the step of processing the block division feature via the composite transformation network to obtain a first feature, processing the bitrate control parameter of the block division feature via the composite transformation network to obtain a second feature, and generating a third feature based on the first and second features, and determining a block division reconstructed image block corresponding to the block division feature based on the third feature. The step of inputting the reconstruction features of the stage subblock and the bitrate control parameters of the stage subblock into the composite transformation network to obtain the block-partitioned reconstructed image block includes: processing the reconstruction features of the stage subblock via the composite transformation network to obtain a first feature; processing the bitrate control parameters of the stage subblock via the composite transformation network to obtain a second feature; generating a third feature based on the first and second features, and determining the block-partitioned reconstructed image block corresponding to the stage subblock based on the third feature. The method according to the present invention, characterized by the present invention.
17. The steps include inputting the current image block into an analysis and transformation network to obtain a feature block corresponding to the current image block, The steps include dividing the feature block into features to be encoded in multiple sub-blocks, For each stage subblock corresponding to the current image block, the steps include obtaining the coefficient hyperparameter features of the stage subblock and encoding the coefficient hyperparameter features of the stage subblock into the first bitstream of the current image block. A step of determining the residual features of the stage subblock based on the features to be encoded of the stage subblock and the mean value features of the stage subblock, The process includes the steps of determining probability distribution parameters based on the coefficient hyperparameter features of the stage subblock, and encoding the residual features of the stage subblock into a second bitstream of the current image block based on the probability distribution parameters, An encoding method characterized by the following.
18. The steps include decoding the first bitstream of the current image block to obtain the coefficient hyperparameter features of the current image block, The steps include determining probability distribution parameters based on the coefficient hyperparameter features, decoding the second bitstream of the current image block based on the probability distribution parameters to obtain residual features of the current image block, and determining reconstruction features of the current image block based on the residual features, The process includes: decoding the auxiliary bitstream corresponding to the current image block to obtain bitrate control parameters corresponding to the current image block; inputting the reconstruction features and the bitrate control parameters into a composite transformation network to obtain a reconstructed image block corresponding to the current image block; A decoding method characterized by the following:
19. The step of inputting the reconstruction features and the bitrate control parameters into the composite transformation network to obtain a reconstructed image block corresponding to the current image block is: The steps include processing the reconstructed features via the aforementioned synthesis transformation network to obtain a first feature, The steps include processing the bitrate control parameters via the aforementioned synthesis and conversion network to obtain a second feature, A step of generating a third feature based on the first and second features, The step of determining a reconstructed image block corresponding to the current image block based on the third characteristic described above, The method according to the present invention, characterized by the present invention.
20. A decoding module for decoding the first bitstream of the current image block to obtain coefficient hyperparameter features of each stage subblock of the current image block, determining probability distribution parameters for each stage subblock based on the coefficient hyperparameter features of the stage subblock, decoding the second bitstream of the current image block based on the probability distribution parameters to obtain residual features of the stage subblock, Includes a decision module for determining the reconstruction features of a stage subblock based on the residual features and mean features of the stage subblock, and for determining the reconstructed image block corresponding to the current image block based on the reconstruction features of each stage subblock, A decoding device characterized by the following features.
21. An acquisition module for inputting the current image block into an analysis and transformation network to obtain a feature block corresponding to the current image block, and for dividing the feature block into features to be encoded in multiple stage subblocks, For each stage subblock corresponding to the current image block, an encoding module for obtaining the coefficient hyperparameter features of the stage subblock and encoding the coefficient hyperparameter features of the stage subblock into the first bitstream of the current image block, Includes a decision module for determining the residual features of the stage subblock based on the features to be encoded of the stage subblock and the mean features of the stage subblock, The encoding module is further used to determine probability distribution parameters based on the coefficient hyperparameter features of the stage subblock, and to encode the residual features of the stage subblock into a second bitstream of the current image block based on the probability distribution parameters. An encoding device characterized by the following features.
22. A decoding device comprising a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor. The processor is used to execute machine-executable instructions and carry out the method according to any one of claims 1 to 16. A decoding device characterized by the following features.
23. An encoding device comprising a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor, The processor is used to execute machine-executable instructions and carry out the method according to claim 17. An encoding device characterized by the following features.
24. A machine-readable storage medium storing a plurality of computer instructions, wherein when the computer instructions are executed by a processor, the method according to any one of claims 1 to 16 is performed, or when the computer instructions are executed by a processor, the method according to claim 17 is performed. A machine-readable storage medium characterized by the following features.
25. When executed by a processor, the method according to any one of claims 1 to 16 is performed, or when executed by a processor, the method according to claim 17 is performed. A computer application characterized by the following features.