A decoding, encoding method, apparatus and device thereof

By enhancing features through encoding parameters at the encoding end and utilizing probability distribution parameters at the decoding end, the problems of poor encoding performance and high complexity in neural network encoding and decoding methods are solved, improving the reconstruction quality and encoding efficiency of video images, and achieving efficient encoding and decoding with low complexity.

CN119011850BActive Publication Date: 2025-11-25HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411375202.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-20
Publication Date
2025-11-25
Estimated Expiration
2043-04-20

AI Technical Summary

Technical Problem

Existing neural network-based encoding and decoding methods suffer from poor encoding performance, poor decoding performance, and high complexity. In particular, they fail to fully utilize advanced prior edge information in video image encoding, resulting in insufficient improvement in image reconstruction quality.

Method used

At the encoding end, encoding enhancement parameters are input into the header information bitstream. At the decoding end, probability distribution parameters are used to enhance features, thereby improving the quality of the reconstructed image. By combining feature domain and image domain enhancement techniques of neural networks, encoding and decoding efficiency is improved.

Benefits of technology

While maintaining low complexity, it effectively improves the reconstruction quality and coding performance of video images, reduces the complexity of encoding and decoding, and achieves better encoding and decoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119011850B_ABST
    Figure CN119011850B_ABST
Patent Text Reader

Abstract

The application provides a decoding method, an encoding method, a device and equipment thereof. The decoding method comprises: decoding a code stream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block; determining a probability distribution parameter based on the coefficient hyperparameter feature; decoding the code stream corresponding to the current image block based on the probability distribution parameter to obtain an initial reconstruction feature corresponding to the current image block; and determining a target reconstruction image block corresponding to the current image block based on the initial reconstruction feature. Through the technical solution of the application, the encoding performance and the decoding performance can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of coding and decoding, and in particular to a decoding and encoding method, device and equipment. BACKGROUND

[0002] In order to save space, video images are transmitted after being encoded. A complete video encoding method can include prediction, transformation, quantization, entropy encoding, filtering and the like. Among them, the prediction process can include intra prediction and inter prediction. Inter prediction uses the correlation in the time domain of the video to predict the current pixel of the current frame image using the pixels of the adjacent coded image, so as to effectively remove the temporal redundancy of the video. Intra prediction uses the correlation in the spatial domain of the video to predict the current pixel using the pixels of the coded block of the current frame image, so as to effectively remove the spatial redundancy of the video.

[0003] With the rapid development of deep learning, deep learning has achieved success in many high-level computer vision problems, such as image classification, object detection, etc. Deep learning has also gradually begun to be applied in the field of coding and decoding, that is, a neural network can be used to encode and decode images. Although the coding and decoding method based on neural network shows great performance potential, the coding and decoding method based on neural network still has problems such as poor coding performance, poor decoding performance and high complexity. SUMMARY

[0004] The present application provides a decoding and encoding method, device and equipment.

[0005] The present application provides a decoding method applied to a decoding end, which comprises: decoding a code stream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block; determining a probability distribution parameter based on the coefficient hyperparameter feature; decoding the code stream corresponding to the current image block based on the probability distribution parameter to obtain an initial reconstruction feature corresponding to the current image block; and determining a target reconstruction image block corresponding to the current image block based on the initial reconstruction feature.

[0006] The present application provides an encoding method applied to an encoding end, which comprises: encoding a coefficient hyperparameter feature corresponding to a current image block to obtain a first code stream corresponding to the current image block; determining a probability distribution parameter based on the coefficient hyperparameter feature; encoding an initial image feature corresponding to the current image block based on the probability distribution parameter to obtain a second code stream corresponding to the current image block; and encoding an important channel identifier to obtain a third code stream corresponding to the current image block.

[0007] The application provides a decoding device, which is applied to a decoding end, and the device comprises: a decoding module, which is used for decoding a code stream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block; determining a probability distribution parameter based on the coefficient hyperparameter feature, decoding the code stream corresponding to the current image block based on the probability distribution parameter to obtain an initial reconstruction feature corresponding to the current image block; and a determining module, which is used for determining a target reconstruction image block corresponding to the current image block based on the initial reconstruction feature.

[0008] The application provides an encoding device, which is applied to an encoding end, and the device comprises: an encoding module, which is used for encoding a coefficient hyperparameter feature corresponding to a current image block to obtain a first code stream corresponding to the current image block; determining a probability distribution parameter based on the coefficient hyperparameter feature; encoding an initial image feature corresponding to the current image block based on the probability distribution parameter to obtain a second code stream corresponding to the current image block; and encoding an important channel identifier to obtain a third code stream corresponding to the current image block.

[0009] The application provides a decoding end device, which comprises: a processor and a machine readable storage medium, wherein the machine readable storage medium stores machine executable instructions which can be executed by the processor; and the processor is used for executing the machine executable instructions to implement the decoding method.

[0010] The application provides an encoding end device, which comprises: a processor and a machine readable storage medium, wherein the machine readable storage medium stores machine executable instructions which can be executed by the processor; and the processor is used for executing the machine executable instructions to implement the encoding method.

[0011] The application provides an electronic device, which comprises: a processor and a machine readable storage medium, wherein the machine readable storage medium stores machine executable instructions which can be executed by the processor; and the processor is used for executing the machine executable instructions to implement the decoding method or the encoding method.

[0012] The application provides a machine readable storage medium, wherein the machine readable storage medium stores a plurality of computer instructions, and the computer instructions are executed by a processor to implement the decoding method or the encoding method.

[0013] The application provides a computer application program, wherein the computer application program is executed by a processor to implement the decoding method or the encoding method. BRIEF DESCRIPTION OF DRAWINGS

[0014] Figure 1is a schematic diagram of a three-dimensional feature matrix in an embodiment of the present application.

[0015] Figure 2 is a flowchart of a decoding method in an embodiment of the present application.

[0016] Figure 3 is a flowchart of an encoding method in an embodiment of the present application.

[0017] Figure 4 is a schematic diagram of a processing procedure of an encoding end in an embodiment of the present application.

[0018] Figure 5 is a schematic diagram of a processing procedure of a decoding end in an embodiment of the present application.

[0019] Figure 6A is a schematic diagram of a position of a feature domain enhancement module in an embodiment of the present application.

[0020] Figure 6B is a schematic diagram of a position of a feature domain enhancement module and an image domain enhancement module in an embodiment of the present application.

[0021] Figure 6C is a schematic diagram of a position of an image domain enhancement module in an embodiment of the present application.

[0022] Figure 7A is a hardware structure diagram of a decoding end device in an embodiment of the present application.

[0023] Figure 7B is a hardware structure diagram of an encoding end device in an embodiment of the present application. DETAILED DESCRIPTION

[0024] The terminology used in the embodiments of the present application merely describes specific embodiments, and is not intended to limit the present application. The singular forms "a," "an," and "the" used in the embodiments of the present application and the claims are intended to include both singular and plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application means any or all possible combinations of one or more associated listed items. It should be understood that although the terms first, second, third, etc. can be used in the embodiments of the present application to describe various information, these terms are not intended to limit the scope of the present application. These terms are only used to distinguish one piece of information from another piece of information of the same type. For example, a first information can also be referred to as a second information, and similarly, a second information can also be referred to as a first information, without departing from the scope of the embodiments of the present application. In addition, depending on the context, the word "if" used in the present application can be interpreted as "when" or "upon determining" or "in response to determining".

[0025] The embodiments of the present application propose a decoding method and an encoding method, which can involve the following concepts: entropy encoding, neural network (NN), convolutional neural network (CNN), deconvolution, generalization ability, feature, rate-distortion optimization (RDO) principle.

[0026] Entropy encoding refers to encoding in an encoding process without losing any information according to the entropy principle. The information entropy is the average amount of information of a source (a measure of uncertainty). The encoding mode of entropy encoding can include, but is not limited to, Shannon encoding, Huffman encoding and arithmetic coding.

[0027] Neural network refers to artificial neural network. The neural network is an operation model composed of a large number of nodes (or neurons) connected with each other. In the neural network, the neuron processing unit can represent different objects, such as features, letters, concepts or some meaningful abstract patterns. The types of processing units in the neural network can be divided into three categories: input units, output units and hidden units. The input units receive signals and data from the external world; the output units output the processing results; the hidden units are between the input units and the output units and cannot be observed from the outside of the system. The connection weight between the neurons reflects the connection strength between the units, and the information representation and processing are embodied in the connection relationship of the processing units. The neural network is a non-programmed, brain-like information processing method. The essence of the neural network is to obtain a parallel distributed information processing function through the transformation and dynamics of the neural network, and to imitate the information processing function of the human brain neural system at different levels and degrees. In the field of video processing, the commonly used neural networks can include, but are not limited to, convolutional neural network (CNN), recurrent neural network (RNN) and fully connected network.

[0028] A convolutional neural network (CNN) is a feed-forward neural network, which is one of the most representative network structures in deep learning technology. The artificial neurons of a convolutional neural network can respond to a portion of the surrounding units within the coverage range, and have excellent performance for large image processing. The basic structure of a convolutional neural network includes two layers: a feature extraction layer (also referred to as a convolution layer) and a feature mapping layer (also referred to as an activation layer). For the feature extraction layer, the input of each neuron is connected to the local receptive field of the previous layer, and the local features are extracted. Once the local features are extracted, the positional relationship between the local features and other features is also determined. For the feature mapping layer, each calculation layer of the neural network is composed of multiple feature mappings, and each feature mapping is a plane, and the weights of all neurons on the plane are equal. The feature mapping structure can use a Sigmoid function, a ReLU function, a Leaky-ReLU function, a PReLU function, a GDN function, etc. as the activation function of the convolutional neural network. In addition, since the neurons on a mapping plane share weights, the number of free parameters of the network is reduced.

[0029] For example, one of the advantages of a convolutional neural network over an image processing algorithm is that it avoids complex pre-processing of images (extracting artificial features, etc.), and can directly input the original image for end-to-end learning. One of the advantages of a convolutional neural network over a general neural network is that the general neural network uses full connection, i.e. all neurons of the input layer are connected to the hidden layer. This will result in a large number of parameters, making the network training time-consuming and even difficult to train. The convolutional neural network avoids this difficulty through local connection and weight sharing.

[0030] The deconvolution layer is also referred to as a transposed convolution layer. The working process of the deconvolution layer is similar to that of the convolution layer, and the main difference is that the deconvolution layer will pad to make the output larger than the input (of course, it can also remain the same). If the stride is 1, the output size is equal to the input size. If the stride is N, the width of the output feature is N times the width of the input feature, and the height of the output feature is N times the height of the input feature.

[0031] The generalization capability can refer to the adaptability of a machine learning algorithm to new samples. The purpose of learning is to learn the rules hidden in the data pairs. The trained network can also give appropriate output for data outside the learning set with the same rules, and this capability can be referred to as generalization capability.

[0032] The feature referred to in the present application is a three-dimensional feature matrix or tensor of C*W*H. Referring to Figure 1 as shown, Figure 1is a schematic diagram of a three-dimensional feature matrix, in which C represents the number of channels, H represents the feature height, and W represents the feature width. The three-dimensional feature matrix can be an input of a neural network or an output of the neural network.

[0033] There are two indicators for evaluating coding efficiency: code rate and PSNR (Peak Signal to Noise Ratio). The smaller the bit stream, the greater the compression rate, and the greater the PSNR, the better the quality of the reconstructed image. In mode selection, the decision formula is essentially a comprehensive evaluation of the two. For example, the cost of the mode:

[0034] J(mode) = D + λ * R.

[0035] where D represents Distortion, which can be measured using the SSE (Sum of Squared Errors) indicator, which is the sum of squares of the differences between the reconstructed image block and the source image. In order to achieve cost consideration, the SAD (Sum of Absolute Differences) indicator can also be used, which is the sum of absolute values of the differences between the reconstructed image block and the source image; λ is the Lagrange multiplier, and R is the actual number of bits required for encoding the image block under the mode, including the total number of bits required for encoding mode information, motion information, and residual. In mode selection, if the rate-distortion optimization principle is used to make a comparative decision on the encoding mode, the best encoding performance can usually be guaranteed.

[0036] A large number of encoding tools are proposed for each module of the encoding end, and each tool often has multiple modes. For different video sequences, the encoding tool that can obtain the optimal encoding performance is often different. Therefore, in the encoding process, the RDO principle is usually used to compare the encoding performance of different tools or modes to select the optimal tool or mode. After determining the optimal tool or mode, the decision information of the tool or mode is transmitted by encoding the flag information in the bit stream. Although this method brings higher encoding complexity, it can adaptively select the optimal mode combination for different contents to obtain the optimal encoding performance. The decoding end can obtain the relevant mode information by directly parsing the flag information, which has little effect on the decoding complexity.

[0037] In the end-to-end image coding general framework, a feature main information part and a hyper-prior side information part are mainly included. The feature main information part includes an analysis network, quantization, normal entropy coding, normal entropy decoding and a synthesis network, and the hyper-prior side information part includes a hyper-prior analysis network, quantization, factorized entropy coding, factorized entropy decoding and a hyper-prior synthesis network. Image components are compressed and coded and reconstructed and recovered in the analysis network and the synthesis network of the feature main information part; the hyper-prior side information part is mainly used for modeling the probability of the feature main information, and guiding the entropy coding and decoding of the feature main information. In the end-to-end image coding general framework, there is a problem that the hyper-prior side information is not fully utilized, and the image reconstruction quality is not fully improved.

[0038] Therefore, in the embodiments of the present application, the characteristics of the end-to-end image coding framework are utilized, the decoding end uses the probability distribution parameters to enhance the quality of the features, and the quality of the reconstructed image is improved. The encoding end does not directly change the main information, but encodes the enhancement parameters into the header information code stream (the third code stream), and the decoding end enhances the main information features through the enhancement parameters.

[0039] The decoding method and the encoding method in the embodiments of the present application are described in detail below in combination with several specific embodiments.

[0040] Embodiment 1: A decoding method is proposed in the embodiments of the present application, as shown in Figure 2 , which is a flowchart of the decoding method. The decoding method can be applied to the decoding end (also referred to as a video decoder), and the method can include steps 201-204.

[0041] In step 201, a first code stream corresponding to a current image block is decoded to obtain a coefficient hyper-parameter feature corresponding to the current image block.

[0042] In step 202, a probability distribution parameter is determined based on the coefficient hyper-parameter feature, and a second code stream corresponding to the current image block is decoded based on the probability distribution parameter to obtain an initial reconstructed feature corresponding to the current image block.

[0043] In step 203, a third code stream corresponding to the current image block is decoded to obtain an enhancement parameter corresponding to the current image block, wherein the third code stream can also be referred to as a header information code stream corresponding to the current image block.

[0044] In step 204, the initial reconstructed feature is enhanced based on the enhancement parameter and the probability distribution parameter to obtain an enhanced reconstructed feature, and a target reconstructed image block corresponding to the current image block is determined based on the enhanced reconstructed feature.

[0045] Exemplarily, the enhancement parameter comprises an important channel identifier. The initial reconstructed feature comprises C feature channel maps. The probability distribution parameter comprises C probability distribution channel maps, and the C probability distribution channel maps correspond to the C feature channel maps. Figure 1 An important feature channel map is selected from the C feature channel maps according to the important channel identifier, and the remaining feature channel maps are selected as non-important feature channel maps. An important probability distribution channel map is selected according to the important feature channel map, and a non-important probability distribution channel map is selected according to the non-important feature channel map.

[0046] Exemplarily, the initial reconstructed feature can comprise C feature channel maps, and the probability distribution parameter can comprise C probability distribution channel maps, and the C probability distribution channel maps correspond to the C feature channel maps. Figure 1 An important feature channel map is selected from the C feature channel maps according to the important channel identifier, and the remaining feature channel maps are selected as non-important feature channel maps. An important probability distribution channel map is selected according to the important feature channel map, and a non-important probability distribution channel map is selected according to the non-important feature channel map.

[0047] Exemplarily, the initial reconstructed feature can comprise an important feature channel map and a non-important feature channel map, and the probability distribution parameter can comprise an important probability distribution channel map corresponding to the important feature channel map and a non-important probability distribution channel map corresponding to the non-important feature channel map. The enhancement of the initial reconstructed feature based on the enhancement parameter and the probability distribution parameter to obtain the enhanced reconstructed feature can comprise but is not limited to: performing feature adaptive edge enhancement on the important feature channel map based on the feature domain enhancement parameter and the important probability distribution channel map to obtain a first reconstructed feature after feature adaptive edge enhancement; performing feature adaptive stretching on the non-important feature channel map based on the feature domain enhancement parameter and the non-important probability distribution channel map to obtain a second reconstructed feature after feature adaptive stretching; and generating the enhanced reconstructed feature based on the first reconstructed feature and the second reconstructed feature.

[0048] Exemplarily, the feature domain enhancement parameter can include a plurality of edge enhancement segment intensity values and a plurality of edge enhancement segment threshold values, the plurality of edge enhancement segment threshold values can constitute a plurality of edge enhancement threshold intervals, and the plurality of edge enhancement threshold intervals correspond to the plurality of edge enhancement segment intensity values one by one. On this basis, the feature adaptive edge enhancement on the important feature channel map based on the feature domain enhancement parameter and the important probability distribution channel map to obtain the first reconstructed feature after the feature adaptive edge enhancement can include but is not limited to: if the important probability distribution channel map includes a plurality of probability distribution values, for each probability distribution value, the edge enhancement segment intensity value corresponding to the probability distribution value can be determined based on the edge enhancement threshold interval corresponding to the probability distribution value; then, the important feature channel map is subjected to feature adaptive edge enhancement based on the edge enhancement segment intensity value corresponding to each probability distribution value to obtain the first reconstructed feature after the feature adaptive edge enhancement.

[0049] Exemplarily, the feature adaptive edge enhancement on the important feature channel map based on the edge enhancement segment intensity value corresponding to each probability distribution value to obtain the first reconstructed feature after the feature adaptive edge enhancement can include but is not limited to: normalizing the important feature channel map to obtain a normalized feature map; generating a high-frequency detail image based on the important feature channel map and the normalized feature map; for each feature value in the high-frequency detail image, the feature value is subjected to edge enhancement based on the edge enhancement segment intensity value corresponding to the probability distribution value corresponding to the feature value to obtain an edge enhanced feature value; determining an edge enhanced feature map based on the edge enhanced feature value corresponding to each feature value; and de-normalizing the edge enhanced feature map to obtain the first reconstructed feature.

[0050] Exemplarily, the feature domain enhancement parameter can include a stretch parameter value, and the feature adaptive stretching on the unimportant feature channel map based on the feature domain enhancement parameter and the unimportant probability distribution channel map to obtain the second reconstructed feature after the feature adaptive stretching can include but is not limited to: if the unimportant feature channel map includes a plurality of feature values, and the unimportant probability distribution channel map includes a plurality of probability distribution values, for each feature value in the unimportant feature channel map, the stretched feature value corresponding to the feature value can be determined based on the feature value, the stretch parameter value and the probability distribution value corresponding to the feature value; then, the second reconstructed feature is determined based on the stretched feature value corresponding to each feature value in the unimportant feature channel map.

[0051] For example, determining the target reconstructed image block corresponding to the current image block based on the enhanced reconstructed feature can include, but is not limited to: inputting the enhanced reconstructed feature into a synthesis transformation network to obtain the target reconstructed image block corresponding to the current image block; or inputting the enhanced reconstructed feature into the synthesis transformation network to obtain an initial reconstructed image block corresponding to the current image block; performing image adaptive edge enhancement on the initial reconstructed image block based on the image domain enhancement parameter corresponding to the current image block and the probability distribution parameter to obtain the target reconstructed image block corresponding to the current image block; wherein the image domain enhancement parameter is obtained by decoding the third code stream corresponding to the current image block, i.e., decoding the third code stream to obtain the image domain enhancement parameter corresponding to the current image block.

[0052] For example, the image domain enhancement parameter can include a plurality of image enhancement segment intensity values and a plurality of image enhancement segment threshold values, the plurality of image enhancement segment threshold values can form a plurality of image enhancement threshold intervals, and the plurality of image enhancement threshold intervals correspond to the plurality of image enhancement segment intensity values one by one. Based on this, performing image adaptive edge enhancement on the initial reconstructed image block based on the image domain enhancement parameter corresponding to the current image block and the probability distribution parameter to obtain the target reconstructed image block corresponding to the current image block can include, but is not limited to: obtaining a target probability distribution channel map based on the probability distribution parameter; if the target probability distribution channel map includes a plurality of probability distribution values, then for each probability distribution value, determining the image enhancement segment intensity value corresponding to the probability distribution value based on the image enhancement threshold interval corresponding to the probability distribution value; and performing image adaptive edge enhancement on the initial reconstructed image block based on the image enhancement segment intensity value corresponding to each probability distribution value to obtain the target reconstructed image block corresponding to the current image block.

[0053] For example, obtaining a target probability distribution channel map based on the probability distribution parameter can include, but is not limited to: if the probability distribution parameter includes an important probability distribution channel map and a non-important probability distribution channel map, upsampling the important probability distribution channel map to obtain the target probability distribution channel map; the size of the target probability distribution channel map is the same as the size of the initial reconstructed image block.

[0054] For example, performing image adaptive edge enhancement on the initial reconstructed image block based on the image enhancement segment intensity value corresponding to each probability distribution value to obtain the target reconstructed image block corresponding to the current image block can include, but is not limited to: generating a high-frequency detail image based on the initial reconstructed image block; for each feature value in the high-frequency detail image, performing edge enhancement on the feature value based on the image enhancement segment intensity value corresponding to the probability distribution value corresponding to the feature value to obtain an image enhancement feature value; and determining the target reconstructed image block based on the image enhancement feature value corresponding to each feature value in the high-frequency detail image.

[0055] For example, the enhancing the initial reconstructed feature based on the enhancement parameter and the probability distribution parameter to obtain an enhanced reconstructed feature can include, but is not limited to, performing feature domain enhancement on the initial reconstructed feature corresponding to the luminance component of the current image block based on the feature domain enhancement parameter and the probability distribution parameter to obtain an enhanced reconstructed feature corresponding to the luminance component. The performing image adaptive edge enhancement on the initial reconstructed image block based on the image domain enhancement parameter and the probability distribution parameter to obtain a target reconstructed image block corresponding to the current image block can include, but is not limited to, performing image adaptive edge enhancement on the initial reconstructed image block corresponding to the luminance component of the current image block based on the image domain enhancement parameter and the probability distribution parameter to obtain a target reconstructed image block corresponding to the luminance component; and performing image adaptive edge enhancement on the initial reconstructed image block corresponding to the chroma component of the current image block based on the image domain enhancement parameter and the probability distribution parameter to obtain a target reconstructed image block corresponding to the chroma component.

[0056] For example, the initial reconstructed feature includes a plurality of feature channel maps, and the probability distribution parameter includes a plurality of probability distribution channel maps, the plurality of probability distribution channel maps correspond to the plurality of feature channel maps in a one-to-one manner. Figure 1 Correspondingly, the method further includes: decoding a bitstream corresponding to the current image block to obtain an important channel identifier; selecting a feature channel map corresponding to the important channel identifier from the plurality of feature channel maps as an important feature channel map according to the important channel identifier, and selecting the remaining feature channel maps as non-important feature channel maps; selecting a probability distribution channel map corresponding to the important feature channel map as the important probability distribution channel map, and selecting probability distribution channel maps corresponding to the non-important feature channel maps as the non-important probability distribution channel maps.

[0057] For example, the initial reconstructed feature includes a plurality of feature channel maps, and the probability distribution parameter includes a plurality of probability distribution channel maps, the plurality of probability distribution channel maps correspond to the plurality of feature channel maps in a one-to-one manner. Figure 1 Correspondingly, the method further includes: for each feature channel map, determining a consumed bit number corresponding to the feature channel map based on a feature value in the feature channel map and a probability distribution value in a probability distribution channel map corresponding to the feature channel map; selecting an important feature channel map from the plurality of feature channel maps based on the consumed bit number corresponding to each feature channel map, and selecting the remaining feature channel maps as non-important feature channel maps; selecting a probability distribution channel map corresponding to the important feature channel map as the important probability distribution channel map, and selecting probability distribution channel maps corresponding to the non-important feature channel maps as the non-important probability distribution channel maps.

[0058] For example, the coefficient hyperparameter feature corresponding to the current image block and the probability distribution parameter are obtained by decoding the first code stream related to the current image block; the initial reconstruction feature corresponding to the current image block is obtained by decoding the second code stream related to the current image block; and the important channel identifier is obtained by decoding the third code stream related to the current image block; wherein the first code stream, the second code stream and the third code stream are code streams in which different information is encoded.

[0059] For example, the above execution order is only an example given for the convenience of description, and in actual application, the execution order between steps can also be changed, and the execution order is not limited. Moreover, in other embodiments, the steps of the corresponding method can not be executed in the order shown and described in the specification, and the steps included in the method can be more or less than those described in the specification. In addition, a single step described in the specification can be divided into multiple steps for description in other embodiments; multiple steps described in the specification can also be combined into a single step for description in other embodiments.

[0060] As can be seen from the above technical solutions, in the embodiments of the present application, after obtaining the initial reconstruction feature corresponding to the current image block, the initial reconstruction feature can be enhanced based on the enhancement parameter and the probability distribution parameter to obtain the enhanced reconstruction feature, and the target reconstruction image block corresponding to the current image block is determined based on the enhanced reconstruction feature, thereby proposing an end-to-end video image compression method, which can realize encoding and decoding of video images based on a neural network, and achieve the purpose of improving coding efficiency and decoding efficiency by combining the enhancement parameter and the probability distribution parameter. By combining the network structure design and the header information code stream (such as the third code stream), the neural network can effectively ensure the quality of the reconstructed image block while maintaining low complexity, achieve the purpose of improving coding performance and decoding performance, and reduce complexity. The quality of the feature is enhanced by using the enhancement parameter and the probability distribution parameter, and the feature information is not directly changed at the encoding end, but the enhancement parameter (such as the important channel identifier) is encoded into the header information code stream, and the feature is enhanced by the enhancement parameter at the decoding end to improve the coding performance and improve the quality of the reconstructed image.

[0061] Embodiment 2: In the embodiments of the present application, an encoding method is proposed, as shown in Figure 3 The encoding method can be applied to the encoding end (also referred to as a video encoder), and the method can include steps 301-304.

[0062] In step 301, the coefficient hyperparameter feature corresponding to the current image block is encoded to obtain the first code stream corresponding to the current image block.

[0063] At step 302, a probability distribution parameter is determined based on the coefficient hyperparameter feature, and an initial image feature corresponding to the current image block is encoded based on the probability distribution parameter to obtain a second code stream corresponding to the current image block.

[0064] At step 303, for each candidate enhancement parameter, the initial reconstructed feature is enhanced based on the candidate enhancement parameter and the probability distribution parameter to obtain an enhanced reconstructed feature, a target reconstructed image block is determined based on the enhanced reconstructed feature, and an objective value corresponding to the candidate enhancement parameter is determined based on the target reconstructed image block.

[0065] At step 304, the enhancement parameter corresponding to the current image block is selected from all candidate enhancement parameters based on the objective value corresponding to each candidate enhancement parameter, and the enhancement parameter is encoded to obtain a third code stream corresponding to the current image block. The third code stream can also be referred to as a header information code stream corresponding to the current image block.

[0066] For example, the initial reconstructed feature includes an important feature channel graph and a non-important feature channel graph, and the probability distribution parameter includes an important probability distribution channel graph corresponding to the important feature channel graph and a non-important probability distribution channel graph corresponding to the non-important feature channel graph. Enhancing the initial reconstructed feature based on the candidate enhancement parameter and the probability distribution parameter to obtain the enhanced reconstructed feature can include but is not limited to: performing feature adaptive edge enhancement on the important feature channel graph based on the candidate feature domain enhancement parameter and the important probability distribution channel graph to obtain a first reconstructed feature after feature adaptive edge enhancement; performing feature adaptive stretching on the non-important feature channel graph based on the candidate feature domain enhancement parameter and the non-important probability distribution channel graph to obtain a second reconstructed feature after feature adaptive stretching; and generating the enhanced reconstructed feature based on the first reconstructed feature and the second reconstructed feature.

[0067] For example, determining the target reconstructed image block based on the enhanced reconstructed feature can include but is not limited to: inputting the enhanced reconstructed feature into a synthesis transformation network to obtain an initial reconstructed image block corresponding to the current image block; for each candidate image domain enhancement parameter, performing image adaptive edge enhancement on the initial reconstructed image block based on the candidate image domain enhancement parameter and the probability distribution parameter to obtain a target reconstructed image block corresponding to the current image block. On this basis, for each candidate image domain enhancement parameter, the objective value corresponding to the candidate image domain enhancement parameter can also be determined based on the target reconstructed image block; the image domain enhancement parameter corresponding to the current image block is selected from all candidate image domain enhancement parameters based on the objective value corresponding to each candidate image domain enhancement parameter, and the image domain enhancement parameter is encoded to obtain a third code stream corresponding to the current image block.

[0068] Exemplarily, the processing procedure at the encoding end is similar to the processing procedure at the decoding end, and the same parts are not repeated. The processing procedure at the decoding end can be applied to the encoding end, i.e., the same processing mode is adopted at the encoding end.

[0069] Exemplarily, the above execution sequence is only an example given for the convenience of description, and the execution sequence between steps can be changed in actual application, and the execution sequence is not limited. Moreover, in other embodiments, the steps of the corresponding method can not be executed in the order shown and described in the specification, and the steps included in the method can be more or less than those described in the specification. In addition, a single step described in the specification can be divided into multiple steps for description in other embodiments, and multiple steps described in the specification can also be combined into a single step for description in other embodiments.

[0070] From the above technical solutions, it can be seen that in the embodiments of the present application, after obtaining the initial reconstructed feature corresponding to the current image block, the initial reconstructed feature can be enhanced based on the enhancement parameter and the probability distribution parameter to obtain an enhanced reconstructed feature, and the target reconstructed image block corresponding to the current image block is determined based on the enhanced reconstructed feature, thereby proposing an end-to-end video image compression method, which can realize encoding and decoding of video images based on a neural network, and achieve the purpose of improving coding efficiency and decoding efficiency by combining the enhancement parameter and the probability distribution parameter. Combined with the network structure design and the header information code stream (such as the third code stream), the neural network can effectively ensure the quality of the reconstructed image block while maintaining low complexity, achieve the purpose of improving coding performance and decoding performance, and reduce complexity. The quality of the feature is enhanced by using the enhancement parameter and the probability distribution parameter, the feature information is not directly changed at the encoding end, but the enhancement parameter (such as the important channel identifier) is encoded into the header information code stream, and the feature is enhanced by the enhancement parameter at the decoding end to improve the coding performance and improve the quality of the reconstructed image.

[0071] Embodiment 3: For embodiments 1 and 2, regarding the processing procedure at the encoding end, please refer to Figure 4 of course, Figure 4 which is only an example of the processing procedure at the encoding end, and the processing procedure at the encoding end is not limited.

[0072] After obtaining the current image block x (the current image block x can be the original image block x, i.e., the input image block), the encoding end can analyze and transform the current image block x by using the analysis and transformation network (i.e., the neural network) to obtain the image feature y corresponding to the current image block x. Wherein, the analysis and transformation of the current image block x by the analysis and transformation network means that the current image block x is transformed into the image feature y in the latent domain, so as to facilitate the operation of all subsequent processes in the latent domain.

[0073] The image can be divided into one image block or multiple image blocks. If the image is divided into one image block, the current image block x can also be the image, that is, the coding process for the image block can also be directly used for the image.

[0074] After obtaining the image feature y, the encoding end performs coefficient hyperparameter feature transformation on the image feature y to obtain the coefficient hyperparameter feature z. For example, the image feature y can be input to a hyperparameter encoding network (i.e., a neural network), and the hyperparameter encoding network performs coefficient hyperparameter feature transformation on the image feature y to obtain the coefficient hyperparameter feature z. The hyperparameter encoding network can be a trained neural network, and the training process of the hyperparameter encoding network is not limited, as long as the image feature y can be transformed into the coefficient hyperparameter feature z. After the image feature y in the latent domain is input to the hyperparameter encoding network, the latent information z is obtained.

[0075] After obtaining the coefficient hyperparameter feature z, the encoding end can quantize the coefficient hyperparameter feature z to obtain the hyperparameter quantization feature corresponding to the coefficient hyperparameter feature z, that is, Figure 4 Q in the above formula represents the quantization process. After obtaining the hyperparameter quantization feature corresponding to the coefficient hyperparameter feature z, the hyperparameter quantization feature is encoded to obtain the Bitstream #1 (i.e., the first code stream) corresponding to the current image block, that is, Figure 4 AE in the above formula represents the encoding process, such as an entropy encoding process. Alternatively, the encoding end can directly encode the coefficient hyperparameter feature z to obtain the Bitstream #1 corresponding to the current image block. The hyperparameter quantization feature or the coefficient hyperparameter feature z carried in the Bitstream #1 is mainly used to obtain the parameters of the mean and the probability distribution model.

[0076] After obtaining the Bitstream #1 corresponding to the current image block, the encoding end can send the Bitstream #1 corresponding to the current image block to the decoding end. For the processing process of the decoding end for the Bitstream #1 corresponding to the current image block, see the subsequent embodiments.

[0077] After obtaining the Bitstream #1 corresponding to the current image block, the encoding end can also decode the Bitstream #1 to obtain the hyperparameter quantization feature, that is, Figure 4 AD in the above formula represents the decoding process. Then, the encoding end can dequantize the hyperparameter quantization feature to obtain the coefficient hyperparameter feature z_hat, and the coefficient hyperparameter feature z_hat can be the same as or different from the coefficient hyperparameter feature z. Figure 4The IQ in the figure represents the inverse quantization process. Alternatively, the encoder can decode Bitstream#1 to obtain the coefficient hyper-parameter feature z_hat after obtaining Bitstream#1 corresponding to the current image block, without involving the inverse quantization process of the coefficient hyper-parameter feature z_hat.

[0078] For the encoding process of Bitstream#1, a fixed probability density model encoding method can be used, and for the decoding process of Bitstream#1, a fixed probability density model decoding method can be used, and the encoding and decoding processes are not limited.

[0079] After the encoder obtains the coefficient hyper-parameter feature z_hat, the coefficient hyper-parameter feature z_hat of the current image block and the reconstructed feature y_hat of the previous image block (the determination process of the reconstructed feature y_hat is described in subsequent embodiments) can be used to perform context-based prediction to obtain the prediction value mu (i.e., the mean value mu) corresponding to the current image block. For example, the coefficient hyper-parameter feature z_hat and the reconstructed feature y_hat are input into the mean value prediction network, and the mean value prediction network determines the prediction value mu based on the coefficient hyper-parameter feature z_hat and the reconstructed feature y_hat. The prediction process is not limited. For the context-based prediction process, the input includes the coefficient hyper-parameter feature z_hat and the decoded reconstructed feature y_hat, and the two are jointly input to obtain a more accurate prediction value mu. The prediction value mu is used to obtain the residual feature r by subtracting the original feature and to obtain the reconstructed feature y_hat by adding the decoded residual feature r_hat.

[0080] It should be noted that the mean value prediction network is an optional neural network, i.e., there can be no mean value prediction network. That is, the prediction value mu does not need to be determined by the mean value prediction network, Figure 4 The dashed box in the figure indicates that the mean value prediction network is optional.

[0081] After the encoder obtains the image feature y, the image feature y and the prediction value mu can be used to determine the residual feature r, such as the difference between the image feature y and the prediction value mu as the residual feature r. Then, the residual feature r is processed to obtain the image feature s, and the feature processing process is not limited and can be any feature processing method. In this case, the mean value prediction network needs to be deployed to provide the prediction value mu. Alternatively, after the encoder obtains the image feature y, the image feature y can be processed to obtain the image feature s, and the feature processing process is not limited and can be any feature processing method. In this case, the mean value prediction network is not required, and the residual process is an optional process as indicated by the dashed box.

[0082] After obtaining the image feature s, the encoding end can quantize the image feature s to obtain an image quantized feature corresponding to the image feature s, i.e. Figure 4 Q in the formula (1) represents a quantization process. After obtaining the image quantized feature corresponding to the image feature s, the encoding end can encode the image quantized feature to obtain a Bitstream#2 (i.e., a second code stream) corresponding to the current image block, i.e. Figure 4 AE in the formula (2) represents an encoding process, such as an entropy encoding process. Alternatively, the encoding end can directly encode the image feature s to obtain the Bitstream#2 corresponding to the current image block without involving the quantization process of the image feature s.

[0083] After obtaining the Bitstream#2 corresponding to the current image block, the encoding end can send the Bitstream#2 corresponding to the current image block to the decoding end. For the processing process of the decoding end for the Bitstream#2 corresponding to the current image block, refer to subsequent embodiments.

[0084] After obtaining the Bitstream#2 corresponding to the current image block, the encoding end can also decode the Bitstream#2 to obtain an image quantized feature, i.e. Figure 4 AD in the formula (3) represents a decoding process. Then, the encoding end can dequantize the image quantized feature to obtain an image feature s', which can be the same as or different from the image feature s, Figure 4 IQ in the formula (4) represents a dequantization process. Alternatively, after obtaining the Bitstream#2 corresponding to the current image block, the encoding end can also decode the Bitstream#2 to obtain the image feature s' without involving the dequantization process of the image quantized feature.

[0085] After obtaining the image feature s', the encoding end can perform feature restoration (i.e., an inverse process of feature processing) on the image feature s' to obtain a residual feature r_hat. The feature restoration process is not limited and can be any feature restoration manner. The residual feature r_hat can be the same as or different from the residual feature r. After obtaining the residual feature r_hat, the encoding end determines an image feature y_hat (i.e., a reconstructed feature) based on the residual feature r_hat and a prediction value mu, such as taking a sum of the residual feature r_hat and the prediction value mu as the image feature y_hat. The image feature y_hat can be the same as or different from the image feature y. In this case, a mean prediction network needs to be deployed to provide the prediction value mu. Alternatively, after obtaining the image feature s', the encoding end performs feature restoration (i.e., an inverse process of feature processing) on the image feature s' to obtain an image feature y_hat. The image feature y_hat can be the same as or different from the image feature y. In this case, the mean prediction network does not need to be deployed, and the residual process is optional, which is represented by a dashed box.

[0086] After obtaining the image feature y_hat, the encoding end can perform a synthesis transformation on the image feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the image feature y_hat is input to a synthesis transformation network, and the synthesis transformation network performs a synthesis transformation on the image feature y_hat to obtain the reconstructed image block x_hat. At this point, the image reconstruction process is completed.

[0087] In a possible implementation, when the encoding end encodes the image quantization feature or the image feature s to obtain the Bitstream#2 corresponding to the current image block, the encoding end needs to first determine the probability distribution model, and then encodes the image quantization feature or the image feature s based on the probability distribution model. In addition, when the encoding end decodes the Bitstream#2, the encoding end also needs to first determine the probability distribution model, and then decodes the Bitstream#2 based on the probability distribution model.

[0088] To obtain the probability distribution model, continue to refer to FIG. 6. Figure 5 As shown in FIG. 6, after obtaining the coefficient hyperparameter feature z_hat, the encoding end can perform a coefficient hyperparameter feature inverse transformation on the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter. For example, the coefficient hyperparameter feature z_hat is input to a probability hyperparameter decoding network, and the probability hyperparameter decoding network performs a coefficient hyperparameter feature inverse transformation on the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter sigma. After obtaining the probability distribution parameter, the probability distribution model can be generated based on the probability distribution parameter. The probability hyperparameter decoding network can be a trained neural network, and the training process of the probability hyperparameter decoding network is not limited. The probability hyperparameter decoding network can only perform a coefficient hyperparameter feature inverse transformation on the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter.

[0089] In a possible implementation, the processing process of the encoding end described above can be performed by a deep learning model or a neural network model, so as to realize an end-to-end image compression and encoding process, and the implementation is not limited in this regard.

[0090] For the processing process of the decoding end of embodiments 1 and 2, refer to FIG. 8. Figure 5 Of course, Figure 5 The processing process of the decoding end is only an example, and the processing process of the decoding end is not limited in this regard.

[0091] After obtaining the Bitstream#1 corresponding to the current image block, the decoding end can also decode the Bitstream#1 to obtain the hyperparameter quantization feature, that is, Figure 5AD in the figure represents a decoding process. Then, the decoding end can perform inverse quantization on the hyper-quantized feature to obtain the coefficient hyper-parameter feature z_hat, which can be the same as or different from the coefficient hyper-parameter feature z, Figure 5 IQ in the figure represents an inverse quantization process. Alternatively, the decoding end can decode Bitstream#1 after obtaining the Bitstream#1 corresponding to the current image block, to obtain the coefficient hyper-parameter feature z_hat, without involving the inverse quantization process of the coefficient hyper-parameter feature z_hat.

[0092] The decoding process of Bitstream#1 can adopt a decoding method of a fixed probability density model, which is not limited.

[0093] The image can be divided into one image block or multiple image blocks. If the image is divided into one image block, the current image block x can also be the image, that is, the decoding process of the image block can also be directly used for the image.

[0094] After the decoding end obtains the coefficient hyper-parameter feature z_hat, the decoding end can perform context-based prediction based on the coefficient hyper-parameter feature z_hat of the current image block and the reconstructed feature y_hat of the previous image block (the determination process of the reconstructed feature y_hat is described in subsequent embodiments), to obtain the prediction value mu (i.e., the mean value mu) corresponding to the current image block. For example, the coefficient hyper-parameter feature z_hat and the reconstructed feature y_hat are input into a mean value prediction network, and the mean value prediction network determines the prediction value mu based on the coefficient hyper-parameter feature z_hat and the reconstructed feature y_hat, which is not limited to the prediction process. For the context-based prediction process, the input includes the coefficient hyper-parameter feature z_hat and the decoded reconstructed feature y_hat, and the two are jointly input to obtain a more accurate prediction value mu.

[0095] It should be noted that the mean value prediction network is an optional neural network, that is, there can be no mean value prediction network. That is, the prediction value mu does not need to be determined by the mean value prediction network, Figure 5 The dashed box in the figure indicates that the mean value prediction network is optional.

[0096] After the decoding end obtains the Bitstream#2 corresponding to the current image block, the decoding end can decode Bitstream#2 to obtain the image quantized feature, that is, Figure 5 AD in the figure represents a decoding process. Then, the decoding end can perform inverse quantization on the hyper-quantized feature to obtain the coefficient hyper-parameter feature z_hat, which can be the same as or different from the coefficient hyper-parameter feature z, Figure 5In the equation, IQ represents the inverse quantization process. Alternatively, the decoding end can decode Bitstream#2 to obtain the image feature s’ after obtaining Bitstream#2 corresponding to the current image block, without involving the inverse quantization process of the image quantized feature.

[0097] After the decoding end obtains the image feature s’, the decoding end can perform feature restoration (i.e., the inverse process of feature processing) on the image feature s’ to obtain the residual feature r_hat, which can be the same as or different from the residual feature r. After the decoding end obtains the residual feature r_hat, the decoding end determines the image feature y_hat (i.e., the reconstructed feature) based on the residual feature r_hat and the prediction value mu, such as taking the sum of the residual feature r_hat and the prediction value mu as the image feature y_hat, which can be the same as or different from the image feature y. In this case, the mean prediction network needs to be deployed to provide the prediction value mu. Alternatively, after the decoding end obtains the image feature s’, the decoding end performs feature restoration on the image feature s’ to obtain the image feature y_hat, which can be the same as or different from the image feature y. In this case, the mean prediction network does not need to be deployed, and the residual process is indicated by the dashed box as an optional process.

[0098] After the decoding end obtains the image feature y_hat, the decoding end can perform synthesis transformation on the image feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the image feature y_hat is input into the synthesis transformation network, and the synthesis transformation network performs synthesis transformation on the image feature y_hat to obtain the reconstructed image block x_hat. At this point, the image reconstruction process is completed.

[0099] In a possible implementation, when the decoding end decodes Bitstream#2, the decoding end needs to first determine the probability distribution model and then decode Bitstream#2 based on the probability distribution model. To obtain the probability distribution model, continue to refer to FIG. 6. Figure 6A As shown in FIG. 6, after the decoding end obtains the coefficient hyperparameter feature z_hat, the decoding end can perform coefficient hyperparameter feature inverse transformation on the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter sigma. For example, the coefficient hyperparameter feature z_hat is input into the probability hyperparameter decoding network, and the probability hyperparameter decoding network performs coefficient hyperparameter feature inverse transformation on the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter. After obtaining the probability distribution parameter, the probability distribution model can be generated based on the probability distribution parameter. The probability hyperparameter decoding network can be a trained neural network, and the training process of the probability hyperparameter decoding network is not limited in this regard. The probability hyperparameter decoding network can only perform coefficient hyperparameter feature inverse transformation on the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter.

[0100] In a possible implementation, the decoding-end processing process described above can be performed by a deep learning model or a neural network model, so as to realize an end-to-end image decoding process, which is not limited.

[0101] In embodiment 5, based on embodiment 3 and embodiment 4, a feature domain enhancement module can be added before the synthesis transformation network, as shown in Figure 6B The input feature of the feature domain enhancement module can be the image feature y_hat (hereinafter referred to as the initial reconstruction feature y_hat), and the output feature of the feature domain enhancement module can be the enhanced reconstruction feature y_hat_enhanced. The enhanced reconstruction feature y_hat_enhanced is input to the synthesis transformation network, and the synthesis transformation network performs synthesis transformation on the enhanced reconstruction feature y_hat_enhanced to obtain the target reconstruction image block x_hat_enhanced corresponding to the current image block x.

[0102] For the encoding end: after obtaining the current image block x, the current image block x is analyzed and transformed by the analysis transformation network to obtain the image feature y corresponding to the current image block x. The image feature y is transformed into the coefficient hyperparameter feature z by the hyperparameter encoding network. The coefficient hyperparameter feature corresponding to the current image block (which can be the coefficient hyperparameter feature z itself or the hyperparameter quantization feature of the coefficient hyperparameter feature z) is encoded to obtain the first code stream corresponding to the current image block.

[0103] The first code stream corresponding to the current image block is decoded to obtain the coefficient hyperparameter feature z_hat corresponding to the current image block (for example, the coefficient hyperparameter feature z_hat itself is decoded from the first code stream, or the hyperparameter quantization feature is decoded from the first code stream, and the coefficient hyperparameter feature z_hat is obtained by inverse quantization of the hyperparameter quantization feature). Then, the coefficient hyperparameter feature z_hat is inversely transformed into the probability distribution parameter sigma by the probability hyperparameter decoding network.

[0104] The initial image feature corresponding to the current image block is encoded based on the probability distribution parameter sigma to obtain the second code stream corresponding to the current image block. The initial image feature can be the image feature y, the residual feature r corresponding to the image feature y, the image feature s obtained by processing the image feature y or the residual feature r, or the image quantization feature corresponding to the image feature y, the residual feature r, or the image feature s, which is not limited.

[0105] The second code stream corresponding to the current image block is decoded based on the probability distribution parameter sigma to obtain the initial reconstructed feature y_hat corresponding to the current image block. For example, if the initial image feature is the image feature y, the initial reconstructed feature y_hat is decoded from the second code stream. If the initial image feature is the image quantized feature corresponding to the image feature y, the image quantized feature is decoded from the second code stream, and the initial reconstructed feature y_hat is obtained by dequantizing the image quantized feature. For another example, if the initial image feature is the residual feature r corresponding to the image feature y, the residual feature r_hat is decoded from the second code stream, and the initial reconstructed feature y_hat is determined based on the residual feature r_hat and the prediction value mu. If the initial image feature is the image quantized feature corresponding to the residual feature r, the image quantized feature is decoded from the second code stream, and the initial reconstructed feature y_hat is obtained by dequantizing the image quantized feature based on the residual feature r_hat and the prediction value mu. If the initial image feature is the image feature s corresponding to the image feature y or the residual feature r, the image feature s' is decoded from the second code stream, the initial reconstructed feature y_hat or the residual feature r_hat is obtained by performing feature restoration on the image feature s', and if the residual feature r_hat is obtained, the initial reconstructed feature y_hat can also be determined based on the residual feature r_hat and the prediction value mu. If the initial image feature is the image quantized feature corresponding to the image feature s, the image quantized feature is decoded from the second code stream, and the image feature s' is obtained by dequantizing the image quantized feature, the initial reconstructed feature y_hat or the residual feature r_hat is obtained by performing feature restoration on the image feature s', and if the residual feature r_hat is obtained, the initial reconstructed feature y_hat can also be determined based on the residual feature r_hat and the prediction value mu.

[0106] After obtaining the initial reconstructed feature y_hat, the initial reconstructed feature y_hat can be input to the feature domain enhancement module, the feature domain enhancement module performs feature domain enhancement on the initial reconstructed feature y_hat to obtain the enhanced reconstructed feature y_hat_enhanced, and the enhanced reconstructed feature y_hat_enhanced is input to the synthesis transformation network, the synthesis transformation network performs synthesis transformation on the enhanced reconstructed feature y_hat_enhanced to obtain the target reconstructed image block x_hat_enhanced corresponding to the current image block x.

[0107] For the decoding end, the first code stream corresponding to the current image block is decoded to obtain the coefficient hyperparameter feature z_hat corresponding to the current image block (for example, the coefficient hyperparameter feature z_hat is decoded from the first code stream itself, or the hyperparameter quantization feature is decoded from the first code stream, and the coefficient hyperparameter feature z_hat is obtained by inverse quantization on the hyperparameter quantization feature). Then, the coefficient hyperparameter feature z_hat is subjected to coefficient hyperparameter feature inverse transformation through the probability hyperparameter decoding network to obtain the probability distribution parameter sigma.

[0108] The second code stream corresponding to the current image block is decoded based on the probability distribution parameter sigma to obtain the initial reconstruction feature y_hat corresponding to the current image block. For example, the initial reconstruction feature y_hat is decoded from the second code stream. Or, the image quantization feature is decoded from the second code stream, and the initial reconstruction feature y_hat is obtained by inverse quantization on the image quantization feature. For another example, the residual feature r_hat is decoded from the second code stream, and the initial reconstruction feature y_hat is determined based on the residual feature r_hat and the prediction value mu. Or, the image quantization feature is decoded from the second code stream, and the residual feature r_hat is obtained by inverse quantization on the image quantization feature, and the initial reconstruction feature y_hat is determined based on the residual feature r_hat and the prediction value mu. For another example, the image feature s' is decoded from the second code stream, the initial reconstruction feature y_hat or the residual feature r_hat is obtained by feature recovery on the image feature s', and if the residual feature r_hat is obtained, the initial reconstruction feature y_hat can also be determined based on the residual feature r_hat and the prediction value mu. Or, the image quantization feature is decoded from the second code stream, and the image feature s' is obtained by inverse quantization on the image quantization feature, the initial reconstruction feature y_hat or the residual feature r_hat is obtained by feature recovery on the image feature s', and if the residual feature r_hat is obtained, the initial reconstruction feature y_hat can also be determined based on the residual feature r_hat and the prediction value mu.

[0109] After obtaining the initial reconstruction feature y_hat, the initial reconstruction feature y_hat can be input to the feature domain enhancement module, the feature domain enhancement module performs feature domain enhancement on the initial reconstruction feature y_hat to obtain the enhanced reconstruction feature y_hat_enhanced, and the enhanced reconstruction feature y_hat_enhanced is input to the synthesis transformation network, and the synthesis transformation network performs synthesis transformation on the enhanced reconstruction feature y_hat_enhanced to obtain the target reconstruction image block x_hat_enhanced corresponding to the current image block x.

[0110] Embodiment 6: On the basis of Embodiment 3 and Embodiment 4, a feature domain enhancement module can be added before the synthesis transformation network, and an image domain enhancement module can be added after the synthesis transformation network, see Figure 6CThe input feature of the feature domain enhancement module is the initial reconstruction feature y_hat, and the output feature of the feature domain enhancement module can be an enhanced reconstruction feature y_hat_enhanced. The enhanced reconstruction feature y_hat_enhanced can be input to the synthesis transformation network, and the synthesis transformation network is used to perform synthesis transformation on the enhanced reconstruction feature y_hat_enhanced to obtain the initial reconstruction image block x_hat corresponding to the current image block x. The input feature of the image domain enhancement module can be the initial reconstruction image block x_hat, and the output feature of the image domain enhancement module can be the target reconstruction image block x_hat_enhanced. That is, the image domain enhancement module can perform image domain enhancement on the initial reconstruction image block x_hat to obtain the target reconstruction image block x_hat_enhanced corresponding to the current image block x.

[0111] For the encoding end: the analysis transformation network is used to perform analysis transformation on the current image block x to obtain the image feature y corresponding to the current image block x. The coefficient hyperparameter feature transformation is performed on the image feature y by the hyperparameter encoding network to obtain the coefficient hyperparameter feature z. The coefficient hyperparameter feature corresponding to the current image block is encoded to obtain the first code stream corresponding to the current image block. The coefficient hyperparameter feature z_hat corresponding to the current image block is obtained by decoding the first code stream corresponding to the current image block, and the coefficient hyperparameter feature z_hat is inversely transformed by the probability hyperparameter decoding network to obtain the probability distribution parameter sigma. The initial image feature corresponding to the current image block is encoded based on the probability distribution parameter sigma to obtain the second code stream corresponding to the current image block. The second code stream corresponding to the current image block is decoded based on the probability distribution parameter sigma to obtain the initial reconstruction feature y_hat corresponding to the current image block. The above process can be referred to in Embodiment 5 and will not be repeated here.

[0112] After obtaining the initial reconstruction feature y_hat, the initial reconstruction feature y_hat can be input to the feature domain enhancement module, the feature domain enhancement module is used to perform feature domain enhancement on the initial reconstruction feature y_hat to obtain the enhanced reconstruction feature y_hat_enhanced, and the enhanced reconstruction feature y_hat_enhanced is input to the synthesis transformation network, and the synthesis transformation network is used to perform synthesis transformation on the enhanced reconstruction feature y_hat_enhanced to obtain the initial reconstruction image block x_hat corresponding to the current image block x.

[0113] After obtaining the initial reconstruction image block x_hat corresponding to the current image block x, the image domain enhancement module can perform image domain enhancement on the initial reconstruction image block x_hat to obtain the target reconstruction image block x_hat_enhanced corresponding to the current image block x.

[0114] For the decoding end: the first bitstream corresponding to the current image block is decoded to obtain the coefficient hyperparameter feature z_hat corresponding to the current image block. The coefficient hyperparameter feature z_hat is then subjected to inverse transformation using a probabilistic hyperparameter decoding network to obtain the probability distribution parameter sigma. Based on the probability distribution parameter sigma, the second bitstream corresponding to the current image block is decoded to obtain the initial reconstructed feature y_hat corresponding to the current image block. The above process can be found in Example 5 and will not be repeated here.

[0115] After obtaining the initial reconstructed feature y_hat, the initial reconstructed feature y_hat can be input into the feature domain enhancement module, which performs feature domain enhancement on the initial reconstructed feature y_hat to obtain the enhanced reconstructed feature y_hat_enhanced. The enhanced reconstructed feature y_hat_enhanced is then input into the synthesis transformation network, which performs a synthesis transformation on the enhanced reconstructed feature y_hat_enhanced to obtain the initial reconstructed image block x_hat corresponding to the current image block x.

[0116] After obtaining the initial reconstructed image block x_hat corresponding to the current image block x, the image domain enhancement module can perform image domain enhancement on the initial reconstructed image block x_hat to obtain the target reconstructed image block x_hat_enhanced corresponding to the current image block x.

[0117] Example 7: Based on Examples 3 and 4, an image domain enhancement module can be added after the synthetic transform network. See [link to example 7]. Figure 1 As shown, the input feature of the image domain enhancement module is the initial reconstructed image block x_hat, and the output feature of the image domain enhancement module is the target reconstructed image block x_hat_enhanced. In other words, the image domain enhancement module can perform image domain enhancement on the initial reconstructed image block x_hat to obtain the target reconstructed image block x_hat_enhanced corresponding to the current image block x.

[0118] For the encoding end: the current image block x is analyzed and transformed by the analysis transformation network to obtain the image feature y corresponding to the current image block x. The image feature y is subjected to coefficient hyperparameter feature transformation by the hyperparameter encoding network to obtain the coefficient hyperparameter feature z. The coefficient hyperparameter feature corresponding to the current image block is encoded to obtain the first code stream corresponding to the current image block. The coefficient hyperparameter feature z_hat corresponding to the current image block is obtained by decoding the first code stream corresponding to the current image block, and the coefficient hyperparameter feature z_hat is subjected to coefficient hyperparameter feature inverse transformation by the probability hyperparameter decoding network to obtain the probability distribution parameter sigma. The initial image feature corresponding to the current image block is encoded based on the probability distribution parameter sigma to obtain the second code stream corresponding to the current image block. The initial reconstruction feature y_hat corresponding to the current image block is obtained by decoding the second code stream corresponding to the current image block based on the probability distribution parameter sigma. The above process can be referred to in Embodiment 5 and will not be repeated here.

[0119] After obtaining the initial reconstruction feature y_hat, the initial reconstruction feature y_hat can be input to the synthesis transformation network, and the initial reconstruction feature y_hat is subjected to synthesis transformation by the synthesis transformation network to obtain the initial reconstruction image block x_hat corresponding to the current image block x. After obtaining the initial reconstruction image block x_hat corresponding to the current image block x, the image domain enhancement module can perform image domain enhancement on the initial reconstruction image block x_hat to obtain the target reconstruction image block x_hat_enhanced corresponding to the current image block x.

[0120] For the decoding end: the coefficient hyperparameter feature z_hat corresponding to the current image block is obtained by decoding the first code stream corresponding to the current image block, and the coefficient hyperparameter feature z_hat is subjected to coefficient hyperparameter feature inverse transformation by the probability hyperparameter decoding network to obtain the probability distribution parameter sigma. The initial reconstruction feature y_hat corresponding to the current image block is obtained by decoding the second code stream corresponding to the current image block based on the probability distribution parameter sigma. The above process can be referred to in Embodiment 5 and will not be repeated here.

[0121] After obtaining the initial reconstruction feature y_hat, the initial reconstruction feature y_hat can be input to the synthesis transformation network, and the initial reconstruction feature y_hat is subjected to synthesis transformation by the synthesis transformation network to obtain the initial reconstruction image block x_hat corresponding to the current image block x. After obtaining the initial reconstruction image block x_hat corresponding to the current image block x, the image domain enhancement module can perform image domain enhancement on the initial reconstruction image block x_hat to obtain the target reconstruction image block x_hat_enhanced corresponding to the current image block x.

[0122] In Embodiment 8, in Embodiment 5 and Embodiment 6, the initial reconstructed feature y_hat is input to the feature domain enhancement module, and the initial reconstructed feature y_hat is subjected to feature domain enhancement by the feature domain enhancement module to obtain an enhanced reconstructed feature y_hat_enhanced. For example, the initial reconstructed feature y_hat can be subjected to feature domain enhancement based on the feature domain enhancement parameter and the probability distribution parameter sigma corresponding to the current image block to obtain the enhanced reconstructed feature y_hat_enhanced. The feature domain enhancement can include feature adaptive edge enhancement and feature adaptive stretching, and the probability distribution parameter sigma is used for auxiliary calculation to determine the channels that need to be subjected to feature adaptive edge enhancement and feature adaptive stretching, and to determine the stretching strength. The process is described below.

[0123] The initial reconstructed feature y_hat is a three-dimensional tensor of C L x H L x W L , C L is the number of channels of the initial reconstructed feature y_hat (i.e., the number of channels of the feature domain or the number of channels of the Latent domain), H L is the feature height of the initial reconstructed feature y_hat (i.e., the feature height of the feature domain or the feature height of the Latent domain), and W L is the feature width of the initial reconstructed feature y_hat (i.e., the feature width of the feature domain or the feature width of the Latent domain). For each channel of the initial reconstructed feature y_hat, a channel ch is taken as an example for description, ch can be any value in [1, 2, …, C L ], and the two-dimensional tensor y_hat_ch of shape H L x W L corresponding to the channel ch is called a feature channel map. Therefore, the initial reconstructed feature y_hat can include C L feature channel maps.

[0124] The probability distribution parameter sigma is also a three-dimensional tensor of C L x H L x W L , C L is the number of channels of the probability distribution parameter sigma, H L is the feature height of the probability distribution parameter sigma, and W L is the feature width of the probability distribution parameter sigma. For each channel of the probability distribution parameter sigma, a channel ch is taken as an example for description, ch can be any value in [1, 2, …, C L ], and the two-dimensional tensor y_hat_ch of shape H L x W LThe two-dimensional tensor sigma_ch of the probability distribution parameter sigma is called a probability distribution channel map. Therefore, the probability distribution parameter sigma can include C L probability distribution channel maps, and the C L probability distribution channel maps correspond to the C L feature channel maps. Figure 1 A correspondence exists, such as the first probability distribution channel map corresponding to the first feature channel map, the second probability distribution channel map corresponding to the second feature channel map, and so on, and the C L probability distribution channel map corresponding to the C L feature channel map.

[0125] For the C L feature channel maps, the C L feature channel maps can be divided into an important feature channel map and non-important feature channel maps. The important feature channel map can be at least one, and the non-important feature channel maps can be multiple. For example, at least one feature channel map can be selected from the C L feature channel maps as the important feature channel map, and the remaining feature channel maps are selected as the non-important feature channel maps. Since the C L probability distribution channel maps correspond to the C L feature channel Figure 1 maps, the probability distribution channel map corresponding to the important feature channel map can be selected as the important probability distribution channel map, and the probability distribution channel maps corresponding to the non-important feature channel maps can be selected as the non-important probability distribution channel maps. For example, if the first feature channel map is the important feature channel map, the first probability distribution channel map corresponding to the first feature channel map is selected as the important probability distribution channel map.

[0126] In summary, the initial reconstructed feature y_hat can include the important feature channel map and the non-important feature channel maps, and the probability distribution parameter sigma can include the important probability distribution channel map corresponding to the important feature channel map and the non-important probability distribution channel maps corresponding to the non-important feature channel maps. For the important feature channel map, feature adaptive edge enhancement can be performed on the important feature channel map based on the feature domain enhancement parameter and the important probability distribution channel map to obtain a first reconstructed feature after feature adaptive edge enhancement. For the non-important feature channel map, feature adaptive stretching can be performed on the non-important feature channel map based on the feature domain enhancement parameter and the non-important probability distribution channel map to obtain a second reconstructed feature after feature adaptive stretching. On this basis, an enhanced reconstructed feature y_hat_enhanced can be generated based on the first reconstructed feature and the second reconstructed feature.

[0127] For example, it is assumed that the feature channel Figure 2 is the important feature channel map, and the feature channel Figure 4 to the feature channel Figure 1The feature channel graph is a non-important feature channel graph, and the probability distribution channel is a probability distribution channel Figure 2 The feature channel graph is an important probability distribution channel graph, and the probability distribution channel is a probability distribution channel Figure 4 to the probability distribution channel Figure 1 The feature channel graph is a non-important probability distribution channel graph, and then: the feature domain enhancement parameter and the probability distribution channel can be used for Figure 1 The feature channel is subjected to feature adaptive edge enhancement Figure 2 to obtain the first reconstructed feature 1 (i.e., the reconstructed feature of the first channel) after feature adaptive edge enhancement. The feature domain enhancement parameter and the probability distribution channel can be used for Figure 2 The feature channel is subjected to feature adaptive stretching Figure 3 to obtain the second reconstructed feature 2 (i.e., the reconstructed feature of the second channel) after feature adaptive stretching. The feature domain enhancement parameter and the probability distribution channel can be used for Figure 3 The feature channel is subjected to feature adaptive stretching Figure 4 to obtain the second reconstructed feature 3 (i.e., the reconstructed feature of the third channel) after feature adaptive stretching. The feature domain enhancement parameter and the probability distribution channel can be used for Figure 4 The feature channel is subjected to feature adaptive stretching Edge enhancement threshold interval to obtain the second reconstructed feature 4 (i.e., the reconstructed feature of the fourth channel) after feature adaptive stretching. Then, the first reconstructed feature 1, the second reconstructed feature 2, the second reconstructed feature 3, and the second reconstructed feature 4 can be combined (e.g., spliced) according to the channel, that is, the reconstructed features of the four channels are combined to obtain the enhanced reconstructed feature y_hat_enhanced.

[0128] In embodiment 8, C L feature channel graphs need to be divided into important feature channel graphs and non-important feature channel graphs. For example, the C L feature channel graphs are divided into important feature channel graphs and non-important feature channel graphs in the following manner.

[0129] Manner 1: The encoding end encodes the important channel identifier in the third code stream corresponding to the current image block, and the decoding end decodes the important channel identifier from the third code stream corresponding to the current image block.

[0130] For example, the encoding end divides the C L feature channel graphs into important feature channel graphs and non-important feature channel graphs, and does not limit this process, and encodes the important channel identifier (also referred to as the important channel number, denoted as important_channel) in the third code stream corresponding to the current image block. After receiving the third code stream, the decoding end decodes the important channel identifier from the third code stream, and then divides the C LThe important feature channel graph is selected from the C feature channel graphs based on the feature values in each feature channel graph and the probability distribution values in each probability distribution channel graph, and the remaining feature channel graphs are selected as non-important feature channel graphs.

[0131] wherein the encoding end selects the C L When the C feature channel graphs are divided into important feature channel graphs and non-important feature channel graphs, the important feature channel graph can be selected from the C L feature channel graphs based on the feature values in each feature channel graph and the probability distribution values in each probability distribution channel graph, and the remaining feature channel graphs are selected as non-important feature channel graphs.

[0132] In mode 2, the important feature channel graph is selected from the C L feature channel graphs based on the feature values in each feature channel graph and the probability distribution values in each probability distribution channel graph.

[0133] For example, for each feature channel graph, the number of consumed bits corresponding to the feature channel graph can be determined based on the feature values in the feature channel graph and the probability distribution values in the probability distribution channel graph corresponding to the feature channel graph. For example, the number of consumed bits corresponding to the feature channel graph can be determined by using the following expression, of course, the following expression is only an example.

[0134]

[0135] In the above expression, bits_per_ch is used to represent the number of consumed bits corresponding to the feature channel graph ch, y_hat_ch(i,j) is used to represent the feature value of the feature point (i,j) in the feature channel graph ch, sigma_ch(i,j) is used to represent the probability distribution value of the feature point (i,j) in the probability distribution channel graph ch, and Φ(.) is a standard normal distribution function.

[0136] After the above processing is performed on each feature channel graph, the number of consumed bits corresponding to each feature channel graph can be obtained. Based on the number of consumed bits corresponding to each feature channel graph, the important feature channel graph can be selected from the C L feature channel graphs. For example, the feature channel graph with the largest number of consumed bits is selected as the important feature channel graph, or the K feature channel graphs with large numbers of consumed bits are selected as the important feature channel graphs, and K can be a positive integer greater than 1.

[0137] After the important feature channel graph is selected from the C LAfter selecting the important feature channel graph from the C feature channel graphs, the remaining feature channel graphs can be selected as non-important feature channel graphs, the probability distribution channel graph corresponding to the important feature channel graph can be selected as an important probability distribution channel graph, and the probability distribution channel graph corresponding to the non-important feature channel graph can be selected as a non-important probability distribution channel graph.

[0138] Method 3, for the encoding end and the decoding end, based on the consumed code rate corresponding to each feature channel graph, from the C L feature channel graphs, select the important feature channel graph. For example, based on the consumed code rate corresponding to each feature channel graph, the C L feature channel graphs can be sorted in descending order of consumed code rate, and the top K feature channel graphs in the order can be selected as the important feature channel graphs. Alternatively, based on the consumed code rate corresponding to each feature channel graph, the C L feature channel graphs can be sorted in ascending order of consumed code rate, and the last K feature channel graphs in the order can be selected as the important feature channel graphs.

[0139] After selecting the important feature channel graph from the C L feature channel graphs, the remaining feature channel graphs can be selected as non-important feature channel graphs, the probability distribution channel graph corresponding to the important feature channel graph can be selected as an important probability distribution channel graph, and the probability distribution channel graph corresponding to the non-important feature channel graph can be selected as a non-important probability distribution channel graph.

[0140] Method 4, for the encoding end and the decoding end, the default feature channel graph (i.e., the fixed feature channel graph) is taken as the important feature channel graph. For example, the first feature channel graph is pre-agreed to be the important feature channel graph, or the sixth feature channel graph is pre-agreed to be the important feature channel graph, or the tenth feature channel graph is pre-agreed to be the important feature channel graph. Of course, any feature channel graph can be taken as the important feature channel graph, which is not limited.

[0141] Embodiment 10: In embodiment 8, the important feature channel graph is subjected to feature adaptive edge enhancement based on the feature domain enhancement parameter and the important probability distribution channel graph, to obtain the first reconstructed feature. The feature adaptive edge enhancement process is described below.

[0142] For example, the encoding end encodes the plurality of edge enhancement segmentation intensity values and the plurality of edge enhancement segmentation threshold values in the third code stream corresponding to the current image block, the decoding end decodes the plurality of edge enhancement segmentation intensity values and the plurality of edge enhancement segmentation threshold values from the third code stream corresponding to the current image block, and the plurality of edge enhancement segmentation intensity values and the plurality of edge enhancement segmentation threshold values are taken as the feature domain enhancement parameter.

[0143] The number of edge enhancement segment intensity values and the number of edge enhancement segment threshold values can be the same or different. Taking the same number as an example, the encoder can also encode the number of edge enhancement segment intensity values (or edge enhancement segment threshold values) in the third code stream corresponding to the current image block, and the decoder decodes the number of edge enhancement segment intensity values from the third code stream corresponding to the current image block. Taking the number of edge enhancement segment intensity values as n as an example, n edge enhancement segment intensity values can be recorded as magl-1, magl-2, …, magl-n, and n edge enhancement segment threshold values can be recorded as thrl-1, thrl-2, …, thrl-n.

[0144] For example, the plurality of edge enhancement segment threshold values can constitute a plurality of edge enhancement threshold intervals, and the plurality of edge enhancement threshold intervals correspond one-to-one to the plurality of edge enhancement segment intensity values. See Table 1 for an example of the correspondence.

[0145] Table 1

[0146] Edge enhancement segment intensity value Less than thrl-1 1 (i.e. no edge enhancement) [thrl-1, thrl-2) magl-1 [thrl-2, thrl-3) magl-2 [thrl-3, thrl-4) magl-3 [thrl-(n-1), thrl-n) … … magl-(n-1) Greater than or equal to thrl-n magl-n Image enhancement threshold interval

[0147] For example, the important probability distribution channel map can include a plurality of probability distribution values. For each probability distribution value, the edge enhancement threshold interval corresponding to the probability distribution value is first determined, and the edge enhancement segment intensity value corresponding to the probability distribution value is determined based on the edge enhancement threshold interval. For example, if the probability distribution value is located in [thrl-2, thrl-3), the edge enhancement segment intensity value corresponding to the probability distribution value is magl-2, and if the probability distribution value is located in [thrl-3, thrl-4), the edge enhancement segment intensity value corresponding to the probability distribution value is magl-3, and so on. After obtaining the edge enhancement segment intensity value corresponding to each probability distribution value, the important feature channel map can be subjected to feature adaptive edge enhancement based on the edge enhancement segment intensity value corresponding to each probability distribution value to obtain the first reconstructed feature after feature adaptive edge enhancement.

[0148] For example, the important feature channel map can be subjected to feature adaptive edge enhancement by using the following steps S11-S15.

[0149] In step S11, the important feature channel map is normalized to obtain a normalized feature map.

[0150] Exemplarily, the normalized feature map can be obtained based on the important feature channel map, the mean feature corresponding to the important feature channel map, and the variance feature corresponding to the important feature channel map. For example, the important feature channel map can be subtracted by the mean feature, and then divided by the variance feature to obtain the normalized feature map. For another example, the important feature channel map can be subtracted by the mean feature, and then divided by the variance feature to obtain an intermediate feature, and then the intermediate feature is converted to obtain the normalized feature map, such as multiplying the intermediate feature by 0.1 and adding 0.5 (0.1 and 0.5 are examples), and then limiting the feature value to be between 0 and 1. Of course, the above is only several examples of normalizing the important feature channel map, and the normalization manner is not limited.

[0151] In step S12, the high-frequency detail image is generated based on the important feature channel map and the normalized feature map.

[0152] Exemplarily, the normalized feature map can be convolved (such as two-dimensional convolution) with a Gaussian blur convolution kernel (which can be a 3*3 convolution kernel, or other size convolution kernel, and the convolution kernel is not limited) to obtain a Gaussian blur image, and then the important feature channel map is subtracted by the Gaussian blur image to obtain the high-frequency detail image. Of course, the above is only an example of generating the high-frequency detail image, as long as the high-frequency detail of the important feature channel map can be obtained, and the limitation is not made.

[0153] In step S13, for each feature value in the high-frequency detail image, the feature value is edge enhanced based on the edge enhancement segmented intensity value corresponding to the probability distribution value corresponding to the feature value to obtain an edge enhanced feature value.

[0154] Exemplarily, the important feature channel map includes a plurality of feature values, the important probability distribution channel map includes a plurality of probability distribution values, and the plurality of probability distribution values correspond one-to-one to the plurality of feature values. In addition, the high-frequency detail image includes a plurality of feature values (corresponding one-to-one to the plurality of feature values of the important feature channel map), and therefore the plurality of probability distribution values of the important probability distribution channel map correspond one-to-one to the plurality of feature values of the high-frequency detail image. Based on this, for each feature value in the high-frequency detail image, the probability distribution value corresponding to the feature value can be determined from the important probability distribution channel map, and the edge enhancement threshold interval corresponding to the probability distribution value is determined, and the edge enhancement segmented intensity value corresponding to the probability distribution value is determined based on the edge enhancement threshold interval.

[0155] After obtaining the edge enhancement segmented intensity value corresponding to the probability distribution value, the feature value can be edge enhanced based on the edge enhancement segmented intensity value to obtain an edge enhanced feature value. For example, the feature value is multiplied by the edge enhancement segmented intensity value (e.g., magl-1, magl-2, etc.) to obtain the edge enhanced feature value. Obviously, if the edge enhancement threshold interval is "less than thrl-1", the feature value is multiplied by 1, that is, the feature value remains unchanged, and the feature value is not edge enhanced. If the edge enhancement threshold interval is [thrl-1, thrl-2), the feature value is multiplied by magl-1, and magl-1 is a value greater than 1, so that the feature value is edge enhanced. Similarly, other edge enhancement segmented intensity values such as magl-2 are greater than 1, which can achieve edge enhancement.

[0156] For each feature value in the high-frequency detail image, the edge enhancement segmented intensity value corresponding to the feature value can be obtained in the above manner, and then the feature value is edge enhanced based on the edge enhancement segmented intensity value to obtain an edge enhanced feature value.

[0157] In step S14, the edge enhanced feature map is determined based on the edge enhanced feature value corresponding to each feature value. For example, the edge enhanced feature values corresponding to all feature values can be combined to obtain the edge enhanced feature map.

[0158] In step S15, the edge enhanced feature map is de-normalized to obtain the first reconstructed feature (i.e., the first reconstructed feature map), which is the reconstructed feature (map) after feature adaptive edge enhancement of the important feature channel map.

[0159] After obtaining the edge enhanced feature map, the edge enhanced feature map can be directly de-normalized to obtain the first reconstructed feature map. Alternatively, after obtaining the edge enhanced feature map, the edge enhanced feature map can be added to the normalized feature map to obtain a corrected edge enhanced feature map, and the corrected edge enhanced feature map is de-normalized to obtain the first reconstructed feature map.

[0160] For example, the edge enhanced normalized feature map can be determined based on the edge enhanced feature map, and the first reconstructed feature map can be determined based on the edge enhanced normalized feature map, the mean feature corresponding to the edge enhanced normalized feature map, and the variance feature corresponding to the edge enhanced normalized feature map. The first reconstructed feature can also be referred to as y_hat_sharp.

[0161] For example, the edge-enhanced feature map can be converted to obtain an edge-enhanced normalized feature map. For example, the feature values in the edge-enhanced feature map are limited to between 0 and 1, and then the feature values in the edge-enhanced feature map are subtracted by 0.5 and divided by 0.1 (0.5 and 0.1 are examples) to obtain the edge-enhanced normalized feature map.

[0162] The edge-enhanced normalized feature map is multiplied by the variance feature, and then the mean feature is added to obtain an edge-enhanced feature map corresponding to the important feature channel map, and the edge-enhanced feature map is denoted as the first reconstructed feature y_hat_sharp.

[0163] Of course, the above is only an example of de-normalization of the edge-enhanced feature map, and the de-normalization manner is not limited.

[0164] In a possible implementation, the encoding end needs to encode multiple edge-enhanced segmentation intensity values and multiple edge-enhanced segmentation thresholds in the third code stream corresponding to the current image block. For this process, the encoding end can use the following manner.

[0165] The encoding end can configure multiple candidate feature domain enhancement parameters. For each candidate feature domain enhancement parameter, the candidate feature domain enhancement parameter can include multiple edge-enhanced segmentation intensity values and multiple edge-enhanced segmentation thresholds. The encoding end can determine the objective value corresponding to each candidate feature domain enhancement parameter. Based on the objective value corresponding to each candidate feature domain enhancement parameter, the feature domain enhancement parameter corresponding to the current image block can be selected from all candidate feature domain enhancement parameters, that is, the candidate feature domain enhancement parameter with the minimum objective value. The encoding end can encode the feature domain enhancement parameter to obtain the third code stream corresponding to the current image block.

[0166] For example, for each candidate feature domain enhancement parameter, the initial reconstructed feature can be feature domain enhanced based on the candidate feature domain enhancement parameter and the probability distribution parameter to obtain an enhanced reconstructed feature. The feature domain enhancement process is described in the above embodiments. The objective value corresponding to the candidate feature domain enhancement parameter is determined based on the enhanced reconstructed feature. The determination manner of the objective value is not limited.

[0167] For example, for each candidate feature field enhancement parameter, the candidate feature field enhancement parameter includes magl-1, magl-2, …, magl-n, thrl-1, thrl-2, …, thrl-n. Based on the candidate feature field enhancement parameter, an enhanced reconstructed feature can be obtained, and then a target reconstructed image block x_hat_enhanced is obtained. The distortion index value of the target reconstructed image block x_hat_enhanced and the current image block x is calculated using the distortion index. After obtaining the distortion index value corresponding to each candidate feature field enhancement parameter, the candidate feature field enhancement parameter corresponding to the smallest distortion index value can be selected as the feature field enhancement parameter corresponding to the current image block, that is, the optimal feature field enhancement parameter, so that the encoding end encodes the feature field enhancement parameter in the third code stream.

[0168] In embodiment 8, the feature-adaptive stretching needs to be performed on the non-important feature channel map based on the feature field enhancement parameter and the non-important probability distribution channel map (that is, the non-important probability distribution channel map corresponding to the non-important feature channel map) to obtain the second reconstructed feature after feature-adaptive stretching. The feature-adaptive stretching process is described below.

[0169] For example, the feature field enhancement parameter can include a stretching parameter value. For example, the encoding end can encode the stretching parameter value in the third code stream corresponding to the current image block, and the decoding end can decode the stretching parameter value from the third code stream corresponding to the current image block, and the stretching parameter value is used as the feature field enhancement parameter. The stretching parameter value can be denoted as p.

[0170] For example, the non-important feature channel map can include a plurality of feature values, and the non-important probability distribution channel map corresponding to the non-important feature channel map can include a plurality of probability distribution values, and the plurality of probability distribution values correspond one-to-one to the plurality of feature values. Based on this, for each feature value in the non-important feature channel map, the stretched feature value corresponding to the feature value can be determined based on the feature value, the stretching parameter value, and the probability distribution value corresponding to the feature value. For example, the stretched feature value corresponding to the feature value can be determined using the following expression, of course, the following expression is only an example, and this is not limited.

[0171] y_hat_scale = y_hat + p * clip3(sigma * y_hat, -0.5, 0.5)

[0172] In the above expression, y_hat_scale represents the stretched eigenvalue, y_hat represents the eigenvalue in the non-important feature channel graph, p represents the stretching parameter value, sigma represents the probability distribution value in the non-important probability distribution channel graph, and the probability distribution value sigma corresponds to the eigenvalue y_hat, and clip3 is a limiting operation for limiting sigma*y_hat to between -0.5 and 0.5, -0.5 and 0.5 are configurable values, and there is no limitation on this, for example, when sigma*y_hat is less than -0.5, sigma*y_hat is limited to -0.5, when sigma*y_hat is greater than 0.5, sigma*y_hat is limited to 0.5, and when sigma*y_hat is greater than or equal to -0.5 and less than or equal to 0.5, the value of sigma*y_hat remains unchanged.

[0173] After the above processing is performed on each eigenvalue in the non-important feature channel graph, the stretched eigenvalue corresponding to each eigenvalue can be obtained, and then the second reconstructed feature (i.e., the second reconstructed feature graph) is determined based on the stretched eigenvalue corresponding to each eigenvalue. The second reconstructed feature (graph) is the reconstructed feature (graph) after the feature adaptive stretching of the non-important feature channel graph, for example, the second reconstructed feature graph is obtained by combining the stretched eigenvalues corresponding to all eigenvalues. The second reconstructed feature is also referred to as y_hat_scale.

[0174] The second reconstructed feature y_hat_scale corresponds to the non-important feature channel graph of the initial reconstructed feature y_hat, that is, the second reconstructed feature y_hat_scale is obtained by stretching the non-important feature channel graph. The first reconstructed feature y_hat_sharp corresponds to the important feature channel graph of the initial reconstructed feature y_hat, that is, the first reconstructed feature y_hat_sharp is obtained by enhancing the important feature channel graph. The second reconstructed feature y_hat_scale and the first reconstructed feature y_hat_sharp are combined to obtain the enhanced reconstructed feature y_hat_enhanced.

[0175] In a possible implementation, the encoding end needs to encode the stretching parameter value in the third code stream corresponding to the current image block. For this process, the encoding end can use the following manner: the encoding end can configure multiple candidate stretching parameter values, can determine the generation value corresponding to each candidate stretching parameter value, and based on the generation value corresponding to each candidate stretching parameter value, the stretching parameter value corresponding to the current image block can be selected from all candidate stretching parameter values, that is, the candidate stretching parameter value with the minimum generation value, and the encoding end can encode the stretching parameter value to obtain the third code stream corresponding to the current image block.

[0176] For example, for each candidate stretching parameter value, the non-important feature channel map can be adaptively stretched based on the candidate stretching parameter value and the non-important probability distribution channel map to obtain a second reconstructed feature after feature adaptive stretching, and then an enhanced reconstructed feature is obtained based on the second reconstructed feature, and the target reconstructed image block x_hat_enhanced is determined based on the enhanced reconstructed feature, and the objective value corresponding to the candidate stretching parameter value is determined based on the target reconstructed image block x_hat_enhanced.

[0177] For example, for each candidate stretching parameter value, the enhanced reconstructed feature can be obtained based on the candidate stretching parameter value, and then the target reconstructed image block x_hat_enhanced is obtained. The distortion index value of the target reconstructed image block x_hat_enhanced and the current image block x is calculated using the distortion index. After obtaining the distortion index value corresponding to each candidate stretching parameter value, the encoding end can select the candidate stretching parameter value corresponding to the smallest distortion index value as the stretching parameter value corresponding to the current image block, that is, the optimal stretching parameter value, and the encoding end can encode the stretching parameter value in the third code stream.

[0178] In another possible implementation, the encoding end configures multiple candidate feature domain enhancement parameters, for each candidate feature domain enhancement parameter, the candidate feature domain enhancement parameter includes multiple edge enhancement segment intensity values, multiple edge enhancement segment threshold values and a stretching parameter value, determines the objective value corresponding to each candidate feature domain enhancement parameter, selects the feature domain enhancement parameter corresponding to the current image block from all candidate feature domain enhancement parameters based on the objective value corresponding to each candidate feature domain enhancement parameter, that is, the candidate feature domain enhancement parameter with the smallest objective value, and the encoding end encodes the feature domain enhancement parameter (that is, multiple edge enhancement segment intensity values, multiple edge enhancement segment threshold values and a stretching parameter value) to obtain the third code stream corresponding to the current image block. For example, for each candidate feature domain enhancement parameter, the initial reconstructed feature is enhanced in the feature domain based on the candidate feature domain enhancement parameter and the probability distribution parameter to obtain an enhanced reconstructed feature, the target reconstructed image block x_hat_enhanced is determined based on the enhanced reconstructed feature, and the objective value corresponding to the candidate feature domain enhancement parameter is determined based on the target reconstructed image block x_hat_enhanced.

[0179] In Embodiment 6 and Embodiment 7, the initial reconstructed image block x_hat is input to the image domain enhancement module, the initial reconstructed image block x_hat is image domain enhanced by the image domain enhancement module, and the target reconstructed image block x_hat_enhanced corresponding to the current image block x is obtained. For example, the initial reconstructed image block x_hat can be image adaptive edge enhanced based on the image domain enhancement parameter and the probability distribution parameter sigma corresponding to the current image block, and the target reconstructed image block x_hat_enhanced corresponding to the current image block x is obtained. The image adaptive edge enhancement process is described below.

[0180] For example, the image domain enhancement parameter can include a plurality of image enhancement segmentation intensity values and a plurality of image enhancement segmentation threshold values. For example, the encoder encodes the plurality of image enhancement segmentation intensity values and the plurality of image enhancement segmentation threshold values in the third code stream corresponding to the current image block, the decoder decodes the plurality of image enhancement segmentation intensity values and the plurality of image enhancement segmentation threshold values from the third code stream corresponding to the current image block, and the plurality of image enhancement segmentation intensity values and the plurality of image enhancement segmentation threshold values are used as the image domain enhancement parameter.

[0181] The number of image enhancement segmentation intensity values and the number of image enhancement segmentation threshold values can be the same or different. For example, the number of image enhancement segmentation intensity values (or image enhancement segmentation threshold values) is encoded by the encoder in the third code stream corresponding to the current image block, and the number of image enhancement segmentation intensity values is decoded by the decoder from the third code stream corresponding to the current image block. For example, the number of image enhancement segmentation intensity values is m, the m image enhancement segmentation intensity values are denoted as magy-1, magy-2, …, magy-m, and the m image enhancement segmentation threshold values are denoted as thry-1, thry-2, …, thry-m.

[0182] For example, the plurality of image enhancement segmentation threshold values can form a plurality of image enhancement threshold intervals, and the plurality of image enhancement threshold intervals correspond one-to-one to the plurality of image enhancement segmentation intensity values. For example, the correspondence is shown in Table 2.

[0183] Table 2

[0184] Image enhancement segment intensity value Less than thry-1 1 (i.e. no edge enhancement) [thry-1, thry-2) magy-1 [thry-2, thry-3) magy-2 [thry-3, thry-4) magy-3 [thry-(m-1), thry-m) … … magy-(m-1) Greater than or equal to thry-m magy-m Figure 1

[0185] For example, the initial reconstructed image block x_hat can be image adaptive edge enhanced by the following steps S21-S23.

[0186] In step S21, the target probability distribution channel map is obtained based on the probability distribution parameter.

[0187] The initial reconstructed image block x_hat is a two-dimensional tensor of HxW, H is the image height of the initial reconstructed image block x_hat, and W is the image width of the initial reconstructed image block x_hat. The probability distribution parameter sigma is a three-dimensional tensor of C L xH L xW L , C L is the number of channels of the probability distribution parameter sigma, H L is the feature height of the probability distribution parameter sigma, and W L is the feature width of the probability distribution parameter sigma. The image height H of the initial reconstructed image block x_hat can be greater than the feature height H L of the probability distribution parameter sigma, such as the image height H being 4 times, 8 times, 16 times, etc. the feature height H L . The image width W of the initial reconstructed image block x_hat can be greater than the feature width W L of the probability distribution parameter sigma, such as the image width W being 4 times, 8 times, 16 times, etc. the feature width W L .

[0188] In order to perform image adaptive edge enhancement on the initial reconstructed image block x_hat using the probability distribution parameter, a target probability distribution channel map needs to be obtained based on the probability distribution parameter, and the target probability distribution channel map is a two-dimensional tensor of HxW.

[0189] For example, one probability distribution channel map can be selected from the C L probability distribution channel maps. Since the probability distribution parameter sigma includes important probability distribution channel maps (at least one) and non-important probability distribution channel maps (multiple), one important probability distribution channel map can be selected from the C L probability distribution channel maps, or one non-important probability distribution channel map can be selected from the C L probability distribution channel maps. In the following, an important probability distribution channel map is selected as an example.

[0190] Then, the important probability distribution channel map is upsampled to obtain the target probability distribution channel map, and the upsampling manner is not limited, and the size of the target probability distribution channel map can be the same as the size of the initial reconstructed image block x_hat.

[0191] For example, the nearest neighbor upscaling method can be used to upscale the important probability distribution channel map to obtain a target probability distribution channel map. For the nearest neighbor upscaling method, a pixel in the original low-resolution image (i.e., the important probability distribution channel map) can be selected as the center point of the corresponding region in the target high-resolution image (i.e., the target probability distribution channel map). The value of the pixel closest to the pixel in the original low-resolution image is taken as the value of the corresponding pixel in the target high-resolution image. The above process is repeated until all the pixels of the target high-resolution image are assigned values.

[0192] As described above, the important probability distribution channel map can be upsampled to the same size as the initial reconstructed image block x_hat, and the upsampled probability distribution channel map is denoted as the target probability distribution channel map sigma_ch_upscale.

[0193] In step S22, if the target probability distribution channel map includes multiple probability distribution values, for each probability distribution value, an image enhancement segmentation intensity value corresponding to the probability distribution value is determined based on the image enhancement threshold interval corresponding to the probability distribution value.

[0194] For example, if the probability distribution value is located in [thry-2, thry-3), the image enhancement segmentation intensity value corresponding to the probability distribution value can be magy-2, if the probability distribution value is located in [thry-3, thry-4), the image enhancement segmentation intensity value corresponding to the probability distribution value can be magy-3, and so on.

[0195] In step S23, image adaptive edge enhancement is performed on the initial reconstructed image block x_hat based on the image enhancement segmentation intensity value corresponding to each probability distribution value to obtain a target reconstructed image block x_hat_enhanced corresponding to the current image block.

[0196] First, a high-frequency detail image can be generated based on the initial reconstructed image block x_hat. For example, a convolution operation (such as a two-dimensional convolution operation) can be performed on the initial reconstructed image block x_hat and a Gaussian blur convolution kernel (which can be a 3*3 convolution kernel or other size convolution kernel, and the convolution kernel is not limited to this). A Gaussian blur image is obtained, and then the initial reconstructed image block x_hat is subtracted from the Gaussian blur image to obtain a high-frequency detail image. Of course, the above is only an example of generating a high-frequency detail image, as long as the high-frequency details of the initial reconstructed image block x_hat can be obtained, and the limitation is not made.

[0197] Then, for each feature value in the high-frequency detail image, the feature value is edge enhanced based on the image enhancement segmented intensity value corresponding to the probability distribution value corresponding to the feature value, to obtain an image enhancement feature value. For example, the initial reconstructed image block x_hat includes a plurality of feature values, the target probability distribution channel image includes a plurality of probability distribution values, and the plurality of probability distribution values correspond one-to-one to the plurality of feature values. In addition, the high-frequency detail image includes a plurality of feature values (corresponding one-to-one to the plurality of feature values of the initial reconstructed image block x_hat), so the plurality of probability distribution values of the target probability distribution channel image correspond one-to-one to the plurality of feature values of the high-frequency detail image. Based on this, for each feature value in the high-frequency detail image, the probability distribution value corresponding to the feature value can be determined from the target probability distribution channel image, and the image enhancement threshold interval corresponding to the probability distribution value can be determined, and then the image enhancement segmented intensity value corresponding to the probability distribution value can be determined based on the image enhancement threshold interval.

[0198] After obtaining the image enhancement segmented intensity value corresponding to the probability distribution value, the feature value can be edge enhanced based on the image enhancement segmented intensity value to obtain an image enhancement feature value. For example, the feature value is multiplied by the image enhancement segmented intensity value (such as magy-1, magy-2, etc.) to obtain an image enhancement feature value. Obviously, if the image enhancement threshold interval is "less than thry-1", the feature value is multiplied by 1, that is, the feature value remains unchanged and is not edge enhanced. If the image enhancement threshold interval is [thry-1, thry-2), the feature value is multiplied by magy-1, and magy-1 is a value greater than 1, so that the feature value is edge enhanced. Similarly, other image enhancement segmented intensity values such as magy-2 are greater than 1, which can achieve edge enhancement.

[0199] Then, the target reconstructed image block x_hat_enhanced is determined based on the image enhancement feature value corresponding to each feature value in the high-frequency detail image. For example, for each feature value in the high-frequency detail image, the image enhancement feature value corresponding to the feature value can be obtained in the above manner, and then the image enhancement feature map is determined based on the image enhancement feature value corresponding to each feature value. For example, the image enhancement feature values corresponding to all feature values can be combined to obtain an image enhancement feature map. Then, the image enhancement feature map is added to the initial reconstructed image block x_hat to obtain a final edge enhanced image, and the edge enhanced image is limited to the value range of the image, thereby obtaining the target reconstructed image block x_hat_enhanced.

[0200] In a possible implementation, the encoding end needs to encode a plurality of image enhancement segmented intensity values and a plurality of image enhancement segmented threshold values in the third code stream corresponding to the current image block. For this process, the encoding end can use the following method.

[0201] The encoding end can configure multiple candidate image domain enhancement parameters. For each candidate image domain enhancement parameter, the candidate image domain enhancement parameter can include multiple image enhancement segment intensity values and multiple image enhancement segment threshold values. The encoding end can determine a generation value corresponding to each candidate image domain enhancement parameter. Based on the generation value corresponding to each candidate image domain enhancement parameter, the encoding end can select, from all candidate image domain enhancement parameters, an image domain enhancement parameter corresponding to the current image block, i.e., a candidate image domain enhancement parameter with the minimum generation value. The encoding end can encode the image domain enhancement parameter to obtain a third code stream corresponding to the current image block.

[0202] For example, for each candidate image domain enhancement parameter, the initial reconstructed image block x_hat is subjected to image adaptive edge enhancement based on the candidate image domain enhancement parameter and the probability distribution parameter to obtain a target reconstructed image block x_hat_enhanced corresponding to the current image block. The image adaptive edge enhancement process is described above and will not be repeated here. The generation value corresponding to the candidate image domain enhancement parameter is determined based on the target reconstructed image block x_hat_enhanced. The determination manner of the generation value is not limited.

[0203] For example, for each candidate image domain enhancement parameter, the candidate image domain enhancement parameter includes magy-1, magy-2, …, magy-m, thry-1, thry-2, …, thry-m. The target reconstructed image block x_hat_enhanced can be obtained based on the candidate image domain enhancement parameter. The distortion index value of the target reconstructed image block x_hat_enhanced and the current image block x is calculated using the distortion index. After obtaining the distortion index value corresponding to each candidate image domain enhancement parameter, the candidate image domain enhancement parameter corresponding to the minimum distortion index value can be selected as the image domain enhancement parameter corresponding to the current image block, i.e., the optimal image domain enhancement parameter. In this way, the encoding end can encode the image domain enhancement parameter in the third code stream.

[0204] Embodiment 13: For embodiments 5-12, if a feature domain enhancement module is added before the synthesis transformation network, the feature domain enhancement module can perform feature domain enhancement on the luminance component, can perform feature domain enhancement on the chrominance component, or can perform feature domain enhancement on both the luminance component and the chrominance component. In this embodiment, the feature domain enhancement on the luminance component is taken as an example. In order to perform feature domain enhancement on the luminance component, the initial reconstructed feature y_hat corresponding to the luminance component of the current image block can be subjected to feature domain enhancement based on the feature domain enhancement parameter and the probability distribution parameter to obtain an enhanced reconstructed feature y_hat_enhanced corresponding to the luminance component. The feature domain enhancement process can refer to embodiments 5-12 and will not be repeated here.

[0205] For Embodiments 5-12, if an image domain enhancement module is added after the synthesis transformation network, the image domain enhancement module can perform image domain enhancement on the luminance component, can perform image domain enhancement on the chrominance component, or can perform image domain enhancement on both the luminance component and the chrominance component. In this embodiment, the image domain enhancement on both the luminance component and the chrominance component is taken as an example. For example, to perform image domain enhancement on the luminance component, the initial reconstructed image block x_hat corresponding to the luminance component of the current image block can be subjected to image adaptive edge enhancement based on the image domain enhancement parameter and the probability distribution parameter to obtain a target reconstructed image block x_hat_enhanced corresponding to the luminance component. The image domain enhancement process can refer to Embodiments 5-12, which are not repeated here. In addition, to perform image domain enhancement on the chrominance component, the initial reconstructed image block x_hat corresponding to the chrominance component of the current image block can be subjected to image adaptive edge enhancement based on the image domain enhancement parameter and the probability distribution parameter to obtain a target reconstructed image block x_hat_enhanced corresponding to the chrominance component.

[0206] When the initial reconstructed image block x_hat corresponding to the chrominance component is subjected to image adaptive edge enhancement, a target probability distribution channel map of the chrominance component can be obtained based on the probability distribution parameter of the luminance component. For example, an important probability distribution channel map is selected from all the probability distribution channel maps of the probability distribution parameter of the luminance component, and the important probability distribution channel map is up-sampled to obtain the target probability distribution channel map of the chrominance component, which has the same size as the initial reconstructed image block x_hat corresponding to the chrominance component. For example, the nearest neighbor up-sampling method is used to up-sample the important probability distribution channel map to obtain the target probability distribution channel map of the chrominance component. For the nearest neighbor up-sampling method, a pixel on the original low-resolution image (i.e., the important probability distribution channel map) is selected as the center point of the corresponding region on the target high-resolution image (i.e., the target probability distribution channel map). The value of the pixel closest to the pixel on the original low-resolution image is taken as the value of the corresponding pixel on the target high-resolution image. The above process is repeated until all the pixels of the target high-resolution image are assigned a value.

[0207] After obtaining the target probability distribution channel map of the chrominance component, for each probability distribution value in the target probability distribution channel map, the image enhancement sub-intensity value corresponding to the probability distribution value is determined based on the image enhancement threshold interval corresponding to the probability distribution value, and the initial reconstructed image block x_hat corresponding to the chrominance component is subjected to image adaptive edge enhancement based on the image enhancement sub-intensity value corresponding to each probability distribution value to obtain a target reconstructed image block x_hat_enhanced corresponding to the chrominance component. The image adaptive edge enhancement process can refer to Embodiment 12, and the relevant operations are performed on the chrominance component.

[0208] In this embodiment, a decoding method is proposed, which can include the following steps S31-S38.

[0209] In step S31, feature domain enhancement parameters and image domain enhancement parameters are decoded from the Bitstream#3 (i.e., the third code stream, which can also be referred to as the feature enhancement header information code stream) corresponding to the current image block. The feature domain enhancement parameters can include important channel identification (important channel number important_channel), a plurality of edge enhancement segment intensity values (denoted as magl-1, magl-2, …, magl-n), a plurality of edge enhancement segment threshold values (denoted as thrl-1, thrl-2, …, thrl-n), and a stretching parameter value ρ. The image domain enhancement parameters can include a plurality of image enhancement segment intensity values (denoted as magy-1, magy-2, …, magy-m) and a plurality of image enhancement segment threshold values (denoted as thry-1, thry-2, …, thry-m). The feature domain enhancement parameters can further include the number n of edge enhancement segment intensity values, and the image domain enhancement parameters can further include the number m of image enhancement segment intensity values.

[0210] In step S32, the input features of the feature domain enhancement module are the initial reconstructed features y_hat and the probability distribution parameter sigma. The initial reconstructed features y_hat are a three-dimensional tensor with a size of C L xH L xW L , where C L is the number of channels of the feature domain (i.e., the Latent domain), H L is the feature height of the Latent domain, and W L is the feature width of the Latent domain. The probability distribution parameter sigma is a three-dimensional tensor with a size of C L xH L xW L . When the channel of the initial reconstructed features y_hat is any value in [1, 2, …, C L ], a two-dimensional tensor y_hat_ch with a shape of H L xW L is called a feature channel map. When the channel of the probability distribution parameter sigma is any value in [1, 2, …, C L ], a two-dimensional tensor sigma_ch with a shape of H L xW L is called a probability distribution channel map.

[0211] In step S33, the initial reconstructed features y_hat have C La probability distribution channel map, C L a probability distribution channel map, C L a feature channel Figure 7A Correspondingly, the probability distribution channel map is denoted as sigma_ch. The feature adaptive edge enhancement is performed on the feature channel map of which the channel ch is identified as important_channel, and the feature adaptive stretching is performed on the feature channel map of which the channel ch is identified as non-important_channel.

[0212] In step S34, for the feature adaptive edge enhancement process, the input data is the important feature channel map y_hat_ch, the important probability distribution channel map sigma_ch, a plurality of edge enhancement segment intensity values (magl-1, magl-2, …, magl-n), a plurality of edge enhancement segment threshold values (thrl-1, thrl-2, …, thrl-n). Based on the above input data, the important feature channel map y_hat_ch can be subjected to feature adaptive edge enhancement to obtain a reconstructed feature map after feature adaptive edge enhancement, which is referred to as the first reconstructed feature y_hat_sharp. The feature adaptive edge enhancement process is described in Embodiment 10 and will not be repeated here.

[0213] In step S35, for the feature adaptive stretching process, the input data is the non-important feature channel map (C L -1) non-important feature channel maps in addition to the important feature channel map, the non-important probability distribution channel map (C L -1) non-important probability distribution channel maps in addition to the important probability distribution channel map), and the stretching parameter value ρ. Based on the above input data, each non-important feature channel map can be subjected to feature adaptive stretching to obtain a reconstructed feature map after feature adaptive stretching, which is referred to as the second reconstructed feature y_hat_scale. The feature adaptive stretching process is described in Embodiment 11 and will not be repeated here. The size of the non-important feature channel map subjected to feature adaptive stretching is (C L -1) x H L x W L , and the size of the corresponding probability distribution parameter sigma is also (C L -1) x H L x W L . Each element of this three-dimensional tensor can be subjected to the following adaptive stretching algorithm to obtain the stretched second reconstructed feature:

[0214] y_hat_scale = y_hat + ρ * clip3(sigma * y_hat, -0.5, 0.5).

[0215] wherein clip3 is a clipping operation.

[0216] In step S36, the second reconstructed feature y_hat_scale corresponds to the non-important channel enhancement of the initial reconstructed feature y_hat, the first reconstructed feature y_hat_sharp corresponds to the important channel enhancement of the initial reconstructed feature y_hat, and the first reconstructed feature y_hat_sharp and the second reconstructed feature y_hat_scale are combined to obtain the enhanced reconstructed feature y_hat_enhanced after feature domain enhancement.

[0217] In step S37, the enhanced reconstructed feature y_hat_enhanced after feature domain enhancement is input into the synthesis transformation network to obtain the initial reconstructed image block x_hat, which is a two-dimensional tensor with a size of HxW. The important probability distribution channel map sigma_ch of the probability distribution parameter sigma is upsampled to the same size as the initial reconstructed image block x_hat, and the upsampled probability distribution channel map is denoted as the target probability distribution channel map sigma_ch_upscale.

[0218] In step S38, for the image domain enhancement process, the input data of the image domain enhancement module are the initial reconstructed image block x_hat, the target probability distribution channel map sigma_ch_upscale, the plurality of image enhancement segment intensity values (magy-1, magy-2, …, magy-m), and the plurality of image enhancement segment threshold values (thry-1, thry-2, …, thry-m). Based on the above input data, the image domain enhancement module can perform image domain enhancement on the initial reconstructed image block x_hat to obtain the target reconstructed image block x_hat_enhanced corresponding to the current image block x. The image domain enhancement process is described in Embodiment 12 and will not be repeated here.

[0219] Embodiment 15: An encoding method is provided in this embodiment, which can include the following steps S41 to S45.

[0220] In step S41, based on C L feature channel maps y_hat and C L probability distribution channel maps sigma, the important channel identifier (also referred to as important channel number important_channel) corresponding to the important feature channel map is determined.

[0221] For example, slice y_hat and sigma along the channel dimension to obtain the current y_hat_ch and sigma_ch, and calculate the bits_per_ch of each feature channel map through the two tensors, as shown in the following expression. The above process is repeated C LSecondly, bits_per_ch of each feature channel map is obtained, and the channel serial number corresponding to the maximum bits_per_ch is selected as the important channel identifier.

[0222]

[0223] In step S42, the stretching parameter value p corresponding to the feature adaptive stretching process is determined. For example, N1 candidate stretching parameter values p are selected, the enhanced reconstructed feature and the target reconstructed image block x_hat_enhanced corresponding to each candidate stretching parameter value p are obtained, the distortion index value of the target reconstructed image block x_hat_enhanced and the current image block x is calculated using the distortion index, and the candidate stretching parameter value p corresponding to the minimum distortion index value is selected as the optimal stretching parameter value p corresponding to the feature adaptive stretching process.

[0224] In step S43, the feature domain enhancement parameter corresponding to the feature adaptive edge enhancement process is determined. The feature domain enhancement parameter includes a plurality of edge enhancement segment intensity values (magl-1, magl-2, …, magl-n) and a plurality of edge enhancement segment threshold values (thrl-1, thrl-2, …, thrl-n). For example, N2 candidate feature domain enhancement parameters are selected, the enhanced reconstructed feature and the target reconstructed image block x_hat_enhanced corresponding to each candidate feature domain enhancement parameter are obtained, the distortion index value of the target reconstructed image block x_hat_enhanced and the current image block x is calculated using the distortion index, and the candidate feature domain enhancement parameter corresponding to the minimum distortion index value is selected as the optimal feature domain enhancement parameter corresponding to the feature adaptive edge enhancement process.

[0225] In step S44, the image domain enhancement parameter corresponding to the image domain enhancement process is determined. The image domain enhancement parameter can include a plurality of image enhancement segment intensity values (magy-1, magy-2, …, magy-m) and a plurality of image enhancement segment threshold values (thry-1, thry-2, …, thry-m). For example, N3 candidate image domain enhancement parameters are selected, the enhanced reconstructed feature and the target reconstructed image block x_hat_enhanced corresponding to each candidate image domain enhancement parameter are obtained, the distortion index value of the target reconstructed image block x_hat_enhanced and the current image block x is calculated using the distortion index, and the candidate image domain enhancement parameter corresponding to the minimum distortion index value is selected as the optimal image domain enhancement parameter corresponding to the image domain enhancement process.

[0226] In step S45, the important channel identifier corresponding to the important feature channel map, the optimal stretching parameter value p, the optimal feature domain enhancement parameter, and the optimal image domain enhancement parameter are encoded into the header information code stream (the third code stream Bitstream#3 corresponding to the current image block). It should be noted that the feature domain enhancement module and the image domain enhancement module do not change Bitstream#1 and Bitstream#2, but the reconstructed image will pass through the enhancement module when the encoding end and the decoding end reconstruct the image, so the reconstructed image will be consistent at the encoding and decoding ends.

[0227] Embodiment 16: A method of image adaptive edge enhancement, i.e., the process of image domain enhancement module performing image domain enhancement on the initial reconstructed image block x_hat to obtain the target reconstructed image block x_hat_enhanced, can be a non-edge enhancement mask edge enhancement algorithm, i.e., a USM edge enhancement algorithm (Unsharp Masking edge enhancement), which includes steps S51 to S54.

[0228] In step S51, a two-dimensional convolution operation is performed on the original reconstructed image and a Gaussian blur convolution kernel to obtain a Gaussian blur image.

[0229] In step S52, the original reconstructed image is subtracted from the Gaussian blur image to obtain a high-frequency detail image.

[0230] In step S53, the high-frequency detail image is multiplied by an edge enhancement coefficient (i.e., an image enhancement segment intensity value corresponding to a probability distribution value corresponding to a feature value), and added to the original reconstructed image to obtain a final edge enhancement image.

[0231] In step S54, the edge enhancement image is limited to the value range of the image.

[0232] For example, the original reconstructed image is the initial reconstructed image block x_hat, and after limiting the edge enhancement image to the value range of the image, the target reconstructed image block x_hat_enhanced can be obtained. The process can be referred to in Embodiment 12.

[0233] Embodiment 17: A method of feature adaptive edge enhancement, i.e., the process of feature domain enhancement module performing feature adaptive edge enhancement on the important feature channel map based on the feature domain enhancement parameter and the important probability distribution channel map to obtain the first reconstructed feature, can be a non-edge enhancement mask edge enhancement algorithm, i.e., a USM edge enhancement algorithm, which includes steps S61 to S68.

[0234] In step S61, the important feature channel map y_hat_ch is subtracted from the mean and divided by the variance to obtain a normalized feature map.

[0235] In step S62, the normalized feature map is multiplied by 0.1 and added by 0.5, and the feature value is limited between 0 and 1.

[0236] In step S63, the normalized feature map is subjected to a two-dimensional convolution operation with a Gaussian blur kernel to obtain a Gaussian blur image.

[0237] In step S64, the original reconstructed image is subtracted by the Gaussian blur image to obtain a high-frequency detail image.

[0238] In step S65, the high-frequency detail image is multiplied by an edge enhancement coefficient (i.e., an edge enhancement segment intensity value corresponding to a probability distribution value corresponding to the feature value), and added to the normalized feature map to obtain an edge-enhanced feature map.

[0239] In step S66, the edge-enhanced feature map is limited between 0 and 1.

[0240] In step S67, the edge-enhanced feature map is subtracted by 0.5 and divided by 0.1 to obtain an edge-enhanced normalized feature map.

[0241] In step S68, the edge-enhanced normalized feature map is multiplied by the variance and added by the mean to obtain an important feature channel edge-enhanced feature map y_hat_sharp corresponding to the important feature channel map y_hat_ch, i.e., the first reconstructed feature.

[0242] For example, the original reconstructed image is the important feature channel map y_hat_ch, and the process can refer to embodiment 10.

[0243] For example, each of embodiments 1-17 can be implemented individually, and at least two of embodiments 1-17 can be implemented in combination.

[0244] For example, in each of the above embodiments, the content of the encoding end can also be applied to the decoding end, i.e., the decoding end can be processed in the same way as the encoding end, and the content of the decoding end can also be applied to the encoding end, i.e., the encoding end can be processed in the same way as the decoding end.

[0245] Based on the same application concept as the above method, the present embodiment also proposes a decoding device, which is applied to a decoding end, and the device comprises: a memory configured to store video data; and a decoder configured to implement the decoding method in embodiments 1-17 above, i.e., the processing flow of the decoding end.

[0246] For example, in a possible implementation, a decoder is configured to implement: decoding a first code stream corresponding to a current image block to obtain a coefficient hyper-parameter feature corresponding to the current image block; determining a probability distribution parameter based on the coefficient hyper-parameter feature, decoding a second code stream corresponding to the current image block based on the probability distribution parameter to obtain an initial reconstruction feature corresponding to the current image block; decoding a third code stream corresponding to the current image block to obtain an enhancement parameter corresponding to the current image block; and enhancing the initial reconstruction feature based on the enhancement parameter and the probability distribution parameter to obtain an enhanced reconstruction feature, and determining a target reconstruction image block corresponding to the current image block based on the enhanced reconstruction feature.

[0247] Based on the same application concept as the above method, an encoding device is further provided in the embodiments of the present application, and the device is applied to an encoding end. The device comprises: a memory configured to store video data; and an encoder configured to implement the encoding method in the above embodiments 1-17, i.e., the processing flow of the encoding end.

[0248] For example, in a possible implementation, an encoder is configured to implement: encoding a coefficient hyper-parameter feature corresponding to a current image block to obtain a first code stream corresponding to the current image block; determining a probability distribution parameter based on the coefficient hyper-parameter feature, and encoding an initial image feature corresponding to the current image block based on the probability distribution parameter to obtain a second code stream corresponding to the current image block; for each candidate enhancement parameter, enhancing an initial reconstruction feature based on the candidate enhancement parameter and the probability distribution parameter to obtain an enhanced reconstruction feature, determining a target reconstruction image block based on the enhanced reconstruction feature, and determining a value of a generation function corresponding to the candidate enhancement parameter based on the target reconstruction image block; selecting an enhancement parameter corresponding to the current image block from all candidate enhancement parameters based on the value of the generation function corresponding to each candidate enhancement parameter, and encoding the enhancement parameter to obtain a third code stream corresponding to the current image block.

[0249] Based on the same application concept as the above method, an encoding end device (which can also be referred to as a video decoder) is provided in the embodiments of the present application. From the hardware layer, the hardware architecture diagram of the encoding end device can be seen from FIG. 7. Figure 7B The encoding end device comprises: a processor 711 and a machine readable storage medium 712, the machine readable storage medium 712 stores machine executable instructions capable of being executed by the processor 711; and the processor 711 is used to execute the machine executable instructions to implement the decoding method in the above embodiments 1-17 of the present application.

[0250] For example, in a possible implementation, the processor 711, when executing the machine executable instructions, is configured to: decode a first code stream corresponding to a current image block to obtain a coefficient hyper-parameter feature corresponding to the current image block; determine a probability distribution parameter based on the coefficient hyper-parameter feature, decode a second code stream corresponding to the current image block based on the probability distribution parameter to obtain an initial reconstruction feature corresponding to the current image block; decode a third code stream corresponding to the current image block to obtain an enhancement parameter corresponding to the current image block; and enhance the initial reconstruction feature based on the enhancement parameter and the probability distribution parameter to obtain an enhanced reconstruction feature, and determine a target reconstruction image block corresponding to the current image block based on the enhanced reconstruction feature.

[0251] Based on the same application concept as the above method, an embodiment of the present application provides an encoding end device (which can also be referred to as a video encoder). In terms of hardware, a hardware architecture diagram of the encoding end device can be seen from FIG. 7. Figure 1 The encoding end device includes a processor 721 and a machine readable storage medium 722. The machine readable storage medium 722 stores machine executable instructions that can be executed by the processor 721. The processor 721 is configured to execute the machine executable instructions to implement the encoding method in the above embodiments 1-17 of the present application.

[0252] For example, in a possible implementation, the processor 721, when executing the machine executable instructions, is configured to: encode a coefficient hyper-parameter feature corresponding to a current image block to obtain a first code stream corresponding to the current image block; determine a probability distribution parameter based on the coefficient hyper-parameter feature, encode an initial image feature corresponding to the current image block based on the probability distribution parameter to obtain a second code stream corresponding to the current image block; for each candidate enhancement parameter, enhance an initial reconstruction feature based on the candidate enhancement parameter and the probability distribution parameter to obtain an enhanced reconstruction feature, determine a target reconstruction image block based on the enhanced reconstruction feature, and determine a value of a generation function corresponding to the candidate enhancement parameter based on the target reconstruction image block; select an enhancement parameter corresponding to the current image block from all candidate enhancement parameters based on the value of the generation function corresponding to each candidate enhancement parameter, and encode the enhancement parameter to obtain a third code stream corresponding to the current image block.

[0253] Based on the same application concept as the above method, an embodiment of the present application provides an electronic device, which includes a processor and a machine readable storage medium. The machine readable storage medium stores machine executable instructions that can be executed by the processor. The processor is configured to execute the machine executable instructions to implement the decoding method or the encoding method in the above embodiments 1-17 of the present application.

[0254] Based on the same application concept as the above method, the embodiments of the present application also provide a machine readable storage medium, wherein a plurality of computer instructions are stored on the machine readable storage medium, and the computer instructions can realize the method disclosed in the above embodiments of the present application, such as the decoding method or the encoding method in the above embodiments, when the computer instructions are executed by a processor.

[0255] Based on the same application concept as the above method, the embodiments of the present application also provide a computer application program, wherein the computer application program can realize the decoding method or the encoding method disclosed in the above embodiments of the present application, when the computer application program is executed by a processor.

[0256] Based on the same application concept as the above method, the embodiments of the present application also provide a decoding device, which can be applied to a decoding end (also referred to as a video decoder), and the decoding device can include: a decoding module, configured to decode a first code stream corresponding to a current image block to obtain a coefficient hyperparameter feature corresponding to the current image block; determine a probability distribution parameter based on the coefficient hyperparameter feature, decode a second code stream corresponding to the current image block based on the probability distribution parameter to obtain an initial reconstruction feature corresponding to the current image block; and decode a third code stream corresponding to the current image block to obtain an enhancement parameter corresponding to the current image block; an enhancement module, configured to enhance the initial reconstruction feature based on the enhancement parameter and the probability distribution parameter to obtain an enhanced reconstruction feature; and a determination module, configured to determine a target reconstruction image block corresponding to the current image block based on the enhanced reconstruction feature.

[0257] For example, the enhancement parameter includes an important channel identifier, the initial reconstruction feature includes C feature channel maps, the probability distribution parameter includes C probability distribution channel maps, and the C probability distribution channel maps correspond to the C feature channel maps Figure 1 For example, the enhancement module is further configured to select a feature channel map corresponding to the important channel identifier from the C feature channel maps as an important feature channel map, and select the remaining feature channel maps as non-important feature channel maps; select a probability distribution channel map corresponding to the important feature channel map as an important probability distribution channel map, and select the probability distribution channel maps corresponding to the non-important feature channel maps as non-important probability distribution channel maps.

[0258] For example, the initial reconstruction feature includes C feature channel maps, the probability distribution parameter includes C probability distribution channel maps, and the C probability distribution channel maps correspond to the C feature channel maps Figure 1Correspondingly, the enhancement module is further configured to determine, for each feature channel map, a number of consumed bits corresponding to the feature channel map based on a feature value in the feature channel map and a probability distribution value in a probability distribution channel map corresponding to the feature channel map; select, from the C feature channel maps, an important feature channel map based on the number of consumed bits corresponding to each feature channel map, and select a remaining feature channel map as a non-important feature channel map; and select a probability distribution channel map corresponding to the important feature channel map as an important probability distribution channel map, and select a probability distribution channel map corresponding to the non-important feature channel map as a non-important probability distribution channel map.

[0259] Illustratively, the initial reconstructed feature includes an important feature channel map and a non-important feature channel map, the probability distribution parameter includes an important probability distribution channel map corresponding to the important feature channel map and a non-important probability distribution channel map corresponding to the non-important feature channel map, and the enhancement module is configured to enhance the initial reconstructed feature based on the enhancement parameter and the probability distribution parameter to obtain an enhanced reconstructed feature, specifically by: performing feature-adaptive edge enhancement on the important feature channel map based on a feature domain enhancement parameter and the important probability distribution channel map to obtain a first reconstructed feature after feature-adaptive edge enhancement; performing feature-adaptive stretching on the non-important feature channel map based on the feature domain enhancement parameter and the non-important probability distribution channel map to obtain a second reconstructed feature after feature-adaptive stretching; and generating the enhanced reconstructed feature based on the first reconstructed feature and the second reconstructed feature.

[0260] Illustratively, the feature domain enhancement parameter includes a plurality of edge enhancement segment intensity values and a plurality of edge enhancement segment threshold values, the plurality of edge enhancement segment threshold values form a plurality of edge enhancement threshold intervals, and the plurality of edge enhancement threshold intervals correspond one-to-one to the plurality of edge enhancement segment intensity values; and the enhancement module is configured to perform feature-adaptive edge enhancement on the important feature channel map based on the feature domain enhancement parameter and the important probability distribution channel map to obtain a first reconstructed feature after feature-adaptive edge enhancement, specifically by: if the important probability distribution channel map includes a plurality of probability distribution values, then for each probability distribution value, determining an edge enhancement segment intensity value corresponding to the probability distribution value based on an edge enhancement threshold interval corresponding to the probability distribution value; and performing feature-adaptive edge enhancement on the important feature channel map based on the edge enhancement segment intensity value corresponding to each probability distribution value to obtain the first reconstructed feature after feature-adaptive edge enhancement.

[0261] Illustratively, the enhancement module performs feature-adaptive edge enhancement on the important feature channel map based on the edge enhancement sub-intensity value corresponding to each probability distribution value to obtain a first reconstructed feature, specifically: normalizing the important feature channel map to obtain a normalized feature map; generating a high-frequency detail image based on the important feature channel map and the normalized feature map; for each feature value in the high-frequency detail image, performing edge enhancement on the feature value based on the edge enhancement sub-intensity value corresponding to the probability distribution value corresponding to the feature value to obtain an edge-enhanced feature value; determining an edge-enhanced feature map based on the edge-enhanced feature value corresponding to each feature value; and de-normalizing the edge-enhanced feature map to obtain the first reconstructed feature.

[0262] Illustratively, the feature domain enhancement parameter includes a stretching parameter value, and the enhancement module performs feature-adaptive stretching on the unimportant feature channel map based on the feature domain enhancement parameter and the unimportant probability distribution channel map to obtain a second reconstructed feature after feature-adaptive stretching, specifically: if the unimportant feature channel map includes a plurality of feature values and the unimportant probability distribution channel map includes a plurality of probability distribution values, then for each feature value in the unimportant feature channel map, determining a stretched feature value corresponding to the feature value based on the feature value, the stretching parameter value, and the probability distribution value corresponding to the feature value; and determining the second reconstructed feature based on the stretched feature value corresponding to each feature value in the unimportant feature channel map.

[0263] Illustratively, the determination module determines the target reconstructed image block corresponding to the current image block based on the enhanced reconstructed feature, specifically: inputting the enhanced reconstructed feature into a synthesis transformation network to obtain the target reconstructed image block corresponding to the current image block; or inputting the enhanced reconstructed feature into a synthesis transformation network to obtain an initial reconstructed image block corresponding to the current image block; performing image-adaptive edge enhancement on the initial reconstructed image block based on the image domain enhancement parameter corresponding to the current image block and the probability distribution parameter to obtain the target reconstructed image block corresponding to the current image block; wherein the image domain enhancement parameter is obtained by decoding a third code stream corresponding to the current image block.

[0264] Exemplarily, the image domain enhancement parameter comprises a plurality of image enhancement segment intensity values and a plurality of image enhancement segment threshold values, the plurality of image enhancement segment threshold values form a plurality of image enhancement threshold intervals, and the plurality of image enhancement threshold intervals correspond to the plurality of image enhancement segment intensity values one by one; when the determination module performs image self-adaptive edge enhancement on the initial reconstructed image block based on the image domain enhancement parameter corresponding to the current image block and the probability distribution parameter to obtain the target reconstructed image block corresponding to the current image block, the determination module is specifically configured to: acquire a target probability distribution channel map based on the probability distribution parameter; if the target probability distribution channel map comprises a plurality of probability distribution values, for each probability distribution value, determine an image enhancement segment intensity value corresponding to the probability distribution value based on an image enhancement threshold interval corresponding to the probability distribution value; and perform image self-adaptive edge enhancement on the initial reconstructed image block based on the image enhancement segment intensity value corresponding to each probability distribution value to obtain the target reconstructed image block corresponding to the current image block.

[0265] Exemplarily, when the determination module acquires a target probability distribution channel map based on the probability distribution parameter, the determination module is specifically configured to: if the probability distribution parameter comprises an important probability distribution channel map and a non-important probability distribution channel map, perform up-sampling on the important probability distribution channel map to obtain the target probability distribution channel map; and a size of the target probability distribution channel map is the same as a size of the initial reconstructed image block.

[0266] Exemplarily, when the determination module performs image self-adaptive edge enhancement on the initial reconstructed image block based on the image enhancement segment intensity value corresponding to each probability distribution value to obtain the target reconstructed image block corresponding to the current image block, the determination module is specifically configured to: generate a high-frequency detail image based on the initial reconstructed image block; for each feature value in the high-frequency detail image, perform edge enhancement on the feature value based on an image enhancement segment intensity value corresponding to a probability distribution value corresponding to the feature value to obtain an image enhancement feature value; and determine the target reconstructed image block based on the image enhancement feature value corresponding to each feature value in the high-frequency detail image.

[0267] Illustratively, the enhancing module enhances the initial reconstructed feature based on the enhancement parameter and the probability distribution parameter to obtain an enhanced reconstructed feature, specifically by: performing feature domain enhancement on the initial reconstructed feature corresponding to the luminance component of the current image block based on a feature domain enhancement parameter and the probability distribution parameter to obtain an enhanced reconstructed feature corresponding to the luminance component; and performing image adaptive edge enhancement on the initial reconstructed image block based on the image domain enhancement parameter and the probability distribution parameter to obtain the target reconstructed image block corresponding to the current image block, specifically by: performing image adaptive edge enhancement on the initial reconstructed image block corresponding to the luminance component of the current image block based on the image domain enhancement parameter and the probability distribution parameter to obtain a target reconstructed image block corresponding to the luminance component; and performing image adaptive edge enhancement on the initial reconstructed image block corresponding to the chroma component of the current image block based on the image domain enhancement parameter and the probability distribution parameter to obtain a target reconstructed image block corresponding to the chroma component.

[0268] Illustratively, the initial reconstructed feature includes a plurality of feature channel maps, and the probability distribution parameter includes a plurality of probability distribution channel maps, the plurality of probability distribution channel maps correspond to the plurality of feature channel maps in a one-to-one manner. Figure 1 Illustratively, the decoding module is further configured to: decode the bitstream corresponding to the current image block to obtain an important channel identifier; and the determining module is further configured to: select, according to the important channel identifier, a feature channel map corresponding to the important channel identifier from the plurality of feature channel maps as an important feature channel map, and select remaining feature channel maps as non-important feature channel maps; select a probability distribution channel map corresponding to the important feature channel map as the important probability distribution channel map, and select probability distribution channel maps corresponding to the non-important feature channel maps as the non-important probability distribution channel maps.

[0269] Illustratively, the initial reconstructed feature includes a plurality of feature channel maps, and the probability distribution parameter includes a plurality of probability distribution channel maps, the plurality of probability distribution channel maps correspond to the plurality of feature channel maps in a one-to-one manner. ​ Illustratively, the determining module is further configured to: for each feature channel map, determine a consumed bit quantity corresponding to the feature channel map based on a feature value in the feature channel map and a probability distribution value in a probability distribution channel map corresponding to the feature channel map; select, based on the consumed bit quantity corresponding to each feature channel map, an important feature channel map from the plurality of feature channel maps, and select remaining feature channel maps as non-important feature channel maps; and select a probability distribution channel map corresponding to the important feature channel map as the important probability distribution channel map, and select probability distribution channel maps corresponding to the non-important feature channel maps as the non-important probability distribution channel maps.

[0270] Illustratively, the coefficient hyper-parameter feature corresponding to the current image block and the probability distribution parameter are obtained by decoding a first code stream related to the current image block; the initial reconstructed feature corresponding to the current image block is obtained by decoding a second code stream related to the current image block; and the important channel identifier is obtained by decoding a third code stream related to the current image block; wherein the first code stream, the second code stream and the third code stream are code streams encoding different information.

[0271] Based on the same application concept as the above method, an encoding device is also proposed in the embodiments of the present application, which is applied to an encoding end (also referred to as a video encoder). The device can include: an encoding module, configured to encode the coefficient hyper-parameter feature corresponding to the current image block to obtain a first code stream corresponding to the current image block; determine a probability distribution parameter based on the coefficient hyper-parameter feature, and encode the initial image feature corresponding to the current image block based on the probability distribution parameter to obtain a second code stream corresponding to the current image block; an enhancement module, configured to, for each candidate enhancement parameter, enhance the initial reconstructed feature based on the candidate enhancement parameter and the probability distribution parameter to obtain an enhanced reconstructed feature; a determination module, configured to determine a target reconstructed image block based on the enhanced reconstructed feature, and determine the objective value of the candidate enhancement parameter based on the target reconstructed image block; select the enhancement parameter corresponding to the current image block from all candidate enhancement parameters based on the objective value of each candidate enhancement parameter; and the encoding module is further configured to encode the enhancement parameter to obtain a third code stream corresponding to the current image block.

[0272] Illustratively, the initial reconstructed feature includes an important feature channel map and a non-important feature channel map, and the probability distribution parameter includes an important probability distribution channel map corresponding to the important feature channel map and a non-important probability distribution channel map corresponding to the non-important feature channel map. When the enhancement module enhances the initial reconstructed feature based on the candidate enhancement parameter and the probability distribution parameter to obtain the enhanced reconstructed feature, it is specifically configured to: perform feature adaptive edge enhancement on the important feature channel map based on the candidate feature domain enhancement parameter and the important probability distribution channel map to obtain a first reconstructed feature after feature adaptive edge enhancement; perform feature adaptive stretching on the non-important feature channel map based on the candidate feature domain enhancement parameter and the non-important probability distribution channel map to obtain a second reconstructed feature after feature adaptive stretching; and generate the enhanced reconstructed feature based on the first reconstructed feature and the second reconstructed feature.

[0273] For example, the determining module is specifically configured to determine the target reconstructed image block based on the enhanced reconstructed feature, by inputting the enhanced reconstructed feature into a synthesis transformation network to obtain an initial reconstructed image block corresponding to the current image block, and performing image adaptive edge enhancement on the initial reconstructed image block based on each candidate image domain enhancement parameter and the probability distribution parameter to obtain a target reconstructed image block corresponding to the current image block.

[0274] The encoding module is further configured to determine, for each candidate image domain enhancement parameter, a value of a cost function corresponding to the candidate image domain enhancement parameter based on the target reconstructed image block, and select, from all the candidate image domain enhancement parameters, an image domain enhancement parameter corresponding to the current image block based on the value of the cost function corresponding to each candidate image domain enhancement parameter, and encode the image domain enhancement parameter to obtain a third code stream corresponding to the current image block.

[0275] Those skilled in the art will understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. The present application can be in the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. The embodiments of the present application can be in the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The above is only an embodiment of the present application and is not intended to limit the present application.

[0276] The present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application shall be included in the scope of the claims of the present application.

Claims

1. An image decoding method, characterized in that, The method includes: The initial reconstruction features corresponding to the current image patch are input into the synthesis transformation network to obtain the initial reconstructed image patch corresponding to the current image patch; Obtain the image domain enhancement parameters corresponding to the current image block; Based on the image domain enhancement parameters and the probability distribution parameters corresponding to the current image block, adaptive edge enhancement is performed on the initial reconstructed image block to obtain the target reconstructed image block corresponding to the current image block; The image domain enhancement parameters include multiple image enhancement segment intensity values ​​and multiple image enhancement segment thresholds. The multiple image enhancement segment thresholds form multiple image enhancement threshold intervals, and the multiple image enhancement threshold intervals correspond one-to-one with the multiple image enhancement segment intensity values. The step of performing adaptive edge enhancement on the initial reconstructed image patch based on the image domain enhancement parameters and the probability distribution parameters corresponding to the current image patch to obtain the target reconstructed image patch corresponding to the current image patch includes: Determine the target probability distribution channel map based on the probability distribution parameters; If the target probability distribution channel map includes multiple probability distribution values, for each probability distribution value, the image enhancement segment intensity value corresponding to the probability distribution value is determined based on the image enhancement threshold interval corresponding to the probability distribution value. Based on the image enhancement segment intensity value corresponding to each probability distribution value, the initial reconstructed image block is subjected to adaptive edge enhancement to obtain the target reconstructed image block corresponding to the current image block.

2. The method according to claim 1, characterized in that, Before inputting the initial reconstructed features into the synthetic transform network to obtain the initial reconstructed image patch corresponding to the current image patch, the method further includes: Decode the bitstream corresponding to the current image block to obtain the coefficient hyperparameter features corresponding to the current image block; Based on the coefficient hyperparameter features, probability distribution parameters are determined, and the bitstream corresponding to the current image block is decoded based on the probability distribution parameters to obtain the initial reconstruction features corresponding to the current image block.

3. The method according to claim 1, characterized in that, The step of performing adaptive edge enhancement on the initial reconstructed image patch based on the image enhancement segment intensity value corresponding to each probability distribution value to obtain the target reconstructed image patch corresponding to the current image patch includes: Generate a high-frequency detail image based on the initial reconstructed image patch; For each feature value in the high-frequency detail image, an image enhancement segment intensity value is determined based on the probability distribution value corresponding to the feature value; the image enhancement segment intensity value is then used to perform edge enhancement on the feature value to obtain the image enhancement feature value. The target reconstructed image block is determined based on the image enhancement feature value corresponding to each feature value in the high-frequency detail image.

4. The method according to claim 2, characterized in that, By decoding the first bitstream corresponding to the current image patch, the coefficient hyperparameter features corresponding to the current image patch are obtained; by decoding the second bitstream corresponding to the current image patch, the initial reconstruction features corresponding to the current image patch are obtained. The first bitstream and the second bitstream are bitstreams that encode different information.

5. An image encoding method, characterized in that, The method includes: The initial reconstruction features corresponding to the current image patch are input into the synthesis transformation network to obtain the initial reconstructed image patch corresponding to the current image patch; Obtain the image domain enhancement parameters corresponding to the current image block; Based on the image domain enhancement parameters and the probability distribution parameters corresponding to the current image block, adaptive edge enhancement is performed on the initial reconstructed image block to obtain the target reconstructed image block corresponding to the current image block; The image domain enhancement parameters include multiple image enhancement segment intensity values ​​and multiple image enhancement segment thresholds. The multiple image enhancement segment thresholds form multiple image enhancement threshold intervals, and the multiple image enhancement threshold intervals correspond one-to-one with the multiple image enhancement segment intensity values. The step of performing adaptive edge enhancement on the initial reconstructed image patch based on the image domain enhancement parameters and the probability distribution parameters corresponding to the current image patch to obtain the target reconstructed image patch corresponding to the current image patch includes: Determine the target probability distribution channel map based on the probability distribution parameters; If the target probability distribution channel map includes multiple probability distribution values, for each probability distribution value, the image enhancement segment intensity value corresponding to the probability distribution value is determined based on the image enhancement threshold interval corresponding to the probability distribution value. Based on the image enhancement segment intensity value corresponding to each probability distribution value, the initial reconstructed image block is subjected to adaptive edge enhancement to obtain the target reconstructed image block corresponding to the current image block.

6. An image decoding device, characterized in that, The device includes: The determination module is used to input the initial reconstruction features corresponding to the current image block into the synthesis transform network to obtain the initial reconstruction image block corresponding to the current image block; The acquisition module is used to acquire the image domain enhancement parameters corresponding to the current image block; The enhancement module is used to perform adaptive edge enhancement on the initial reconstructed image block based on the image domain enhancement parameters and the probability distribution parameters corresponding to the current image block, so as to obtain the target reconstructed image block corresponding to the current image block; The image domain enhancement parameters include multiple image enhancement segment intensity values ​​and multiple image enhancement segment thresholds. These multiple image enhancement segment thresholds form multiple image enhancement threshold intervals, and each of these intervals corresponds one-to-one with the multiple image enhancement segment intensity values. The enhancement module performs adaptive edge enhancement on the initial reconstructed image block based on the image domain enhancement parameters and the probability distribution parameters corresponding to the current image block. Specifically, when obtaining the target reconstructed image block corresponding to the current image block, it is used for: Determine the target probability distribution channel map based on the probability distribution parameters; If the target probability distribution channel map includes multiple probability distribution values, for each probability distribution value, the image enhancement segment intensity value corresponding to the probability distribution value is determined based on the image enhancement threshold interval corresponding to the probability distribution value. Based on the image enhancement segment intensity value corresponding to each probability distribution value, the initial reconstructed image block is subjected to adaptive edge enhancement to obtain the target reconstructed image block corresponding to the current image block.

7. An image encoding device, characterized in that, The device includes: The determination module is used to input the initial reconstruction features corresponding to the current image block into the synthesis transform network to obtain the initial reconstruction image block corresponding to the current image block; The acquisition module is used to acquire the image domain enhancement parameters corresponding to the current image block; The enhancement module is used to perform adaptive edge enhancement on the initial reconstructed image block based on the image domain enhancement parameters and the probability distribution parameters corresponding to the current image block, so as to obtain the target reconstructed image block corresponding to the current image block; The image domain enhancement parameters include multiple image enhancement segment intensity values ​​and multiple image enhancement segment thresholds. These multiple image enhancement segment thresholds form multiple image enhancement threshold intervals, and each of these intervals corresponds one-to-one with the multiple image enhancement segment intensity values. The enhancement module performs adaptive edge enhancement on the initial reconstructed image block based on the image domain enhancement parameters and the probability distribution parameters corresponding to the current image block. Specifically, when obtaining the target reconstructed image block corresponding to the current image block, it is used for: Determine the target probability distribution channel map based on the probability distribution parameters; If the target probability distribution channel map includes multiple probability distribution values, for each probability distribution value, the image enhancement segment intensity value corresponding to the probability distribution value is determined based on the image enhancement threshold interval corresponding to the probability distribution value. Based on the image enhancement segment intensity value corresponding to each probability distribution value, the initial reconstructed image block is subjected to adaptive edge enhancement to obtain the target reconstructed image block corresponding to the current image block.

8. A decoding device, characterized in that, The decoding device includes a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor. The processor is configured to execute machine-executable instructions to implement the method according to any one of claims 1-4.

9. An encoding terminal device, characterized in that, The encoding end device includes: a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; The processor is configured to execute machine-executable instructions to implement the method of claim 5.

10. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores a plurality of computer instructions, which, when executed by a processor, implement the method described in any one of claims 1-4, or, when executed by a processor, implement the method described in claim 5.

11. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method described in any one of claims 1-4, or, when executed by a processor, implements the method described in claim 5.

Citation Information

Patent Citations

  • Coarse-grained context entropy coding method

    CN113347422A

  • Enhancement layer coding and decoding method and device

    CN114615500A