A decoding, encoding method, apparatus and device thereof
By employing an end-to-end video image compression method, and utilizing feature adjustment factors and convolutional layer structure design, the problem of insufficient neural network encoding and decoding performance is solved, thereby improving encoding and decoding performance and reducing complexity.
Patent Information
- Application Number
- CN202411921056.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-10
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2043-01-10
AI Technical Summary
Existing neural network-based encoding and decoding methods suffer from poor encoding performance, poor decoding performance, and high complexity.
The first neural network acquires the features of the current image patch and adjusts them based on the feature adjustment factor. The second neural network is then used to acquire the reconstructed image patch. An end-to-end video image compression method is adopted, and the encoding and decoding performance is improved by using convolutional layer structure design and auxiliary bitstream.
While maintaining low complexity, it effectively ensures the quality of reconstructed image patches, improves encoding and decoding performance, and reduces complexity.
Smart Images

Figure CN119653096B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of encoding and decoding technology, and in particular to a decoding and encoding method, apparatus and device thereof. Background Technology
[0002] To save space, video images are encoded before transmission. Complete video encoding includes processes such as prediction, transform, quantization, entropy coding, and filtering. The prediction process can be divided into intra-frame prediction and inter-frame prediction. Inter-frame prediction utilizes temporal correlation to predict the current pixel using pixels from neighboring encoded images, effectively removing temporal redundancy. Intra-frame prediction utilizes spatial correlation to predict the current pixel using pixels from the encoded blocks of the current frame, removing spatial redundancy.
[0003] With the rapid development of deep learning, it has achieved success in many high-level computer vision problems, such as image classification and object detection. Deep learning is also gradually being applied in the field of encoding and decoding, where neural networks can be used to encode and decode images. Although neural network-based encoding and decoding methods have shown great performance potential, they still suffer from problems such as poor encoding performance, poor decoding performance, and high complexity. Summary of the Invention
[0004] In view of this, this application provides a decoding and encoding method, apparatus and device, to improve encoding and decoding performance.
[0005] This application provides a decoding method applied at a decoding end, the method comprising:
[0006] The first feature corresponding to the current image patch is obtained through a first neural network, and the feature adjustment factor corresponding to the current image patch is obtained. The first neural network includes at least one convolutional layer.
[0007] The target feature is determined based on the first feature and the feature adjustment factor;
[0008] Based on the target features, a reconstructed image block corresponding to the current image block is obtained through a second neural network, wherein the second neural network includes at least one convolutional layer.
[0009] This application provides an encoding method applied to an encoding end, the method comprising: obtaining a first feature corresponding to a current image patch through a first neural network, wherein the first neural network includes at least one convolutional layer;
[0010] Based on the first feature, obtain the feature adjustment factor corresponding to the current image block;
[0011] The feature adjustment factor is encoded in the auxiliary bitstream corresponding to the current image block.
[0012] This application provides a decoding device applied at a decoding end, the device comprising:
[0013] A memory configured to store video data;
[0014] The decoder, which is configured to achieve:
[0015] The first feature corresponding to the current image patch is obtained through a first neural network, and the feature adjustment factor corresponding to the current image patch is obtained. The first neural network includes at least one convolutional layer.
[0016] The target feature is determined based on the first feature and the feature adjustment factor;
[0017] Based on the target features, a reconstructed image block corresponding to the current image block is obtained through a second neural network, wherein the second neural network includes at least one convolutional layer.
[0018] This application provides an encoding device applied at an encoding end, the device comprising:
[0019] A memory configured to store video data;
[0020] An encoder configured to: acquire a first feature corresponding to the current image patch via a first neural network, the first neural network including at least one convolutional layer;
[0021] Based on the first feature, obtain the feature adjustment factor corresponding to the current image block;
[0022] The feature adjustment factor is encoded in the auxiliary bitstream corresponding to the current image block.
[0023] This application provides a decoding end device, including: a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor;
[0024] The processor is used to execute machine-executable instructions to implement the above-described decoding method.
[0025] This application provides an encoding terminal device, including: a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor;
[0026] The processor is used to execute machine-executable instructions to implement the above-described encoding method.
[0027] This application provides an electronic device, including: a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; the processor is configured to execute the machine-executable instructions to implement the above-described decoding method; or, the processor is configured to execute the machine-executable instructions to implement the above-described encoding method.
[0028] This application provides a machine-readable storage medium storing a plurality of computer instructions, which, when executed by a processor, implement the above-described decoding method; or, implement the above-described encoding method.
[0029] As can be seen from the above technical solutions, in this embodiment, a first feature corresponding to the current image block is obtained through a neural network, and the first feature is adjusted based on the feature adjustment factor corresponding to the current image block to obtain the target feature. Based on the target feature, the reconstructed image block corresponding to the current image block is obtained through the neural network, thereby proposing an end-to-end video image compression method. This method can realize the encoding and decoding of video images based on a neural network, and improve the encoding and decoding efficiency by combining the feature adjustment factor. By combining the network structure design and auxiliary bitstream, the neural network can effectively ensure the quality of the reconstructed image block while maintaining low complexity, thereby improving the encoding and decoding performance and reducing complexity. Attached Figure Description
[0030] Figure 1 This is a schematic diagram of a three-dimensional feature matrix in one embodiment of this application;
[0031] Figure 2 This is a flowchart of a decoding method in one embodiment of this application;
[0032] Figure 3 This is a flowchart of an encoding method in one embodiment of this application;
[0033] Figure 4 This is a schematic diagram of the processing procedure at the encoding end in one embodiment of this application;
[0034] Figure 5 This is a schematic diagram of the processing procedure at the decoding end in one embodiment of this application;
[0035] Figures 6A-6H This is a schematic diagram of the network structure in one embodiment of this application;
[0036] Figure 7A This is a schematic diagram of the processing procedure at the decoding end in one embodiment of this application;
[0037] Figure 7B This is a schematic diagram of the processing procedure at the encoding end in one embodiment of this application;
[0038] Figure 8A and Figure 8B This is a schematic diagram of the structure of the decoding sub-network in one embodiment of this application;
[0039] Figure 9A and Figure 9B This is a schematic diagram of the structure of the attention subnetwork in one embodiment of this application;
[0040] Figures 10A-10C This is a schematic diagram showing the adjustment position of the feature adjustment factor in one embodiment of this application;
[0041] Figure 11A This is a hardware structure diagram of the decoding end device in one embodiment of this application;
[0042] Figure 11B This is a hardware structure diagram of the encoding end device in one embodiment of this application. Detailed Implementation
[0043] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the application. The singular forms “a,” “the,” and “the” used in the embodiments and claims of this application are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to any or all possible combinations including one or more of the associated listed items. It should be understood that although the terms first, second, third, etc., may be used to describe various information in the embodiments of this application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first information may also be referred to as second information without departing from the scope of the embodiments of this application, and similarly, second information may also be referred to as first information, depending on the context. Furthermore, the word “if” as used can be interpreted as “when,” “in response to a determination,” or “when…”.
[0044] This application proposes a decoding method and an encoding method, which may involve the following concepts:
[0045] JPEG (Joint Photographic Experts Group): JPEG is a standard for compressing continuous-tone still images. Files with the extension .jpg or .jpeg are commonly used image file formats. JPEG uses a joint coding method combining predictive coding (e.g., Differential Pulse-Code Modulation, DPCM), discrete cosine transform (DCT), and entropy coding to remove redundant image and color data. It is a lossy compression format, capable of compressing images into a very small storage space, but this inevitably causes some damage to the image data. Especially when using excessively high compression ratios, the quality of the decompressed image will be reduced. If high-quality images are desired, JPEG should not use excessively high compression ratios.
[0046] JPEG-AI (Joint Photographic Experts Group Artificial Intelligence): JPEG-AI aims to create a learning-based image coding standard that provides a single-stream, compact compression domain representation, significantly improving compression efficiency compared to commonly used image coding standards while maintaining the same subjective quality. This results in improved performance in image processing and computer vision tasks. JPEG-AI is geared towards a wide range of applications, including cloud storage, vision management, autonomous vehicles and devices, image acquisition, storage and management, real-time management of visual data, and media distribution. The goal of JPEG-AI is to design an encoding and decoding solution that significantly improves compression efficiency while maintaining the same subjective quality, providing efficient compression domain processing for machine learning-based image processing and computer vision tasks. JPEG-AI requires hardware and software-friendly encoding and decoding, support for 8-bit and 10-bit depths, and efficient encoding and progressive decoding of images using text and graphics.
[0047] Entropy coding: Entropy coding is a coding process that follows the principle of entropy without losing any information. Information entropy is the average amount of information in the source (a measure of uncertainty). Entropy coding methods can include, but are not limited to, Shannon coding, Huffman coding, and arithmetic coding.
[0048] Neural Networks (NNs): Neural networks are artificial neural networks, a computational model composed of numerous interconnected nodes (or neurons). In a neural network, neurons can represent different objects, such as features, letters, concepts, or meaningful abstract patterns. The types of processing units in a neural network can be divided into three categories: input units, output units, and hidden units. Input units receive signals and data from the external world; output units output the processed results; hidden units are located between the input and output units and cannot be observed from outside the system. The connection weights between neurons reflect the connection strength between units; the representation and processing of information are reflected in the connections between processing units. Neural networks are a non-programmed, brain-like information processing method. The essence of a neural network is to obtain a parallel and distributed information processing function through the transformations and dynamic behavior of the network, mimicking the information processing function of the human brain's nervous system to varying degrees and levels. In the field of video processing, commonly used neural networks include, but are not limited to, Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), and fully connected networks.
[0049] Convolutional Neural Networks (CNNs) are a type of feedforward neural network and one of the most representative network structures in deep learning. The artificial neurons in a CNN can respond to surrounding units within a certain coverage area, exhibiting excellent performance in large-scale image processing. The basic structure of a CNN consists of two layers: a feature extraction layer (also called a convolutional layer), where the input of each neuron is connected to the local receptive field of the previous layer, extracting features from that local area. Once these local features are extracted, their positional relationship with other features is determined. The second layer is a feature mapping layer (also called an activation layer). Each computational layer of the neural network consists of multiple feature maps, each a plane where all neurons have equal weights. Feature mapping structures can use functions such as the Sigmoid function, ReLU function, Leaky-ReLU function, PReLU function, and GDN function as activation functions for the convolutional network. Furthermore, because neurons on a single feature map share weights, the number of free parameters in the network is reduced.
[0050] For example, one advantage of convolutional neural networks (CNNs) over image processing algorithms is that they avoid complex preprocessing steps (such as extracting artificial features) and can directly input the original image for end-to-end learning. Another advantage of CNNs over ordinary neural networks is that ordinary neural networks use fully connected layers, meaning all neurons from the input layer to the hidden layer are connected. This results in a huge number of parameters, making network training time-consuming or even difficult. CNNs, however, avoid this difficulty through local connectivity and weight sharing.
[0051] Deconvolution: Also known as transposed convolution, deconvolution layers work similarly to convolutional layers. The main difference is that deconvolution layers use padding to make the output larger than the input (though they can also remain the same). If the stride is 1, the output size equals the input size; if the stride is N, the width of the output feature is N times the width of the input feature, and the height of the output feature is N times the height of the input feature.
[0052] Generalization ability: Generalization ability refers to the ability of a machine learning algorithm to adapt to new samples. The goal of learning is to learn the patterns hidden behind data pairs. The trained network can also give appropriate outputs for data outside the learning set that have the same pattern. This ability can be called generalization ability.
[0053] Features: The features involved in this application are a three-dimensional feature matrix of size C*W*H, see [link to relevant documentation]. Figure 1 The diagram shows a schematic of a three-dimensional feature matrix. In the three-dimensional feature matrix, C represents the number of channels, H represents the feature height, and W represents the feature width. The three-dimensional feature matrix can be either the input or the output of a neural network.
[0054] Rate-Distortion Optimized (RDBEM) principle: Two main metrics for evaluating coding efficiency are bitrate and PSNR (Peak Signal-to-Noise Ratio). A smaller bitrate results in a higher compression ratio, and a higher PSNR leads to better reconstructed image quality. In mode selection, the decision formula essentially evaluates both factors. For example, the cost of a mode is: J(mode) = D + λ*R, where D represents Distortion, typically measured using the SSE metric (Sum of Mean Squares of Differences between the Reconstructed Image Patch and the Source Image). Alternatively, the SAD metric (Sum of Absolute Differences between the Reconstructed Image Patch and the Source Image) can be used to consider the cost. λ is a Lagrange multiplier, and R is the actual number of bits required to encode the image patch in that mode, including the total number of bits needed for encoding mode information, motion information, residuals, etc. Using the RDBEM principle to compare and decide on coding modes during mode selection usually ensures optimal coding performance.
[0055] Numerous encoding tools have been proposed for various modules at the encoding end, and each tool often has multiple modes. The optimal encoding tool for different video sequences often differs. Therefore, during encoding, Rate-Distortion Optimization (RDO) is typically used to compare the encoding performance of different tools or modes to select the best mode. After determining the optimal tool or mode, the decision information is transmitted by encoding marker information in the bitstream. Although this method introduces higher encoding complexity, it can adaptively select the optimal mode combination for different content to achieve the best encoding performance. The decoding end can obtain the relevant mode information by directly parsing the marker information, with minimal impact from complexity.
[0056] The decoding and encoding methods in the embodiments of this application will be described in detail below with reference to several specific examples.
[0057] Example 1: This application proposes a decoding method, see [link to example]. Figure 2 The diagram shown illustrates the flowchart of this decoding method, which can be applied to the decoding end (also known as a video decoder). This method may include:
[0058] Step 201: Obtain the first feature corresponding to the current image patch through the first neural network, and obtain the feature adjustment factor corresponding to the current image patch. The first neural network may include at least one convolutional layer.
[0059] In one possible implementation, obtaining the first feature corresponding to the current image patch through the first neural network may include, but is not limited to: obtaining probability distribution parameters based on the first bitstream corresponding to the current image patch; determining a probability distribution model based on the probability distribution parameters, and decoding the second bitstream corresponding to the current image patch based on the probability distribution model to obtain decoded image features; and determining the first feature corresponding to the current image patch based on the decoded image features. In summary, in this implementation, the first neural network is used to achieve the functions of obtaining probability distribution parameters, determining the probability distribution model, decoding the second bitstream corresponding to the current image patch, and determining the first feature corresponding to the current image patch.
[0060] In another possible implementation, obtaining the first feature corresponding to the current image patch through the first neural network may include, but is not limited to: obtaining probability distribution parameters and predicted values based on the first bitstream corresponding to the current image patch; determining a probability distribution model based on the probability distribution parameters, and decoding the second bitstream corresponding to the current image patch based on the probability distribution model to obtain decoded image features; performing residual recovery on the decoded image features to obtain residual features; and determining the first feature corresponding to the current image patch based on the residual features and predicted values. In summary, in this implementation, the first neural network is used to achieve the functions of obtaining probability distribution parameters, obtaining predicted values, determining a probability distribution model, decoding the second bitstream corresponding to the current image patch, performing residual recovery, and determining the first feature corresponding to the current image patch.
[0061] For example, obtaining the feature adjustment factor corresponding to the current image patch may include, but is not limited to: decoding the auxiliary bitstream corresponding to the current image patch to obtain the feature adjustment factor corresponding to the current image patch, that is, parsing the feature adjustment factor corresponding to the current image patch from the auxiliary bitstream. Alternatively, a fixed parameter value may be determined as the feature adjustment factor corresponding to the current image patch; for example, the fixed parameter value 1 may be determined as the feature adjustment factor corresponding to the current image patch.
[0062] Step 202: Determine the target feature based on the first feature and the feature adjustment factor.
[0063] For example, determining the target feature based on the first feature and the feature adjustment factor may include, but is not limited to: enhancing the first feature to obtain the second feature; after obtaining the second feature, adjusting the second feature based on the feature adjustment factor to obtain the third feature; and after obtaining the third feature, determining the target feature based on the third feature.
[0064] In one possible implementation, feature enhancement is performed on the first feature to obtain the second feature, which may include, but is not limited to: determining the initial feature corresponding to the attention subnetwork based on the first feature; and enhancing the initial feature through the attention subnetwork to obtain the second feature. After obtaining the second feature, feature adjustment can be performed on the second feature based on a feature adjustment factor to obtain the third feature (i.e., the feature-adjusted feature is used as the third feature). After obtaining the third feature, a target feature is determined based on the third feature, which may include, but is not limited to: determining the third feature as the target feature.
[0065] For example, the attention subnetwork may include a residual enhancement subnetwork and a weight generation subnetwork. The attention subnetwork enhances the initial feature to obtain a second feature. This may include, but is not limited to: using the residual enhancement subnetwork to enhance the initial feature to obtain an enhanced feature, using the weight generation subnetwork to generate a weight feature corresponding to the initial feature, and generating a second feature based on the initial feature, the enhanced feature, and the weight feature.
[0066] For example, the weight generation subnetwork may include a residual block subnetwork, a convolutional subnetwork, and a feature mapping subnetwork. The weight generation subnetwork generates the weight features corresponding to the initial feature, which may include, but is not limited to: using the residual block subnetwork to perform convolutional activation on the initial feature to obtain the convolutionally activated feature; using the convolutional subnetwork to perform convolution on the convolutionally activated feature to obtain the convolutionally activated feature; and using the feature mapping subnetwork to perform feature mapping on the convolutionally activated feature to obtain the weight features.
[0067] For example, determining the initial feature corresponding to the attention sub-network based on the first feature may include, but is not limited to: using an initial enhancement sub-network to enhance the first feature to obtain the enhanced feature, and using an upsampling convolution sub-network to perform upsampling convolution on the enhanced feature to obtain the initial feature corresponding to the attention sub-network.
[0068] In another possible implementation, feature enhancement is performed on the first feature to obtain the second feature. This may include, but is not limited to: determining an initial feature corresponding to the attention sub-network based on the first feature; enhancing the initial feature through the first sub-network in the attention sub-network to obtain the second feature. After obtaining the second feature, feature adjustment can be performed on the second feature based on a feature adjustment factor to obtain a third feature (i.e., the feature-adjusted feature is used as the third feature). After obtaining the third feature, a target feature is determined based on the third feature. This may include, but is not limited to: processing the third feature through the second sub-network in the attention sub-network to obtain a fourth feature, and determining the fourth feature as the target feature.
[0069] For example, the initial feature is enhanced by the first subnetwork in the attention subnetwork to obtain the second feature. This can include, but is not limited to, using a residual enhancement subnetwork to enhance the initial feature to obtain the enhanced feature, using a weight generation subnetwork to generate the weight feature corresponding to the initial feature, and generating the second feature based on the enhanced feature and the weight feature. The third feature is processed by the second subnetwork in the attention subnetwork to obtain the fourth feature. This can include, but is not limited to, generating the fourth feature based on the initial feature and the third feature.
[0070] The weight generation subnetwork may include a residual block subnetwork, a convolutional subnetwork, and a feature mapping subnetwork. The weight generation subnetwork generates the weight features corresponding to the initial feature, which may include, but is not limited to: using the residual block network to perform convolutional activation on the initial feature to obtain the convolutionally activated feature; using the convolutional subnetwork to perform convolution on the convolutionally activated feature to obtain the convolutionally activated feature; and using the feature mapping subnetwork to perform feature mapping on the convolutionally activated feature to obtain the weight features.
[0071] For example, the initial feature is enhanced by the first sub-network in the attention sub-network to obtain the second feature. This can include, but is not limited to: enhancing the initial feature using a residual block sub-network to obtain the enhanced feature; convolving the enhanced feature using a convolution sub-network to obtain the convolutional feature; and determining the second feature based on the convolutional feature. The third feature is processed by the second sub-network in the attention sub-network to obtain the fourth feature. This can include, but is not limited to: mapping the third feature using a feature mapping sub-network to obtain the weighted feature; enhancing the initial feature using a residual enhancement sub-network to obtain the enhanced feature; and generating the fourth feature based on the initial feature, the enhanced feature, and the weighted feature.
[0072] In the above embodiments, determining the initial feature corresponding to the attention sub-network based on the first feature may include, but is not limited to: using an initial enhancement sub-network to enhance the first feature to obtain the enhanced feature, and using an upsampling convolution sub-network to perform upsampling convolution on the enhanced feature to obtain the initial feature corresponding to the attention sub-network.
[0073] In one possible implementation, the second feature is adjusted based on a feature adjustment factor to obtain the third feature. This adjustment may include, but is not limited to: if the second feature comprises C*H*W feature values, where C represents the number of channels, H represents the feature height, and W represents the feature width, then feature adjustment values corresponding to the C*H*W feature values are determined based on the feature adjustment factor, and the C*H*W feature values are adjusted based on these feature adjustment values to obtain the adjusted feature values corresponding to the C*H*W feature values. Based on this, the third feature can be generated using the adjusted feature values corresponding to the C*H*W feature values.
[0074] For example, determining the feature adjustment values corresponding to C*H*W feature values based on the feature adjustment factor may include, but is not limited to: if the feature adjustment factor includes C*H*W feature adjustment values, then determining the feature adjustment values corresponding to C*H*W feature values based on the C*H*W feature adjustment values; or, if the feature adjustment factor includes H*W feature adjustment values, then determining the feature adjustment values corresponding to C*H*W feature values based on the H*W feature adjustment values; or, if the feature adjustment factor includes C feature adjustment values, then determining the feature adjustment values corresponding to C*H*W feature values based on the C feature adjustment values.
[0075] For example, feature adjustment is performed on C*H*W feature values based on feature adjustment values to obtain the adjusted feature values corresponding to C*H*W feature values. This can include, but is not limited to, cases where one of the C*H*W feature values corresponds to one feature adjustment value. In this case, the adjusted feature value corresponding to the feature value can be determined using the following formula: yr = rf*y; where yr represents the adjusted feature value, rf represents the feature adjustment value, and y represents the feature value. Alternatively, if one of the C*H*W feature values corresponds to N+1 feature adjustment values, where N is a positive integer, the adjusted feature value corresponding to the feature value can be determined using the following formula: yr = rf_0*y0 + rf_1*y1 + rf_2*y2 + ... + rf_N*yN; where yr represents the adjusted feature value, rf_0, rf_1, rf_2, ..., rf_N represent the N+1 feature adjustment values, and y represents the feature value.
[0076] Step 203: Based on the target features, obtain the reconstructed image block corresponding to the current image block (i.e. the final output reconstructed image block) through the second neural network. The second neural network may include at least one convolutional layer.
[0077] For example, the second neural network may include a reconstruction decoding subnetwork. Based on target features, the second neural network obtains the reconstructed image patch corresponding to the current image patch. This may include, but is not limited to, processing the target features through the reconstruction decoding subnetwork to obtain the reconstructed image patch corresponding to the current image patch. For instance, the reconstruction decoding subnetwork may include an upsampling convolution subnetwork, which can perform upsampling convolution on the target features to obtain the reconstructed image patch corresponding to the current image patch. As another example, the reconstruction decoding subnetwork may include an upsampling convolution subnetwork and a color space transformation subnetwork. The upsampling convolution subnetwork can perform upsampling convolution on the target features, and the color space transformation subnetwork can perform color space transformation on the upsampling convolution features to obtain the reconstructed image patch corresponding to the current image patch.
[0078] In one possible implementation, if the current image patch has feature adjustment mode enabled, the feature adjustment factor corresponding to the current image patch is obtained, and the target feature is determined based on the first feature and the feature adjustment factor. Alternatively, if the current image patch does not have feature adjustment mode enabled, the target feature is determined based on the first feature, and it is not necessary to obtain the feature adjustment factor corresponding to the current image patch.
[0079] For example, a feature adjustment flag can be parsed from the bitstream. If the feature adjustment flag allows the current image block to enable feature adjustment mode, then it can be determined that the current image block enables feature adjustment mode. Alternatively, a first range of feature values can be parsed from the bitstream. If the feature value of the second feature is within the first range, then it can be determined that the current image block enables feature adjustment mode. Or, a second range of feature change values can be parsed from the bitstream. If the feature change value corresponding to the second feature is within the second range, then it can be determined that the current image block enables feature adjustment mode.
[0080] For example, if the attribute value corresponding to the second feature is within the configured attribute value range, it can be determined that the current image patch is in feature adjustment mode; wherein, the attribute value may include the variance value.
[0081] For example, the above execution order is merely an example for ease of description. In practical applications, the execution order between steps can be changed, and there is no limitation on this execution order. Moreover, in other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in this specification, and the method may include more or fewer steps than described in this specification. Furthermore, a single step described in this specification may be broken down into multiple steps in other embodiments; multiple steps described in this specification may also be combined into a single step in other embodiments.
[0082] As can be seen from the above technical solutions, in this embodiment, a first feature corresponding to the current image block is obtained through a first neural network, and the first feature is adjusted based on the feature adjustment factor corresponding to the current image block to obtain the target feature. Based on the target feature, a reconstructed image block corresponding to the current image block is obtained through a second neural network, thereby proposing an end-to-end video image compression method. This method can decode video images based on neural networks and improve decoding efficiency by combining feature adjustment factors. By combining network structure design and auxiliary bitstream, the neural network can effectively ensure the quality of reconstructed image blocks while maintaining low complexity, thereby improving decoding performance and reducing complexity.
[0083] Example 2: An encoding method is proposed in this application embodiment, see [link to example]. Figure 3The diagram shown illustrates the flowchart of this encoding method, which can be applied to the encoding end (also known as a video encoder). This method may include:
[0084] Step 301: Obtain the first feature corresponding to the current image patch through the first neural network. The first neural network may include at least one convolutional layer.
[0085] In one possible implementation, obtaining the first feature corresponding to the current image patch through the first neural network may include, but is not limited to: obtaining probability distribution parameters based on the first bitstream corresponding to the current image patch; determining a probability distribution model based on the probability distribution parameters, and decoding the second bitstream corresponding to the current image patch based on the probability distribution model to obtain decoded image features; and determining the first feature corresponding to the current image patch based on the decoded image features. In summary, in this implementation, the first neural network is used to achieve the functions of obtaining probability distribution parameters, determining the probability distribution model, decoding the second bitstream corresponding to the current image patch, and determining the first feature corresponding to the current image patch.
[0086] In another possible implementation, obtaining the first feature corresponding to the current image patch through the first neural network may include, but is not limited to: obtaining probability distribution parameters and predicted values based on the first bitstream corresponding to the current image patch; determining a probability distribution model based on the probability distribution parameters, and decoding the second bitstream corresponding to the current image patch based on the probability distribution model to obtain decoded image features; performing residual recovery on the decoded image features to obtain residual features; and determining the first feature corresponding to the current image patch based on the residual features and predicted values. In summary, in this implementation, the first neural network is used to achieve the functions of obtaining probability distribution parameters, obtaining predicted values, determining a probability distribution model, decoding the second bitstream corresponding to the current image patch, performing residual recovery, and determining the first feature corresponding to the current image patch.
[0087] Step 302: Obtain the feature adjustment factor corresponding to the current image block based on the first feature.
[0088] For example, obtaining the feature adjustment factor corresponding to the current image patch based on the first feature may include, but is not limited to: obtaining at least one candidate feature adjustment factor, and determining the rate-distortion cost corresponding to each candidate feature adjustment factor based on the first feature. Based on the rate-distortion cost corresponding to each candidate feature adjustment factor, one candidate feature adjustment factor can be selected from all candidate feature adjustment factors as the feature adjustment factor corresponding to the current image patch.
[0089] Step 303: Encode the feature adjustment factor in the auxiliary bitstream corresponding to the current image block.
[0090] For example, a fixed parameter value can be determined as the feature adjustment factor corresponding to the current image patch. For instance, a fixed parameter value of 1 can be determined as the feature adjustment factor corresponding to the current image patch. In this case, the feature adjustment factor may not need to be encoded in the auxiliary bitstream corresponding to the current image patch.
[0091] For example, the target feature can also be determined based on the first feature corresponding to the current image patch and the feature adjustment factor corresponding to the current image patch. Based on the target feature, the reconstructed image patch corresponding to the current image patch (i.e. the final output reconstructed image patch) can be obtained through the second neural network. The second neural network may include at least one convolutional layer.
[0092] For example, determining the target feature based on the first feature and the feature adjustment factor may include, but is not limited to: enhancing the first feature to obtain the second feature; after obtaining the second feature, adjusting the second feature based on the feature adjustment factor to obtain the third feature; and after obtaining the third feature, determining the target feature based on the third feature.
[0093] In one possible implementation, feature enhancement is performed on the first feature to obtain the second feature, which may include, but is not limited to: determining the initial feature corresponding to the attention subnetwork based on the first feature; and enhancing the initial feature through the attention subnetwork to obtain the second feature. After obtaining the second feature, feature adjustment can be performed on the second feature based on a feature adjustment factor to obtain the third feature (i.e., the feature-adjusted feature is used as the third feature). After obtaining the third feature, a target feature is determined based on the third feature, which may include, but is not limited to: determining the third feature as the target feature.
[0094] For example, the attention subnetwork may include a residual enhancement subnetwork and a weight generation subnetwork. The attention subnetwork enhances the initial feature to obtain a second feature. This may include, but is not limited to: using the residual enhancement subnetwork to enhance the initial feature to obtain an enhanced feature, using the weight generation subnetwork to generate a weight feature corresponding to the initial feature, and generating a second feature based on the initial feature, the enhanced feature, and the weight feature.
[0095] For example, the weight generation subnetwork may include a residual block subnetwork, a convolutional subnetwork, and a feature mapping subnetwork. The weight generation subnetwork generates the weight features corresponding to the initial feature, which may include, but is not limited to: using the residual block subnetwork to perform convolutional activation on the initial feature to obtain the convolutionally activated feature; using the convolutional subnetwork to perform convolution on the convolutionally activated feature to obtain the convolutionally activated feature; and using the feature mapping subnetwork to perform feature mapping on the convolutionally activated feature to obtain the weight features.
[0096] For example, determining the initial feature corresponding to the attention sub-network based on the first feature may include, but is not limited to: using an initial enhancement sub-network to enhance the first feature to obtain the enhanced feature, and using an upsampling convolution sub-network to perform upsampling convolution on the enhanced feature to obtain the initial feature corresponding to the attention sub-network.
[0097] In another possible implementation, feature enhancement is performed on the first feature to obtain the second feature. This may include, but is not limited to: determining an initial feature corresponding to the attention sub-network based on the first feature; enhancing the initial feature through the first sub-network in the attention sub-network to obtain the second feature. After obtaining the second feature, feature adjustment can be performed on the second feature based on a feature adjustment factor to obtain a third feature (i.e., the feature-adjusted feature is used as the third feature). After obtaining the third feature, a target feature is determined based on the third feature. This may include, but is not limited to: processing the third feature through the second sub-network in the attention sub-network to obtain a fourth feature, and determining the fourth feature as the target feature.
[0098] For example, the initial feature is enhanced by the first subnetwork in the attention subnetwork to obtain the second feature. This can include, but is not limited to, using a residual enhancement subnetwork to enhance the initial feature to obtain the enhanced feature, using a weight generation subnetwork to generate the weight feature corresponding to the initial feature, and generating the second feature based on the enhanced feature and the weight feature. The third feature is processed by the second subnetwork in the attention subnetwork to obtain the fourth feature. This can include, but is not limited to, generating the fourth feature based on the initial feature and the third feature.
[0099] The weight generation subnetwork may include a residual block subnetwork, a convolutional subnetwork, and a feature mapping subnetwork. The weight generation subnetwork generates the weight features corresponding to the initial feature, which may include, but is not limited to: using the residual block network to perform convolutional activation on the initial feature to obtain the convolutionally activated feature; using the convolutional subnetwork to perform convolution on the convolutionally activated feature to obtain the convolutionally activated feature; and using the feature mapping subnetwork to perform feature mapping on the convolutionally activated feature to obtain the weight features.
[0100] For example, the initial feature is enhanced by the first sub-network in the attention sub-network to obtain the second feature. This can include, but is not limited to: enhancing the initial feature using a residual block sub-network to obtain the enhanced feature; convolving the enhanced feature using a convolution sub-network to obtain the convolutional feature; and determining the second feature based on the convolutional feature. The third feature is processed by the second sub-network in the attention sub-network to obtain the fourth feature. This can include, but is not limited to: mapping the third feature using a feature mapping sub-network to obtain the weighted feature; enhancing the initial feature using a residual enhancement sub-network to obtain the enhanced feature; and generating the fourth feature based on the initial feature, the enhanced feature, and the weighted feature.
[0101] In the above embodiments, determining the initial feature corresponding to the attention sub-network based on the first feature may include, but is not limited to: using an initial enhancement sub-network to enhance the first feature to obtain the enhanced feature, and using an upsampling convolution sub-network to perform upsampling convolution on the enhanced feature to obtain the initial feature corresponding to the attention sub-network.
[0102] In one possible implementation, the second feature is adjusted based on a feature adjustment factor to obtain the third feature. This adjustment may include, but is not limited to: if the second feature comprises C*H*W feature values, where C represents the number of channels, H represents the feature height, and W represents the feature width, then feature adjustment values corresponding to the C*H*W feature values are determined based on the feature adjustment factor, and the C*H*W feature values are adjusted based on these feature adjustment values to obtain the adjusted feature values corresponding to the C*H*W feature values. Based on this, the third feature can be generated using the adjusted feature values corresponding to the C*H*W feature values.
[0103] For example, determining the feature adjustment values corresponding to C*H*W feature values based on the feature adjustment factor may include, but is not limited to: if the feature adjustment factor includes C*H*W feature adjustment values, then determining the feature adjustment values corresponding to C*H*W feature values based on the C*H*W feature adjustment values; or, if the feature adjustment factor includes H*W feature adjustment values, then determining the feature adjustment values corresponding to C*H*W feature values based on the H*W feature adjustment values; or, if the feature adjustment factor includes C feature adjustment values, then determining the feature adjustment values corresponding to C*H*W feature values based on the C feature adjustment values.
[0104] For example, feature adjustment is performed on C*H*W feature values based on feature adjustment values to obtain the adjusted feature values corresponding to C*H*W feature values. This can include, but is not limited to: if one feature value in the C*H*W feature values corresponds to one feature adjustment value, the adjusted feature value corresponding to that feature value can be determined by the following formula: yr = rf*y; where yr represents the adjusted feature value, rf represents the feature adjustment value, and y represents the feature value. Alternatively, if one feature value in the C*H*W feature values corresponds to N+1 feature adjustment values, where N is a positive integer, the adjusted feature value corresponding to that feature value can be determined by the following formula: yr = rf_0*y0 + rf_1*y1 + rf_2*y2 + ... + rf_N*yN; where yr represents the adjusted feature value, rf_0, rf_1, rf_2, ..., rf_N represent N+1 feature adjustment values, and y represents the feature value.
[0105] For example, the second neural network may include a reconstruction decoding subnetwork. Based on target features, the second neural network obtains the reconstructed image patch corresponding to the current image patch. This may include, but is not limited to, processing the target features through the reconstruction decoding subnetwork to obtain the reconstructed image patch corresponding to the current image patch. For instance, the reconstruction decoding subnetwork may include an upsampling convolution subnetwork, which can perform upsampling convolution on the target features to obtain the reconstructed image patch corresponding to the current image patch. As another example, the reconstruction decoding subnetwork may include an upsampling convolution subnetwork and a color space transformation subnetwork. The upsampling convolution subnetwork can perform upsampling convolution on the target features, and the color space transformation subnetwork can perform color space transformation on the upsampling convolution features to obtain the reconstructed image patch corresponding to the current image patch.
[0106] In one possible implementation, if the current image patch has feature adjustment mode enabled, then the feature adjustment factor corresponding to the current image patch is obtained based on the first feature, and the feature adjustment factor is encoded in the auxiliary bitstream corresponding to the current image patch. Alternatively, if the current image patch does not have feature adjustment mode enabled, then it is not necessary to obtain the feature adjustment factor corresponding to the current image patch.
[0107] For example, a feature adjustment flag can be encoded in the bitstream, which either allows the current image block to enable feature adjustment mode, or it can disable feature adjustment mode for the current image block. Alternatively, a first range of feature values can be encoded in the bitstream. Or, a second range of feature change values can be encoded in the bitstream.
[0108] For example, the above execution order is merely an example for ease of description. In practical applications, the execution order between steps can be changed, and there is no limitation on this execution order. Moreover, in other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in this specification, and the method may include more or fewer steps than described in this specification. Furthermore, a single step described in this specification may be broken down into multiple steps in other embodiments; multiple steps described in this specification may also be combined into a single step in other embodiments.
[0109] As can be seen from the above technical solutions, this application proposes an end-to-end video image compression method that can encode video images based on neural networks and improve coding efficiency by combining feature adjustment factors. By combining network structure design and auxiliary bitstream (used to carry feature adjustment factors), the neural network can effectively ensure the quality of reconstructed image blocks while maintaining low complexity, thereby improving coding performance and reducing complexity.
[0110] Example 3: For the processing procedures at the encoding end in Examples 1 and 2, please refer to... Figure 4 As shown, of course, Figure 4 This is just one example of the processing procedure at the encoding end, and no restrictions are imposed on this processing procedure.
[0111] After obtaining the current image block x (which can be the original image block x, i.e., the input image block), the encoding end can perform analysis and transformation on the current image block x through an analysis and transformation network (i.e., a neural network) to obtain the image features y corresponding to the current image block x. Specifically, performing feature transformation on the current image block x through the analysis and transformation network means transforming the current image block x to image features y in the latent domain, thereby facilitating all subsequent processes to be performed in the latent domain.
[0112] An image can be divided into one image block or multiple image blocks. If the image is divided into one image block, then the current image block x can also be an image. That is, the encoding and decoding process for image blocks can also be directly applied to the image.
[0113] After obtaining image features y, the encoder performs a coefficient hyperparameter feature transformation on y to obtain coefficient hyperparameter features z. For example, image features y can be input into a hyperparameter coding network (i.e., a neural network), which then performs the coefficient hyperparameter feature transformation on y to obtain coefficient hyperparameter features z. The hyperparameter coding network can be a trained neural network, and its training process is not restricted, as long as it can perform the coefficient hyperparameter feature transformation on image features y. The latent domain image features y, after passing through the hyperparameter coding network, yield the hyper-prior latent information z.
[0114] After obtaining the coefficient hyperparameter feature z, the encoder can quantize the coefficient hyperparameter feature z to obtain the hyperparameter quantized feature corresponding to the coefficient hyperparameter feature z, i.e. Figure 4 The Q-operation in the code represents the quantization process. After obtaining the hyperparameter quantization features corresponding to the coefficient hyperparameter features z, these features are encoded to obtain Bitstream#1 (i.e., the first bitstream) corresponding to the current image patch. Figure 4 The AE operation in the code represents the encoding process, such as entropy encoding. Alternatively, the encoder can directly encode the coefficient hyperparameter feature z to obtain the Bitstream#1 corresponding to the current image patch. The hyperparameter quantization feature or coefficient hyperparameter feature z carried in Bitstream#1 is mainly used to obtain the parameters of the mean and probability distribution model.
[0115] After obtaining the Bitstream#1 corresponding to the current image block, the encoding end can send the Bitstream#1 corresponding to the current image block to the decoding end. For the processing of the Bitstream#1 corresponding to the current image block by the decoding end, please refer to the following embodiments.
[0116] After obtaining Bitstream#1 corresponding to the current image block, the encoding end can also decode Bitstream#1 to obtain the hyperparameter quantization features, i.e. Figure 4 In the diagram, AD represents the decoding process. Then, the hyperparameter quantization features are dequantized to obtain the coefficient hyperparameter features z_hat. The coefficient hyperparameter features z_hat and the coefficient hyperparameter features z can be the same or different. Figure 4 The IQ operation in the code is an inverse quantization process. Alternatively, after obtaining Bitstream#1 corresponding to the current image block, the encoder can decode Bitstream#1 to obtain the coefficient hyperparameter feature z_hat without involving the inverse quantization process of the coefficient hyperparameter feature z_hat.
[0117] For the encoding process of Bitstream#1, a fixed probability density model encoding method can be used, and for the decoding process of Bitstream#1, a fixed probability density model decoding method can be used. There are no restrictions on the encoding and decoding processes.
[0118] After obtaining the coefficient hyperparameter feature z_hat, the encoder can perform context-based prediction based on the coefficient hyperparameter feature z_hat of the current image patch and the residual feature y_hat of the previous image patch (the determination process of the residual feature y_hat is described in subsequent embodiments) to obtain the predicted value mu (i.e., the mean mu) corresponding to the current image patch. For example, the coefficient hyperparameter feature z_hat and the residual feature y_hat can be input into the mean prediction network, which determines the predicted value mu based on the coefficient hyperparameter feature z_hat and the residual feature y_hat. This prediction process is not restricted. Specifically, for the context-based prediction process, the input includes the coefficient hyperparameter feature z_hat and the decoded residual feature y_hat. The two are jointly input to obtain a more accurate predicted value mu. The predicted value mu is used to subtract the original feature to obtain the residual and add it to the decoded residual to obtain the reconstructed y.
[0119] It is important to note that the mean prediction network is an optional neural network; that is, it is not necessary to use a mean prediction network to determine the predicted value mu. Figure 4 The dashed box in the image indicates that the mean prediction network is optional.
[0120] After obtaining image features y, the encoder can determine residual features r based on image features y and predicted values mu, such as using the difference between image features y and predicted values mu as residual features r. Then, feature processing is performed on the residual features r to obtain image features s. This feature processing method is not restricted and can be any method. In this case, a mean prediction network needs to be deployed to provide the predicted values mu. Alternatively, after obtaining image features y, the encoder can perform feature processing on image features y to obtain image features s. This feature processing method is not restricted and can be any method. In this case, a mean prediction network is not needed, and the residual process is indicated by a dashed box as optional.
[0121] After obtaining image features s, the encoding end can quantize image features s to obtain the quantized image features corresponding to image features s, i.e. Figure 4 The Q operation in the code represents the quantization process. After obtaining the quantized image features corresponding to image features s, the encoder can encode these quantized features to obtain Bitstream#2 (i.e., the second bitstream) corresponding to the current image block. Figure 4The AE operation in the code represents the encoding process, such as entropy encoding. Alternatively, the encoding end can directly encode the image features s to obtain the Bitstream#2 corresponding to the current image block, without involving the quantization process of image features s.
[0122] After obtaining the Bitstream#2 corresponding to the current image block, the encoding end can send the Bitstream#2 corresponding to the current image block to the decoding end. For the processing of the Bitstream#2 corresponding to the current image block by the decoding end, please refer to the following embodiments.
[0123] After obtaining Bitstream#2 corresponding to the current image block, the encoding end can also decode Bitstream#2 to obtain the image quantization features, i.e. Figure 4 In this context, AD represents the decoding process. Then, the encoding end can perform inverse quantization on the image quantization features to obtain image features s'. Image features s' can be the same as or different from image features s. Figure 4 The IQ operation in the code is the dequantization process. Alternatively, after obtaining Bitstream#2 corresponding to the current image block, the encoder can decode Bitstream#2 to obtain image features s' without involving the dequantization process of image quantization features.
[0124] After obtaining image features s', the encoder can perform feature recovery (the inverse process of feature processing) on s'. This feature recovery process is not restricted and can be any method, resulting in residual features r_hat. Residual features r_hat and r can be the same or different. After obtaining residual features r_hat, the encoder determines image features y_hat based on residual features r_hat and predicted values mu. Image features y_hat and y can be the same or different; for example, the sum of residual features r_hat and predicted values mu can be used as image features y_hat. In this case, a mean prediction network needs to be deployed to provide the predicted values mu. Alternatively, after obtaining image features s', the encoder can perform feature recovery (the inverse process of feature processing) on s' to obtain image features y_hat. Image features y_hat and y can be the same or different. In this case, a mean prediction network is not needed, and the residual process is indicated by a dashed box as optional.
[0125] After obtaining the image feature y_hat, the encoder can perform a synthetic transformation on the image feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the image feature y_hat can be input into the synthetic transformation network, which will perform a synthetic transformation on the image feature y_hat to obtain the reconstructed image block x_hat. Thus, the image reconstruction process is completed.
[0126] In one possible implementation, when the encoding end encodes the image quantization features or image features s to obtain Bitstream#2 corresponding to the current image block, the encoding end needs to first determine the probability distribution model, and then encode the image quantization features or image features s based on the probability distribution model. Furthermore, when the encoding end decodes Bitstream#2, it also needs to first determine the probability distribution model, and then decode Bitstream#2 based on the probability distribution model.
[0127] To obtain the probability distribution model, please refer to [link / reference]. Figure 4 As shown, after obtaining the coefficient hyperparameter feature z_hat, the encoder can perform an inverse hyperparameter feature transformation on z_hat to obtain the probability distribution parameter p. For example, the coefficient hyperparameter feature z_hat can be input into a probabilistic hyperparameter decoding network, which will then perform an inverse hyperparameter feature transformation on z_hat to obtain the probability distribution parameter p. After obtaining the probability distribution parameter p, a probability distribution model can be generated based on p. The probabilistic hyperparameter decoding network can be a trained neural network; the training process of this network is not restricted, as long as it can perform the inverse hyperparameter feature transformation on z_hat.
[0128] In one possible implementation, the above-mentioned encoding process can be executed by a deep learning model or a neural network model to achieve end-to-end image compression and encoding, without any restrictions on the encoding process.
[0129] Example 4: For the processing procedures at the decoding end in Examples 1 and 2, please refer to... Figure 5 As shown, of course, Figure 5 This is just one example of the processing procedure at the decoding end, and no restrictions are imposed on the processing procedure at the decoding end.
[0130] After obtaining Bitstream#1 corresponding to the current image block, the decoding end can further decode Bitstream#1 to obtain the hyperparameter quantization features, i.e. Figure 5 In the diagram, AD represents the decoding process. Then, the hyperparameter quantization features are dequantized to obtain the coefficient hyperparameter features z_hat. The coefficient hyperparameter features z_hat and the coefficient hyperparameter features z can be the same or different. Figure 5 The IQ operation in the code is an inverse quantization process. Alternatively, after obtaining Bitstream#1 corresponding to the current image block, the decoding end can also decode Bitstream#1 to obtain the coefficient hyperparameter feature z_hat, without involving the inverse quantization process of the coefficient hyperparameter feature z_hat.
[0131] For the decoding process of Bitstream#1, a decoding method with a fixed probability density model can be used, and there are no restrictions on this.
[0132] An image can be divided into one image block or multiple image blocks. If the image is divided into one image block, then the current image block x can also be an image. That is, the decoding process for the image block can also be directly applied to the image.
[0133] After obtaining the coefficient hyperparameter feature z_hat, the decoder can perform context-based prediction based on the coefficient hyperparameter feature z_hat of the current image patch and the residual feature y_hat of the previous image patch (the determination process of the residual feature y_hat is described in subsequent embodiments) to obtain the predicted value mu (i.e., the mean mu) corresponding to the current image patch. For example, the coefficient hyperparameter feature z_hat and the residual feature y_hat can be input into the mean prediction network, which determines the predicted value mu based on the coefficient hyperparameter feature z_hat and the residual feature y_hat. This prediction process is not restricted. Specifically, for the context-based prediction process, the input includes the coefficient hyperparameter feature z_hat and the decoded residual feature y_hat; the combined input of these two features yields a more accurate predicted value mu.
[0134] It is important to note that the mean prediction network is an optional neural network; that is, it is not necessary to use a mean prediction network to determine the predicted value mu. Figure 5 The dashed box in the image indicates that the mean prediction network is optional.
[0135] After obtaining Bitstream#2 corresponding to the current image block, the decoding end can further decode Bitstream#2 to obtain the image quantization features, i.e. Figure 5 In this context, AD represents the decoding process. Then, the decoding end can perform inverse quantization on the image quantization features to obtain image features s'. Image features s' can be the same as or different from image features s. Figure 5 The IQ operation in the code is an inverse quantization process. Alternatively, after obtaining Bitstream#2 corresponding to the current image block, the decoding end can also decode Bitstream#2 to obtain image features s' without involving the inverse quantization process of image quantization features.
[0136] After obtaining image features s', the decoder can perform feature recovery (the inverse process of feature processing) on image features s' to obtain residual features r_hat. Residual features r_hat may be the same as or different from residual features r. After obtaining residual features r_hat, the decoder determines image features y_hat based on residual features r_hat and predicted values mu. Image features y_hat may be the same as or different from image features y. For example, the sum of residual features r_hat and predicted values mu can be used as image features y_hat. In this case, a mean prediction network needs to be deployed to provide the predicted values mu. Alternatively, after obtaining image features s', the decoder can perform feature recovery on image features s' to obtain image features y_hat. Image features y_hat may be the same as or different from image features y. In this case, a mean prediction network is not needed, and the residual process is indicated by a dashed box as an optional process.
[0137] After obtaining the image feature y_hat, the decoding end can perform a synthetic transformation on the image feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the image feature y_hat can be input into the synthetic transformation network, and the synthetic transformation network can perform a synthetic transformation on the image feature y_hat to obtain the reconstructed image block x_hat. Thus, the image reconstruction process is completed.
[0138] In one possible implementation, when decoding Bitstream#2, the decoding end needs to first determine the probability distribution model, and then decode Bitstream#2 based on that probability distribution model. To obtain the probability distribution model, see [link to documentation]. Figure 5 As shown, after obtaining the hyperparameter feature z_hat, the decoder can perform an inverse hyperparameter feature transformation on z_hat to obtain the probability distribution parameter p. For example, the hyperparameter feature z_hat can be input into a probabilistic hyperparameter decoding network, which will then perform an inverse hyperparameter feature transformation on z_hat to obtain the probability distribution parameter p. After obtaining the probability distribution parameter p, a probability distribution model can be generated based on p. The probabilistic hyperparameter decoding network can be a trained neural network; the training process of this network is not restricted, as long as it can perform an inverse hyperparameter feature transformation on z_hat to obtain the probability distribution parameter p.
[0139] In one possible implementation, the above-mentioned decoding process can be executed by a deep learning model or a neural network model to achieve end-to-end image compression and encoding, without limiting the decoding process.
[0140] Example 5: For Examples 1, 2, 3, and 4, an analytical transform network is involved at the encoding end, and a synthetic transform network is involved at both the encoding and decoding ends. Both the analytical transform network and the synthetic transform network are neural networks. In one possible implementation, a schematic diagram of the analytical transform network can be found [link to schematic diagram]. Figure 6A As shown, a schematic diagram of the synthetic transformation network can be found in [reference needed]. Figure 6B As shown, of course, Figure 6A and Figure 6B This is merely an example; the structure of the analytical transform network and the synthetic transform network in this embodiment is not limited, as long as they can achieve the analytical transform function and the synthetic transform function.
[0141] The analysis transform network consists of, in sequence, a Padding layer, a Conv layer, a ResAU layer, a Padding layer, an RNAB layer, a Conv layer, a ResAU layer, a Padding layer, a Conv layer, a ResAU layer, a Padding layer, a Conv layer, and another Conv layer. The Padding layer is used to expand the edges of the features. The Conv layer is a convolutional layer used to perform convolution operations on the features. C*3*3 indicates that a 3*3 convolutional kernel (C channels) is used for the feature convolution operation, and C*1*1 indicates that a 1*1 convolutional kernel (C channels) is used for the feature convolution operation. The down arrow 2 indicates that the features are downsampled by a factor of 2.
[0142] The ResAU layer is the activation layer. A schematic diagram of the ResAU layer can be found in [reference needed]. Figure 6C As shown, LeakyReLU represents the activation operation, Conv 1*1 represents the convolution operation of features using a 1*1 convolution kernel, and tanh represents the hyperbolic tangent operation.
[0143] The RNAB layer is a residual nonlocal attention block. A structural diagram of the RNAB layer can be found in [reference needed]. Figure 6D As shown, the RNAB layer can include an RB layer, a Conv layer, and a sigmoid layer. In the Conv layer, the up arrow indicating a 2x upsampling of the features. The sigmoid layer is used for feature mapping, such as using the sigmoid function to map features to a range between 0 and 1.
[0144] The RB (Residual Blocks) layer consists of residual blocks. A structural diagram of the RB layer can be found in [reference needed]. Figure 6E As shown, the RB layer can include a Conv layer, LeakyReLU, and a Conv layer, with a 3*3 convolution kernel used in the Conv layer.
[0145] The synthetic transformation network consists of the following layers in sequence: ResBlock layer (RB layer), ResBlock layer, Conv layer, Cropping layer, ResAU layer, Conv layer, Cropping layer, ResAU layer, Conv layer, RNAB layer, Cropping layer, ResAU layer, Conv layer, and Cropping layer. The ResBlock layer is a residual block; a schematic diagram of the ResBlock layer can be found in [reference needed]. Figure 6E As shown. The Conv layer is a convolutional layer used to perform convolution operations on features. C*3*3 indicates that a 3*3 convolutional kernel (C channels) is used to perform the convolution operation on features. The upward arrow 2 indicates that the features are upsampled by a factor of 2. The Cropping layer is a clipping layer used to perform feature clipping operations, which is the inverse operation of the Padding operation. The ResAU layer is an activation layer. A schematic diagram of the ResAU layer can be found in [reference needed]. Figure 6C As shown. The RNAB layer is a residual nonlocal attention block. A schematic diagram of the RNAB layer structure can be found in [reference needed]. Figure 6D As shown, the RNAB layer may include the RB layer, the Conv layer, and the sigmoid layer.
[0146] Example 6: Regarding Example 5, from Figure 6B As can be seen, the synthesis transformation network mainly includes four upsampling 3*3 convolutional layers, two ResBlock layers (each ResBlock layer includes two 3*3 convolutional layers), three activation layers (ResAU layers, each ResAU layer includes one 1*1 convolutional layer), and one RNAB layer (the RNAB layer includes nine RB blocks and three 3*3 convolutional layers, totaling 21 3*3 convolutional layers). In summary, the synthesis transformation network has a total of 29 3*3 convolutional layers and three 1*1 convolutional layers, with the RNAB layer accounting for nearly 70% of the complexity, indicating significant room for optimization in its complexity. Therefore, in this embodiment, the synthesis transformation network can be optimized to design a low-complexity, high-performance synthesis transformation network. Alternatively, the analysis transformation network can also be optimized to design a low-complexity, high-performance analysis transformation network. Of course, other high-complexity network structures can also be optimized to design low-complexity, high-performance network structures; this embodiment does not impose any limitations on this. For ease of description, this embodiment uses the optimization of the synthesis transformation network as an example. For instance, complex networks similar to RNAB layers in the synthesis transformation network can be optimized. Of course, other network aspects in the synthesis transformation network can also be optimized; there are no restrictions on this, as long as the complexity of the synthesis transformation network can be simplified.
[0147] In one possible implementation, for Figure 6DThe RNAB layer shown can consist of 9 RB layers, 3 Conv layers, and 1 sigmoid layer. Based on this, some network layers can be removed to obtain an optimized RNAB layer. See [link to documentation]. Figure 6F As shown, three RB layers and two Conv layers can be removed to obtain the optimized RNAB layer. Of course, other network layers can also be removed to obtain the optimized RNAB layer; there are no restrictions on the method of removing these network layers.
[0148] For example, the optimized RNAB layer can be seen in [reference needed]. Figure 6G As shown, the RNAB layer may include an RB layer (composed of at least one RB block, such as three consecutive RB blocks, of course, the number of RB blocks may be more or less), an RB layer (composed of at least one RB block, such as three consecutive RB blocks, of course, the number of RB blocks may be more or less), a Conv layer (such as a 3*3 Conv layer) and a sigmoid layer (i.e., a feature mapping layer).
[0149] For example, a 3x3 Conv layer can be replaced with a 1x1 Conv layer or a 5x5 Conv layer. Taking a 1x1 Conv layer as an example, the optimized RNAB layer can be found in [reference needed]. Figure 6H As shown, the RNAB layer may include an RB layer, an RB layer, a Conv layer (such as a 1*1 Conv layer), and a sigmoid layer (i.e., a feature mapping layer).
[0150] For example, after obtaining the optimized RNAB layer, the optimized RNAB layer can be substituted into... Figure 6B The optimized synthetic transformation network is obtained. Based on the optimized synthetic transformation network, the image feature y_hat can be input into the synthetic transformation network, and the synthetic transformation network performs synthetic transformation on the image feature y_hat to obtain the reconstructed image patch x_hat.
[0151] Example 7: Based on Example 5, the synthetic transformation network can be optimized to design a low-complexity, high-performance synthetic transformation network. Alternatively, the analytical transformation network can also be optimized to design a low-complexity, high-performance analytical transformation network. Of course, other high-complexity network structures can also be optimized to design low-complexity, high-performance network structures. In this example, a method of "simplified network structure + adaptive adjustment factor" can be used for encoding and decoding.
[0152] First, the complex network is simplified to obtain a simplified network. For example, the synthesis transformation network can be simplified to obtain a simplified synthesis transformation network. For instance, the RNAB layer (i.e., the complex network) in the synthesis transformation network can be simplified to obtain a simplified synthesis transformation network. Of course, other networks can also be simplified, and there are no restrictions on this.
[0153] Then, adjustments are made to at least one layer of features (such as key features) of the simplified network. For example, the simplified network is trained on a large number of images, and the features it generates have general effectiveness. However, for a specific frame of an image, its features are often not optimal. Therefore, customized features need to be designed for the image. To obtain these customized features, a feature adjustment factor needs to be introduced at some point in the network process. The features are then adjusted based on the feature adjustment factor to obtain the adjusted customized features. For example, the encoder adaptively selects the feature adjustment factor for the current image patch and encodes the feature adjustment factor in the bitstream, while the decoder can parse the feature adjustment factor from the bitstream.
[0154] In one possible implementation, for the decoding end, the decoding method can be found in [reference needed]. Figure 7A As shown, the decoder can receive the main bitstream (i.e., the first bitstream Bitstream #1 and the second bitstream Bitstream #2) for the current image patch. After passing through coefficient decoding and the decoding sub-network, the main bitstream yields the first feature corresponding to the current image patch. The decoder can decode the feature adjustment factor corresponding to the current image patch from the auxiliary bitstream. After obtaining the first feature corresponding to the current image patch, the feature enhancement network can enhance the first feature to obtain the second feature, and then adjust the second feature using the feature adjustment factor corresponding to the current image patch to obtain the third feature. After obtaining the third feature, the target feature corresponding to the current image patch can be determined based on the third feature, and this target feature is input into the reconstruction decoding network. The reconstruction decoding network processes the target feature to obtain the reconstructed image patch corresponding to the current image patch.
[0155] For example, the feature adjustment factor can be decoded from the auxiliary bitstream by the decoding end, or it can be a fixed parameter value. When the feature adjustment factor is a fixed parameter value, such as when the feature adjustment factor is a fixed parameter value of 1, it is equivalent to not adjusting the second feature, that is, not involving the feature adjustment process, which is equivalent to the scheme of embodiment 6.
[0156] For example, the decoder can also decode a feature adjustment factor for the decoding sub-network (which may be the same as or different from the feature adjustment factor for the second feature) from the auxiliary bitstream. After obtaining the feature adjustment factor for the decoding sub-network, a feature in the decoding sub-network can be adjusted based on this feature adjustment factor. For instance, after obtaining feature A (feature A is any feature in the decoding sub-network), the decoding sub-network can use this feature adjustment factor to adjust feature A, resulting in the adjusted feature B. The decoding sub-network continues processing based on feature B to finally obtain the first feature. By using the feature adjustment factor in certain decoding processes of the decoding sub-network, a first feature with less distortion can be generated.
[0157] In one possible implementation, the encoding method for the encoding end can be found in [reference needed]. Figure 7B As shown, after the current image block is encoded by the encoding network and coefficients, the main bitstream corresponding to the current image block (i.e., the first bitstream Bitstream#1 and the second bitstream Bitstream#2) can be obtained. The encoding end can send the main bitstream of the current image block to the decoding end.
[0158] After the main bitstream passes through coefficient decoding and the decoding sub-network, the first feature corresponding to the current image block can be obtained. Based on the first feature, the feature adjustment factor corresponding to the current image block can be obtained and encoded in the auxiliary bitstream corresponding to the current image block, so that the decoding end can decode the feature adjustment factor corresponding to the current image block from the auxiliary bitstream.
[0159] For example, the encoder can obtain at least one candidate feature adjustment factor (for instance, a list of candidate feature adjustment factors can be pre-constructed, and all feature adjustment factors in the list can be used as candidate feature adjustment factors; or, an algorithm can be used to generate at least one candidate feature adjustment factor, without limitation). The rate-distortion cost corresponding to each candidate feature adjustment factor can be determined based on a first feature. For example, for each candidate feature adjustment factor, after obtaining the first feature corresponding to the current image patch, the first feature can be enhanced using a feature enhancement network to obtain a second feature, and then the second feature can be adjusted using the candidate feature adjustment factor to obtain a third feature. After obtaining the third feature, the target feature corresponding to the current image patch can be determined based on the third feature, and this target feature can be input into a reconstruction decoding network. The reconstruction decoding network processes the target feature to obtain the reconstructed image patch corresponding to the current image patch. After obtaining the reconstructed image patch, the rate-distortion cost corresponding to the candidate feature adjustment factor can be determined based on the loss between the reconstructed image patch and the current image patch, thus obtaining the rate-distortion cost corresponding to each candidate feature adjustment factor. After obtaining the rate-distortion cost corresponding to each candidate feature adjustment factor, based on the rate-distortion cost corresponding to each candidate feature adjustment factor, select one candidate feature adjustment factor from all candidate feature adjustment factors as the feature adjustment factor corresponding to the current image patch, such as selecting the candidate feature adjustment factor with the smallest rate-distortion cost as the feature adjustment factor corresponding to the current image patch.
[0160] For example, the encoder can also obtain a feature adjustment factor for the decoding sub-network and perform feature adjustment on a certain feature in the decoding sub-network based on this feature adjustment factor. For instance, after obtaining feature A (feature A is any feature in the decoding sub-network), the decoding sub-network can use this feature adjustment factor to adjust feature A, obtaining the adjusted feature B. The decoding sub-network continues processing based on feature B to finally obtain the first feature. Furthermore, the encoder can also encode the feature adjustment factor for the decoding sub-network (which may be the same as or different from the feature adjustment factor for the second feature) in the auxiliary bitstream, so that the decoder can decode the feature adjustment factor for the decoding sub-network from the auxiliary bitstream.
[0161] Example 8: For Examples 1-7, for the encoding end and the decoding end, the first feature corresponding to the current image block can be obtained through the first neural network. For example, the probability distribution parameters are obtained based on the first bitstream corresponding to the current image block, and the probability distribution model is determined based on the probability distribution parameters. The second bitstream corresponding to the current image block is decoded based on the probability distribution model to obtain the decoded image features. The first feature corresponding to the current image block is determined based on the decoded image features.
[0162] For example, the first neural network can be a decoding sub-network. The following explanation uses a decoding sub-network as an example. Figure 8A As shown, the network within the dashed box is the decoding sub-network, and y_hat is the first feature output by the decoding sub-network.
[0163] For example, after obtaining Bitstream#1 corresponding to the current image block, the encoding or decoding end can decode Bitstream#1 to obtain the hyperparameter quantization feature, and then dequantize the hyperparameter quantization feature to obtain the coefficient hyperparameter feature z_hat. Alternatively, after obtaining Bitstream#1 corresponding to the current image block, the encoding or decoding end can also decode Bitstream#1 to obtain the coefficient hyperparameter feature z_hat without involving the dequantization process of the coefficient hyperparameter feature z_hat.
[0164] After obtaining the hyperparameter features z_hat, the encoder or decoder can perform an inverse hyperparameter feature transform on z_hat to obtain the probability distribution parameters p. For example, the hyperparameter features z_hat can be input into a probabilistic hyperparameter decoding network, which will then perform an inverse hyperparameter feature transform on z_hat to obtain the probability distribution parameters p. After obtaining the probability distribution parameters p, a probability distribution model can be generated based on p.
[0165] After obtaining Bitstream#2 corresponding to the current image block, the encoding or decoding end can decode Bitstream#2 to obtain image quantization features, and then dequantize these features to obtain image features s'. Alternatively, the encoding or decoding end can decode Bitstream#2 to obtain image features s' without involving the dequantization process of the image quantization features. For example, when decoding Bitstream#2, the encoding or decoding end can decode it based on a probability distribution model.
[0166] After obtaining the image feature s', the encoder or decoder can perform feature recovery on the image feature s' to obtain the image feature y_hat, which can be used as the first feature, i.e., the first feature y_hat is output by the decoding sub-network.
[0167] Example 9: For Examples 1-7, for the encoding and decoding ends, the first feature corresponding to the current image block can be obtained through the first neural network. For example, the probability distribution parameters and predicted values (such as the mean) are obtained based on the first bitstream corresponding to the current image block; the probability distribution model is determined based on the probability distribution parameters, and the second bitstream corresponding to the current image block is decoded based on the probability distribution model to obtain the decoded image features; residual recovery is performed on the decoded image features to obtain residual features; based on the residual features and the predicted value, the first feature corresponding to the current image block is determined.
[0168] For example, the first neural network can be a decoding sub-network. The following explanation uses a decoding sub-network as an example. Figure 8B As shown, the network within the dashed box is the decoding sub-network, and y_hat is the first feature output by the decoding sub-network.
[0169] For example, after obtaining Bitstream#1 corresponding to the current image block, the encoding or decoding end can decode Bitstream#1 to obtain the hyperparameter quantization feature, and then dequantize the hyperparameter quantization feature to obtain the coefficient hyperparameter feature z_hat. Alternatively, after obtaining Bitstream#1 corresponding to the current image block, the encoding or decoding end can also decode Bitstream#1 to obtain the coefficient hyperparameter feature z_hat without involving the dequantization process of the coefficient hyperparameter feature z_hat.
[0170] After obtaining the hyperparameter features z_hat, the encoder or decoder can perform an inverse hyperparameter feature transform on z_hat to obtain the probability distribution parameters p. For example, the hyperparameter features z_hat can be input into a probabilistic hyperparameter decoding network, which will then perform an inverse hyperparameter feature transform on z_hat to obtain the probability distribution parameters p. After obtaining the probability distribution parameters p, a probability distribution model can be generated based on p.
[0171] After obtaining the hyperparameter feature z_hat, the encoder or decoder can perform context-based prediction based on the hyperparameter feature z_hat of the current image patch and the residual feature y_hat of the previous image patch to obtain the predicted value mu (i.e., the mean mu) for the current image patch. For example, the hyperparameter feature z_hat and the residual feature y_hat can be input into the mean prediction network, which then determines the predicted value mu based on these two features. This prediction process is not restricted. Specifically, for the context-based prediction process, the input to the mean prediction network can include the hyperparameter feature z_hat and the decoded residual feature y_hat; combining these two inputs yields a more accurate predicted value mu.
[0172] After obtaining Bitstream#2 corresponding to the current image block, the encoding or decoding end can decode Bitstream#2 to obtain image quantization features, and then dequantize these features to obtain image features s'. Alternatively, the encoding or decoding end can decode Bitstream#2 to obtain image features s' without involving the dequantization process of the image quantization features. For example, when decoding Bitstream#2, the encoding or decoding end can decode it based on a probability distribution model.
[0173] After obtaining image features s', the encoder or decoder performs feature recovery (i.e., residual recovery, the inverse process of residual processing) on image features s' to obtain residual features r_hat. After obtaining residual features r_hat, the encoder or decoder determines image features y_hat based on residual features r_hat and predicted values mu. For example, the sum of residual features r_hat and predicted values mu can be used as image features y_hat. Image features y_hat can be used as the first feature, that is, the decoding sub-network outputs the first feature y_hat.
[0174] For example, Figure 8B and Figure 8A In contrast, in the decoding sub-network, besides obtaining the probability distribution parameters for the second bitstream based on the first bitstream, a predicted value mu for the first feature can also be generated based on the first bitstream. The feature residual is obtained by decoding the second bitstream, and then the residual r_hat for the first feature is obtained through residual recovery. The first feature can be obtained based on mu and r_hat.
[0175] Example 10: Referring to Examples 1-9, for both the encoder and decoder ends, after obtaining the first feature y_hat, a synthetic transformation network can be used to perform a synthetic transformation on the first feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x. During the synthetic transformation process, feature enhancement can be performed on the first feature y_hat to obtain the second feature; feature adjustment can be performed on the second feature based on a feature adjustment factor to obtain the third feature; after obtaining the third feature, the target feature is determined based on the third feature. Then, the reconstructed image block x_hat corresponding to the current image block x is obtained based on the target feature.
[0176] See Figure 6BThe diagram shown illustrates the structure of a synthetic transform network. This network can sequentially include a ResBlock layer (RB layer), a ResBlock layer, a Conv layer, a Cropping layer, a ResAU layer, a Conv layer, a Cropping layer, a ResAU layer, a Conv layer, an RNAB layer, a Cropping layer, a ResAU layer, a Conv layer, and a Cropping layer. A schematic diagram of the RB layer can be found in [reference needed]. Figure 6E As shown, a schematic diagram of the ResAU layer structure can be found in [reference needed]. Figure 6C As shown.
[0177] The RNAB layer is a residual nonlocal attention block. A structural diagram of the RNAB layer can be found in [reference needed]. Figure 6D As shown, or, can be removed. Figure 6D From a portion of the network layers, the optimized RNAB layer was obtained; see [link to documentation]. Figure 6F As shown, three RB layers and two Conv layers can be removed to obtain the optimized RNAB layer. Therefore, a schematic diagram of the RNAB layer structure can be found in [reference needed]. Figure 6G As shown, or, a schematic diagram of the RNAB layer structure can be found in [reference needed]. Figure 6H As shown. In Figure 6G and Figure 6H In this model, the RNAB layer can include an RB layer (composed of at least one RB block, such as three consecutive RB blocks), an RB layer (composed of at least one RB block, such as three consecutive RB blocks), a Conv layer, and a sigmoid layer (i.e., a feature mapping layer). The Conv layer can be a 3x3 Conv layer, a 5x5 Conv layer, a 1x1 Conv layer, or other sizes such as a 7x7 Conv layer or a 9x9 Conv layer; there are no restrictions on the size of the Conv layer.
[0178] The synthetic transformation network described above can be divided into a feature enhancement subnetwork and a reconstruction / decoding subnetwork. The feature enhancement subnetwork may include an initial enhancement subnetwork, a first upsampling convolutional subnetwork, and an attention subnetwork. The reconstruction / decoding subnetwork may include a second upsampling convolutional subnetwork and a color space transformation subnetwork.
[0179] For example, all networks preceding the RNAB layer in the synthetic transformation network can be used as the initial enhancement subnetwork and the first upsampling convolutional subnetwork. The initial enhancement subnetwork may include two RB networks, and the first upsampling convolutional subnetwork may include at least one upsampling convolutional network (e.g., three upsampling convolutional networks). See, for example, [link to documentation]. Figure 6BAs shown, all the networks before the RNAB layer are, in sequence, RB layer, RB layer, Conv layer, Cropping layer, ResAU layer, Conv layer, Cropping layer, ResAU layer, and Conv layer. Based on this, the initial enhancement sub-network can include RB layer and RB layer, and the first upsampling convolutional sub-network can include Conv layer, Cropping layer, ResAU layer, Conv layer, Cropping layer, ResAU layer, and Conv layer, which involves 3 upsampling convolutional networks (Conv layers). Here, we take a 3*3 Conv layer as an example.
[0180] For example, the RNAB layer in the synthetic transformation network can be used as an attention subnetwork. The attention subnetwork effectively reduces complexity by removing the upsampling and downsampling convolutional networks. See [link to relevant documentation]. Figure 6G and Figure 6H The diagram shows the structure of the attention subnetwork, which may include an RB layer (composed of at least one RB block), an RB layer (composed of at least one RB block), a Conv layer, and a sigmoid layer (i.e., a feature mapping layer).
[0181] The attention subnetwork can include a residual skip subnetwork, a residual enhancement subnetwork, and a weight generation subnetwork. For the residual skip subnetwork, x_out0 = x_in, where x_in is the input feature of the residual skip subnetwork, and x_out0 is the output feature. For the residual enhancement subnetwork, x_out1 = RB(x_in), where x_in is the input feature of the residual enhancement subnetwork, and x_out1 is the output feature. The weight generation subnetwork is used to obtain the weight feature k of the output feature x_out1 of the residual enhancement subnetwork. Therefore, the final output of the attention subnetwork is: x_out = x_in + k * x_out1.
[0182] See Figure 9A and Figure 9BAs shown, the residual skip subnetwork is used to superimpose the input feature x_in (i.e., x_out0) onto the final output feature of the attention subnetwork. The residual enhancement subnetwork includes at least one RB block (such as one RB block or three RB blocks, etc.), and is used to process the input feature x_in based on the RB blocks to obtain the output feature x_out1. The weight generation subnetwork may include at least one RB block (such as one or three RB blocks), one Conv layer (such as a 3*3 Conv layer, a 5*5 Conv layer, a 1*1 Conv layer, etc.), and a feature mapping layer (such as using a sigmoid process to implement feature mapping). The weight generation subnetwork is used to process the input feature x_in based on the RB block, and then the Conv layer is used to perform convolution processing on the feature after the RB block is processed. Then, the feature mapping layer is used to perform feature mapping on the feature after the Conv layer is processed to obtain the weight feature k. The dimension of the weight feature k is the same as the dimension of the output feature x_out1 of the residual enhancement subnetwork. After multiplying the output feature x_out1 with the weight feature k, it is added to the output feature x_out0 of the residual skip subnetwork to obtain the final output feature of the attention subnetwork: x_out = x_in + k * x_out1.
[0183] For example, all networks after the RNAB layer in the synthesis transform network can be used as reconstruction decoding subnetworks. These reconstruction decoding subnetworks can include a second upsampling convolutional subnetwork and a color space transformation subnetwork, meaning all networks after the RNAB layer can be used as the second upsampling convolutional subnetwork and the color space transformation subnetwork. The second upsampling convolutional subnetwork can include at least one upsampling convolutional network (e.g., a single upsampling convolutional network), and the color space transformation subnetwork is used to implement the image transformation process from the YUV domain to RGB, or it can be a filtering process from the YUV domain to the YUV domain.
[0184] For example, see Figure 6B As shown, all the networks following the RNAB layer are, in sequence, a Cropping layer, a ResAU layer, a Conv layer, and another Cropping layer. Therefore, the second upsampling convolutional sub-network can include a Cropping layer, a ResAU layer, a Conv layer, and another Cropping layer, which involves one upsampling convolutional network (Conv layer). Here, we take a 3*3 Conv layer as an example. If there is a color space transformation requirement, a color space transformation sub-network can be included (…). Figure 6B (Not shown in the image) The image transformation process from YUV domain to RGB is performed by the color space transformation subnetwork, or the filtering process from YUV domain to YUV domain is performed by the color space transformation subnetwork. If there is no need for color space transformation, the color space transformation subnetwork can be omitted.
[0185] After dividing the synthesis transformation network into an initial enhancement sub-network, a first upsampling convolutional sub-network, an attention sub-network, and a reconstruction decoding sub-network (the reconstruction decoding sub-network includes a second upsampling convolutional network and a color space transformation sub-network), the first feature y_hat can be synthesized and transformed based on the initial enhancement sub-network, the first upsampling convolutional network, the attention sub-network, and the reconstruction decoding sub-network to obtain the reconstructed image block x_hat corresponding to the current image block x.
[0186] Example 11: In Example 10, the first feature y_hat can be synthesized and transformed based on the initial enhancement subnetwork, the first upsampling convolution subnetwork, the attention subnetwork, and the reconstruction decoding subnetwork to obtain the reconstructed image block x_hat corresponding to the current image block x. In one possible implementation, a feature adjustment factor can be used to adjust the output features of the feature enhancement subnetwork. The adjustment position of the feature adjustment factor can be found in [reference needed]. Figure 10A As shown, under the adjustment method of this feature adjustment factor, the first feature y_hat can be synthesized and transformed in the following way to obtain the reconstructed image block x_hat.
[0187] An initial enhancement subnetwork is used to enhance the first feature y_hat, resulting in the enhanced feature. For example, the first feature y_hat can be input into the initial enhancement subnetwork, which then enhances it to obtain the enhanced feature. For instance, the initial enhancement subnetwork may include two RB layers; therefore, the first feature y_hat can be enhanced using two RB layers to obtain the enhanced feature. There are no restrictions on this process.
[0188] After obtaining the enhanced features, a first upsampling convolutional sub-network is used to perform upsampling convolution on the enhanced features to obtain the initial features corresponding to the attention sub-network, i.e., the input features x_in of the attention sub-network. For example, the enhanced features (i.e., the output features of the initial enhanced sub-network) can be input into the first upsampling convolutional sub-network, which then performs upsampling convolution on the enhanced features to obtain the initial features corresponding to the attention sub-network. For example, the first upsampling convolutional sub-network may include Conv layers, Cropping layers, ResAU layers, and so on. Upsampling convolution can be performed on the enhanced features through these network layers; this process is not limited.
[0189] After obtaining the initial feature x_in corresponding to the attention sub-network, the attention sub-network can be used to enhance the initial feature x_in to obtain the second feature. For example, the initial feature x_in can be input into the residual skip sub-network, the residual enhancement sub-network, and the weight generation sub-network respectively. After obtaining the initial feature x_in, the residual skip sub-network superimposes the initial feature x_in (i.e., x_out0) onto the final output feature of the attention sub-network. After obtaining the initial feature x_in, the residual enhancement sub-network enhances the initial feature x_in to obtain the enhanced feature. For example, the residual enhancement sub-network enhances the initial feature x_in based on RB blocks to obtain the enhanced feature x_out1. After obtaining the initial feature x_in, the weight generation sub-network generates the weight feature corresponding to the initial feature x_in, that is, the weight feature k of the output feature x_out1 of the residual enhancement sub-network.
[0190] For example, the weight generation subnetwork can include a residual block subnetwork, a convolutional subnetwork, and a feature mapping subnetwork. The residual block network can be used to perform convolutional activation on the initial feature x_in to obtain the convolutionally activated feature. The convolutional subnetwork can then be used to convolve the convolutionally activated feature to obtain the convolutionally generated feature. Finally, the feature mapping subnetwork can be used to map the convolutionally generated feature to obtain the weight feature k. See also Figure 10A As shown, the weight generation subnetwork may include at least one RB block (such as one or three RB blocks, i.e., a residual block subnetwork), one Conv layer (such as a 3*3 Conv layer, a 5*5 Conv layer, a 1*1 Conv layer, etc., i.e., a convolutional subnetwork), and a feature mapping layer (such as using a sigmoid process to implement feature mapping, i.e., a feature mapping subnetwork). Based on this, the initial feature x_in can be convolved and activated based on the RB block to obtain the convolved activated feature. Then, the Conv layer can be used to convolve the convolved activated feature to obtain the convolved feature. Finally, the feature mapping layer can be used to perform feature mapping on the convolved feature to obtain the weight feature k.
[0191] See Figure 10A As shown, after obtaining the enhanced feature x_out1, the weight feature k, and the initial feature x_in (i.e., x_out0), a second feature can be generated based on the initial feature x_in, the enhanced feature x_out1, and the weight feature k. The second feature can be the final output feature of the attention sub-network. For example, multiplying the enhanced feature x_out1 by the weight feature k and then adding it to the output feature x_out0 of the residual skip sub-network, the second feature is: x_out = x_in + k * x_out1.
[0192] After obtaining the second feature, it can be adjusted based on a feature adjustment factor, and the adjusted feature can be used as the third feature. After obtaining the third feature, it can be determined as the target feature.
[0193] After obtaining the target features, the reconstruction decoding subnetwork can process these features to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the target features can be input into the reconstruction decoding subnetwork, which processes them to obtain the reconstructed image block x_hat corresponding to the current image block x. For instance, the reconstruction decoding subnetwork may include a second upsampling convolution subnetwork, which can perform upsampling convolution on the target features to obtain the reconstructed image block x_hat corresponding to the current image block x. Alternatively, the reconstruction decoding subnetwork may include a second upsampling convolution subnetwork and a color space transformation subnetwork. The second upsampling convolution subnetwork can perform upsampling convolution on the target features, and the color space transformation subnetwork can perform color space transformation on the upsampling convolution features to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the second upsampling convolution subnetwork may include Cropping layers, ResAU layers, Conv layers, and Cropping layers; these network layers can be used to perform upsampling convolution on the target features, and this process is not limited. The color space transformation subnetwork is used to perform image transformation from the YUV domain to RGB, or it is used to perform filtering processing from the YUV domain to the YUV domain. No restrictions are placed on this process.
[0194] Example 12: In Example 10, the first feature y_hat can be synthesized and transformed based on the initial enhancement subnetwork, the first upsampling convolutional subnetwork, the attention subnetwork, and the reconstruction decoding subnetwork to obtain the reconstructed image block x_hat corresponding to the current image block x. In one possible implementation, a feature adjustment factor can be used to adjust the residual enhancement output feature of the attention subnetwork. The adjustment position of the feature adjustment factor can be found in [reference needed]. Figure 10B As shown, under the adjustment method of this feature adjustment factor, the first feature y_hat can be synthesized and transformed in the following way to obtain the reconstructed image block x_hat.
[0195] An initial enhancement subnetwork is used to enhance the first feature y_hat, resulting in the enhanced feature. For example, the first feature y_hat can be input into the initial enhancement subnetwork, which then enhances it to obtain the enhanced feature. For instance, the initial enhancement subnetwork may include two RB layers; therefore, the first feature y_hat can be enhanced using two RB layers to obtain the enhanced feature. There are no restrictions on this process.
[0196] After obtaining the enhanced features, a first upsampling convolutional sub-network is used to perform upsampling convolution on the enhanced features to obtain the initial features corresponding to the attention sub-network, i.e., the input features x_in of the attention sub-network. For example, the enhanced features (i.e., the output features of the initial enhanced sub-network) can be input into the first upsampling convolutional sub-network, which then performs upsampling convolution on the enhanced features to obtain the initial features corresponding to the attention sub-network. For example, the first upsampling convolutional sub-network may include Conv layers, Cropping layers, ResAU layers, and so on. Upsampling convolution can be performed on the enhanced features through these network layers; this process is not limited.
[0197] After obtaining the initial feature x_in corresponding to the attention sub-network, the initial feature x_in is input into the residual skip sub-network, the residual enhancement sub-network, and the weight generation sub-network, respectively. The residual skip sub-network, after obtaining the initial feature x_in, superimposes the initial feature x_in (i.e., x_out0) onto the final output feature of the attention sub-network. The residual enhancement sub-network, after obtaining the initial feature x_in, performs feature enhancement on the initial feature x_in to obtain the enhanced feature; for example, the residual enhancement sub-network enhances the initial feature x_in based on RB blocks to obtain the enhanced feature x_out1. The weight generation sub-network, after obtaining the initial feature x_in, generates the weight feature corresponding to the initial feature x_in, which is the weight feature k of the output feature x_out1 of the residual enhancement sub-network.
[0198] For example, the weight generation subnetwork includes a residual block subnetwork, a convolutional subnetwork, and a feature mapping subnetwork. The residual block subnetwork performs convolutional activation on the initial feature x_in to obtain the convolutionally activated feature. The convolutional subnetwork then convolves the convolutionally activated feature to obtain the convolutionally generated feature. Finally, the feature mapping subnetwork maps the convolutionally generated feature to obtain the weight feature k. See also... Figure 10B As shown, the weight generation subnetwork includes at least one RB block (i.e., residual block subnetwork), one Conv layer (i.e., convolutional subnetwork), and a feature mapping layer (i.e., feature mapping subnetwork). Based on this, the initial feature x_in is convolved and activated based on the RB block to obtain the convolved activated feature. The convolved activated feature is then convolved using the Conv layer to obtain the convolved feature. Finally, the convolved feature is mapped using the feature mapping layer to obtain the weight feature k.
[0199] See Figure 10BAs shown, after obtaining the enhanced feature x_out1 and the weight feature k, a second feature can be generated based on the enhanced feature x_out1 and the weight feature k. For example, the enhanced feature x_out1 can be multiplied with the weight feature k, and the feature after multiplication can be used as the second feature. That is, the second feature can be: k*x_out1.
[0200] After obtaining the second feature, it can be adjusted based on the feature adjustment factor, and the adjusted feature can be used as the third feature. See [link to relevant documentation]. Figure 10B As shown, the third feature can be feature rd.
[0201] After obtaining the third feature rd, a fourth feature can be generated based on the initial feature x_in (i.e., the output feature x_out0 of the residual skip sub-network) and the third feature rd. The fourth feature can be the final output feature of the attention sub-network. For example, adding the third feature rd to the output feature x_out0 of the residual skip sub-network yields the fourth feature: x_out = x_in + rd.
[0202] In summary, the initial feature x_in can be enhanced using the first sub-network within the attention sub-network (such as the residual enhancement sub-network and the weight generation sub-network, which can include the residual block sub-network, convolutional sub-network, and feature mapping sub-network) to obtain the second feature. After obtaining the second feature, it can be adjusted based on a feature adjustment factor to obtain the third feature rd. After obtaining the third feature rd, it can be processed using the second sub-network within the attention sub-network (such as the residual skip sub-network) to obtain the fourth feature x_out. Finally, the fourth feature x_out can be designated as the target feature.
[0203] After obtaining the target features, the reconstruction decoding subnetwork can process these features to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the target features can be input into the reconstruction decoding subnetwork, which processes them to obtain the reconstructed image block x_hat corresponding to the current image block x. For instance, the reconstruction decoding subnetwork may include a second upsampling convolution subnetwork, which can perform upsampling convolution on the target features to obtain the reconstructed image block x_hat corresponding to the current image block x. Alternatively, the reconstruction decoding subnetwork may include a second upsampling convolution subnetwork and a color space transformation subnetwork. The second upsampling convolution subnetwork can perform upsampling convolution on the target features, and the color space transformation subnetwork can perform color space transformation on the upsampling convolution features to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the second upsampling convolution subnetwork may include Cropping layers, ResAU layers, Conv layers, and Cropping layers; these network layers can be used to perform upsampling convolution on the target features, and this process is not limited. The color space transformation subnetwork is used to perform image transformation from the YUV domain to RGB, or it is used to perform filtering processing from the YUV domain to the YUV domain. No restrictions are placed on this process.
[0204] Example 13: In Example 10, the first feature y_hat can be synthesized and transformed based on the initial enhancement subnetwork, the first upsampling convolutional subnetwork, the attention subnetwork, and the reconstruction decoding subnetwork to obtain the reconstructed image block x_hat corresponding to the current image block x. In one possible implementation, a feature adjustment factor can be used to adjust the weight output features of the attention subnetwork. The adjustment position of the feature adjustment factor can be found in [reference needed]. Figure 10C As shown, under the adjustment method of this feature adjustment factor, the first feature y_hat can be synthesized and transformed in the following way to obtain the reconstructed image block x_hat.
[0205] An initial enhancement subnetwork is used to enhance the first feature y_hat, resulting in the enhanced feature. For example, the first feature y_hat can be input into the initial enhancement subnetwork, which then enhances it to obtain the enhanced feature. For instance, the initial enhancement subnetwork may include two RB layers; therefore, the first feature y_hat can be enhanced using two RB layers to obtain the enhanced feature. There are no restrictions on this process.
[0206] After obtaining the enhanced features, a first upsampling convolutional sub-network is used to perform upsampling convolution on the enhanced features to obtain the initial features corresponding to the attention sub-network, i.e., the input features x_in of the attention sub-network. For example, the enhanced features (i.e., the output features of the initial enhanced sub-network) can be input into the first upsampling convolutional sub-network, which then performs upsampling convolution on the enhanced features to obtain the initial features corresponding to the attention sub-network. For example, the first upsampling convolutional sub-network may include Conv layers, Cropping layers, ResAU layers, and so on. Upsampling convolution can be performed on the enhanced features through these network layers; this process is not limited.
[0207] After obtaining the initial feature x_in corresponding to the attention sub-network, the initial feature x_in is input into the residual skip sub-network, the residual enhancement sub-network, and the weight generation sub-network, respectively. For example, after obtaining the initial feature x_in, the residual skip sub-network superimposes the initial feature x_in (i.e., x_out0) onto the final output feature of the attention sub-network. After obtaining the initial feature x_in, the residual enhancement sub-network enhances the initial feature x_in to obtain the enhanced feature; for example, the residual enhancement sub-network enhances the initial feature x_in based on RB blocks to obtain the enhanced feature x_out1.
[0208] The weight generation subnetwork may include a residual block subnetwork, a convolutional subnetwork, and a feature map subnetwork. See [link to documentation]. Figure 10C As shown, the weight generation subnetwork can include at least one RB block (i.e., residual block subnetwork), one Conv layer (i.e., convolutional subnetwork), and a feature mapping layer (i.e., feature mapping subnetwork). Based on this, the initial feature x_in can be input into the residual block subnetwork, and the residual block subnetwork can perform convolutional activation on the initial feature x_in to obtain the convolutionally activated feature. For example, the residual block subnetwork performs convolutional activation on the initial feature x_in based on the RB block to obtain the convolutionally activated feature.
[0209] Then, a convolutional sub-network can be used to convolve the features activated by the convolution to obtain convolutional features. For example, the convolutional sub-network can use Conv layers to convolve the features activated by the convolution to obtain convolutional features. After obtaining the convolutional features, a second feature can be determined based on the convolutional features. For example, the convolutional features can be used as the second feature.
[0210] In summary, a residual block sub-network can be used to enhance the initial feature x_in, resulting in enhanced features. A convolutional sub-network is then used to convolve these enhanced features, yielding the convolutional features, i.e., the second feature. After obtaining the second feature, it can be adjusted based on a feature adjustment factor, and the adjusted feature can be used as the third feature.
[0211] See Figure 10C As shown, after obtaining the third feature, a feature mapping sub-network can be used to perform feature mapping on the third feature to obtain the weight feature k. After obtaining the enhanced feature x_out1, the weight feature k, and the initial feature x_in (i.e., x_out0), a fourth feature can be generated based on the initial feature x_in, the enhanced feature x_out1, and the weight feature k. The fourth feature can be the final output feature of the attention sub-network. For example, multiplying the enhanced feature x_out1 by the weight feature k, and then adding it to the output feature x_out0 of the residual skip sub-network, the fourth feature is: x_out = x_in + k * x_out1.
[0212] In summary, we can see that the initial feature x_in can be enhanced using the first sub-network in the attention sub-network (such as the residual block sub-network and convolutional sub-network in the weight generation sub-network) to obtain the second feature. After obtaining the second feature, it can be adjusted based on the feature adjustment factor to obtain the third feature. After obtaining the third feature, it can be processed using the second sub-network in the attention sub-network (such as the residual skip sub-network, residual enhancement sub-network, and feature mapping sub-network in the weight generation sub-network) to obtain the fourth feature x_out. For example, the feature mapping sub-network can be used to perform feature mapping on the third feature to obtain the weight feature k; the residual enhancement sub-network can be used to enhance the initial feature x_in to obtain the enhanced feature x_out1; and the fourth feature x_out can be generated based on the initial feature x_in, the enhanced feature x_out1, and the weight feature k. After obtaining the fourth feature x_out, it can be determined as the target feature.
[0213] After obtaining the target features, the reconstruction decoding subnetwork can process these features to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the target features can be input into the reconstruction decoding subnetwork, which processes them to obtain the reconstructed image block x_hat corresponding to the current image block x. For instance, the reconstruction decoding subnetwork may include a second upsampling convolution subnetwork, which can perform upsampling convolution on the target features to obtain the reconstructed image block x_hat corresponding to the current image block x. Alternatively, the reconstruction decoding subnetwork may include a second upsampling convolution subnetwork and a color space transformation subnetwork. The second upsampling convolution subnetwork can perform upsampling convolution on the target features, and the color space transformation subnetwork can perform color space transformation on the upsampling convolution features to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the second upsampling convolution subnetwork may include Cropping layers, ResAU layers, Conv layers, and Cropping layers; these network layers can be used to perform upsampling convolution on the target features, and this process is not limited. The color space transformation subnetwork is used to perform image transformation from the YUV domain to RGB, or it is used to perform filtering processing from the YUV domain to the YUV domain. No restrictions are placed on this process.
[0214] Example 14: For Examples 1-13, at both the encoding and decoding ends, the second feature can be adjusted based on a feature adjustment factor to obtain the third feature. For instance, if the second feature includes C*H*W feature values, where C represents the number of channels, H represents the feature height, and W represents the feature width, then the feature adjustment value corresponding to each feature value can be determined based on the feature adjustment factor. For each feature value, the feature value is adjusted based on its corresponding feature adjustment value to obtain the adjusted feature value. The third feature is then generated based on the adjusted feature values corresponding to each feature value.
[0215] For example, if the second feature includes C*H*W feature values, then the feature adjustment value corresponding to each feature value of the second feature can be determined based on the feature adjustment factor. For this process, the following can be used:
[0216] Case 1: The dimension of the feature adjustment factor is the same as the dimension of the second feature, meaning that each feature value of the second feature has a different feature adjustment value. For example, if the feature adjustment factor can include C*H*W feature adjustment values, then the feature adjustment values corresponding to the C*H*W feature values can be determined based on the C*H*W feature adjustment values.
[0217] For example, for the feature value of the c-th channel in the second feature with position h*w, the feature adjustment value corresponding to this feature value can be: the feature adjustment value of the c-th channel with position h*w in the feature adjustment factor.
[0218] Scenario 2: The W and H dimensions of the feature adjustment factor are the same as the dimensions of the second feature, but the C dimension of the feature adjustment factor is 1. That is, the feature values at the same position in each channel of the second feature correspond to the same feature adjustment value. For example, if the feature adjustment factor can include H*W feature adjustment values, then the feature adjustment values corresponding to C*H*W feature values can be determined based on the H*W feature adjustment values (i.e., the C dimension of the feature adjustment factor is 1).
[0219] For example, for the feature value at position h*w of the c-th channel (i.e., each channel in the second feature), the feature adjustment value corresponding to this feature value can be: the feature adjustment value at position h*w in the feature adjustment factor.
[0220] Scenario 3: The C dimension of the feature adjustment factor is the same as the C dimension of the second feature, but the W and H dimensions of the feature adjustment factor are both 1. That is, the W*H feature values of the same channel of the second feature correspond to the same feature adjustment value. For example, if the feature adjustment factor can include C feature adjustment values, then the feature adjustment values corresponding to C*H*W feature values can be determined based on the C feature adjustment values (i.e., the W and H dimensions of the feature adjustment factor are both 1).
[0221] For example, for all feature values (i.e., feature values at any position) of the c-th channel in the second feature, the feature adjustment value corresponding to that feature value can be: the feature adjustment value of the c-th channel in the feature adjustment factor.
[0222] For example, after obtaining the feature adjustment value corresponding to each feature value, for each feature value, the feature value can be adjusted based on the feature adjustment value corresponding to that feature value to obtain the adjusted feature value corresponding to that feature value.
[0223] For example, if the eigenvalue corresponds to only one eigenvalue adjustment value, the adjusted eigenvalue can be determined using the following formula: yr = rf * y; where yr represents the adjusted eigenvalue, rf represents the eigenvalue adjustment value, and y represents the eigenvalue. For instance, for a eigenvalue y(c, w, h) with a spatial location (w, h) in channel c, after determining the eigenvalue adjustment value rf(c, w, h) corresponding to eigenvalue y(c, w, h), the adjusted eigenvalue yr(c, w, h) corresponding to eigenvalue y(c, w, h) is: yr(c, w, h) = rf(c, w, h) * y(c, w, h).
[0224] If the eigenvalue corresponds to N+1 eigenvalue adjustment values, where N is a positive integer (meaning at least two eigenvalue adjustment values), then the adjusted eigenvalue corresponding to the eigenvalue can be determined using the following formula: yr=rf_0*y 0 +rf_1*y 1 +rf_2*y 2 +…+rf_N*y N Where yr represents the adjusted eigenvalue, rf_0, rf_1, rf_2, ..., rf_N represent N+1 eigenvalue adjustment values, and y represents the eigenvalue. For example, for a eigenvalue y(c, w, h) with a spatial location (w, h) in channel c, after determining the corresponding eigenvalue adjustment value rf_i(c, w, h), where i = 0, 1, 2...N, the adjusted eigenvalue yr(c, w, h) corresponding to eigenvalue y(c, w, h) is: yr(c, w, h) = rf_0(c, w, h) * y 0 (c, w, h) + rf_1(c, w, h) * y 1 (c, w, h) + rf_2(c, w, h) * y 2 (c,w,h)+…+rf_N(c,w,h)*y N (c, w, h). When N is 1, yr(c,w,h)=rf_0(c,w,h)+rf_1(c,w,h)*y(c,w,h). When N is 2, yr(c,w,h)=rf_0(c,w,h)+rf_1(c,w,h)*y(c,w,h)+rf_2(c,w,h)*y 2 (c, w, h). And so on, the form is similar when N takes other values, and will not be repeated here. In the above formula, y k (c, w, h) represents y raised to the power of k.
[0225] For example, after obtaining the adjusted feature value corresponding to each feature value, a third feature can be generated based on the adjusted feature value corresponding to each feature value. The third feature may include the adjusted feature value corresponding to each feature value.
[0226] Example 15: For Examples 1-14, at the encoding end, it is also necessary to determine whether the current image block has feature adjustment mode enabled. If the current image block has feature adjustment mode enabled, the feature adjustment factor corresponding to the current image block is obtained and encoded in the auxiliary bitstream corresponding to the current image block. If the current image block does not have feature adjustment mode enabled, it is not necessary to obtain the feature adjustment factor corresponding to the current image block, nor is it necessary to encode the feature adjustment factor in the auxiliary bitstream corresponding to the current image block. At the decoding end, it is also necessary to determine whether the current image block has feature adjustment mode enabled. If the current image block has feature adjustment mode enabled, the feature adjustment factor corresponding to the current image block is decoded from the auxiliary bitstream corresponding to the current image block, and the target feature is determined based on the first feature and the feature adjustment factor. If the current image block does not have feature adjustment mode enabled, it is not necessary to decode the feature adjustment factor corresponding to the current image block from the auxiliary bitstream corresponding to the current image block.
[0227] For example, to determine whether feature adjustment mode is enabled for the current image patch, the following method can be used:
[0228] Method 1: The encoder encodes a feature adjustment flag in the bitstream, and the decoder decodes the feature adjustment flag from the bitstream. This feature adjustment flag either allows the current image block to use feature adjustment mode, or it disables the current image block from using feature adjustment mode. If the feature adjustment flag allows the current image block to use feature adjustment mode, the decoder determines that the current image block uses feature adjustment mode. Otherwise, if the feature adjustment flag disables the current image block from using feature adjustment mode, the decoder determines that the current image block disables feature adjustment mode. For example, if the feature adjustment flag has the first value, it allows the current image block to use feature adjustment mode; if it has the second value, it disables the current image block from using feature adjustment mode.
[0229] For example, the feature adjustment flag can be a sequence-level feature adjustment flag, meaning that the feature adjustment flag corresponds to all image blocks in the sequence; or, the feature adjustment flag can be an image-level feature adjustment flag, meaning that the feature adjustment flag corresponds to all image blocks in the image; or, the feature adjustment flag can be a slice-level feature adjustment flag, meaning that the feature adjustment flag corresponds to all image blocks in the slice, without any limitation.
[0230] Method 2: The encoder encodes a first range of feature values in the bitstream, and the decoder decodes the first range of feature values from the bitstream. Based on this, if the feature values of the second feature (e.g., all feature values) fall within the first range, the decoder determines that the current image block is in feature adjustment mode. Otherwise, if the feature values of the second feature (e.g., any feature value) do not fall within the first range, the decoder determines that the current image block is in feature adjustment mode. For example, after obtaining the second feature, the decoder can determine whether the feature values of the second feature fall within the first range.
[0231] For example, the first value range can be a sequence-level first value range, that is, the first value range corresponds to all image blocks in the sequence; or, the first value range can be an image-level first value range, that is, the first value range corresponds to all image blocks in the image; or, the first value range can be a slice-level first value range, that is, the first value range corresponds to all image blocks in the slice, and there is no limitation on this first value range.
[0232] Method 3: The encoder encodes the second range of feature change values in the bitstream, and the decoder decodes the second range of feature change values from the bitstream. Based on this, if the feature change value corresponding to the second feature (such as the feature change value corresponding to all feature values, i.e., each feature value corresponds to one feature change value) is within the second range, the decoder determines that the current image block is enabled in feature adjustment mode. Otherwise, if the feature change value corresponding to the second feature (such as any feature change value) is not within the second range, the decoder determines that the current image block is disabled in feature adjustment mode. For example, after obtaining the second feature, the decoder can also determine the feature change value corresponding to the second feature, i.e., determine the feature change value corresponding to each feature value in the second feature. The feature change value can represent the range of changes in the spatial domain or channel domain, i.e., determine the changes of the second feature in the spatial domain or channel domain. For example, the feature change value can include, but is not limited to, gradient values.
[0233] For example, the second value range can be a sequence-level second value range, that is, the second value range corresponds to all image blocks in the sequence; or, the second value range can be an image-level second value range, that is, the second value range corresponds to all image blocks in the image; or, the second value range can be a slice-level second value range, that is, the second value range corresponds to all image blocks in the slice, and there is no limitation on this second value range.
[0234] Method 4: The encoding end encodes the feature adjustment flag and the first value range in the bitstream, and the decoding end decodes the feature adjustment flag and the first value range from the bitstream. If the feature adjustment flag allows the current image block to enable feature adjustment mode, and the feature values of the second feature (such as all feature values) are within the first value range, then it is determined that the current image block enables feature adjustment mode. Otherwise, if the feature adjustment flag prohibits the current image block from enabling feature adjustment mode, and / or, the feature values of the second feature (such as any feature value) are not within the first value range, then it is determined that the current image block prohibits feature adjustment mode.
[0235] Method 5: The encoder encodes the feature adjustment flag and the second value range in the bitstream, and the decoder decodes the feature adjustment flag and the second value range from the bitstream. If the feature adjustment flag allows the current image block to enable feature adjustment mode, and the feature change value corresponding to the second feature (such as the feature change value corresponding to all feature values, i.e., each feature value corresponds to one feature change value) is within the second value range, then the decoder determines that the current image block enables feature adjustment mode. Otherwise, if the feature adjustment flag prohibits the current image block from enabling feature adjustment mode, and / or, the feature change value corresponding to the second feature (such as any feature change value) is not within the second value range, then the current image block is determined to be prohibited from enabling feature adjustment mode.
[0236] Method 6: The encoding end encodes a first value range and a second value range in the bitstream, and the decoding end decodes the first value range and the second value range from the bitstream. If the feature value of the second feature is within the first value range, and the feature change value corresponding to the second feature is within the second value range, then the decoding end can determine that the feature adjustment mode is enabled for the current image block. Otherwise, if the feature value of the second feature is not within the first value range, and / or the feature change value corresponding to the second feature is not within the second value range, then the decoding end can determine that the feature adjustment mode is disabled for the current image block.
[0237] Method 7: The encoding end encodes the feature adjustment flag, the first value range, and the second value range in the bitstream. The decoding end decodes the feature adjustment flag, the first value range, and the second value range from the bitstream. If the feature adjustment flag allows the current image block to enable feature adjustment mode, and the feature value of the second feature is within the first value range, and the feature change value corresponding to the second feature is within the second value range, then it is determined that the current image block enables feature adjustment mode. Otherwise, if the feature adjustment flag prohibits the current image block from enabling feature adjustment mode, the feature value of the second feature is not within the first value range, and / or the feature change value corresponding to the second feature is not within the second value range, then it is determined that the current image block prohibits feature adjustment mode.
[0238] Method 8: If the attribute value corresponding to the second feature falls within the configured attribute value range, the decoder can determine that the current image block is in feature adjustment mode. Otherwise, if the attribute value corresponding to the second feature does not fall within the configured attribute value range, the decoder can determine that the current image block is in feature adjustment mode disabled. In Method 8, the decoder does not need to parse information related to the feature adjustment mode from the bitstream; instead, it determines whether the current image block is in feature adjustment mode based on the configured attribute value range. In Method 8, the attribute value corresponding to the second feature may include, but is not limited to, the variance value corresponding to the feature value in the second feature. The variance value is used to represent the magnitude of the drastic change in the position of the feature value of the second feature.
[0239] Example 16: In Example 14, for each feature value in the second feature, the feature value can be adjusted based on the feature adjustment value corresponding to that feature value to obtain the adjusted feature value. For the above feature adjustment process, it is necessary to first determine whether the conditions are met. If the conditions are met, the feature adjustment process is performed (that is, the feature value can be adjusted based on the feature adjustment value corresponding to that feature value to obtain the adjusted feature value).
[0240] For example, these conditions may include, but are not limited to: 1. Determining whether to adjust the feature value based on additional bitstream parsing information. This information may include: flags indicating whether the region where the current feature value is located (which could be the feature itself, the channel of the feature, or a segment of the channel) needs adjustment; range information of the size of the feature value to be adjusted (determining whether the current feature value is within this range to determine whether adjustment is needed); or range of the degree of change of the feature value in the spatial or channel domain (calculating the change of the current feature value in the spatial or channel domain (e.g., gradient value) and determining whether the change is within this range to determine whether adjustment is needed). 2. Making a judgment based on a certain attribute of the feature value, such as the variance corresponding to the current feature value (characterizing the degree of drastic change in the location of the feature value).
[0241] For example, for each feature value in the second feature, in order to determine whether to perform feature adjustment on the feature value based on the feature adjustment value corresponding to that feature value to obtain the adjusted feature value, the following method can be used:
[0242] Method 1: The encoding end encodes the feature adjustment flag corresponding to the feature value in the bitstream. The decoding end decodes the feature adjustment flag corresponding to the feature value from the bitstream. This feature adjustment flag indicates whether feature adjustment is performed on the feature value, or whether no feature adjustment is performed. If the feature adjustment flag indicates feature adjustment, the decoding end adjusts the feature value based on the corresponding feature adjustment value to obtain the adjusted feature value. Otherwise, if the feature adjustment flag indicates no feature adjustment, the decoding end does not adjust the feature value based on the corresponding feature adjustment value. For example, if the feature adjustment flag is the first value, it indicates feature adjustment; if it is the second value, it indicates no feature adjustment.
[0243] Method 2: The encoding end encodes a first range of feature values in the bitstream, and the decoding end decodes the first range of feature values from the bitstream. Based on this, if the feature value falls within the first range, the decoding end adjusts the feature value based on the corresponding feature adjustment value to obtain the adjusted feature value. Otherwise, if the feature value does not fall within the first range, the decoding end does not adjust the feature value based on the corresponding feature adjustment value.
[0244] Method 3: The encoder encodes a second range of feature change values in the bitstream, and the decoder decodes the second range of feature change values from the bitstream. Based on this, if the feature change value (e.g., gradient value) corresponding to the feature value is within the second range, the decoder adjusts the feature value based on the corresponding feature adjustment value to obtain the adjusted feature value. Otherwise, if the feature change value (e.g., gradient value) corresponding to the feature value is not within the second range, the decoder does not adjust the feature value based on the corresponding feature adjustment value.
[0245] Method 4: The encoding end encodes the feature adjustment flag and the first value range in the bitstream, and the decoding end decodes the feature adjustment flag and the first value range from the bitstream. If the feature adjustment flag indicates that feature adjustment is to be performed on the feature value, and the feature value is within the first value range, then the decoding end performs feature adjustment on the feature value based on the corresponding feature adjustment value. Otherwise, the decoding end does not perform feature adjustment on the feature value based on the corresponding feature adjustment value.
[0246] Method 5: The encoding end encodes the feature adjustment flag and the second value range in the bitstream, and the decoding end decodes the feature adjustment flag and the second value range from the bitstream. If the feature adjustment flag indicates that feature adjustment is to be performed on the feature value, and the feature change value corresponding to the feature value is within the second value range, then the decoding end performs feature adjustment on the feature value based on the feature adjustment value corresponding to the feature value. Otherwise, it does not perform feature adjustment on the feature value based on the feature adjustment value corresponding to the feature value.
[0247] Method 6: The encoding end encodes a first value range and a second value range in the bitstream, and the decoding end decodes the first value range and the second value range from the bitstream. If the feature value is within the first value range and the feature change value corresponding to the feature value is within the second value range, then the decoding end adjusts the feature value based on the feature adjustment value corresponding to the feature value. Otherwise, the decoding end does not adjust the feature value based on the feature adjustment value corresponding to the feature value.
[0248] Method 7: The encoding end encodes a feature adjustment flag, a first value range, and a second value range in the bitstream. The decoding end decodes the feature adjustment flag, the first value range, and the second value range from the bitstream. If the feature adjustment flag indicates that feature adjustment is to be performed on the feature value, and the feature value is within the first value range, and the feature change value corresponding to the feature value is within the second value range, then the decoding end performs feature adjustment on the feature value based on the feature adjustment value corresponding to the feature value. Otherwise, the decoding end does not perform feature adjustment on the feature value based on the feature adjustment value corresponding to the feature value.
[0249] Method 8: If the attribute value corresponding to the feature value falls within the configured attribute value range, the decoder adjusts the feature value based on the corresponding feature adjustment value. Otherwise, the decoder does not adjust the feature value based on the corresponding feature adjustment value. In Method 8, the decoder does not need to parse information related to the feature adjustment mode from the bitstream; instead, it determines whether to adjust the feature value based on the corresponding feature adjustment value based on the configured attribute value range. In Method 8, the attribute value corresponding to the feature value may include, but is not limited to, the variance value corresponding to the feature value. The variance value is used to represent the magnitude of the drastic change in the location of the feature value.
[0250] Example 17: Referring to Example 7, the encoder can obtain a feature adjustment factor for the decoding sub-network, and perform feature adjustment on a certain feature in the decoding sub-network based on this feature adjustment factor, and encode the feature adjustment factor for the decoding sub-network in the auxiliary bitstream. The decoder can decode the feature adjustment factor for the decoding sub-network from the auxiliary bitstream, and perform feature adjustment on a certain feature in the decoding sub-network based on this feature adjustment factor. For example, after obtaining feature A (feature A is any feature in the decoding sub-network), the decoding sub-network can use this feature adjustment factor to adjust feature A, obtaining the adjusted feature B. The decoding sub-network continues processing based on feature B, and finally obtains the first feature.
[0251] In one possible implementation, see Example 8. Figure 8A As shown, the feature adjustment factor for the decoding sub-network can be used in the feature recovery process of the decoding sub-network. That is, in the feature recovery process, the feature adjustment factor is used to adjust a certain feature. The feature adjustment process can be referred to the feature adjustment process of the second feature, which will not be repeated here.
[0252] In one possible implementation, see Example 9. Figure 8B As shown, the feature adjustment factor for the decoding subnetwork can be used in the residual recovery process of the decoding subnetwork. That is, in the residual recovery process, the feature adjustment factor is used to adjust a certain feature. The feature adjustment process can be referred to the feature adjustment process of the second feature, which will not be repeated here.
[0253] For example, the above embodiments can be implemented individually or in combination. For instance, each of embodiments 1-17 can be implemented individually, and at least two embodiments 1-17 can be implemented in combination.
[0254] For example, in the above embodiments, the content of the encoding end can also be applied to the decoding end, that is, the decoding end can be processed in the same way, and the content of the decoding end can also be applied to the encoding end, that is, the encoding end can be processed in the same way.
[0255] Based on the same application concept as the above method, this application also proposes a decoding device, which is applied to the decoding end. The device includes: a memory configured to store video data; and a decoder configured to implement the decoding methods in embodiments 1-17 above, i.e., the processing flow of the decoding end.
[0256] For example, in one possible implementation, the decoder is configured to:
[0257] The first feature corresponding to the current image patch is obtained through a first neural network, and the feature adjustment factor corresponding to the current image patch is obtained. The first neural network includes at least one convolutional layer.
[0258] The target feature is determined based on the first feature and the feature adjustment factor;
[0259] Based on the target features, a reconstructed image block corresponding to the current image block is obtained through a second neural network, wherein the second neural network includes at least one convolutional layer.
[0260] Based on the same application concept as the above method, this application also proposes an encoding device, which is applied to the encoding end. The device includes: a memory configured to store video data; and an encoder configured to implement the encoding methods in embodiments 1-17 above, i.e., the processing flow of the encoding end.
[0261] For example, in one possible implementation, the encoder is configured to:
[0262] A first feature corresponding to the current image patch is obtained through a first neural network, the first neural network including at least one convolutional layer; a feature adjustment factor corresponding to the current image patch is obtained based on the first feature.
[0263] The feature adjustment factor is encoded in the auxiliary bitstream corresponding to the current image block.
[0264] Based on the same concept as the above method, the decoding device (also known as a video decoder) provided in this application embodiment, from a hardware perspective, its hardware architecture diagram can be found in [reference needed]. Figure 11A As shown, it includes: a processor 1101 and a machine-readable storage medium 1102, the machine-readable storage medium 1102 storing machine-executable instructions that can be executed by the processor 1101; the processor 1101 is used to execute the machine-executable instructions to implement the decoding methods of embodiments 1-17 of this application described above. For example, in one possible implementation, the decoding end device is used to implement:
[0265] The first feature corresponding to the current image patch is obtained through a first neural network, and the feature adjustment factor corresponding to the current image patch is obtained. The first neural network includes at least one convolutional layer.
[0266] The target feature is determined based on the first feature and the feature adjustment factor;
[0267] Based on the target features, a reconstructed image block corresponding to the current image block is obtained through a second neural network, wherein the second neural network includes at least one convolutional layer.
[0268] Based on the same concept as the above method, the encoding end device (also known as a video encoder) provided in this application embodiment, from a hardware perspective, its hardware architecture diagram can be found in [reference needed]. Figure 11B As shown, it includes: a processor 1111 and a machine-readable storage medium 1112, the machine-readable storage medium 1112 storing machine-executable instructions that can be executed by the processor 1111; the processor 1111 is used to execute the machine-executable instructions to implement the encoding methods of embodiments 1-17 of this application described above. For example, in one possible implementation, the encoding end device is used to implement:
[0269] A first feature corresponding to the current image patch is obtained through a first neural network, the first neural network including at least one convolutional layer; a feature adjustment factor corresponding to the current image patch is obtained based on the first feature.
[0270] The feature adjustment factor is encoded in the auxiliary bitstream corresponding to the current image block.
[0271] Based on the same application concept as the methods described above, this application provides an electronic device. It includes a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; the processor executes the machine-executable instructions to implement the decoding or encoding methods of embodiments 1-17 of this application described above.
[0272] Based on the same application concept as the above methods, embodiments of this application also provide a machine-readable storage medium storing a plurality of computer instructions. When the computer instructions are executed by a processor, they can implement the methods disclosed in the above examples of this application, such as the decoding method or encoding method in the above embodiments.
[0273] Based on the same application concept as the above method, this application embodiment also provides a computer application that, when executed by a processor, can implement the decoding method or encoding method disclosed in the above examples of this application.
[0274] Based on the same application concept as the above method, this application also proposes a decoding device, which can be applied to a decoding end. The decoding device may include: an acquisition module, used to acquire a first feature corresponding to the current image patch through a first neural network, and acquire a feature adjustment factor corresponding to the current image patch, wherein the first neural network includes at least one convolutional layer; a determination module, used to determine a target feature based on the first feature and the feature adjustment factor; the acquisition module is further used to acquire a reconstructed image patch corresponding to the current image patch through a second neural network based on the target feature, wherein the second neural network includes at least one convolutional layer.
[0275] For example, when the acquisition module acquires the first feature corresponding to the current image block through the first neural network, it is specifically used to: acquire probability distribution parameters based on the first bitstream corresponding to the current image block; determine a probability distribution model based on the probability distribution parameters, and decode the second bitstream corresponding to the current image block based on the probability distribution model to obtain the decoded image features; and determine the first feature corresponding to the current image block based on the decoded image features.
[0276] For example, when the acquisition module acquires the first feature corresponding to the current image block through the first neural network, it is specifically used to: acquire probability distribution parameters and predicted values based on the first bitstream corresponding to the current image block; determine a probability distribution model based on the probability distribution parameters, and decode the second bitstream corresponding to the current image block based on the probability distribution model to obtain the decoded image features; perform residual recovery on the decoded image features to obtain residual features; and determine the first feature corresponding to the current image block based on the residual features and the predicted values.
[0277] For example, when the acquisition module acquires the feature adjustment factor corresponding to the current image block, it is specifically used to: decode the auxiliary bitstream corresponding to the current image block to obtain the feature adjustment factor corresponding to the current image block; or, determine a fixed parameter value as the feature adjustment factor corresponding to the current image block.
[0278] For example, when the determining module determines the target feature based on the first feature and the feature adjustment factor, it is specifically used to: enhance the first feature to obtain a second feature; adjust the second feature based on the feature adjustment factor to obtain a third feature; and determine the target feature based on the third feature.
[0279] For example, when the determining module performs feature enhancement on the first feature to obtain the second feature, it is specifically used to: determine the initial feature corresponding to the attention sub-network based on the first feature; and perform feature enhancement on the initial feature through the attention sub-network to obtain the second feature; for example, when the determining module determines the target feature based on the third feature, it is specifically used to: determine the third feature as the target feature.
[0280] The attention subnetwork includes a residual enhancement subnetwork and a weight generation subnetwork. When the determining module enhances the initial feature through the attention subnetwork to obtain the second feature, it specifically performs the following steps: using the residual enhancement subnetwork to enhance the initial feature to obtain the enhanced feature; using the weight generation subnetwork to generate the weight feature corresponding to the initial feature; and generating the second feature based on the initial feature, the enhanced feature, and the weight feature.
[0281] For example, the weight generation subnetwork may include a residual block subnetwork, a convolutional subnetwork, and a feature mapping subnetwork. When the determining module uses the weight generation subnetwork to generate the weight features corresponding to the initial features, it specifically performs the following steps: using the residual block subnetwork to perform convolutional activation on the initial features to obtain convolutionally activated features; using the convolutional subnetwork to perform convolution on the convolutionally activated features to obtain convolutionally activated features; and using the feature mapping subnetwork to perform feature mapping on the convolutionally activated features to obtain the weight features corresponding to the initial features.
[0282] For example, when the determining module performs feature enhancement on the first feature to obtain the second feature, it is specifically used to: determine the initial feature corresponding to the attention sub-network based on the first feature; perform feature enhancement on the initial feature through the first sub-network in the attention sub-network to obtain the second feature; when the determining module determines the target feature based on the third feature, it is specifically used to: process the third feature through the second sub-network in the attention sub-network to obtain the fourth feature, and determine the fourth feature as the target feature.
[0283] For example, when the determining module performs feature enhancement on the initial feature through the first sub-network in the attention sub-network to obtain the second feature, it specifically performs the following: it performs feature enhancement on the initial feature using a residual enhancement sub-network to obtain the enhanced feature; it generates weight features corresponding to the initial feature using a weight generation sub-network; and it generates the second feature based on the enhanced feature and the weight features. When the determining module processes the third feature through the second sub-network in the attention sub-network to obtain the fourth feature, it specifically performs the following: it generates the fourth feature based on the initial feature and the third feature.
[0284] For example, when the determining module performs feature enhancement on the initial feature through the first sub-network in the attention sub-network to obtain the second feature, it specifically performs the following steps: using a residual block sub-network to perform feature enhancement on the initial feature to obtain enhanced features; using a convolution sub-network to perform convolution on the enhanced features to obtain convolutional features; and determining the second feature based on the convolutional features. When the determining module processes the third feature through the second sub-network in the attention sub-network to obtain the fourth feature, it specifically performs the following steps: using a feature mapping sub-network to perform feature mapping on the third feature to obtain weighted features; using a residual enhancement sub-network to perform feature enhancement on the initial feature to obtain enhanced features; and generating the fourth feature based on the initial feature, the enhanced features, and the weighted features.
[0285] For example, when the determining module determines the initial feature corresponding to the attention sub-network based on the first feature, it is specifically used to: perform feature enhancement on the first feature using an initial enhancement sub-network to obtain enhanced features; and perform upsampling convolution on the enhanced features using an upsampling convolution sub-network to obtain initial features.
[0286] For example, when the determining module adjusts the second feature based on the feature adjustment factor to obtain the third feature, it is specifically used as follows: if the second feature includes C*H*W feature values, where C represents the number of channels, H represents the feature height, and W represents the feature width, then the feature adjustment value corresponding to each feature value is determined based on the feature adjustment factor, and the feature value is adjusted based on the feature adjustment value corresponding to the feature value to obtain the adjusted feature value corresponding to the feature value; the third feature is generated based on the adjusted feature value corresponding to each feature value.
[0287] For example, when the determining module determines the feature adjustment value corresponding to each feature value based on the feature adjustment factor, it is specifically used to: if the feature adjustment factor includes C*H*W feature adjustment values, then determine the feature adjustment values corresponding to the C*H*W feature values based on the C*H*W feature adjustment values; or, if the feature adjustment factor includes H*W feature adjustment values, then determine the feature adjustment values corresponding to the C*H*W feature values based on the H*W feature adjustment values; or, if the feature adjustment factor includes C feature adjustment values, then determine the feature adjustment values corresponding to the C*H*W feature values based on the C feature adjustment values.
[0288] For example, when the determining module adjusts the feature value based on the feature adjustment value corresponding to the feature value to obtain the adjusted feature value, it is specifically used as follows: If the feature value corresponds to only one feature adjustment value, the adjusted feature value is determined using the following formula: yr = rf * y; where yr represents the adjusted feature value, rf represents the feature adjustment value, and y represents the feature value; or, if the feature value corresponds to N+1 feature adjustment values, where N is a positive integer, the adjusted feature value is determined using the following formula: yr = rf_0 * y 0 +rf_1*y 1 +rf_2*y 2 +…+rf_N*y N Where yr represents the adjusted feature value, rf_0, rf_1, rf_2, ..., rf_N represent N+1 feature adjustment values, and y represents the feature value.
[0289] For example, the second neural network includes a reconstruction decoding sub-network. When the acquisition module obtains the reconstructed image block corresponding to the current image block based on the target features through the second neural network, it is specifically used to: process the target features through the reconstruction decoding sub-network to obtain the reconstructed image block; wherein, the reconstruction decoding sub-network includes an upsampling convolution sub-network, which performs upsampling convolution on the target features to obtain the reconstructed image block; or, the reconstruction decoding sub-network includes an upsampling convolution sub-network and a color space transformation sub-network, which performs upsampling convolution on the target features and performs color space transformation on the upsampling convolution features through the color space transformation sub-network to obtain the reconstructed image block.
[0290] Based on the same application concept as the above method, this application also proposes an encoding device, which is applied to the encoding end. The device includes: an acquisition module, used to acquire a first feature corresponding to the current image block through a first neural network, the first neural network including at least one convolutional layer; and to acquire a feature adjustment factor corresponding to the current image block based on the first feature; and an encoding module, used to encode the feature adjustment factor in the auxiliary bitstream corresponding to the current image block.
[0291] For example, when the acquisition module acquires the feature adjustment factor corresponding to the current image block based on the first feature, it is specifically used to: acquire at least one candidate feature adjustment factor; determine the rate distortion cost corresponding to each candidate feature adjustment factor based on the first feature; and select one candidate feature adjustment factor from all candidate feature adjustment factors as the feature adjustment factor corresponding to the current image block based on the rate distortion cost corresponding to each candidate feature adjustment factor.
[0292] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. This application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Embodiments of this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The above descriptions are merely embodiments of this application and are not intended to limit this application.
[0293] Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. An image decoding method, characterized in that, The method includes: The first feature corresponding to the current image block is obtained through the first neural network, and the feature adjustment factor corresponding to the current image block is obtained; Based on the first feature, an initial feature corresponding to the attention sub-network is determined; the initial feature is enhanced using a residual enhancement sub-network to obtain an enhanced feature; a weight generation sub-network is used to generate a weight feature corresponding to the initial feature; and a second feature is generated based on the enhanced feature and the weight feature. The second feature is adjusted based on the feature adjustment factor to obtain the third feature; wherein, if the second feature includes C×H×W feature values, where C represents the number of channels, H represents the feature height, and W represents the feature width, then the feature adjustment values corresponding to the C×H×W feature values are determined based on the feature adjustment factor, and the C×H×W feature values are adjusted based on the feature adjustment values to obtain the adjusted feature values corresponding to the C×H×W feature values; the third feature is generated based on the adjusted feature values corresponding to the C×H×W feature values. The target feature is determined based on the third feature; Based on the target features, the reconstructed image block corresponding to the current image block is obtained through a second neural network.
2. The method according to claim 1, characterized in that, The step of determining the feature adjustment values corresponding to the C×H×W feature values based on the feature adjustment factor includes: If the feature adjustment factor includes C×H×W feature adjustment values, then the feature adjustment values corresponding to the C×H×W feature values are determined based on the C×H×W feature adjustment values; or, If the feature adjustment factor includes H×W feature adjustment values, then the feature adjustment values corresponding to the C×H×W feature values are determined based on the H×W feature adjustment values; or, If the feature adjustment factor includes C feature adjustment values, then the feature adjustment values corresponding to the C×H×W feature values are determined based on the C feature adjustment values.
3. The method according to claim 1, characterized in that, The first neural network includes a decoding sub-network, and the step of obtaining the first feature corresponding to the current image patch through the first neural network includes: The probability distribution parameters are obtained based on the bitstream corresponding to the current image block; a probability distribution model is determined based on the probability distribution parameters, and the bitstream corresponding to the current image block is decoded based on the probability distribution model to obtain the decoded image features; a first feature corresponding to the current image block is determined based on the decoded image features. Alternatively, obtain probability distribution parameters and predicted values based on the bitstream corresponding to the current image block; determine a probability distribution model based on the probability distribution parameters, and decode the bitstream corresponding to the current image block based on the probability distribution model to obtain decoded image features; perform residual recovery on the decoded image features to obtain residual features; and determine the first feature corresponding to the current image block based on the residual features and the predicted values.
4. The method according to claim 1, characterized in that, The step of obtaining the feature adjustment factor corresponding to the current image patch includes: Decode the bitstream corresponding to the current image block to obtain the feature adjustment factor corresponding to the current image block; or, The fixed parameter value is determined as the feature adjustment factor corresponding to the current image block.
5. The method according to claim 1, characterized in that, The second neural network includes a reconstruction decoding sub-network. The step of obtaining the reconstructed image patch corresponding to the current image patch through the second neural network based on the target features includes: The target features are processed by a reconstruction decoding subnetwork to obtain the reconstructed image patch; The reconstruction decoding subnetwork includes an upsampling convolution subnetwork, which performs upsampling convolution on the target features to obtain the reconstructed image patch; or, the reconstruction decoding subnetwork includes an upsampling convolution subnetwork and a color space transformation subnetwork, which performs upsampling convolution on the target features and performs color space transformation on the upsampling convolution features through the color space transformation subnetwork to obtain the reconstructed image patch.
6. The method according to claim 1, characterized in that, The method further includes: If the current image patch is in feature adjustment mode, then the feature adjustment factor corresponding to the current image patch is obtained; wherein, the process of determining that the current image patch is in feature adjustment mode includes: Parse the feature adjustment flag from the bitstream. If the feature adjustment flag allows the current image block to enable feature adjustment mode, then determine that the current image block is enabled in feature adjustment mode; or... The first range of feature values is parsed from the bitstream. If the feature value of the second feature is within the first range, then the feature adjustment mode is enabled for the current image block.
7. An image encoding method, characterized in that, The method includes: A first feature corresponding to the current image patch is obtained through a first neural network, and a feature adjustment factor corresponding to the current image patch is obtained based on the first feature; the feature adjustment factor is encoded in the bitstream corresponding to the current image patch. Based on the first feature, an initial feature corresponding to the attention sub-network is determined; the initial feature is enhanced using a residual enhancement sub-network to obtain an enhanced feature; a weight generation sub-network is used to generate a weight feature corresponding to the initial feature; and a second feature is generated based on the enhanced feature and the weight feature. The second feature is adjusted based on the feature adjustment factor to obtain the third feature; wherein, if the second feature includes C×H×W feature values, where C represents the number of channels, H represents the feature height, and W represents the feature width, then the feature adjustment values corresponding to the C×H×W feature values are determined based on the feature adjustment factor, and the C×H×W feature values are adjusted based on the feature adjustment values to obtain the adjusted feature values corresponding to the C×H×W feature values; the third feature is generated based on the adjusted feature values corresponding to the C×H×W feature values. The target feature is determined based on the third feature; Based on the target features, the reconstructed image block corresponding to the current image block is obtained through a second neural network.
8. An image decoding device, characterized in that, The device includes: The acquisition module is used to acquire the first feature corresponding to the current image block through the first neural network, and to acquire the feature adjustment factor corresponding to the current image block; A determination module is used to: determine an initial feature corresponding to an attention subnetwork based on the first feature; enhance the initial feature using a residual enhancement subnetwork to obtain an enhanced feature; generate a weighted feature corresponding to the initial feature using a weight generation subnetwork; generate a second feature based on the enhanced feature and the weighted feature; adjust the second feature based on the feature adjustment factor to obtain a third feature; wherein, if the second feature includes C×H×W feature values, where C represents the number of channels, H represents the feature height, and W represents the feature width, then the feature adjustment value corresponding to the C×H×W feature values is determined based on the feature adjustment factor, and the C×H×W feature values are adjusted based on the feature adjustment value to obtain the adjusted feature value corresponding to the C×H×W feature values; generate the third feature based on the adjusted feature value corresponding to the C×H×W feature values; and determine a target feature based on the third feature. The acquisition module is further configured to acquire the reconstructed image block corresponding to the current image block through a second neural network based on the target features.
9. An image encoding device, characterized in that, The device includes: The acquisition module is used to acquire the first feature corresponding to the current image patch through the first neural network, and to acquire the feature adjustment factor corresponding to the current image patch based on the first feature; The encoding module is used to encode the feature adjustment factor in the bitstream corresponding to the current image block; A determination module is used to: determine an initial feature corresponding to an attention subnetwork based on the first feature; enhance the initial feature using a residual enhancement subnetwork to obtain an enhanced feature; generate a weighted feature corresponding to the initial feature using a weight generation subnetwork; generate a second feature based on the enhanced feature and the weighted feature; adjust the second feature based on the feature adjustment factor to obtain a third feature; wherein, if the second feature includes C×H×W feature values, where C represents the number of channels, H represents the feature height, and W represents the feature width, then the feature adjustment value corresponding to the C×H×W feature values is determined based on the feature adjustment factor, and the C×H×W feature values are adjusted based on the feature adjustment value to obtain the adjusted feature value corresponding to the C×H×W feature values; generate the third feature based on the adjusted feature value corresponding to the C×H×W feature values; and determine a target feature based on the third feature. The acquisition module is further configured to acquire the reconstructed image block corresponding to the current image block through a second neural network based on the target features.
10. An image decoding device, characterized in that, include: A processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; The processor is configured to execute machine-executable instructions to implement the method according to any one of claims 1-6.
11. An image encoding device, characterized in that, include: A processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; The processor is configured to execute machine-executable instructions to implement the method of claim 7.
12. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores a plurality of computer instructions, which, when executed by a processor, implement the method described in any one of claims 1-6, or, when executed by a processor, implement the method described in claim 7.
13. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-6, or, when executed by a processor, implements the method according to claim 7.
Citation Information
Patent Citations
Video decoding method, video encoding method, video decoder and video encoder
CN111800629A