A decoding, encoding method, apparatus and device thereof
By using an end-to-end neural network approach and feature adjustment factors, the encoding and decoding performance is improved, solving the problems of poor encoding performance and high complexity in existing technologies, and achieving efficient video image compression.
Patent Information
- Application Number
- CN202411924186.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-10
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2043-01-10
AI Technical Summary
Existing neural network-based encoding and decoding methods suffer from poor encoding performance, poor decoding performance, and high complexity.
The first neural network acquires the features of the image patch and adjusts them based on the feature adjustment factor. The second neural network is then used to acquire the reconstructed image patch. An end-to-end video image compression method is adopted, and the encoding and decoding performance is improved by using convolutional layer structure design and auxiliary bitstream.
While maintaining low complexity, it effectively improves encoding and decoding efficiency, ensures the quality of reconstructed image patches, and reduces complexity.
Smart Images

Figure CN119450065B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of coding and decoding, in particular to a decoding and encoding method, device and equipment. BACKGROUND
[0002] In order to save space, video images are transmitted after being encoded. Complete video encoding can include prediction, transformation, quantization, entropy encoding, filtering and the like. For the prediction process, the prediction process can include intra prediction and inter prediction. Inter prediction refers to using the correlation in the time domain of a video to predict the current pixel using the pixels of the adjacent encoded image, so as to effectively remove the temporal redundancy of the video. Intra prediction refers to using the correlation in the spatial domain of a video to predict the current pixel using the pixels of the encoded block of the current frame image, so as to remove the spatial redundancy of the video.
[0003] With the rapid development of deep learning, deep learning has achieved success in many high-level computer vision problems such as image classification and object detection. Deep learning has also gradually begun to be applied in the field of coding and decoding, that is, a neural network can be used to encode and decode images. Although the neural network-based coding and decoding method has shown great performance potential, the neural network-based coding and decoding method still has problems such as poor coding performance, poor decoding performance and high complexity. SUMMARY
[0004] Therefore, the present application provides a decoding and encoding method, device and equipment to improve the coding and decoding performance.
[0005] The present application provides a decoding method applied to a decoding end, which comprises:
[0006] obtaining a first feature corresponding to a current image block through a first neural network, and obtaining a feature adjustment factor corresponding to the current image block, wherein the first neural network comprises at least one convolutional layer;
[0007] determining a target feature based on the first feature and the feature adjustment factor;
[0008] obtaining a reconstructed image block corresponding to the current image block through a second neural network based on the target feature, wherein the second neural network comprises at least one convolutional layer.
[0009] The present application provides an encoding method applied to an encoding end, which comprises: obtaining a first feature corresponding to a current image block through a first neural network, wherein the first neural network comprises at least one convolutional layer;
[0010] obtaining a feature adjustment factor corresponding to the current image block based on the first feature;
[0011] encode the feature adjustment factor in a side stream corresponding to the current image block.
[0012] The application provides a decoding device applied to a decoding end, the device comprising:
[0013] a memory configured to store video data;
[0014] a decoder configured to implement:
[0015] obtaining a first feature corresponding to a current image block through a first neural network, the first neural network comprising at least one convolutional layer;
[0016] determining a target feature based on the first feature and a feature adjustment factor corresponding to the current image block;
[0017] obtaining a reconstructed image block corresponding to the current image block through a second neural network based on the target feature, the second neural network comprising at least one convolutional layer.
[0018] The application provides an encoding device applied to an encoding end, the device comprising:
[0019] a memory configured to store video data;
[0020] an encoder configured to implement:
[0021] obtaining a first feature corresponding to a current image block through a first neural network, the first neural network comprising at least one convolutional layer;
[0022] obtaining a feature adjustment factor corresponding to the current image block based on the first feature;
[0023] encoding the feature adjustment factor in a side stream corresponding to the current image block.
[0024] The processor is configured to execute the machine-executable instructions to implement the decoding method.
[0025] The application provides an encoding end device, comprising: a processor and a machine readable storage medium, the machine readable storage medium stores machine executable instructions capable of being executed by the processor;
[0026] The processor is configured to execute the machine-executable instructions to implement the encoding method.
[0027] The application provides an electronic device, comprising a processor and a machine readable storage medium, the machine readable storage medium stores machine executable instructions capable of being executed by the processor; the processor is used for executing the machine executable instructions to realize the decoding method; or the processor is used for executing the machine executable instructions to realize the encoding method.
[0028] The application provides a machine readable storage medium, the machine readable storage medium stores a plurality of computer instructions, the computer instructions are executed by the processor to realize the decoding method; or realize the encoding method.
[0029] From the above technical solutions, in the embodiment of the application, the first feature corresponding to the current image block is obtained through the neural network, and the first feature is adjusted based on the feature adjustment factor corresponding to the current image block to obtain the target feature, and based on the target feature, the reconstruction image block corresponding to the current image block is obtained through the neural network, thereby an end-to-end video image compression method is proposed, which can realize the encoding and decoding of video images based on the neural network, and the purpose of improving the encoding efficiency and decoding efficiency is achieved by combining the feature adjustment factor. The network structure design and auxiliary code stream are combined to effectively ensure the quality of the reconstructed image block while maintaining low complexity, to improve the encoding performance and decoding performance, and to reduce complexity. BRIEF DESCRIPTION OF DRAWINGS
[0030] Figure 1 is a schematic diagram of a three-dimensional feature matrix in an embodiment of the application;
[0031] Figure 2 is a flowchart of the decoding method in an embodiment of the application;
[0032] Figure 3 is a flowchart of the encoding method in an embodiment of the application;
[0033] Figure 4 is a schematic diagram of the processing process of the encoding end in an embodiment of the application;
[0034] Figure 5 is a schematic diagram of the processing process of the decoding end in an embodiment of the application;
[0035] Figures 6A-6H is a schematic diagram of the network structure in an embodiment of the application;
[0036] Figure 7A is a schematic diagram of the processing process of the decoding end in an embodiment of the application;
[0037] Figure 7B is a schematic diagram of the processing process of the encoding end in an embodiment of the application;
[0038] Figure 8A and Figure 8B is a structural schematic diagram of a decoding subnetwork in an embodiment of the present application;
[0039] Figure 9A and Figure 9B is a structural schematic diagram of an attention subnetwork in an embodiment of the present application;
[0040] Figures 10A-10C is an adjustment position schematic diagram of a feature adjustment factor in an embodiment of the present application;
[0041] Figure 11A is a hardware structure diagram of a decoding end device in an embodiment of the present application;
[0042] Figure 11B is a hardware structure diagram of an encoding end device in an embodiment of the present application. DETAILED DESCRIPTION
[0043] The terminology used in the present application merely for the purpose of describing particular embodiments, and is not intended to limit the present application. The singular forms "a", "an", and "the" used in the embodiments and claims of the present application are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to any or all possible combinations of one or more associated listed items. It should be understood that although the terms first, second, third, etc. can be used in the embodiments of the present application to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information, without departing from the scope of the embodiments of the present application, depending on the context. In addition, the word "if" used can be interpreted as "when", or "upon", or "in response to determining".
[0044] The embodiments of the present application propose a decoding method and an encoding method, which can involve the following concepts:
[0045] JPEG (Joint Photographic Experts Group): JPEG is a standard for compressing continuous-tone still images, and the file suffix can be.jpg or.jpeg. It is a commonly used image file format. JPEG is a joint coding method that uses predictive coding (such as DPCM, Differential Pulse-Code Modulation), discrete cosine transform (DCT), and entropy coding to remove redundant image and color data. It is a lossy compression format that can compress images into a very small storage space, but will cause damage to image data to some extent. Especially when using too high a compression ratio, the quality of the image recovered after final decompression will be reduced. If high-quality images are pursued, JPEG should not use too high a compression ratio.
[0046] JPEG-AI (Joint Photographic Experts Group Artificial Intelligence): The scope of JPEG-AI is to create a learning-based image coding standard that provides a single-stream, compact compression domain representation, significantly improves compression efficiency compared to commonly used image coding standards, and effectively improves performance in image processing and computer vision tasks. JPEG-AI is aimed at a wide range of applications, such as cloud storage, visual management, autonomous driving cars and devices, image acquisition, storage and management, real-time management of visual data, and media distribution. The goal of JPEG-AI is to design a coding solution that significantly improves compression efficiency with the same subjective quality, providing efficient compression domain processing for machine learning-based image processing and computer vision tasks. JPEG-AI needs to support 8-bit and 10-bit depth, and use text and graphics to efficiently encode and progressively decode images.
[0047] Entropy Encoding: Entropy encoding is a coding process that does not lose any information according to the entropy principle. The information entropy is the average amount of information of the source (a measure of uncertainty). The encoding method of entropy encoding can include but is not limited to: Shannon coding, Huffman coding, and arithmetic coding.
[0048] Neural Network (NN): Neural network refers to artificial neural network. The neural network is an operation model composed of a large number of nodes (or called neurons) connected with each other. In the neural network, the neuron processing unit can represent different objects, such as features, letters, concepts, or some meaningful abstract patterns. The types of processing units in the neural network can be divided into three categories: input units, output units and hidden units. The input unit accepts the signals and data of the external world; the output unit realizes the output of the processing result; the hidden unit is the unit between the input and output units, which cannot be observed from the outside of the system. The connection weight between neurons reflects the connection strength between units, and the information representation and processing are embodied in the connection relationship of processing units. Neural network is a non-programmed, brain-like information processing method. The essence of neural network is to obtain a parallel distributed information processing function through the transformation and dynamics of neural network, and to imitate the information processing function of the human brain neural system at different levels and levels. In the field of video processing, the commonly used neural network can include but is not limited to: convolutional neural network (CNN), recurrent neural network (RNN), fully connected network, etc.
[0049] Convolutional Neural Network (CNN): Convolutional neural network is a kind of feedforward neural network, which is one of the most representative network structures in deep learning technology. The artificial neuron of convolutional neural network can respond to a part of the surrounding units in the coverage range, and has excellent performance for large image processing. The basic structure of convolutional neural network includes two layers. One is the feature extraction layer (also called convolution layer), the input of each neuron is connected with the local receptive field of the previous layer, and the local feature is extracted. Once the local feature is extracted, the positional relationship between the local feature and other features is also determined. The second is the feature mapping layer (also called activation layer). Each calculation layer of neural network is composed of multiple feature mappings. Each feature mapping is a plane, and all the weights of the neurons on the plane are equal. The feature mapping structure can use Sigmoid function, ReLU function, Leaky-ReLU function, PReLU function, GDN function, etc. as the activation function of convolutional network. In addition, since the neurons on a mapping plane share weights, the number of free parameters of the network is reduced.
[0050] For example, one advantage of convolutional neural networks (CNNs) over image processing algorithms is that they avoid complex preprocessing steps (such as extracting artificial features) and can directly input the original image for end-to-end learning. Another advantage of CNNs over ordinary neural networks is that ordinary neural networks use fully connected layers, meaning all neurons from the input layer to the hidden layer are connected. This results in a huge number of parameters, making network training time-consuming or even difficult. CNNs, however, avoid this difficulty through local connectivity and weight sharing.
[0051] Deconvolution: Also known as transposed convolution, deconvolution layers work similarly to convolutional layers. The main difference is that deconvolution layers use padding to make the output larger than the input (though they can also remain the same). If the stride is 1, the output size equals the input size; if the stride is N, the width of the output feature is N times the width of the input feature, and the height of the output feature is N times the height of the input feature.
[0052] Generalization ability: Generalization ability refers to the ability of a machine learning algorithm to adapt to new samples. The goal of learning is to learn the patterns hidden behind data pairs. The trained network can also give appropriate outputs for data outside the learning set that have the same pattern. This ability can be called generalization ability.
[0053] Features: The features involved in this application are a three-dimensional feature matrix of size C*W*H, see [link to relevant documentation]. Figure 1 The diagram shows a schematic of a three-dimensional feature matrix. In the three-dimensional feature matrix, C represents the number of channels, H represents the feature height, and W represents the feature width. The three-dimensional feature matrix can be either the input or the output of a neural network.
[0054] Rate-Distortion Optimized: There are two indicators to evaluate the coding efficiency: code rate and PSNR (Peak Signal to Noise Ratio). The smaller the bit stream, the greater the compression rate, and the greater the PSNR, the better the quality of the reconstructed image. In mode selection, the evaluation formula is essentially a comprehensive evaluation of the two. For example, the cost of a mode: J (mode) = D + λ * R, where D represents Distortion (distortion), which can usually be measured using the SSE indicator, which is the sum of the squares of the differences between the reconstructed image block and the source image. In order to achieve cost consideration, the SAD indicator can also be used, which is the sum of the absolute values of the differences between the reconstructed image block and the source image. λ is the Lagrange multiplier, and R is the actual number of bits required for image block coding under this mode, including the total number of bits required for coding mode information, motion information, and residual error. In mode selection, if the rate-distortion principle is used to compare and decide the coding mode, the best coding performance can usually be guaranteed.
[0055] A large number of encoding tools are proposed for each module of the encoding end, and each tool often has multiple modes. For different video sequences, the encoding tool that can achieve the best encoding performance is often different. Therefore, in the encoding process, RDO (Rate-Distortion Optimize) is usually used to compare the encoding performance of different tools or modes to select the best mode. After determining the optimal tool or mode, the decision information of the tool or mode is transmitted by encoding the marker information in the bit stream. This method, although it brings higher encoding complexity, can adaptively select the optimal mode combination for different content and achieve the best encoding performance. The decoding end can obtain the relevant mode information by directly parsing the flag information, and the complexity is less affected.
[0056] The decoding method and the encoding method in the embodiments of the present application will be described in detail in combination with several specific embodiments.
[0057] Embodiment 1: A decoding method is proposed in the embodiments of the present application, as shown in Figure 2 The method can be applied to the decoding end (also referred to as a video decoder), and the method can include:
[0058] Step 201: obtaining a first feature corresponding to a current image block through a first neural network, and obtaining a feature adjustment factor corresponding to the current image block, the first neural network can include at least one convolutional layer.
[0059] In a possible implementation, the first feature corresponding to the current image block is acquired by the first neural network, which can include but is not limited to: acquiring a probability distribution parameter based on the first code stream corresponding to the current image block; determining a probability distribution model based on the probability distribution parameter, and decoding the second code stream corresponding to the current image block based on the probability distribution model to obtain a decoded image feature; and determining the first feature corresponding to the current image block based on the decoded image feature. As can be seen from the above, in this implementation, the first neural network is used to realize the functions of acquiring a probability distribution parameter, determining a probability distribution model, decoding the second code stream corresponding to the current image block, and determining the first feature corresponding to the current image block.
[0060] In another possible implementation, the first feature corresponding to the current image block is acquired by the first neural network, which can include but is not limited to: acquiring a probability distribution parameter and a prediction value based on the first code stream corresponding to the current image block; determining a probability distribution model based on the probability distribution parameter, and decoding the second code stream corresponding to the current image block based on the probability distribution model to obtain a decoded image feature; performing residual recovery on the decoded image feature to obtain a residual feature; and determining the first feature corresponding to the current image block based on the residual feature and the prediction value. As can be seen from the above, in this implementation, the first neural network is used to realize the functions of acquiring a probability distribution parameter, acquiring a prediction value, determining a probability distribution model, decoding the second code stream corresponding to the current image block, performing residual recovery, and determining the first feature corresponding to the current image block.
[0061] For example, the feature adjustment factor corresponding to the current image block is acquired, which can include but is not limited to: decoding the auxiliary code stream corresponding to the current image block to obtain the feature adjustment factor corresponding to the current image block, i.e., parsing the feature adjustment factor corresponding to the current image block from the auxiliary code stream. Alternatively, a fixed parameter value is determined as the feature adjustment factor corresponding to the current image block, for example, the fixed parameter value 1 can be determined as the feature adjustment factor corresponding to the current image block.
[0062] In step 202, the target feature is determined based on the first feature and the feature adjustment factor.
[0063] For example, the target feature is determined based on the first feature and the feature adjustment factor, which can include but is not limited to: performing feature enhancement on the first feature to obtain a second feature; after obtaining the second feature, performing feature adjustment on the second feature based on the feature adjustment factor to obtain a third feature; and after obtaining the third feature, determining the target feature based on the third feature.
[0064] In a possible implementation, the feature enhancement on the first feature to obtain the second feature can include but is not limited to: determining an initial feature corresponding to the attention subnetwork based on the first feature; and performing feature enhancement on the initial feature through the attention subnetwork to obtain the second feature. After obtaining the second feature, the feature adjustment on the second feature based on the feature adjustment factor to obtain the third feature (i.e., the feature-adjusted feature as the third feature) can be performed. After obtaining the third feature, the determination of the target feature based on the third feature can include but is not limited to: determining the third feature as the target feature.
[0065] For example, the attention subnetwork can include a residual enhancement subnetwork and a weight generation subnetwork. The feature enhancement on the initial feature through the attention subnetwork to obtain the second feature can include but is not limited to: performing feature enhancement on the initial feature by using the residual enhancement subnetwork to obtain an enhanced feature, generating a weight feature corresponding to the initial feature by using the weight generation subnetwork, and generating the second feature based on the initial feature, the enhanced feature, and the weight feature.
[0066] For example, the weight generation subnetwork can include a residual block subnetwork, a convolution subnetwork, and a feature mapping subnetwork. The generation of the weight feature corresponding to the initial feature by using the weight generation subnetwork can include but is not limited to: performing convolution activation on the initial feature by using the residual block subnetwork to obtain a convolution-activated feature; performing convolution on the convolution-activated feature by using the convolution subnetwork to obtain a convolved feature; and performing feature mapping on the convolved feature by using the feature mapping subnetwork to obtain the weight feature.
[0067] For example, the determination of the initial feature corresponding to the attention subnetwork based on the first feature can include but is not limited to: performing feature enhancement on the first feature by using an initial enhancement subnetwork to obtain an enhanced feature, and performing up-sampling convolution on the enhanced feature by using an up-sampling convolution subnetwork to obtain the initial feature corresponding to the attention subnetwork.
[0068] In another possible implementation, the feature enhancement on the first feature to obtain the second feature can include but is not limited to: determining an initial feature corresponding to the attention subnetwork based on the first feature; and performing feature enhancement on the initial feature through a first subnetwork in the attention subnetwork to obtain the second feature. After obtaining the second feature, the feature adjustment on the second feature based on the feature adjustment factor to obtain the third feature (i.e., the feature-adjusted feature as the third feature) can be performed. After obtaining the third feature, the determination of the target feature based on the third feature can include but is not limited to: performing processing on the third feature through a second subnetwork in the attention subnetwork to obtain a fourth feature, and determining the fourth feature as the target feature.
[0069] For example, the feature enhancement on the initial feature by the first sub-network in the attention sub-network to obtain the second feature can include but is not limited to: performing feature enhancement on the initial feature by a residual enhancement sub-network to obtain an enhanced feature, generating a weight feature corresponding to the initial feature by a weight generation sub-network, and generating the second feature based on the enhanced feature and the weight feature. The processing on the third feature by the second sub-network in the attention sub-network to obtain the fourth feature can include but is not limited to: generating the fourth feature based on the initial feature and the third feature.
[0070] The weight generation sub-network can include a residual block sub-network, a convolution sub-network and a feature mapping sub-network, and the generation of the weight feature corresponding to the initial feature by the weight generation sub-network can include but is not limited to: performing convolution activation on the initial feature by the residual block sub-network to obtain a convolution activated feature; performing convolution on the convolution activated feature by the convolution sub-network to obtain a convolution feature; and performing feature mapping on the convolution feature by the feature mapping sub-network to obtain the weight feature.
[0071] For example, the feature enhancement on the initial feature by the first sub-network in the attention sub-network to obtain the second feature can include but is not limited to: performing feature enhancement on the initial feature by a residual block sub-network to obtain an enhanced feature, performing convolution on the enhanced feature by a convolution sub-network to obtain a convolution feature, and determining the second feature based on the convolution feature. The processing on the third feature by the second sub-network in the attention sub-network to obtain the fourth feature can include but is not limited to: performing feature mapping on the third feature by a feature mapping sub-network to obtain a weight feature; performing feature enhancement on the initial feature by a residual enhancement sub-network to obtain an enhanced feature; and generating the fourth feature based on the initial feature, the enhanced feature and the weight feature.
[0072] In the above embodiment, the determination of the initial feature corresponding to the attention sub-network based on the first feature can include but is not limited to: performing feature enhancement on the first feature by an initial enhancement sub-network to obtain an enhanced feature, and performing up-sampling convolution on the enhanced feature by an up-sampling convolution sub-network to obtain the initial feature corresponding to the attention sub-network.
[0073] In a possible implementation, the feature adjustment on the second feature based on the feature adjustment factor to obtain the third feature can include but is not limited to: if the second feature includes C*H*W feature values, C represents the number of channels, H represents the feature height, and W represents the feature width, determining feature adjustment values corresponding to the C*H*W feature values based on the feature adjustment factor, and performing feature adjustment on the C*H*W feature values based on the feature adjustment values to obtain adjusted feature values corresponding to the C*H*W feature values. On this basis, the third feature can be generated based on the adjusted feature values corresponding to the C*H*W feature values.
[0074] For example, determining the feature adjustment value corresponding to each of the C*H*W feature values based on the feature adjustment factor can include, but is not limited to: if the feature adjustment factor includes C*H*W feature adjustment values, determining the feature adjustment value corresponding to each of the C*H*W feature values based on the C*H*W feature adjustment values; or, if the feature adjustment factor includes H*W feature adjustment values, determining the feature adjustment value corresponding to each of the C*H*W feature values based on the H*W feature adjustment values; or, if the feature adjustment factor includes C feature adjustment values, determining the feature adjustment value corresponding to each of the C*H*W feature values based on the C feature adjustment values.
[0075] For example, adjusting the C*H*W feature values based on the feature adjustment value to obtain the adjusted feature value corresponding to each of the C*H*W feature values can include, but is not limited to: if one of the C*H*W feature values corresponds to one feature adjustment value, the adjusted feature value corresponding to the feature value can be determined by the following formula: yr=rf*y; wherein, yr represents the adjusted feature value, rf represents the feature adjustment value, and y represents the feature value. Or, if one of the C*H*W feature values corresponds to N+1 feature adjustment values, N is a positive integer, the adjusted feature value corresponding to the feature value can be determined by the following formula: yr=rf_0*y0+rf_1*y1+rf_2*y2+…+rf_N*yN; wherein, yr represents the adjusted feature value, rf_0, rf_1, rf_2, …, rf_N represent N+1 feature adjustment values, and y represents the feature value.
[0076] In step 203, a reconstruction image block corresponding to the current image block is obtained based on the target feature through a second neural network. The second neural network can include at least one convolutional layer.
[0077] For example, the second neural network can include a reconstruction decoding subnetwork. Obtaining the reconstruction image block corresponding to the current image block based on the target feature through the second neural network can include, but is not limited to: processing the target feature through the reconstruction decoding subnetwork to obtain the reconstruction image block corresponding to the current image block. For example, the reconstruction decoding subnetwork can include an up-sampling convolutional subnetwork. The target feature can be up-sampled and convolved through the up-sampling convolutional subnetwork to obtain the reconstruction image block corresponding to the current image block. For another example, the reconstruction decoding subnetwork can include an up-sampling convolutional subnetwork and a color space transformation subnetwork. The target feature can be up-sampled and convolved through the up-sampling convolutional subnetwork, and the feature after up-sampling and convolution can be color space transformed through the color space transformation subnetwork to obtain the reconstruction image block corresponding to the current image block.
[0078] In a possible implementation, if the current image block enables the feature adjustment mode, a feature adjustment factor corresponding to the current image block is obtained, and the target feature is determined based on the first feature and the feature adjustment factor. Alternatively, if the current image block does not enable the feature adjustment mode, the target feature is determined based on the first feature, and the feature adjustment factor corresponding to the current image block does not need to be obtained.
[0079] For example, a feature adjustment flag can be parsed from the bitstream, and if the feature adjustment flag allows the current image block to enable the feature adjustment mode, it can be determined that the current image block enables the feature adjustment mode. Alternatively, a first value range of a feature value can be parsed from the bitstream, and if the feature value of the second feature is in the first value range, it can be determined that the current image block enables the feature adjustment mode. Alternatively, a second value range of a feature change value can be parsed from the bitstream, and if the feature change value corresponding to the second feature is in the second value range, it can be determined that the current image block enables the feature adjustment mode.
[0080] For example, if the attribute value corresponding to the second feature is in the configured attribute value range, it can be determined that the current image block enables the feature adjustment mode; wherein the attribute value can include a variance value.
[0081] For example, the above execution order is only an example given for convenience of description, and in actual application, the execution order between steps can also be changed, and the execution order is not limited. Moreover, in other embodiments, the steps of the corresponding method can not necessarily be executed in the order shown and described in the specification, and the steps included in the method can be more or less than those described in the specification. In addition, a single step described in the specification can be divided into multiple steps for description in other embodiments; multiple steps described in the specification can also be combined into a single step for description.
[0082] From the above technical solutions, in the embodiments of the present application, the first feature corresponding to the current image block is obtained through the first neural network, and the first feature is adjusted based on the feature adjustment factor corresponding to the current image block to obtain the target feature. Based on the target feature, the reconstructed image block corresponding to the current image block is obtained through the second neural network, thereby proposing an end-to-end video image compression method, which can realize decoding of video images based on neural networks, and improve decoding efficiency by combining feature adjustment factors. By combining network structure design and auxiliary bitstream, the neural network can maintain low complexity while effectively ensuring the quality of the reconstructed image block, improving decoding performance and reducing complexity.
[0083] Embodiment 2: In the embodiments of the present application, an encoding method is proposed, which is described with reference to Figure 3As shown, a flowchart of the encoding method is shown. The method can be applied to an encoding end (also referred to as a video encoder). The method can include the following steps:
[0084] In step 301, a first feature corresponding to a current image block is obtained by using a first neural network. The first neural network can include at least one convolutional layer.
[0085] In a possible implementation, the first feature corresponding to the current image block is obtained by using the first neural network, which can include but is not limited to the following steps: obtaining a probability distribution parameter based on a first code stream corresponding to the current image block; determining a probability distribution model based on the probability distribution parameter, and decoding a second code stream corresponding to the current image block based on the probability distribution model to obtain a decoded image feature; and determining the first feature corresponding to the current image block based on the decoded image feature. As can be seen from the above, in this implementation, the first neural network is used to realize the functions of obtaining the probability distribution parameter, determining the probability distribution model, decoding the second code stream corresponding to the current image block, and determining the first feature corresponding to the current image block.
[0086] In another possible implementation, the first feature corresponding to the current image block is obtained by using the first neural network, which can include but is not limited to the following steps: obtaining a probability distribution parameter and a prediction value based on the first code stream corresponding to the current image block; determining a probability distribution model based on the probability distribution parameter, and decoding the second code stream corresponding to the current image block based on the probability distribution model to obtain a decoded image feature; performing residual recovery on the decoded image feature to obtain a residual feature; and determining the first feature corresponding to the current image block based on the residual feature and the prediction value. As can be seen from the above, in this implementation, the first neural network is used to realize the functions of obtaining the probability distribution parameter, obtaining the prediction value, determining the probability distribution model, decoding the second code stream corresponding to the current image block, performing the residual recovery, and determining the first feature corresponding to the current image block.
[0087] In step 302, a feature adjustment factor corresponding to the current image block is obtained based on the first feature.
[0088] For example, the feature adjustment factor corresponding to the current image block is obtained based on the first feature, which can include but is not limited to the following steps: obtaining at least one candidate feature adjustment factor, and determining a rate-distortion cost corresponding to each candidate feature adjustment factor based on the first feature. Based on the rate-distortion cost corresponding to each candidate feature adjustment factor, one candidate feature adjustment factor can be selected from all candidate feature adjustment factors as the feature adjustment factor corresponding to the current image block.
[0089] In step 303, the feature adjustment factor is encoded in an auxiliary code stream corresponding to the current image block.
[0090] Exemplarily, the fixed parameter value can also be determined as a feature adjustment factor corresponding to the current image block. For example, the fixed parameter value 1 can be determined as the feature adjustment factor corresponding to the current image block. In this case, the feature adjustment factor corresponding to the current image block can also not be encoded in the auxiliary code stream.
[0091] Exemplarily, the target feature can also be determined based on the first feature corresponding to the current image block and the feature adjustment factor corresponding to the current image block. Based on the target feature, the reconstructed image block corresponding to the current image block (i.e., the final output reconstructed image block) can be obtained through the second neural network. The second neural network can include at least one convolutional layer.
[0092] Exemplarily, determining the target feature based on the first feature and the feature adjustment factor can include but is not limited to: performing feature enhancement on the first feature to obtain a second feature; after obtaining the second feature, performing feature adjustment on the second feature based on the feature adjustment factor to obtain a third feature; and after obtaining the third feature, determining the target feature based on the third feature.
[0093] In a possible implementation, performing feature enhancement on the first feature to obtain a second feature can include but is not limited to: determining an initial feature corresponding to an attention subnetwork based on the first feature; and performing feature enhancement on the initial feature through the attention subnetwork to obtain the second feature. After obtaining the second feature, the second feature can be adjusted based on the feature adjustment factor to obtain a third feature (i.e., the feature adjusted feature as the third feature). After obtaining the third feature, determining the target feature based on the third feature can include but is not limited to: determining the third feature as the target feature.
[0094] Exemplarily, the attention subnetwork can include a residual enhancement subnetwork and a weight generation subnetwork. Performing feature enhancement on the initial feature through the attention subnetwork to obtain the second feature can include but is not limited to: performing feature enhancement on the initial feature through the residual enhancement subnetwork to obtain an enhanced feature, generating a weight feature corresponding to the initial feature through the weight generation subnetwork, and generating the second feature based on the initial feature, the enhanced feature, and the weight feature.
[0095] Exemplarily, the weight generation subnetwork can include a residual block subnetwork, a convolutional subnetwork, and a feature mapping subnetwork. Generating the weight feature corresponding to the initial feature through the weight generation subnetwork can include but is not limited to: performing convolutional activation on the initial feature through the residual block subnetwork to obtain a convolutional activation feature; performing convolution on the convolutional activation feature through the convolutional subnetwork to obtain a convolution feature; and performing feature mapping on the convolution feature through the feature mapping subnetwork to obtain the weight feature.
[0096] For example, determining the initial feature corresponding to the attention subnetwork based on the first feature can include, but is not limited to: performing feature enhancement on the first feature by using an initial enhancement subnetwork to obtain an enhanced feature, and performing up-sampling convolution on the enhanced feature by using an up-sampling convolution subnetwork to obtain the initial feature corresponding to the attention subnetwork.
[0097] In another possible implementation, performing feature enhancement on the first feature to obtain a second feature can include, but is not limited to: determining the initial feature corresponding to the attention subnetwork based on the first feature; and performing feature enhancement on the initial feature by using a first subnetwork in the attention subnetwork to obtain the second feature. After obtaining the second feature, performing feature adjustment on the second feature based on a feature adjustment factor to obtain a third feature (i.e., the feature-adjusted feature as the third feature) can be performed. After obtaining the third feature, determining the target feature based on the third feature can include, but is not limited to: processing the third feature by using a second subnetwork in the attention subnetwork to obtain a fourth feature, and determining the fourth feature as the target feature.
[0098] For example, performing feature enhancement on the initial feature by using the first subnetwork in the attention subnetwork to obtain the second feature can include, but is not limited to: performing feature enhancement on the initial feature by using a residual enhancement subnetwork to obtain an enhanced feature, and generating a weight feature corresponding to the initial feature by using a weight generation subnetwork, and generating the second feature based on the enhanced feature and the weight feature. Processing the third feature by using the second subnetwork in the attention subnetwork to obtain the fourth feature can include, but is not limited to: generating the fourth feature based on the initial feature and the third feature.
[0099] The weight generation subnetwork can include a residual block subnetwork, a convolution subnetwork, and a feature mapping subnetwork. Generating the weight feature corresponding to the initial feature by using the weight generation subnetwork can include, but is not limited to: performing convolution activation on the initial feature by using the residual block subnetwork to obtain a convolution-activated feature; performing convolution on the convolution-activated feature by using the convolution subnetwork to obtain a convolutional feature; and performing feature mapping on the convolutional feature by using the feature mapping subnetwork to obtain the weight feature.
[0100] For example, the initial feature is enhanced by a first subnetwork in the attention subnetwork to obtain a second feature, which can include but is not limited to: the initial feature is enhanced by a residual block subnetwork to obtain an enhanced feature, the enhanced feature is convoluted by a convolution subnetwork to obtain a convoluted feature, and the second feature is determined based on the convoluted feature. The third feature is processed by a second subnetwork in the attention subnetwork to obtain a fourth feature, which can include but is not limited to: the third feature is mapped by a feature mapping subnetwork to obtain a weight feature; the initial feature is enhanced by a residual enhancement subnetwork to obtain an enhanced feature; and the fourth feature is generated based on the initial feature, the enhanced feature, and the weight feature.
[0101] In the above embodiment, the initial feature corresponding to the attention subnetwork is determined based on the first feature, which can include but is not limited to: the first feature is enhanced by an initial enhancement subnetwork to obtain an enhanced feature, and the enhanced feature is up-sampled and convoluted by an up-sampling convolution subnetwork to obtain the initial feature corresponding to the attention subnetwork.
[0102] In a possible implementation, the second feature is adjusted based on the feature adjustment factor to obtain a third feature, which can include but is not limited to: if the second feature includes C*H*W feature values, C represents the number of channels, H represents the feature height, and W represents the feature width, then the feature adjustment values corresponding to the C*H*W feature values are determined based on the feature adjustment factor, and the C*H*W feature values are adjusted based on the feature adjustment values to obtain adjusted feature values corresponding to the C*H*W feature values. On this basis, the third feature can be generated based on the adjusted feature values corresponding to the C*H*W feature values.
[0103] For example, the feature adjustment values corresponding to the C*H*W feature values are determined based on the feature adjustment factor, which can include but is not limited to: if the feature adjustment factor includes C*H*W feature adjustment values, then the feature adjustment values corresponding to the C*H*W feature values are determined based on the C*H*W feature adjustment values; or, if the feature adjustment factor includes H*W feature adjustment values, then the feature adjustment values corresponding to the C*H*W feature values are determined based on the H*W feature adjustment values; or, if the feature adjustment factor includes C feature adjustment values, then the feature adjustment values corresponding to the C*H*W feature values are determined based on the C feature adjustment values.
[0104] For example, the feature adjustment based on the feature adjustment value on the C*H*W feature values to obtain the adjusted feature value corresponding to the C*H*W feature values can include but is not limited to: if one of the C*H*W feature values corresponds to one feature adjustment value, the adjusted feature value corresponding to the feature value can be determined by the following formula: yr=rf*y; wherein, yr represents the adjusted feature value, rf represents the feature adjustment value, and y represents the feature value. Or, if one of the C*H*W feature values corresponds to N+1 feature adjustment values, N is a positive integer, the adjusted feature value corresponding to the feature value can be determined by the following formula: yr=rf_0*y0+rf_1*y1+rf_2*y2+…+rf_N*yN; wherein, yr represents the adjusted feature value, rf_0, rf_1, rf_2, …, rf_N represent N+1 feature adjustment values, and y represents the feature value.
[0105] For example, the second neural network can include a reconstruction decoding subnetwork, and the current image block corresponding to the target feature can be obtained by the second neural network based on the target feature, which can include but is not limited to: the target feature is processed by the reconstruction decoding subnetwork to obtain the current image block corresponding to the reconstruction image block. For example, the reconstruction decoding subnetwork can include an up-sampling convolution subnetwork, and the current image block corresponding to the reconstruction image block can be obtained by up-sampling convolution of the target feature through the up-sampling convolution subnetwork. For another example, the reconstruction decoding subnetwork can include an up-sampling convolution subnetwork and a color space transformation subnetwork, and the current image block corresponding to the reconstruction image block can be obtained by up-sampling convolution of the target feature through the up-sampling convolution subnetwork and color space transformation of the feature after up-sampling convolution through the color space transformation subnetwork.
[0106] In a possible implementation, if the current image block enables the feature adjustment mode, the feature adjustment factor corresponding to the current image block is obtained based on the first feature, and the feature adjustment factor is encoded in the auxiliary code stream corresponding to the current image block. Or, if the current image block does not enable the feature adjustment mode, the feature adjustment factor corresponding to the current image block does not need to be obtained.
[0107] For example, the feature adjustment flag can be encoded in the code stream, which allows the current image block to enable the feature adjustment mode, or the feature adjustment flag prohibits the current image block to enable the feature adjustment mode. Or, the first value range of the feature value can be encoded in the code stream. Or, the second value range of the feature change value can be encoded in the code stream.
[0108] Exemplarily, the execution sequence described above is only an example given for the convenience of description, and in actual application, the execution sequence between steps can also be changed, and the execution sequence is not limited. Moreover, in other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in the specification, and the steps included in the method can be more or less than those described in the specification. In addition, a single step described in the specification can be divided into multiple steps for description in other embodiments; multiple steps described in the specification can also be combined into a single step for description in other embodiments.
[0109] From the above technical solutions, in the embodiments of the present application, an end-to-end video image compression method is proposed, which can realize the encoding of video images based on a neural network, and achieve the purpose of improving the encoding efficiency by combining a feature adjustment factor. By combining network structure design and auxiliary code streams (used to carry the feature adjustment factor), the neural network can effectively guarantee the quality of the reconstructed image block while maintaining low complexity, achieve the purpose of improving the encoding performance, and reduce the complexity.
[0110] Embodiment 3: For embodiments 1 and 2, regarding the processing process at the encoding end, please refer to Figure 4 of course, Figure 4 is only an example of the processing process at the encoding end, and the processing process at the encoding end is not limited.
[0111] After obtaining the current image block x (the current image block x can be the original image block x, that is, the input image block), the encoding end can analyze and transform the current image block x through an analysis transformation network (that is, a neural network) to obtain the image feature y corresponding to the current image block x. Wherein, the feature transformation of the current image block x through the analysis transformation network refers to: transforming the current image block x into the image feature y in the latent domain, so as to facilitate the operation of all subsequent processes in the latent domain.
[0112] The image can be divided into one image block, or can be divided into multiple image blocks, if the image is divided into one image block, the current image block x can also be the image, that is, the encoding and decoding process of the image block can also be directly used for the image.
[0113] After obtaining the image feature y, the encoding end performs coefficient hyperparameter feature transformation on the image feature y to obtain a coefficient hyperparameter feature z. For example, the image feature y can be input to a hyperparameter encoding network (i.e., a neural network), and the hyperparameter encoding network performs coefficient hyperparameter feature transformation on the image feature y to obtain the coefficient hyperparameter feature z. The hyperparameter encoding network can be a trained neural network, and the training process of the hyperparameter encoding network is not limited. The hyperparameter encoding network can perform coefficient hyperparameter feature transformation on the image feature y. After the image feature y in the latent domain is input to the hyperparameter encoding network, the hyperprior latent information z is obtained.
[0114] After obtaining the coefficient hyperparameter feature z, the encoding end can quantize the coefficient hyperparameter feature z to obtain a hyperparameter quantized feature corresponding to the coefficient hyperparameter feature z, i.e., Q(z). Figure 4 The Q operation in the above formula represents a quantization process. After obtaining the hyperparameter quantized feature corresponding to the coefficient hyperparameter feature z, the hyperparameter quantized feature is encoded to obtain a Bitstream #1 (i.e., a first code stream) corresponding to the current image block, i.e., AE(Q(z)). Figure 4 The AE operation in the above formula represents an encoding process, such as an entropy encoding process. Alternatively, the encoding end can directly encode the coefficient hyperparameter feature z to obtain the Bitstream #1 corresponding to the current image block. The hyperparameter quantized feature or the coefficient hyperparameter feature z carried in the Bitstream #1 is mainly used to obtain parameters of the mean and the probability distribution model.
[0115] After obtaining the Bitstream #1 corresponding to the current image block, the encoding end can send the Bitstream #1 corresponding to the current image block to the decoding end. The processing process of the decoding end on the Bitstream #1 corresponding to the current image block is described in subsequent embodiments.
[0116] After obtaining the Bitstream #1 corresponding to the current image block, the encoding end can also decode the Bitstream #1 to obtain a hyperparameter quantized feature, i.e., AD(Bitstream #1). Figure 4 The AD in the above formula represents a decoding process. Then, the hyperparameter quantized feature is dequantized to obtain a coefficient hyperparameter feature z_hat. The coefficient hyperparameter feature z_hat can be the same as or different from the coefficient hyperparameter feature z. Figure 4 The IQ operation in the above formula represents a dequantization process. Alternatively, after obtaining the Bitstream #1 corresponding to the current image block, the encoding end can also decode the Bitstream #1 to obtain the coefficient hyperparameter feature z_hat, without involving the dequantization process of the coefficient hyperparameter feature z_hat.
[0117] For the encoding process of Bitstream#1, a fixed probability density model encoding method can be used, and for the decoding process of Bitstream#1, a fixed probability density model decoding method can be used, and the encoding and decoding processes are not limited.
[0118] After obtaining the coefficient hyperparameter feature z_hat, the encoding end can perform context-based prediction based on the coefficient hyperparameter feature z_hat of the current image block and the residual feature y_hat of the previous image block (the determination process of the residual feature y_hat is described in subsequent embodiments) to obtain the prediction value mu (i.e., the mean mu) corresponding to the current image block. For example, the coefficient hyperparameter feature z_hat and the residual feature y_hat are input into the mean prediction network, and the mean prediction network determines the prediction value mu based on the coefficient hyperparameter feature z_hat and the residual feature y_hat. The prediction process is not limited. For the context-based prediction process, the input includes the coefficient hyperparameter feature z_hat and the decoded residual feature y_hat, which are jointly input to obtain a more accurate prediction value mu. The prediction value mu is used to obtain the residual by subtracting the original feature and to obtain the reconstruction y by adding the decoded residual.
[0119] It should be noted that the mean prediction network is an optional neural network, i.e., there can be no mean prediction network, i.e., the prediction value mu does not need to be determined by the mean prediction network. Figure 4 The dashed box in the above formula indicates that the mean prediction network is optional.
[0120] After obtaining the image feature y, the encoding end can determine the residual feature r based on the image feature y and the prediction value mu, such as the difference between the image feature y and the prediction value mu as the residual feature r. Then, the residual feature r is processed to obtain the image feature s, and the feature processing process is not limited and can be any feature processing method. In this case, the mean prediction network needs to be deployed to provide the prediction value mu. Alternatively, after obtaining the image feature y, the encoding end can process the image feature y to obtain the image feature s, and the feature processing process is not limited and can be any feature processing method. In this case, the mean prediction network is not needed, and the residual process is optional, as indicated by the dashed box.
[0121] After obtaining the image feature s, the encoding end can quantize the image feature s to obtain the image quantization feature corresponding to the image feature s, i.e., Q(s) in the above formula. Figure 4 The Q operation in the above formula is a quantization process. After obtaining the image quantization feature corresponding to the image feature s, the encoding end can encode the image quantization feature to obtain the Bitstream#2 (i.e., the second code stream) corresponding to the current image block, i.e., Bitstream#2 in the above formula. Figure 4The AE operation in the above formula represents an encoding process, such as an entropy encoding process. Alternatively, the encoding end can directly encode the image feature s to obtain Bitstream #2 corresponding to the current image block without involving the quantization process of the image feature s.
[0122] After obtaining the Bitstream #2 corresponding to the current image block, the encoding end can send the Bitstream #2 corresponding to the current image block to the decoding end. For the processing process of the decoding end for the Bitstream #2 corresponding to the current image block, see the subsequent embodiments.
[0123] After obtaining the Bitstream #2 corresponding to the current image block, the encoding end can also decode the Bitstream #2 to obtain the image quantized feature, that is, Figure 4 The AD in the above formula represents a decoding process. Then, the encoding end can dequantize the image quantized feature to obtain the image feature s', which can be the same as or different from the image feature s, Figure 4 The IQ operation in the above formula is a dequantization process. Alternatively, after obtaining the Bitstream #2 corresponding to the current image block, the encoding end can also decode the Bitstream #2 to obtain the image feature s' without involving the dequantization process of the image quantized feature.
[0124] After obtaining the image feature s', the encoding end can perform feature restoration (i.e., the inverse process of feature processing) on the image feature s' without limitation, which can be any feature restoration manner, to obtain the residual feature r_hat, which can be the same as or different from the residual feature r. After obtaining the residual feature r_hat, the encoding end determines the image feature y_hat based on the residual feature r_hat and the prediction value mu, which can be the same as or different from the image feature y, such as taking the sum of the residual feature r_hat and the prediction value mu as the image feature y_hat. In this case, a mean prediction network needs to be deployed to provide the prediction value mu. Alternatively, after obtaining the image feature s', the encoding end can perform feature restoration (i.e., the inverse process of feature processing) on the image feature s' to obtain the image feature y_hat, which can be the same as or different from the image feature y. In this case, the mean prediction network does not need to be deployed, and the residual process is optional, which is represented by the dashed box.
[0125] After obtaining the image feature y_hat, the encoding end can perform synthesis transformation on the image feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x, such as inputting the image feature y_hat to a synthesis transformation network to perform synthesis transformation on the image feature y_hat by the synthesis transformation network to obtain the reconstructed image block x_hat, thereby completing the image reconstruction process.
[0126] In one possible implementation, when the encoding end encodes the image quantization features or image features s to obtain Bitstream#2 corresponding to the current image block, the encoding end needs to first determine the probability distribution model, and then encode the image quantization features or image features s based on the probability distribution model. Furthermore, when the encoding end decodes Bitstream#2, it also needs to first determine the probability distribution model, and then decode Bitstream#2 based on the probability distribution model.
[0127] To obtain the probability distribution model, please refer to [link / reference]. Figure 4 As shown, after obtaining the coefficient hyperparameter feature z_hat, the encoder can perform an inverse hyperparameter feature transformation on z_hat to obtain the probability distribution parameter p. For example, the coefficient hyperparameter feature z_hat can be input into a probabilistic hyperparameter decoding network, which will then perform an inverse hyperparameter feature transformation on z_hat to obtain the probability distribution parameter p. After obtaining the probability distribution parameter p, a probability distribution model can be generated based on p. The probabilistic hyperparameter decoding network can be a trained neural network; the training process of this network is not restricted, as long as it can perform the inverse hyperparameter feature transformation on z_hat.
[0128] In one possible implementation, the above-mentioned encoding process can be executed by a deep learning model or a neural network model to achieve end-to-end image compression and encoding, without any restrictions on the encoding process.
[0129] Example 4: For the processing procedures at the decoding end in Examples 1 and 2, please refer to... Figure 5 As shown, of course, Figure 5 This is just one example of the processing procedure at the decoding end, and no restrictions are imposed on the processing procedure at the decoding end.
[0130] After obtaining Bitstream#1 corresponding to the current image block, the decoding end can further decode Bitstream#1 to obtain the hyperparameter quantization features, i.e. Figure 5 In the diagram, AD represents the decoding process. Then, the hyperparameter quantization features are dequantized to obtain the coefficient hyperparameter features z_hat. The coefficient hyperparameter features z_hat and the coefficient hyperparameter features z can be the same or different. Figure 5 The IQ operation in the code is an inverse quantization process. Alternatively, after obtaining Bitstream#1 corresponding to the current image block, the decoding end can also decode Bitstream#1 to obtain the coefficient hyperparameter feature z_hat, without involving the inverse quantization process of the coefficient hyperparameter feature z_hat.
[0131] For the decoding process of Bitstream#1, a decoding method of fixed probability density model can be adopted, and no limitation is made.
[0132] The image can be divided into one image block or multiple image blocks. If the image is divided into one image block, the current image block x can also be the image, that is, the decoding process of the image block can also be directly used for the image.
[0133] After obtaining the coefficient hyperparameter feature z_hat, the decoding end can perform context-based prediction based on the coefficient hyperparameter feature z_hat of the current image block and the residual feature y_hat of the previous image block (the determination process of the residual feature y_hat is described in subsequent embodiments), to obtain the prediction value mu (i.e., the mean mu) corresponding to the current image block. For example, the coefficient hyperparameter feature z_hat and the residual feature y_hat are input into the mean prediction network, and the mean prediction network determines the prediction value mu based on the coefficient hyperparameter feature z_hat and the residual feature y_hat. The prediction process is not limited. For the context-based prediction process, the input includes the coefficient hyperparameter feature z_hat and the decoded residual feature y_hat, and the two are jointly input to obtain a more accurate prediction value mu.
[0134] It should be noted that the mean prediction network is an optional neural network, that is, there can be no mean prediction network, that is, the prediction value mu does not need to be determined by the mean prediction network. Figure 5 The dashed box in the above formula indicates that the mean prediction network is optional.
[0135] After obtaining the Bitstream#2 corresponding to the current image block, the decoding end can also decode the Bitstream#2 to obtain the image quantization feature, that is, Figure 5 AD in the above formula indicates the decoding process. Then, the decoding end can perform inverse quantization on the image quantization feature to obtain the image feature s', which can be the same as or different from the image feature s, Figure 5 The IQ operation in the above formula is the inverse quantization process. Alternatively, after obtaining the Bitstream#2 corresponding to the current image block, the decoding end can also decode the Bitstream#2 to obtain the image feature s', without involving the inverse quantization process of the image quantization feature.
[0136] After obtaining the image feature s', the decoding end can perform feature restoration (i.e., the inverse process of feature processing) on the image feature s' to obtain a residual feature r_hat, which is the same as or different from the residual feature r. After obtaining the residual feature r_hat, the decoding end determines the image feature y_hat based on the residual feature r_hat and the prediction value mu, which is the same as or different from the image feature y, such as taking the sum of the residual feature r_hat and the prediction value mu as the image feature y_hat. In this case, a mean prediction network needs to be deployed to provide the prediction value mu. Alternatively, after obtaining the image feature s', the decoding end can perform feature restoration on the image feature s' to obtain the image feature y_hat, which can be the same as or different from the image feature y. In this case, a mean prediction network does not need to be deployed, and the residual process is optional, as indicated by the dashed box.
[0137] After obtaining the image feature y_hat, the decoding end can perform synthesis transformation on the image feature y_hat to obtain the reconstructed image block x_hat corresponding to the current image block x, such as inputting the image feature y_hat to a synthesis transformation network to perform synthesis transformation on the image feature y_hat by the synthesis transformation network to obtain the reconstructed image block x_hat, thereby completing the image reconstruction process.
[0138] In a possible implementation, when decoding Bitstream#2, the decoding end needs to first determine a probability distribution model, and then decode Bitstream#2 based on the probability distribution model. To obtain the probability distribution model, continue to refer to FIG. 8. Figure 5 As shown in FIG. 8, after obtaining the coefficient hyperparameter feature z_hat, the decoding end can perform coefficient hyperparameter feature inverse transformation on the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter p, such as inputting the coefficient hyperparameter feature z_hat to a probability hyperparameter decoding network to perform coefficient hyperparameter feature inverse transformation on the coefficient hyperparameter feature z_hat by the probability hyperparameter decoding network to obtain the probability distribution parameter p, and after obtaining the probability distribution parameter p, the probability distribution model can be generated based on the probability distribution parameter p. The probability hyperparameter decoding network can be a trained neural network, and the training process of the probability hyperparameter decoding network is not limited. The probability hyperparameter decoding network can only perform coefficient hyperparameter feature inverse transformation on the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter p.
[0139] In a possible implementation, the processing process of the decoding end described above can be performed by a deep learning model or a neural network model, thereby realizing an end-to-end image compression and encoding process, and the decoding process is not limited.
[0140] Embodiment 5: For Embodiment 1, Embodiment 2, Embodiment 3 and Embodiment 4, the analysis transform network is involved at the encoding end, the synthesis transform network is involved at the encoding end and the decoding end, and the analysis transform network and the synthesis transform network are both neural networks. In a possible implementation, a structural diagram of the analysis transform network can refer to FIG. 5A. Figure 6A A structural diagram of the synthesis transform network can refer to FIG. 5B as shown in the figure. Of course, Figure 6B Figure 6A Figure 6B This analysis transform network and the synthesis transform network in the embodiment are only examples, and the structure of the analysis transform network and the synthesis transform network is not limited in the embodiment, as long as the analysis transform function and the synthesis transform function can be implemented.
[0141] For the analysis transform network, the Padding layer, the Conv layer, the ResAU layer, the Padding layer, the RNAB layer, the Conv layer, the ResAU layer, the Padding layer, the Conv layer, the ResAU layer, the Padding layer, the Conv layer, and the Conv layer are sequentially included. The Padding layer is a padding layer, and is used to implement an edge expansion operation of a feature. The Conv layer is a convolution layer, and is used to implement a convolution operation of the feature. C*3*3 represents that a 3*3 convolution kernel (C channels) is used to implement the convolution operation of the feature. C*1*1 represents that a 1*1 convolution kernel (C channels) is used to implement the convolution operation of the feature. The down arrow of 2 represents that the feature is down-sampled by 2.
[0142] The ResAU layer is an activation layer. A structural diagram of the ResAU layer can refer to FIG. 6 as shown in the figure. LeakyReLU represents an activation operation. Conv 1*1 represents that a 1*1 convolution kernel is used to implement a convolution operation of a feature. Tanh represents a hyperbolic tangent operation. Figure 6C
[0143] The RNAB layer is a residual non-local attention block. A structural diagram of the RNAB layer can refer to FIG. 7 as shown in the figure. The RNAB layer can include the RB layer, the Conv layer, and the sigmoid layer. In the Conv layer, the up arrow of 2 represents that the feature is up-sampled by 2. The sigmoid layer is used to perform feature mapping on the feature, for example, the sigmoid function is used to map the feature to between 0 and 1. Figure 6D
[0144] The RB (Residual Blocks) layer is a residual block. A structural diagram of the RB layer can refer to FIG. 8 as shown in the figure. The RB layer can include the Conv layer, LeakyReLU, and the Conv layer. In the Conv layer, a 3*3 convolution kernel is used. Figure 6E
[0145] For the synthesis transformation network, it sequentially includes a ResBlock layer (i.e., an RB layer), a ResBlock layer, a Conv layer, a Cropping layer, a ResAU layer, a Conv layer, a Cropping layer, a ResAU layer, a Conv layer, an RNAB layer, a Cropping layer, a ResAU layer, a Conv layer, and a Cropping layer. The ResBlock layer is a residual block, and a structure diagram of the ResBlock layer can be seen from FIG. 6. The Conv layer is a convolution layer, which is used to implement a convolution operation of a feature, and C*3*3 indicates that a 3*3 convolution kernel (C channels) is used to implement the convolution operation of the feature, and the upper arrow of 2 indicates that the feature is up-sampled by 2 times. The Cropping layer is a cropping layer, which is used to implement a cropping operation of a feature, i.e., an inverse operation of a padding operation. The ResAU layer is an activation layer, and a structure diagram of the ResAU layer can be seen from FIG. 7. The RNAB layer is a residual non-local attention block, and a structure diagram of the RNAB layer can be seen from FIG. 8. The RNAB layer can include an RB layer, a Conv layer, and a sigmoid layer. Figure 6E Figure 6C Figure 6D
[0146] Embodiment 6: For embodiment 5, from Figure 6B It can be seen that the synthesis transformation network mainly includes four times of 3*3 convolution layers for up-sampling, two ResBlock layers (the ResBlock layer includes two 3*3 convolution layers), three activation layers ResAU (the ResAU includes one 1*1 convolution layer), and one RNAB layer (the RNAB layer includes nine RB blocks and three 3*3 convolution layers, i.e., a total of 21 3*3 convolution layers). As can be seen from the above, the synthesis transformation network has a total of 29 3*3 convolution layers and three 1*1 convolution layers, and the RNAB layer accounts for nearly 70% of the complexity, i.e., the complexity of the RNAB layer has a large optimization space. Based on this, in this embodiment, the synthesis transformation network can be optimized, and a low-complexity and high-performance synthesis transformation network is designed. Alternatively, the analysis transformation network can also be optimized, and a low-complexity and high-performance analysis transformation network is designed. Of course, other high-complexity network structures can also be optimized, and a low-complexity and high-performance network structure is designed, which is not limited in this embodiment. For the convenience of description, the synthesis transformation network is taken as an example in this embodiment, for example, the complex network similar to the RNAB layer in the synthesis transformation network can be optimized. Of course, other network sides in the synthesis transformation network can also be optimized, which is not limited, as long as the complexity of the synthesis transformation network can be simplified.
[0147] In a possible implementation, for the synthesis transformation network, the RNAB layer is replaced by a low-complexity and high-performance network structure. Figure 6D As shown in the RNAB layer, the RNAB layer can include 9 RB layers, 3 Conv layers and 1 sigmoid layer, on the basis of which, part of the network layer can be removed to obtain an optimized RNAB layer, see Figure 6F As shown, 3 RB layers and 2 Conv layers can be removed to obtain an optimized RNAB layer. Of course, other network layers can also be removed to obtain an optimized RNAB layer, and the removal of the network layer is not limited.
[0148] For example, the optimized RNAB layer can refer to Figure 6G As shown, the RNAB layer can include an RB layer (composed of at least one RB block, such as composed of 3 consecutive RB blocks, and of course, the number of RB blocks can be more or less), an RB layer (composed of at least one RB block, such as composed of 3 consecutive RB blocks, and of course, the number of RB blocks can be more or less), 1 Conv layer (such as a 3*3 Conv layer) and 1 sigmoid layer (i.e. feature mapping layer).
[0149] For example, for the 3*3 Conv layer, it can also be replaced by a 1*1 Conv layer or a 5*5 Conv layer. Taking the 1*1 Conv layer as an example, the optimized RNAB layer can refer to Figure 6H As shown, the RNAB layer can include an RB layer, an RB layer, 1 Conv layer (such as a 1*1 Conv layer) and 1 sigmoid layer (i.e. feature mapping layer).
[0150] For example, after obtaining the optimized RNAB layer, the optimized RNAB layer can be substituted into Figure 6B to obtain an optimized synthesis transformation network. Based on the optimized synthesis transformation network, the image feature y_hat can be input to the synthesis transformation network, and the synthesis transformation network can be used to synthesize and transform the image feature y_hat to obtain the reconstructed image block x_hat.
[0151] In embodiment 7, the synthesis transformation network can be optimized to design a low-complexity and high-performance synthesis transformation network. Alternatively, the analysis transformation network can also be optimized to design a low-complexity and high-performance analysis transformation network. Of course, other high-complexity network structures can also be optimized to design a low-complexity and high-performance network structure. In this embodiment, the method of “simplified network structure + adaptive adjustment factor” can be used for encoding and decoding.
[0152] Firstly, the complex network is simplified to obtain a simplified network. For example, the synthesis transformation network can be simplified to obtain a simplified synthesis transformation network. For example, the RNAB layer in the synthesis transformation network (i.e., the complex network) is simplified to obtain a simplified synthesis transformation network. Of course, other networks can also be simplified, and this is not limited.
[0153] Then, at least one layer feature (e.g., a key feature) of the simplified network is adjusted. For example, the simplified network is trained based on a large number of images, and the generated features have a general effect. However, for a specific image frame, the features are often not optimal, and therefore, customized features need to be designed for the image. In order to obtain the customized features, a feature adjustment factor needs to be introduced in a certain process of the network, and the feature is adjusted based on the feature adjustment factor to obtain adjusted customized features. For example, the encoding end adaptively selects a feature adjustment factor for the current image block, and encodes the feature adjustment factor in the code stream. The decoding end can parse the feature adjustment factor from the code stream.
[0154] In a possible implementation, for the decoding end, the decoding method of the decoding end can refer to Figure 7A As shown in the figure, the decoding end can receive the main code stream (i.e., the first code stream Bitstream#1 and the second code stream Bitstream#2) for the current image block. After the main code stream is subjected to coefficient decoding and decoding sub-network, the first feature corresponding to the current image block can be obtained. The decoding end can decode the feature adjustment factor corresponding to the current image block from the auxiliary code stream. After obtaining the first feature corresponding to the current image block, the first feature can be subjected to feature enhancement by the feature enhancement network to obtain the second feature, and the second feature is subjected to feature adjustment by the feature adjustment factor corresponding to the current image block to obtain the third feature. After obtaining the third feature, the target feature corresponding to the current image block can be determined based on the third feature, and the target feature is input to the reconstruction decoding network. The target feature is processed by the reconstruction decoding network to obtain the reconstructed image block corresponding to the current image block.
[0155] For example, the feature adjustment factor can be decoded by the decoding end from the auxiliary code stream. The feature adjustment factor can also be a fixed parameter value. When the feature adjustment factor is a fixed parameter value, such as a fixed parameter value 1, it is equivalent to not adjusting the second feature, i.e., not involving the feature adjustment process, which is equivalent to the scheme of embodiment 6.
[0156] For example, the decoding end can also decode a feature adjustment factor (which can be the same as or different from the feature adjustment factor for the second feature) for the decoding sub-network from the auxiliary code stream. After obtaining the feature adjustment factor for the decoding sub-network, the decoding sub-network can perform feature adjustment on a certain feature in the decoding sub-network based on the feature adjustment factor. For example, after obtaining feature A (which is any feature in the decoding sub-network), the decoding sub-network can perform feature adjustment on feature A using the feature adjustment factor to obtain adjusted feature B. The decoding sub-network continues to process based on feature B, and finally obtains the first feature. By using the feature adjustment factor in the decoding process of the decoding sub-network, the first feature with smaller distortion can be generated.
[0157] In a possible implementation, for the encoding end, the encoding method of the encoding end can refer to Figure 7B As shown in the figure, after the current image block is processed by the encoding network and the coefficient encoding, the main code stream (i.e., the first code stream Bitstream#1 and the second code stream Bitstream#2) corresponding to the current image block can be obtained. The encoding end can send the main code stream of the current image block to the decoding end.
[0158] After the main code stream is processed by the coefficient decoding and the decoding sub-network, the first feature corresponding to the current image block can be obtained. The feature adjustment factor corresponding to the current image block can be obtained based on the first feature, and the feature adjustment factor can be encoded in the auxiliary code stream corresponding to the current image block, so that the decoding end can decode the feature adjustment factor corresponding to the current image block from the auxiliary code stream.
[0159] For example, the encoding end can obtain at least one candidate feature adjustment factor (for example, a candidate feature adjustment factor list can be constructed in advance, all feature adjustment factors in the candidate feature adjustment factor list can be taken as candidate feature adjustment factors, or at least one candidate feature adjustment factor can be generated by using an algorithm, and no limitation is made to this). The rate-distortion cost corresponding to each candidate feature adjustment factor can be determined based on the first feature. For example, for each candidate feature adjustment factor, after the first feature corresponding to the current image block is obtained, the first feature can be enhanced by the feature enhancement network to obtain a second feature, and the second feature can be adjusted by using the candidate feature adjustment factor to obtain a third feature. After the third feature is obtained, the target feature corresponding to the current image block can be determined based on the third feature, and the target feature is input into the reconstruction decoding network, and the target feature is processed by the reconstruction decoding network to obtain the reconstructed image block corresponding to the current image block. After the reconstructed image block is obtained, the rate-distortion cost corresponding to the candidate feature adjustment factor can be determined based on the loss between the reconstructed image block and the current image block, so as to obtain the rate-distortion cost corresponding to each candidate feature adjustment factor. After the rate-distortion cost corresponding to each candidate feature adjustment factor is obtained, a candidate feature adjustment factor is selected from all candidate feature adjustment factors as the feature adjustment factor corresponding to the current image block based on the rate-distortion cost corresponding to each candidate feature adjustment factor, for example, the candidate feature adjustment factor with the minimum rate-distortion cost is taken as the feature adjustment factor corresponding to the current image block.
[0160] For example, the encoding end can also obtain a feature adjustment factor for the decoding subnetwork, and adjust a certain feature in the decoding subnetwork based on the feature adjustment factor. For example, after the decoding subnetwork obtains feature A (feature A is any feature in the decoding subnetwork), the feature adjustment factor is used to adjust feature A to obtain adjusted feature B. The decoding subnetwork continues to process based on feature B, and finally obtains the first feature. On this basis, the encoding end can also encode the feature adjustment factor for the decoding subnetwork (which is the same as or different from the feature adjustment factor for the second feature) in the auxiliary code stream, so that the decoding end decodes the feature adjustment factor for the decoding subnetwork from the auxiliary code stream.
[0161] Embodiment 8: For embodiments 1-7, for the encoding end and the decoding end, the first feature corresponding to the current image block can be obtained by the first neural network, for example, the probability distribution parameter corresponding to the current image block is obtained based on the first code stream, the probability distribution model is determined based on the probability distribution parameter, the decoded image feature is obtained by decoding the second code stream corresponding to the current image block based on the probability distribution model, and the first feature corresponding to the current image block is determined based on the decoded image feature.
[0162] For example, the first neural network can be a decoding subnetwork. For example, the decoding subnetwork is illustrated in FIG. 3. As shown in FIG. 3, the network in the dashed box is the decoding subnetwork, and y_hat is the first feature output by the decoding subnetwork. Figure 8A
[0163] For example, after obtaining the Bitstream #1 corresponding to the current image block, the encoding end or the decoding end can decode the Bitstream #1 to obtain the hyperparameter quantized feature, and then dequantize the hyperparameter quantized feature to obtain the coefficient hyperparameter feature z_hat. Alternatively, after obtaining the Bitstream #1 corresponding to the current image block, the encoding end or the decoding end can decode the Bitstream #1 to obtain the coefficient hyperparameter feature z_hat, without involving the dequantization process of the coefficient hyperparameter feature z_hat.
[0164] After obtaining the coefficient hyperparameter feature z_hat, the encoding end or the decoding end can perform coefficient hyperparameter feature inverse transformation on the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter p. For example, the coefficient hyperparameter feature z_hat is input into the probability hyperparameter decoding network, and the probability hyperparameter decoding network performs coefficient hyperparameter feature inverse transformation on the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter p. After obtaining the probability distribution parameter p, the probability distribution model can be generated based on the probability distribution parameter p.
[0165] For example, after obtaining the Bitstream #2 corresponding to the current image block, the encoding end or the decoding end can decode the Bitstream #2 to obtain the image quantized feature, and then dequantize the image quantized feature to obtain the image feature s’. Alternatively, after obtaining the Bitstream #2 corresponding to the current image block, the encoding end or the decoding end can decode the Bitstream #2 to obtain the image feature s’, without involving the dequantization process of the image quantized feature. For example, when decoding the Bitstream #2, the encoding end or the decoding end can decode the Bitstream #2 based on the probability distribution model.
[0166] After obtaining the image feature s’, the encoding end or the decoding end can perform feature restoration on the image feature s’ to obtain the image feature y_hat, and the image feature y_hat can be used as the first feature, i.e., the first feature y_hat output by the decoding subnetwork.
[0167] Embodiment 9: For the encoder and the decoder, the first feature corresponding to the current image block can be obtained by the first neural network, for example, based on the first bitstream corresponding to the current image block to obtain the probability distribution parameters and the prediction value (such as the mean value); the probability distribution model is determined based on the probability distribution parameters, and the second bitstream corresponding to the current image block is decoded based on the probability distribution model to obtain the decoded image feature; the residual feature is obtained by performing residual recovery on the decoded image feature; and the first feature corresponding to the current image block is determined based on the residual feature and the prediction value.
[0168] For example, the first neural network can be a decoding subnetwork. Taking the decoding subnetwork as an example, referring to FIG. 8, the network in the dashed box is the decoding subnetwork, and y_hat is the first feature output by the decoding subnetwork. Figure 8B
[0169] For example, after obtaining the Bitstream#1 corresponding to the current image block, the encoder or the decoder can decode the Bitstream#1 to obtain the hyperparameter quantized feature, and perform inverse quantization on the hyperparameter quantized feature to obtain the coefficient hyperparameter feature z_hat. Alternatively, after obtaining the Bitstream#1 corresponding to the current image block, the encoder or the decoder can also decode the Bitstream#1 to obtain the coefficient hyperparameter feature z_hat, without involving the inverse quantization process of the coefficient hyperparameter feature z_hat.
[0170] After obtaining the coefficient hyperparameter feature z_hat, the encoder or the decoder can perform coefficient hyperparameter feature inverse transformation on the coefficient hyperparameter feature z_hat to obtain the probability distribution parameter p, for example, inputting the coefficient hyperparameter feature z_hat to the probability hyperparameter decoding network to perform coefficient hyperparameter feature inverse transformation on the coefficient hyperparameter feature z_hat by the probability hyperparameter decoding network to obtain the probability distribution parameter p. After obtaining the probability distribution parameter p, the probability distribution model can be generated based on the probability distribution parameter p.
[0171] After obtaining the coefficient hyperparameter feature z_hat, the encoder or the decoder can perform context-based prediction based on the coefficient hyperparameter feature z_hat of the current image block and the residual feature y_hat of the previous image block to obtain the prediction value mu (i.e., the mean value mu), for example, inputting the coefficient hyperparameter feature z_hat and the residual feature y_hat to the mean value prediction network to determine the prediction value mu based on the coefficient hyperparameter feature z_hat and the residual feature y_hat by the mean value prediction network, and the prediction process is not limited. For the context-based prediction process, the input of the mean value prediction network can include the coefficient hyperparameter feature z_hat and the decoded residual feature y_hat, and the two are jointly input to obtain a more accurate prediction value mu.
[0172] The encoding end or the decoding end can further decode the Bitstream #2 to obtain the image quantization feature after obtaining the Bitstream #2 corresponding to the current image block. The image quantization feature is dequantized to obtain the image feature s'. Alternatively, the encoding end or the decoding end can decode the Bitstream #2 to obtain the image feature s' after obtaining the Bitstream #2 corresponding to the current image block, without involving the dequantization process of the image quantization feature. Illustratively, the encoding end or the decoding end can decode the Bitstream #2 based on the probability distribution model when decoding the Bitstream #2.
[0173] The encoding end or the decoding end can further decode the Bitstream #2 to obtain the image quantization feature after obtaining the Bitstream #2 corresponding to the current image block. The image quantization feature is dequantized to obtain the image feature s'. Alternatively, the encoding end or the decoding end can decode the Bitstream #2 to obtain the image feature s' after obtaining the Bitstream #2 corresponding to the current image block, without involving the dequantization process of the image quantization feature. Illustratively, the encoding end or the decoding end can decode the Bitstream #2 based on the probability distribution model when decoding the Bitstream #2.
[0174] Illustratively, Figure 8B With Figure 8A In the decoding sub-network, in addition to obtaining the probability distribution parameters for the second code stream based on the first code stream, the prediction value mu for the first feature can also be generated based on the first code stream. The feature residual is obtained by decoding the second code stream, and the residual r_hat of the first feature is obtained by residual recovery. The first feature can be obtained based on mu and r_hat.
[0175] Embodiment 10: For embodiments 1-9, for the encoding end and the decoding end, after obtaining the first feature y_hat, the first feature y_hat can be synthesized and transformed by a synthesis transformation network to obtain the reconstructed image block x_hat corresponding to the current image block x. In the synthesis transformation process, the first feature y_hat can be enhanced to obtain a second feature; the second feature is adjusted based on a feature adjustment factor to obtain a third feature; after obtaining the third feature, a target feature is determined based on the third feature. Then, the reconstructed image block x_hat corresponding to the current image block x is obtained based on the target feature.
[0176] See Figure 6BAs shown, the structure diagram of the ResAU layer can be referred to as Figure 6E As shown, the structure diagram of the ResAU layer can be referred to as Figure 6C As shown.
[0177] The RNAB layer is a residual non-local attention block. The structure diagram of the RNAB layer can be referred to as Figure 6D As shown, or, part of the network layer of Figure 6D may be removed to obtain an optimized RNAB layer. As shown, Figure 6F 3 RB layers and 2 Conv layers can be removed to obtain an optimized RNAB layer. Therefore, the structure diagram of the RNAB layer can be referred to as Figure 6G As shown, or, the structure diagram of the RNAB layer can be referred to as Figure 6H As shown. In Figure 6G and Figure 6H , the RNAB layer can include an RB layer (composed of at least one RB block, such as composed of three consecutive RB blocks), an RB layer (composed of at least one RB block, such as composed of three consecutive RB blocks), one Conv layer, and one sigmoid layer (i.e., a feature mapping layer). For the Conv layer, it can be a 3*3 Conv layer, a 5*5 Conv layer, a 1*1 Conv layer, and also other sizes of Conv layers, such as a 7*7 Conv layer, a 9*9 Conv layer, and the size of the Conv layer is not limited.
[0178] For the above-mentioned synthesis transformation network, the synthesis transformation network can be divided into a feature enhancement sub-network and a reconstruction decoding sub-network. The feature enhancement sub-network can include an initial enhancement sub-network, a first up-sampling convolution sub-network, and an attention sub-network. The reconstruction decoding sub-network can include a second up-sampling convolution sub-network and a color space transformation sub-network.
[0179] For example, all the networks before the RNAB layer in the synthesis transformation network can be regarded as the initial enhancement sub-network and the first up-sampling convolution sub-network. The initial enhancement sub-network can include two RB networks, and the first up-sampling convolution sub-network can include at least one up-sampling convolution network (such as three up-sampling convolution networks). For example, as shown in Figure 6BAs shown, all the networks before the RNAB layer are RB layer, RB layer, Conv layer, Cropping layer, ResAU layer, Conv layer, Cropping layer, ResAU layer, Conv layer in turn, based on which, the initial enhancement subnetwork can include the RB layer and the RB layer, and the first up-sampling convolution subnetwork can include the Conv layer, the Cropping layer, the ResAU layer, the Conv layer, the Cropping layer, the ResAU layer, and the Conv layer, that is, 3 up-sampling convolution networks (Conv layers) are involved, and here, a 3*3 Conv layer is taken as an example.
[0180] As an example, the RNAB layer in the synthesis transformation network can be taken as an attention subnetwork (i.e., an attention subnetwork), and the attention subnetwork can effectively reduce the complexity by removing the up-sampling and down-sampling convolution networks, as shown in Figure 6G and Figure 6H As shown, the structure diagram of the attention subnetwork, the attention subnetwork can include the RB layer (composed of at least one RB block), the RB layer (composed of at least one RB block), 1 Conv layer, and 1 sigmoid layer (i.e., a feature mapping layer).
[0181] The attention subnetwork can include a residual skip subnetwork, a residual enhancement subnetwork, and a weight generation subnetwork. For the residual skip subnetwork, x_out0=x_in, x_in is the input feature of the residual skip subnetwork, and x_out0 is the output feature of the residual skip subnetwork. For the residual enhancement subnetwork, x_out1=RB(x_in), x_in is the input feature of the residual enhancement subnetwork, and x_out1 is the output feature of the residual enhancement subnetwork. For the weight generation subnetwork, the weight feature k of the output feature x_out1 of the residual enhancement subnetwork is obtained. Therefore, the final output of the attention subnetwork is: x_out=x_in+k*x_out1.
[0182] As shown in Figure 9A and Figure 9BAs shown, the residual skip subnetwork is used to superimpose the input feature x_in (i.e., x_out0) onto the final output feature of the attention subnetwork. The residual enhancement subnetwork includes at least one RB block (such as one RB block or three RB blocks, etc.), and the residual enhancement subnetwork is used to process the input feature x_in based on the RB block to obtain the output feature x_out1. The weight generation subnetwork can include at least one RB block (such as one RB block or three RB blocks, etc.), one Conv layer (such as a 3*3 Conv layer, a 5*5 Conv layer, a 1*1 Conv layer, etc.), and a feature mapping layer (such as using a sigmoid processing process to realize feature mapping), the weight generation subnetwork is used to process the input feature x_in based on the RB block, then the Conv layer is used to convolve the feature processed by the RB block, then the feature mapping layer is used to map the feature processed by the Conv layer, to obtain the weight feature k, the dimension of the weight feature k is the same as the dimension of the output feature x_out1 of the residual enhancement subnetwork, after multiplying the output feature x_out1 with the weight feature k, and then adding the output feature x_out0 of the residual skip subnetwork, the final output feature of the attention subnetwork is obtained as: x_out = x_in + k*x_out1.
[0183] For example, all the networks after the RNAB layer in the synthesis transformation network can be taken as the reconstruction decoding subnetwork, and the reconstruction decoding subnetwork can include a second up-sampling convolution subnetwork and a color space transformation subnetwork, that is, all the networks after the RNAB layer are taken as the second up-sampling convolution subnetwork and the color space transformation subnetwork. The second up-sampling convolution subnetwork can include at least one up-sampling convolution network (such as one up-sampling convolution network), and the color space transformation subnetwork is used to realize the image transformation process from the YUV domain to the RBG or the filtering process from the YUV domain to the YUV domain.
[0184] For example, referring to Figure 6B As shown, all the networks after the RNAB layer are in turn a Cropping layer, a ResAU layer, a Conv layer, and a Cropping layer, and therefore, the second up-sampling convolution subnetwork can include a Cropping layer, a ResAU layer, a Conv layer, and a Cropping layer, that is, it involves one up-sampling convolution network (Conv layer), and here a 3*3 Conv layer is taken as an example. If there is a color space transformation requirement, a color space transformation subnetwork (not shown in the figure) can be included, which realizes the image transformation process from the YUV domain to the RBG or the filtering process from the YUV domain to the YUV domain. If there is no color space transformation requirement, the color space transformation subnetwork can not be included. Figure 6B
[0185] After the synthesis transformation network is divided into the initial enhancement sub-network, the first up-sampling convolution sub-network, the attention sub-network, and the reconstruction decoding sub-network (the reconstruction decoding sub-network includes the second up-sampling convolution sub-network and the color space transformation sub-network), the first feature y_hat can be subjected to synthesis transformation based on the initial enhancement sub-network, the first up-sampling convolution sub-network, the attention sub-network, and the reconstruction decoding sub-network to obtain the reconstructed image block x_hat corresponding to the current image block x.
[0186] In embodiment 11, the first feature y_hat can be subjected to synthesis transformation based on the initial enhancement sub-network, the first up-sampling convolution sub-network, the attention sub-network, and the reconstruction decoding sub-network to obtain the reconstructed image block x_hat corresponding to the current image block x. In a possible implementation, the output feature of the feature enhancement sub-network can be adjusted by using a feature adjustment factor. The adjustment position of the feature adjustment factor can be seen from Figure 10A In the adjustment manner of the feature adjustment factor, the first feature y_hat can be subjected to synthesis transformation to obtain the reconstructed image block x_hat in the following manner.
[0187] The first feature y_hat is subjected to feature enhancement by using the initial enhancement sub-network to obtain an enhanced feature. For example, the first feature y_hat can be input into the initial enhancement sub-network, and the first feature y_hat is subjected to feature enhancement by the initial enhancement sub-network to obtain an enhanced feature. For example, the initial enhancement sub-network can include an RB layer and an RB layer, and thus the first feature y_hat can be subjected to feature enhancement by the two RB layers to obtain an enhanced feature, and the process is not limited in this regard.
[0188] After the enhanced feature is obtained, the enhanced feature is subjected to up-sampling convolution by using the first up-sampling convolution sub-network to obtain an initial feature corresponding to the attention sub-network, i.e., an input feature x_in corresponding to the attention sub-network. For example, the enhanced feature (i.e., the output feature of the initial enhancement sub-network) can be input into the first up-sampling convolution sub-network, and the enhanced feature is subjected to up-sampling convolution by the first up-sampling convolution sub-network to obtain an initial feature corresponding to the attention sub-network. For example, the first up-sampling convolution sub-network can include a Conv layer, a Cropping layer, a ResAU layer, a Conv layer, a Cropping layer, a ResAU layer, and a Conv layer, and the enhanced feature can be subjected to up-sampling convolution by using these network layers, and the process is not limited in this regard.
[0189] After obtaining the initial feature x_in corresponding to the attention sub-network, the attention sub-network can be used to enhance the initial feature x_in to obtain the second feature. For example, the initial feature x_in can be input into the residual skip sub-network, the residual enhancement sub-network, and the weight generation sub-network respectively. After obtaining the initial feature x_in, the residual skip sub-network superimposes the initial feature x_in (i.e., x_out0) onto the final output feature of the attention sub-network. After obtaining the initial feature x_in, the residual enhancement sub-network enhances the initial feature x_in to obtain the enhanced feature. For example, the residual enhancement sub-network enhances the initial feature x_in based on RB blocks to obtain the enhanced feature x_out1. After obtaining the initial feature x_in, the weight generation sub-network generates the weight feature corresponding to the initial feature x_in, that is, the weight feature k of the output feature x_out1 of the residual enhancement sub-network.
[0190] For example, the weight generation subnetwork can include a residual block subnetwork, a convolutional subnetwork, and a feature mapping subnetwork. The residual block network can be used to perform convolutional activation on the initial feature x_in to obtain the convolutionally activated feature. The convolutional subnetwork can then be used to convolve the convolutionally activated feature to obtain the convolutionally generated feature. Finally, the feature mapping subnetwork can be used to map the convolutionally generated feature to obtain the weight feature k. See also Figure 10A As shown, the weight generation subnetwork may include at least one RB block (such as one or three RB blocks, i.e., a residual block subnetwork), one Conv layer (such as a 3*3 Conv layer, a 5*5 Conv layer, a 1*1 Conv layer, etc., i.e., a convolutional subnetwork), and a feature mapping layer (such as using a sigmoid process to implement feature mapping, i.e., a feature mapping subnetwork). Based on this, the initial feature x_in can be convolved and activated based on the RB block to obtain the convolved activated feature. Then, the Conv layer can be used to convolve the convolved activated feature to obtain the convolved feature. Finally, the feature mapping layer can be used to perform feature mapping on the convolved feature to obtain the weight feature k.
[0191] See Figure 10A As shown, after obtaining the enhanced feature x_out1, the weight feature k, and the initial feature x_in (i.e., x_out0), a second feature can be generated based on the initial feature x_in, the enhanced feature x_out1, and the weight feature k. The second feature can be the final output feature of the attention sub-network. For example, multiplying the enhanced feature x_out1 by the weight feature k and then adding it to the output feature x_out0 of the residual skip sub-network, the second feature is: x_out = x_in + k * x_out1.
[0192] After obtaining the second feature, the second feature can be feature adjusted based on a feature adjustment factor, and the feature adjusted second feature is taken as a third feature. After obtaining the third feature, the third feature can be determined as the target feature.
[0193] After obtaining the target feature, the target feature can be processed by a reconstruction decoding subnetwork to obtain a reconstructed image block x hat corresponding to the current image block x. For example, the target feature can be input into the reconstruction decoding subnetwork, and the target feature is processed by the reconstruction decoding subnetwork to obtain the reconstructed image block x hat corresponding to the current image block x. For example, the reconstruction decoding subnetwork can include a second up-sampling convolution subnetwork, and the target feature can be up-sampled and convolved by the second up-sampling convolution subnetwork to obtain the reconstructed image block x hat corresponding to the current image block x. Alternatively, the reconstruction decoding subnetwork can include a second up-sampling convolution subnetwork and a color space transformation subnetwork, and the target feature can be up-sampled and convolved by the second up-sampling convolution subnetwork, and the up-sampled and convolved feature can be color space transformed by the color space transformation subnetwork to obtain the reconstructed image block x hat corresponding to the current image block x. For example, the second up-sampling convolution subnetwork can include a Cropping layer, a ResAU layer, a Conv layer, and a Cropping layer, and the target feature can be up-sampled and convolved by these network layers, and the process is not limited. The color space transformation subnetwork is used for YUV domain to RBG image transformation process, or the color space transformation subnetwork is used for YUV domain to YUV domain filtering process, and the process is not limited.
[0194] In embodiment 10, the first feature y hat can be synthetically transformed based on the initial enhancement subnetwork, the first up-sampling convolution subnetwork, the attention subnetwork, and the reconstruction decoding subnetwork to obtain the reconstructed image block x hat corresponding to the current image block x. In a possible implementation, the residual enhancement output feature of the attention subnetwork can be adjusted by a feature adjustment factor, and the adjustment position of the feature adjustment factor can be seen from Figure 10B In the adjustment manner of the feature adjustment factor, the first feature y hat can be synthetically transformed to obtain the reconstructed image block x hat in the following manner.
[0195] The first feature y hat is feature enhanced by the initial enhancement subnetwork to obtain an enhanced feature. For example, the first feature y hat can be input into the initial enhancement subnetwork, and the first feature y hat is feature enhanced by the initial enhancement subnetwork to obtain an enhanced feature. For example, the initial enhancement subnetwork can include an RB layer and an RB layer, and thus the first feature y hat can be feature enhanced by the two RB layers to obtain an enhanced feature, and the process is not limited.
[0196] After obtaining the enhanced feature, the first up-sampling convolution subnetwork is adopted to perform up-sampling convolution on the enhanced feature to obtain the initial feature corresponding to the attention subnetwork, i.e., the input feature x_in of the attention subnetwork. For example, the enhanced feature (i.e., the output feature of the initial enhancement subnetwork) can be input to the first up-sampling convolution subnetwork, and the first up-sampling convolution subnetwork is adopted to perform up-sampling convolution on the enhanced feature to obtain the initial feature corresponding to the attention subnetwork. For example, the first up-sampling convolution subnetwork can include a Conv layer, a Cropping layer, a ResAU layer, a Conv layer, a Cropping layer, a ResAU layer, and a Conv layer, and the enhanced feature can be up-sampled and convolved through these network layers, and the process is not limited thereto.
[0197] After obtaining the initial feature x_in corresponding to the attention subnetwork, the initial feature x_in is input to the residual skip subnetwork, the residual enhancement subnetwork, and the weight generation subnetwork, respectively. After obtaining the initial feature x_in, the residual skip subnetwork superimposes the initial feature x_in (i.e., x_out0) to the final output feature of the attention subnetwork. After obtaining the initial feature x_in, the residual enhancement subnetwork performs feature enhancement on the initial feature x_in to obtain an enhanced feature, for example, the residual enhancement subnetwork performs feature enhancement on the initial feature x_in based on the RB block to obtain the enhanced feature x_out1. After obtaining the initial feature x_in, the weight generation subnetwork generates the weight feature corresponding to the initial feature x_in, i.e., the weight feature k of the output feature x_out1 of the residual enhancement subnetwork.
[0198] For example, the weight generation subnetwork includes a residual block subnetwork, a convolution subnetwork, and a feature mapping subnetwork, the residual block subnetwork is adopted to perform convolution and activation on the initial feature x_in to obtain a convolution and activation feature, the convolution subnetwork is adopted to perform convolution on the convolution and activation feature to obtain a convolution feature, and the feature mapping subnetwork is adopted to perform feature mapping on the convolution feature to obtain the weight feature k. As shown in FIG. 6, the weight generation subnetwork includes at least one RB block (i.e., the residual block subnetwork), one Conv layer (i.e., the convolution subnetwork), and a feature mapping layer (i.e., the feature mapping subnetwork), and on this basis, the initial feature x_in is convolved and activated based on the RB block to obtain a convolution and activation feature, the Conv layer is adopted to perform convolution on the convolution and activation feature to obtain a convolution feature, and the feature mapping layer is adopted to perform feature mapping on the convolution feature to obtain the weight feature k. Figure 10B As shown in FIG. 6, the weight generation subnetwork includes at least one RB block (i.e., the residual block subnetwork), one Conv layer (i.e., the convolution subnetwork), and a feature mapping layer (i.e., the feature mapping subnetwork), and on this basis, the initial feature x_in is convolved and activated based on the RB block to obtain a convolution and activation feature, the Conv layer is adopted to perform convolution on the convolution and activation feature to obtain a convolution feature, and the feature mapping layer is adopted to perform feature mapping on the convolution feature to obtain the weight feature k.
[0199] As shown in FIG. 6, the weight generation subnetwork includes at least one RB block (i.e., the residual block subnetwork), one Conv layer (i.e., the convolution subnetwork), and a feature mapping layer (i.e., the feature mapping subnetwork), and on this basis, the initial feature x_in is convolved and activated based on the RB block to obtain a convolution and activation feature, the Conv layer is adopted to perform convolution on the convolution and activation feature to obtain a convolution feature, and the feature mapping layer is adopted to perform feature mapping on the convolution feature to obtain the weight feature k. Figure 10BAs shown, after the enhanced feature x_out1 and the weight feature k are obtained, the second feature can be generated based on the enhanced feature x_out1 and the weight feature k. For example, the enhanced feature x_out1 can be multiplied by the weight feature k, and the multiplied feature can be taken as the second feature, that is, the second feature can be k*x_out1.
[0200] After the second feature is obtained, the second feature can be feature-adjusted based on the feature adjustment factor, and the feature-adjusted feature can be taken as the third feature, as shown in Figure 10B As shown, the third feature can be the feature rd.
[0201] After the third feature rd is obtained, the fourth feature can be generated based on the initial feature x_in (that is, the output feature x_out0 of the residual skip subnetwork) and the third feature rd, and the fourth feature can be the final output feature of the attention subnetwork. For example, the third feature rd can be added to the output feature x_out0 of the residual skip subnetwork, and the fourth feature x_out can be obtained as follows: x_out = x_in + rd.
[0202] As can be seen from the above, the initial feature x_in can be feature-enhanced by the first subnetwork (such as the residual enhancement subnetwork and the weight generation subnetwork, which can include the residual block subnetwork, the convolution subnetwork, and the feature mapping subnetwork) in the attention subnetwork to obtain the second feature. After the second feature is obtained, the second feature can be feature-adjusted based on the feature adjustment factor to obtain the third feature rd. After the third feature rd is obtained, the third feature rd can be processed by the second subnetwork (such as the residual skip subnetwork) in the attention subnetwork to obtain the fourth feature x_out. After the fourth feature x_out is obtained, the fourth feature x_out can be determined as the target feature.
[0203] After obtaining the target feature, the target feature can be processed by a reconstruction decoding sub-network to obtain a reconstructed image block x_hat corresponding to the current image block x. For example, the target feature can be input into the reconstruction decoding sub-network, and the target feature is processed by the reconstruction decoding sub-network to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the reconstruction decoding sub-network can include a second up-sampling convolution sub-network, and the target feature can be up-sampled and convolved by the second up-sampling convolution sub-network to obtain the reconstructed image block x_hat corresponding to the current image block x. Alternatively, the reconstruction decoding sub-network can include a second up-sampling convolution sub-network and a color space transformation sub-network, and the target feature can be up-sampled and convolved by the second up-sampling convolution sub-network, and the color space of the up-sampled and convolved feature can be transformed by the color space transformation sub-network to obtain the reconstructed image block x_hat corresponding to the current image block x. For example, the second up-sampling convolution sub-network can include a Cropping layer, a ResAU layer, a Conv layer, and a Cropping layer, and the target feature can be up-sampled and convolved by these network layers, and the process is not limited. The color space transformation sub-network is used for YUV domain to RBG image transformation process, or the color space transformation sub-network is used for YUV domain to YUV domain filtering process, and the process is not limited.
[0204] In embodiment 10, the first feature y_hat can be synthesized and transformed based on the initial enhancement sub-network, the first up-sampling convolution sub-network, the attention sub-network, and the reconstruction decoding sub-network to obtain the reconstructed image block x_hat corresponding to the current image block x. In a possible implementation, the weight output feature of the attention sub-network can be adjusted by a feature adjustment factor, and the adjustment position of the feature adjustment factor can be seen from Figure 10C The first feature y_hat can be synthesized and transformed to obtain the reconstructed image block x_hat in the following manner.
[0205] The first feature y_hat is feature-enhanced by the initial enhancement sub-network to obtain an enhanced feature. For example, the first feature y_hat can be input into the initial enhancement sub-network, and the first feature y_hat is feature-enhanced by the initial enhancement sub-network to obtain an enhanced feature. For example, the initial enhancement sub-network can include an RB layer and an RB layer, and thus the first feature y_hat can be feature-enhanced by the two RB layers to obtain an enhanced feature, and the process is not limited.
[0206] After obtaining the enhanced feature, the first up-sampling convolution sub-network is adopted to perform up-sampling convolution on the enhanced feature to obtain the initial feature corresponding to the attention sub-network, i.e., the input feature x_in of the attention sub-network. For example, the enhanced feature (i.e., the output feature of the initial enhancement sub-network) can be input to the first up-sampling convolution sub-network, and the first up-sampling convolution sub-network is adopted to perform up-sampling convolution on the enhanced feature to obtain the initial feature corresponding to the attention sub-network. For example, the first up-sampling convolution sub-network can include a Conv layer, a Cropping layer, a ResAU layer, a Conv layer, a Cropping layer, a ResAU layer, and a Conv layer, and the enhanced feature can be up-sampled and convolved through these network layers, and the process is not limited thereto.
[0207] After obtaining the initial feature x_in corresponding to the attention sub-network, the initial feature x_in is input to the residual skip sub-network, the residual enhancement sub-network, and the weight generation sub-network, respectively. For example, after obtaining the initial feature x_in, the residual skip sub-network superimposes the initial feature x_in (i.e., x_out0) to the final output feature of the attention sub-network. After obtaining the initial feature x_in, the residual enhancement sub-network performs feature enhancement on the initial feature x_in to obtain an enhanced feature. For example, the residual enhancement sub-network performs feature enhancement on the initial feature x_in based on the RB block to obtain the enhanced feature x_out1.
[0208] The weight generation sub-network can include a residual block sub-network, a convolution sub-network, and a feature mapping sub-network, as shown in FIG. 2. Figure 10C As shown in FIG. 2, the weight generation sub-network can include at least one RB block (i.e., the residual block sub-network), one Conv layer (i.e., the convolution sub-network), and a feature mapping layer (i.e., the feature mapping sub-network). On this basis, the initial feature x_in can be input to the residual block sub-network, and the residual block sub-network is adopted to perform convolution activation on the initial feature x_in to obtain a convolution-activated feature. For example, the residual block sub-network performs convolution activation on the initial feature x_in based on the RB block to obtain the convolution-activated feature.
[0209] Then, the convolution sub-network can be adopted to perform convolution on the convolution-activated feature to obtain a convolution feature. For example, the convolution sub-network can perform convolution on the convolution-activated feature using the Conv layer to obtain the convolution feature. After obtaining the convolution feature, the second feature can be determined based on the convolution feature. For example, the convolution feature can be taken as the second feature.
[0210] In summary, a residual block sub-network can be used to enhance the initial feature x_in, resulting in enhanced features. A convolutional sub-network is then used to convolve these enhanced features, yielding the convolutional features, i.e., the second feature. After obtaining the second feature, it can be adjusted based on a feature adjustment factor, and the adjusted feature can be used as the third feature.
[0211] See Figure 10C As shown, after obtaining the third feature, a feature mapping sub-network can be used to perform feature mapping on the third feature to obtain the weight feature k. After obtaining the enhanced feature x_out1, the weight feature k, and the initial feature x_in (i.e., x_out0), a fourth feature can be generated based on the initial feature x_in, the enhanced feature x_out1, and the weight feature k. The fourth feature can be the final output feature of the attention sub-network. For example, multiplying the enhanced feature x_out1 by the weight feature k, and then adding it to the output feature x_out0 of the residual skip sub-network, the fourth feature is: x_out = x_in + k * x_out1.
[0212] In summary, we can see that the initial feature x_in can be enhanced using the first sub-network in the attention sub-network (such as the residual block sub-network and convolutional sub-network in the weight generation sub-network) to obtain the second feature. After obtaining the second feature, it can be adjusted based on the feature adjustment factor to obtain the third feature. After obtaining the third feature, it can be processed using the second sub-network in the attention sub-network (such as the residual skip sub-network, residual enhancement sub-network, and feature mapping sub-network in the weight generation sub-network) to obtain the fourth feature x_out. For example, the feature mapping sub-network can be used to perform feature mapping on the third feature to obtain the weight feature k; the residual enhancement sub-network can be used to enhance the initial feature x_in to obtain the enhanced feature x_out1; and the fourth feature x_out can be generated based on the initial feature x_in, the enhanced feature x_out1, and the weight feature k. After obtaining the fourth feature x_out, it can be determined as the target feature.
[0213] After obtaining the target feature, the target feature can be processed by the reconstruction decoding sub-network to obtain the reconstructed image block x hat corresponding to the current image block x. For example, the target feature can be input to the reconstruction decoding sub-network, and the target feature is processed by the reconstruction decoding sub-network to obtain the reconstructed image block x hat corresponding to the current image block x. For example, the reconstruction decoding sub-network can include a second up-sampling convolution sub-network, and the target feature can be up-sampled and convolved by the second up-sampling convolution sub-network to obtain the reconstructed image block x hat corresponding to the current image block x. Alternatively, the reconstruction decoding sub-network can include a second up-sampling convolution sub-network and a color space transformation sub-network, and the target feature can be up-sampled and convolved by the second up-sampling convolution sub-network, and the color space transformation sub-network can be used to perform color space transformation on the up-sampled and convolved feature to obtain the reconstructed image block x hat corresponding to the current image block x. For example, the second up-sampling convolution sub-network can include a Cropping layer, a ResAU layer, a Conv layer, and a Cropping layer, and the target feature can be up-sampled and convolved by these network layers, and this process is not limited. The color space transformation sub-network is used to perform YUV domain to RBG image transformation process, or the color space transformation sub-network is used to perform YUV domain to YUV domain filtering process, and this process is not limited.
[0214] Embodiment 14: For embodiments 1-13, for the encoding end and the decoding end, the second feature can be feature adjusted based on the feature adjustment factor to obtain the third feature. For example, if the second feature includes C*H*W feature values, C represents the number of channels, H represents the feature height, and W represents the feature width, the feature adjustment value corresponding to each feature value can be determined based on the feature adjustment factor. For each feature value, the feature adjustment value corresponding to the feature value is used to adjust the feature value to obtain the adjusted feature value corresponding to the feature value. The third feature is generated based on the adjusted feature value corresponding to each feature value.
[0215] For example, if the second feature includes C*H*W feature values, the feature adjustment value corresponding to each feature value of the second feature can be determined based on the feature adjustment factor, and the following cases can be used for this process:
[0216] Case 1: The dimension of the feature adjustment factor is the same as the dimension of the second feature, that is, each feature value of the second feature has a different feature adjustment value. For example, if the feature adjustment factor can include C*H*W feature adjustment values, the feature adjustment value corresponding to each of the C*H*W feature values can be determined based on the C*H*W feature adjustment values.
[0217] For example, if the feature value at the position of h*w in the cth channel in the second feature, the feature adjustment value corresponding to the feature value can be the feature adjustment value at the position of h*w in the cth channel in the feature adjustment factor.
[0218] Case 2: The W dimension and the H dimension of the feature adjustment factor are the same as the dimensions of the second feature, but the C dimension of the feature adjustment factor is 1, that is, the feature values at the same position of each channel of the second feature correspond to the same feature adjustment value. For example, if the feature adjustment factor can include H*W feature adjustment values, then the feature adjustment values corresponding to the C*H*W feature values can be determined based on the H*W feature adjustment values (i.e., the C dimension of the feature adjustment factor is 1).
[0219] For example, if the feature value at the position of h*w in the cth channel (i.e., each channel in the second feature), the feature adjustment value corresponding to the feature value can be the feature adjustment value at the position of h*w in the feature adjustment factor.
[0220] Case 3: The C dimension of the feature adjustment factor is the same as the C dimension of the second feature, but the W dimension and the H dimension of the feature adjustment factor are both 1, that is, the W*H feature values of the same channel of the second feature correspond to the same feature adjustment value. For example, if the feature adjustment factor can include C feature adjustment values, then the feature adjustment values corresponding to the C*H*W feature values can be determined based on the C feature adjustment values (i.e., the W dimension and the H dimension of the feature adjustment factor are both 1).
[0221] For example, if the feature value at the position of h*w in the cth channel (i.e., each channel in the second feature), the feature adjustment value corresponding to the feature value can be the feature adjustment value at the position of h*w in the feature adjustment factor.
[0222] For example, if the feature value at the position of h*w in the cth channel (i.e., each channel in the second feature), the feature adjustment value corresponding to the feature value can be the feature adjustment value at the position of h*w in the feature adjustment factor.
[0223] For example, if the feature value at the position of h*w in the cth channel (i.e., each channel in the second feature), the feature adjustment value corresponding to the feature value can be the feature adjustment value at the position of h*w in the feature adjustment factor.
[0224] If the feature value corresponds to N+1 feature adjustment values, N is a positive integer, i.e., corresponds to at least two feature adjustment values, the following formula can be used to determine the adjusted feature value corresponding to the feature value: yr=rf_0*y 0 +rf_1*y 1 +rf_2*y 2 +…+rf_N*y N ; wherein, yr represents the adjusted feature value, rf_0, rf_1, rf_2, …, rf_N represent the N+1 feature adjustment values, and y represents the feature value. For example, for a feature value y(c, w, h) of a c-th channel spatial position (w, h), after the feature adjustment value rf_i(c, w, h) corresponding to the feature value y(c, w, h) is determined, i=0, 1, 2…N, the adjusted feature value yr(c, w, h) corresponding to the feature value y(c, w, h) is: yr(c, w, h)=rf_0(c, w, h)*y 0 (c, w, h)+rf_1(c, w, h)*y 1 (c, w, h)+rf_2(c, w, h)*y 2 (c, w, h)+…+rf_N(c, w, h)*y N (c, w, h). When N is 1, yr(c, w, h)=rf_0(c, w, h)+rf_1(c, w, h)*y(c, w, h). When N is 2, yr(c, w, h)=rf_0(c, w, h)+rf_1(c, w, h)*y(c, w, h)+rf_2(c, w, h)*y 2 (c, w, h). Similarly, when N is other values, the form is similar, which is not described here. In the above formula, y k (c, w, h) represents the k-th power of y.
[0225] For example, after the adjusted feature value corresponding to each feature value is obtained, the third feature can be generated based on the adjusted feature value corresponding to each feature value, and the third feature can include the adjusted feature value corresponding to each feature value.
[0226] For the encoding end, it is also needed to determine whether the current image block enables the feature adjustment mode, if the current image block enables the feature adjustment mode, the feature adjustment factor corresponding to the current image block is obtained and encoded in the auxiliary code stream corresponding to the current image block. If the current image block does not enable the feature adjustment mode, the feature adjustment factor corresponding to the current image block does not need to be obtained and encoded in the auxiliary code stream corresponding to the current image block. For the decoding end, it is also needed to determine whether the current image block enables the feature adjustment mode, if the current image block enables the feature adjustment mode, the feature adjustment factor corresponding to the current image block is decoded from the auxiliary code stream corresponding to the current image block, and the target feature is determined based on the first feature and the feature adjustment factor. If the current image block does not enable the feature adjustment mode, the feature adjustment factor corresponding to the current image block does not need to be decoded from the auxiliary code stream corresponding to the current image block.
[0227] For example, in order to determine whether the current image block enables the feature adjustment mode, the following method can be used:
[0228] Method 1: the encoding end encodes a feature adjustment flag bit in the code stream, and the decoding end decodes the feature adjustment flag bit from the code stream. The feature adjustment flag bit allows the current image block to enable the feature adjustment mode, or the feature adjustment flag bit prohibits the current image block to enable the feature adjustment mode. If the feature adjustment flag bit allows the current image block to enable the feature adjustment mode, the decoding end determines that the current image block enables the feature adjustment mode. Otherwise, if the feature adjustment flag bit prohibits the current image block to enable the feature adjustment mode, the decoding end determines that the current image block does not enable the feature adjustment mode. For example, if the feature adjustment flag bit is a first value, the feature adjustment flag bit allows the current image block to enable the feature adjustment mode, and if the feature adjustment flag bit is a second value, the feature adjustment flag bit prohibits the current image block to enable the feature adjustment mode.
[0229] For example, the feature adjustment flag bit can be a sequence-level feature adjustment flag bit, i.e., the feature adjustment flag bit corresponds to all image blocks in a sequence; or the feature adjustment flag bit can be an image-level feature adjustment flag bit, i.e., the feature adjustment flag bit corresponds to all image blocks in an image; or the feature adjustment flag bit can be a slice-level feature adjustment flag bit, i.e., the feature adjustment flag bit corresponds to all image blocks in a slice, without limitation.
[0230] Manner 2, the first value range of the feature value is encoded by the encoding end in the code stream, and the first value range of the feature value is decoded by the decoding end from the code stream. Based on this, if the feature value of the second feature (such as all feature values) is located in the first value range, the decoding end determines that the current image block enables the feature adjustment mode. Otherwise, if the feature value of the second feature (such as any feature value) is not located in the first value range, the decoding end determines that the current image block disables the feature adjustment mode. For example, after obtaining the second feature, the decoding end can determine whether the feature value of the second feature is located in the first value range.
[0231] For example, the first value range can be a sequence-level first value range, that is, the first value range corresponds to all image blocks in the sequence; or the first value range can be an image-level first value range, that is, the first value range corresponds to all image blocks in the image; or the first value range can be a slice-level first value range, that is, the first value range corresponds to all image blocks in the slice, and the first value range is not limited.
[0232] Manner 3, the second value range of the feature change value is encoded by the encoding end in the code stream, and the second value range of the feature change value is decoded by the decoding end from the code stream. Based on this, if the feature change value corresponding to the second feature (such as the feature change value corresponding to all feature values, that is, each feature value corresponds to a feature change value) is located in the second value range, the decoding end determines that the current image block enables the feature adjustment mode. Otherwise, if the feature change value corresponding to the second feature (such as any feature change value) is not located in the second value range, the decoding end determines that the current image block disables the feature adjustment mode. For example, after obtaining the second feature, the decoding end can also determine the feature change value corresponding to the second feature, that is, determine the feature change value corresponding to each feature value in the second feature. The feature change value can represent the change degree range in the spatial domain or the channel domain, that is, determine the change of the second feature in the spatial domain or the channel domain. For example, the feature change value can include but is not limited to gradient value and the like.
[0233] For example, the second value range can be a sequence-level second value range, that is, the second value range corresponds to all image blocks in the sequence; or the second value range can be an image-level second value range, that is, the second value range corresponds to all image blocks in the image; or the second value range can be a slice-level second value range, that is, the second value range corresponds to all image blocks in the slice, and the second value range is not limited.
[0234] Mode 4, the encoding end encodes the feature adjustment flag and the first value range in the bitstream, and the decoding end decodes the feature adjustment flag and the first value range from the bitstream. If the feature adjustment flag allows the current image block to enable the feature adjustment mode, and the feature value of the second feature (such as all feature values) is in the first value range, it is determined that the current image block enables the feature adjustment mode. Otherwise, if the feature adjustment flag prohibits the current image block to enable the feature adjustment mode, and / or the feature value of the second feature (such as any feature value) is not in the first value range, it is determined that the current image block does not enable the feature adjustment mode.
[0235] Mode 5, the encoding end encodes the feature adjustment flag and the second value range in the bitstream, and the decoding end decodes the feature adjustment flag and the second value range from the bitstream. If the feature adjustment flag allows the current image block to enable the feature adjustment mode, and the feature variation value corresponding to the second feature (such as the feature variation value corresponding to all feature values, i.e., each feature value corresponds to a feature variation value) is in the second value range, the decoding end determines that the current image block enables the feature adjustment mode. Otherwise, if the feature adjustment flag prohibits the current image block to enable the feature adjustment mode, and / or the feature variation value corresponding to the second feature (such as any feature variation value) is not in the second value range, it is determined that the current image block does not enable the feature adjustment mode.
[0236] Mode 6, the encoding end encodes the first value range and the second value range in the bitstream, and the decoding end decodes the first value range and the second value range from the bitstream. If the feature value of the second feature is in the first value range, and the feature variation value corresponding to the second feature is in the second value range, the decoding end can determine that the current image block enables the feature adjustment mode. Otherwise, if the feature value of the second feature is not in the first value range, and / or the feature variation value corresponding to the second feature is not in the second value range, the decoding end can determine that the current image block does not enable the feature adjustment mode.
[0237] Mode 7, the encoding end encodes the feature adjustment flag, the first value range and the second value range in the bitstream, and the decoding end decodes the feature adjustment flag, the first value range and the second value range from the bitstream. If the feature adjustment flag allows the current image block to enable the feature adjustment mode, and the feature value of the second feature is in the first value range, and the feature variation value corresponding to the second feature is in the second value range, it is determined that the current image block enables the feature adjustment mode. Otherwise, if the feature adjustment flag prohibits the current image block to enable the feature adjustment mode, the feature value of the second feature is not in the first value range, and / or the feature variation value corresponding to the second feature is not in the second value range, it is determined that the current image block does not enable the feature adjustment mode.
[0238] In the manner 8, the decoding end can determine that the current image block enables the feature adjustment mode if the attribute value corresponding to the second feature is located in the configured attribute value range, otherwise, the decoding end can determine that the current image block disables the feature adjustment mode if the attribute value corresponding to the second feature is not located in the configured attribute value range. In the manner 8, the decoding end does not need to parse the information related to the feature adjustment mode from the code stream, but determines whether the current image block enables the feature adjustment mode based on the configured attribute value range. In the manner 8, the attribute value corresponding to the second feature can include but is not limited to the variance value corresponding to the feature value in the second feature, and the variance value is used to represent the magnitude of the degree of change of the position of the feature value in the second feature.
[0239] In the embodiment 16, for each feature value in the second feature, the feature adjustment can be performed on the feature value based on the feature adjustment value corresponding to the feature value to obtain the adjusted feature value corresponding to the feature value.
[0240] For example, the conditions can include but are not limited to: 1, determining whether to adjust the feature value based on the additional code stream parsing related information, which can be the flag information indicating whether the region (which can be the feature or the channel of the feature, or the slice of the channel) where the current feature value is located needs to be adjusted, or the range information of the size of the feature value that needs to be adjusted (determining whether the current feature value is in the range to determine whether adjustment is needed), or the range of the degree of change of the spatial domain or channel domain of the feature value that needs to be adjusted (calculating the change (such as the gradient value) of the spatial domain or channel domain of the current feature value to determine whether the change is in the range to determine whether adjustment is needed); 2, determining based on the attribute of the feature value, such as the variance (representing the degree of change of the position of the feature value) corresponding to the current feature value.
[0241] For example, for each feature value in the second feature, to determine whether to perform feature adjustment on the feature value based on the feature adjustment value corresponding to the feature value to obtain the adjusted feature value corresponding to the feature value, the following manners can be used:
[0242] Manner 1: The encoding end encodes a feature adjustment flag corresponding to the feature value in the bitstream, and the decoding end decodes the feature adjustment flag corresponding to the feature value from the bitstream. The feature adjustment flag indicates whether the feature value is adjusted or not. If the feature adjustment flag indicates that the feature value is adjusted, the decoding end adjusts the feature value based on the feature adjustment value corresponding to the feature value to obtain the adjusted feature value corresponding to the feature value. Otherwise, if the feature adjustment flag indicates that the feature value is not adjusted, the decoding end does not adjust the feature value based on the feature adjustment value corresponding to the feature value. For example, if the feature adjustment flag is a first value, the feature adjustment flag indicates that the feature value is adjusted, and if the feature adjustment flag is a second value, the feature adjustment flag indicates that the feature value is not adjusted.
[0243] Manner 2: The encoding end encodes a first value range of the feature value in the bitstream, and the decoding end decodes the first value range of the feature value from the bitstream. Based on this, if the feature value is in the first value range, the decoding end adjusts the feature value based on the feature adjustment value corresponding to the feature value to obtain the adjusted feature value corresponding to the feature value. Otherwise, if the feature value is not in the first value range, the decoding end does not adjust the feature value based on the feature adjustment value corresponding to the feature value.
[0244] Manner 3: The encoding end encodes a second value range of the feature change value in the bitstream, and the decoding end decodes the second value range of the feature change value from the bitstream. Based on this, if the feature change value (such as the gradient value) corresponding to the feature value is in the second value range, the decoding end adjusts the feature value based on the feature adjustment value corresponding to the feature value to obtain the adjusted feature value corresponding to the feature value. Otherwise, if the feature change value (such as the gradient value) corresponding to the feature value is not in the second value range, the decoding end does not adjust the feature value based on the feature adjustment value corresponding to the feature value.
[0245] Manner 4: The encoding end encodes a feature adjustment flag and a first value range in the bitstream, and the decoding end decodes the feature adjustment flag and the first value range from the bitstream. If the feature adjustment flag indicates that the feature value is adjusted and the feature value is in the first value range, the decoding end adjusts the feature value based on the feature adjustment value corresponding to the feature value. Otherwise, the decoding end does not adjust the feature value based on the feature adjustment value corresponding to the feature value.
[0246] Manner 5, the encoding end encodes the feature adjustment flag and the second value range in the bitstream, and the decoding end decodes the feature adjustment flag and the second value range from the bitstream. If the feature adjustment flag indicates that the feature value is to be adjusted, and the feature variation value corresponding to the feature value is in the second value range, the decoding end adjusts the feature value based on the feature adjustment value corresponding to the feature value. Otherwise, the decoding end does not adjust the feature value based on the feature adjustment value corresponding to the feature value.
[0247] Manner 6, the encoding end encodes the first value range and the second value range in the bitstream, and the decoding end decodes the first value range and the second value range from the bitstream. If the feature value is in the first value range, and the feature variation value corresponding to the feature value is in the second value range, the decoding end adjusts the feature value based on the feature adjustment value corresponding to the feature value. Otherwise, the decoding end does not adjust the feature value based on the feature adjustment value corresponding to the feature value.
[0248] Manner 7, the encoding end encodes the feature adjustment flag, the first value range and the second value range in the bitstream, and the decoding end decodes the feature adjustment flag, the first value range and the second value range from the bitstream. If the feature adjustment flag indicates that the feature value is to be adjusted, and the feature value is in the first value range, and the feature variation value corresponding to the feature value is in the second value range, the decoding end adjusts the feature value based on the feature adjustment value corresponding to the feature value. Otherwise, the decoding end does not adjust the feature value based on the feature adjustment value corresponding to the feature value.
[0249] Manner 8, if the attribute value corresponding to the feature value is in the configured attribute value range, the decoding end adjusts the feature value based on the feature adjustment value corresponding to the feature value. Otherwise, the decoding end does not adjust the feature value based on the feature adjustment value corresponding to the feature value. In manner 8, the decoding end does not need to parse information related to the feature adjustment mode from the bitstream, but determines whether to adjust the feature value based on the feature adjustment value corresponding to the feature value based on the configured attribute value range. In manner 8, the attribute value corresponding to the feature value can include but is not limited to the variance value corresponding to the feature value, and the variance value is used to represent the magnitude of the degree of change in the position of the feature value.
[0250] Embodiment 17: Refer to Embodiment 7, the encoding end can obtain the feature adjustment factor for the decoding sub-network, and perform feature adjustment on a certain feature in the decoding sub-network based on the feature adjustment factor, and encode the feature adjustment factor for the decoding sub-network in the auxiliary code stream. The decoding end can decode the feature adjustment factor for the decoding sub-network from the auxiliary code stream, and perform feature adjustment on a certain feature in the decoding sub-network based on the feature adjustment factor. For example, after obtaining feature A (feature A is any feature in the decoding sub-network), the decoding sub-network can perform feature adjustment on feature A using the feature adjustment factor to obtain adjusted feature B. The decoding sub-network continues to process based on feature B, and finally obtains the first feature.
[0251] In a possible implementation, refer to Embodiment 8 Figure 8A As shown in the figure, the feature adjustment factor for the decoding sub-network can be used in the feature recovery process in the decoding sub-network, that is, in the feature recovery process, a certain feature is adjusted using the feature adjustment factor. The feature adjustment process can refer to the feature adjustment process of the second feature, which will not be repeated here.
[0252] In a possible implementation, refer to Embodiment 9 Figure 8B As shown in the figure, the feature adjustment factor for the decoding sub-network can be used in the residual recovery process in the decoding sub-network, that is, in the residual recovery process, a certain feature is adjusted using the feature adjustment factor. The feature adjustment process can refer to the feature adjustment process of the second feature, which will not be repeated here.
[0253] For example, each of Embodiments 1-17 can be implemented alone, and at least two of Embodiments 1-17 can be combined for implementation.
[0254] For example, the content of the encoding end in each of the above embodiments can also be applied to the decoding end, that is, the decoding end can be processed in the same way, and the content of the decoding end can also be applied to the encoding end, that is, the encoding end can be processed in the same way.
[0255] Based on the same application idea as the above method, the present embodiment also proposes a decoding device, which is applied to a decoding end, and includes a memory configured to store video data, and a decoder configured to implement the decoding method in Embodiments 1-17 above, that is, the processing flow of the decoding end.
[0256] For example, in a possible implementation, the decoder is configured to implement:
[0257] acquire a first feature corresponding to the current image block by a first neural network, the first neural network comprising at least one convolutional layer; acquire a feature adjustment factor corresponding to the current image block based on the first feature;
[0258] determine a target feature based on the first feature and the feature adjustment factor;
[0259] acquire a reconstructed image block corresponding to the current image block based on the target feature by a second neural network, the second neural network comprising at least one convolutional layer.
[0260] Based on the same application concept as the above method, the embodiment of the present application also proposes an encoding device, which is applied to an encoding end, and the device comprises: a memory configured to store video data; and an encoder configured to implement the encoding method in the above embodiments 1-17, i.e., the processing flow of the encoding end.
[0261] For example, in a possible implementation, the encoder is configured to implement:
[0262] acquire a first feature corresponding to the current image block by a first neural network, the first neural network comprising at least one convolutional layer; acquire a feature adjustment factor corresponding to the current image block based on the first feature;
[0263] encode the feature adjustment factor in a secondary code stream corresponding to the current image block.
[0264] Based on the same application concept as the above method, the decoding end device (which can also be referred to as a video decoder) provided in the embodiment of the present application can be specifically referred to from the hardware level in terms of its hardware architecture diagram as shown in Figure 11A The decoding end device comprises: a processor 1101 and a machine readable storage medium 1102, the machine readable storage medium 1102 stores machine executable instructions that can be executed by the processor 1101; and the processor 1101 is used to execute the machine executable instructions to implement the decoding method in the above embodiments 1-17 of the present application. For example, in a possible implementation, the decoding end device is used to implement:
[0265] acquire a first feature corresponding to the current image block by a first neural network, the first neural network comprising at least one convolutional layer; acquire a feature adjustment factor corresponding to the current image block based on the first feature;
[0266] determine a target feature based on the first feature and the feature adjustment factor;
[0267] acquire a reconstructed image block corresponding to the current image block based on the target feature by a second neural network, the second neural network comprising at least one convolutional layer.
[0268] Based on the same application concept as the above method, the encoding end device (which can also be referred to as a video encoder) provided in the embodiments of the present application can be specifically referred to the hardware architecture diagram shown in the figure Figure 11B The encoding end device can include a processor 1111 and a machine readable storage medium 1112, and the machine readable storage medium 1112 stores machine executable instructions that can be executed by the processor 1111; the processor 1111 is configured to execute the machine executable instructions to implement the encoding method of the above embodiments 1-17 of the present application. For example, in a possible implementation, the encoding end device is configured to implement the following steps:
[0269] obtaining a first feature corresponding to the current image block through a first neural network, the first neural network including at least one convolutional layer; obtaining a feature adjustment factor corresponding to the current image block based on the first feature;
[0270] encoding the feature adjustment factor in the auxiliary code stream corresponding to the current image block.
[0271] Based on the same application concept as the above method, the encoding end device (which can also be referred to as a video encoder) provided in the embodiments of the present application can be specifically referred to the hardware architecture diagram shown in the figure
[0272] Based on the same application concept as the above method, the encoding end device (which can also be referred to as a video encoder) provided in the embodiments of the present application can be specifically referred to the hardware architecture diagram shown in the figure
[0273] Based on the same application concept as the above method, the encoding end device (which can also be referred to as a video encoder) provided in the embodiments of the present application can be specifically referred to the hardware architecture diagram shown in the figure
[0274] Based on the same application concept as the above method, the encoding end device (which can also be referred to as a video encoder) provided in the embodiments of the present application can be specifically referred to the hardware architecture diagram shown in the figure
[0275] Illustratively, the obtaining module obtains the first feature corresponding to the current image block through the first neural network, and specifically is configured to: obtain a probability distribution parameter based on the first code stream corresponding to the current image block; determine a probability distribution model based on the probability distribution parameter, and decode the second code stream corresponding to the current image block based on the probability distribution model to obtain a decoded image feature; and determine the first feature corresponding to the current image block based on the decoded image feature.
[0276] Illustratively, the obtaining module obtains the first feature corresponding to the current image block through the first neural network, and specifically is configured to: obtain a probability distribution parameter and a prediction value based on the first code stream corresponding to the current image block; determine a probability distribution model based on the probability distribution parameter, and decode the second code stream corresponding to the current image block based on the probability distribution model to obtain a decoded image feature; perform residual recovery on the decoded image feature to obtain a residual feature; and determine the first feature corresponding to the current image block based on the residual feature and the prediction value.
[0277] Illustratively, the obtaining module obtains the feature adjustment factor corresponding to the current image block, and specifically is configured to: decode the auxiliary code stream corresponding to the current image block to obtain the feature adjustment factor corresponding to the current image block; or determine a fixed parameter value as the feature adjustment factor corresponding to the current image block.
[0278] Illustratively, the determining module determines the target feature based on the first feature and the feature adjustment factor, and specifically is configured to: perform feature enhancement on the first feature to obtain a second feature; perform feature adjustment on the second feature based on the feature adjustment factor to obtain a third feature; and determine the target feature based on the third feature.
[0279] Illustratively, the determining module performs feature enhancement on the first feature to obtain a second feature, and specifically is configured to: determine an initial feature corresponding to an attention subnetwork based on the first feature; and perform feature enhancement on the initial feature through the attention subnetwork to obtain the second feature. Illustratively, the determining module determines the target feature based on the third feature, and specifically is configured to: determine the third feature as the target feature.
[0280] The attention subnetwork includes a residual enhancement subnetwork and a weight generation subnetwork. Illustratively, the determining module performs feature enhancement on the initial feature through the attention subnetwork to obtain the second feature, and specifically is configured to: perform feature enhancement on the initial feature through the residual enhancement subnetwork to obtain an enhanced feature; generate a weight feature corresponding to the initial feature through the weight generation subnetwork; and generate the second feature based on the initial feature, the enhanced feature, and the weight feature.
[0281] In an example, the weight generation subnetwork can include a residual block subnetwork, a convolution subnetwork, and a feature mapping subnetwork, and the determination module is specifically configured to: adopt the residual block subnetwork to perform convolution activation on the initial feature to obtain a convolution-activated feature; adopt the convolution subnetwork to perform convolution on the convolution-activated feature to obtain a convolutional feature; and adopt the feature mapping subnetwork to perform feature mapping on the convolutional feature to obtain the weight feature corresponding to the initial feature.
[0282] In an example, the determination module is specifically configured to: determine an initial feature corresponding to the attention subnetwork based on the first feature; perform feature enhancement on the initial feature through a first subnetwork in the attention subnetwork to obtain the second feature; and determine the target feature based on the third feature, and specifically configured to: perform processing on the third feature through a second subnetwork in the attention subnetwork to obtain a fourth feature, and determine the fourth feature as the target feature.
[0283] In an example, the determination module is specifically configured to: perform feature enhancement on the initial feature through a first subnetwork in the attention subnetwork to obtain the second feature, and specifically configured to: adopt a residual enhancement subnetwork to perform feature enhancement on the initial feature to obtain an enhanced feature; adopt a weight generation subnetwork to generate a weight feature corresponding to the initial feature; and generate the second feature based on the enhanced feature and the weight feature; and the determination module is specifically configured to perform processing on the third feature through a second subnetwork in the attention subnetwork to obtain a fourth feature, and specifically configured to: generate the fourth feature based on the initial feature and the third feature.
[0284] In an example, the determination module is specifically configured to: perform feature enhancement on the initial feature through a first subnetwork in the attention subnetwork to obtain the second feature, and specifically configured to: adopt a residual block subnetwork to perform feature enhancement on the initial feature to obtain an enhanced feature; adopt a convolution subnetwork to perform convolution on the enhanced feature to obtain a convolutional feature, and determine the second feature based on the convolutional feature; and the determination module is specifically configured to perform processing on the third feature through a second subnetwork in the attention subnetwork to obtain a fourth feature, and specifically configured to: adopt a feature mapping subnetwork to perform feature mapping on the third feature to obtain a weight feature; adopt a residual enhancement subnetwork to perform feature enhancement on the initial feature to obtain an enhanced feature; and generate the fourth feature based on the initial feature, the enhanced feature, and the weight feature.
[0285] Illustratively, the determining module is specifically configured to determine the initial feature corresponding to the attention sub-network based on the first feature by: performing feature enhancement on the first feature by using an initial enhancement sub-network to obtain an enhanced feature; and performing up-sampling convolution on the enhanced feature by using an up-sampling convolution sub-network to obtain the initial feature.
[0286] Illustratively, the determining module is specifically configured to perform feature adjustment on the second feature based on the feature adjustment factor to obtain a third feature by: if the second feature comprises C*H*W feature values, C represents a channel number, H represents a feature height, and W represents a feature width, determining a feature adjustment value corresponding to each feature value based on the feature adjustment factor, and performing feature adjustment on the feature value based on the feature adjustment value corresponding to the feature value to obtain an adjusted feature value corresponding to the feature value; and generating the third feature based on the adjusted feature value corresponding to each feature value.
[0287] Illustratively, the determining module is specifically configured to determine the feature adjustment value corresponding to each feature value based on the feature adjustment factor by: if the feature adjustment factor comprises C*H*W feature adjustment values, determining the feature adjustment value corresponding to each of the C*H*W feature values based on the C*H*W feature adjustment values; or if the feature adjustment factor comprises H*W feature adjustment values, determining the feature adjustment value corresponding to each of the C*H*W feature values based on the H*W feature adjustment values; or if the feature adjustment factor comprises C feature adjustment values, determining the feature adjustment value corresponding to each of the C*H*W feature values based on the C feature adjustment values.
[0288] Illustratively, the determining module is specifically configured to perform feature adjustment on the feature value based on the feature adjustment value corresponding to the feature value to obtain an adjusted feature value corresponding to the feature value by: if the feature value corresponds to only one feature adjustment value, determining the adjusted feature value by using the following formula: yr=rf*y; wherein yr represents the adjusted feature value, rf represents the feature adjustment value, and y represents the feature value; or if the feature value corresponds to N+1 feature adjustment values, N being a positive integer, determining the adjusted feature value by using the following formula: yr=rf_0*y 0 +rf_1*y 1 +rf_2*y 2 +…+rf_N*y N ; wherein yr represents the adjusted feature value, rf_0, rf_1, rf_2, …, and rf_N represent the N+1 feature adjustment values, and y represents the feature value.
[0289] Illustratively, the second neural network comprises a reconstruction decoding subnetwork, and the obtaining module is specifically configured to: obtain the reconstruction image block corresponding to the current image block by processing the target feature through the reconstruction decoding subnetwork, when obtaining the reconstruction image block corresponding to the current image block through the second neural network based on the target feature; wherein the reconstruction decoding subnetwork comprises an up-sampling convolution subnetwork, and the target feature is up-sampled and convolved through the up-sampling convolution subnetwork to obtain the reconstruction image block; or the reconstruction decoding subnetwork comprises an up-sampling convolution subnetwork and a color space transformation subnetwork, the target feature is up-sampled and convolved through the up-sampling convolution subnetwork, and the feature after the up-sampling and convolution is color space transformed through the color space transformation subnetwork to obtain the reconstruction image block.
[0290] Based on the same application concept as the above method, an encoding device is further provided in the embodiments of the present application, and the device is applied to an encoding end, and the device comprises: an obtaining module, configured to obtain a first feature corresponding to a current image block through a first neural network, wherein the first neural network comprises at least one convolution layer; and obtain a feature adjustment factor corresponding to the current image block based on the first feature; and an encoding module, configured to encode the feature adjustment factor in an auxiliary code stream corresponding to the current image block.
[0291] Illustratively, the obtaining module is specifically configured to: obtain at least one candidate feature adjustment factor, when obtaining the feature adjustment factor corresponding to the current image block based on the first feature; determine a rate-distortion cost corresponding to each candidate feature adjustment factor based on the first feature; and select one candidate feature adjustment factor from all candidate feature adjustment factors as the feature adjustment factor corresponding to the current image block based on the rate-distortion cost corresponding to each candidate feature adjustment factor.
[0292] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. The present application can be in the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. The embodiments of the present application can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The above is only an embodiment of the present application and is not intended to limit the present application.
[0293] The present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of the present application shall be included in the scope of the claims of the present application.
Claims
1. An image decoding method characterized by, The method comprises: obtaining a probability distribution parameter corresponding to the code stream of the current image block; determining a probability distribution model based on the probability distribution parameter, and decoding the code stream corresponding to the current image block based on the probability distribution model to obtain a decoded image feature; determining a first feature corresponding to the current image block based on the decoded image feature; obtaining a feature adjustment factor corresponding to the current image block; performing a synthetic transformation on the first feature through a synthetic transformation network to obtain a reconstructed image block corresponding to the current image block; wherein, in the process of synthetic transformation, the first feature is enhanced to obtain a second feature; the second feature is adjusted based on the feature adjustment factor to obtain a third feature; a target feature is determined based on the third feature; and the reconstructed image block corresponding to the current image block is determined based on the target feature.
2. The method of claim 1, wherein The method further comprises: obtaining a prediction value based on the code stream corresponding to the current image block; The determination of the first feature corresponding to the current image block based on the decoded image feature comprises: residual recovery is performed on the decoded image feature to obtain a residual feature; The first feature corresponding to the current image block is determined based on the residual feature and the prediction value.
3. The method of claim 1, wherein The feature adjustment factor corresponding to the current image block is obtained by: decoding the code stream corresponding to the current image block to obtain the feature adjustment factor corresponding to the current image block; or, determining a fixed parameter value as the feature adjustment factor corresponding to the current image block.
4. The method of claim 1, wherein The determination of the reconstructed image block corresponding to the current image block based on the target feature comprises: processing the target feature through a reconstruction decoding sub-network to obtain the reconstructed image block; The reconstruction decoding sub-network comprises an up-sampling convolution sub-network, and the target feature is up-sampled and convolved through the up-sampling convolution sub-network to obtain the reconstructed image block; or, the reconstruction decoding sub-network comprises an up-sampling convolution sub-network and a color space transformation sub-network, and the target feature is up-sampled and convolved through the up-sampling convolution sub-network, and the feature after up-sampling and convolution is color space transformed through the color space transformation sub-network to obtain the reconstructed image block.
5. The method of claim 1, wherein, The method further comprises: if the current image block enables the feature adjustment mode, obtaining the feature adjustment factor corresponding to the current image block; wherein, the determination process of the current image block enabling the feature adjustment mode comprises: parsing a feature adjustment flag from the code stream, and if the feature adjustment flag allows the current image block to enable the feature adjustment mode, determining that the current image block enables the feature adjustment mode; or, parsing a first value range of a feature value from the code stream, and if the feature value of the second feature is located in the first value range, determining that the current image block enables the feature adjustment mode.
6. An image coding method characterized by, The method comprises: acquire a probability distribution parameter based on the code stream corresponding to the current image block; determine a probability distribution model based on the probability distribution parameter, and decode the code stream corresponding to the current image block based on the probability distribution model to obtain a decoded image feature; determine a first feature corresponding to the current image block based on the decoded image feature; acquire a feature adjustment factor corresponding to the current image block based on the first feature; and encode the feature adjustment factor in the code stream corresponding to the current image block; perform synthetic transformation on the first feature through a synthetic transformation network to obtain a reconstructed image block corresponding to the current image block; wherein, in the process of synthetic transformation, the first feature is enhanced to obtain a second feature; the second feature is adjusted based on the feature adjustment factor to obtain a third feature; a target feature is determined based on the third feature; and the reconstructed image block corresponding to the current image block is determined based on the target feature.
7. An image decoding apparatus characterized by comprising: The apparatus comprises: an acquisition module configured to acquire a probability distribution parameter based on the code stream corresponding to the current image block; determine a probability distribution model based on the probability distribution parameter, and decode the code stream corresponding to the current image block based on the probability distribution model to obtain a decoded image feature; determine a first feature corresponding to the current image block based on the decoded image feature; and acquire a feature adjustment factor corresponding to the current image block based on the first feature; a determination module configured to perform synthetic transformation on the first feature through a synthetic transformation network to obtain a reconstructed image block corresponding to the current image block; wherein, in the process of synthetic transformation, the first feature is enhanced to obtain a second feature; the second feature is adjusted based on the feature adjustment factor to obtain a third feature; a target feature is determined based on the third feature; and the reconstructed image block corresponding to the current image block is determined based on the target feature.
8. An image coding apparatus characterized by comprising: The apparatus comprises: an acquisition module configured to acquire a probability distribution parameter based on the code stream corresponding to the current image block; determine a probability distribution model based on the probability distribution parameter, and decode the code stream corresponding to the current image block based on the probability distribution model to obtain a decoded image feature; determine a first feature corresponding to the current image block based on the decoded image feature; and acquire a feature adjustment factor corresponding to the current image block based on the first feature; an encoding module configured to encode the feature adjustment factor in the code stream corresponding to the current image block; a determination module configured to perform synthetic transformation on the first feature through a synthetic transformation network to obtain a reconstructed image block corresponding to the current image block; wherein, in the process of synthetic transformation, the first feature is enhanced to obtain a second feature; the second feature is adjusted based on the feature adjustment factor to obtain a third feature; a target feature is determined based on the third feature; and the reconstructed image block corresponding to the current image block is determined based on the target feature.
9. An image decoding apparatus characterized by comprising: comprise: a processor and a machine readable storage medium having stored machine executable instructions executable by the processor; The processor is configured to execute machine executable instructions to implement the method of any one of claims 1-5.
10. An image coding apparatus characterized by comprising: comprising: a processor and a machine readable storage medium storing machine executable instructions capable of being executed by the processor; The processor is configured to execute machine executable instructions to implement the method of claim 6.
11. A machine-readable storage medium, characterized in that, The machine readable storage medium stores a plurality of computer instructions, which, when executed by a processor, implement the method of any one of claims 1-5, or, when executed by a processor, implement the method of claim 6.
12. A computer program product, characterised in that, The computer program product comprises a computer program, which, when executed by a processor, implements the method of any one of claims 1-5, or, when executed by a processor, implements the method of claim 6.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and storage medium
CN110543900A
Image encoding and decoding, video encoding and decoding: methods, systems and training methods
WO2022084702A1