A decoding, encoding method, apparatus and device thereof
Patent Information
- Application Number
- CN202410391864.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-01
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2044-04-01
AI Technical Summary
虽然基于神经网络的编解码方法展现出巨大性能潜力,但是,基于神经网络的编解码方法仍然存在编码性能较差、解码性能较差和复杂度较高等问题
[0027]本申请提供一种机器可读存储介质,所述机器可读存储介质上存储有若干计算机指令,所述计算机指令被处理器执行时,实现上述的解码方法;或者,实现上述的编码方法。
Smart Images

Figure CN120786076B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of encoding and decoding technology, and in particular to a decoding and encoding method, apparatus and device thereof. Background Technology
[0002] To save space, video images are encoded before transmission. Complete video encoding includes processes such as prediction, transform, quantization, entropy coding, and filtering. The prediction process can be divided into intra-frame prediction and inter-frame prediction. Inter-frame prediction utilizes temporal correlation to predict the current pixel using pixels from neighboring encoded images, effectively removing temporal redundancy. Intra-frame prediction utilizes spatial correlation to predict the current pixel using pixels from the encoded blocks of the current frame, removing spatial redundancy.
[0003] With the rapid development of deep learning, it has achieved success in many high-level computer vision problems, such as image classification and object detection. Deep learning is also gradually being applied in the field of encoding and decoding, where neural networks can be used to encode and decode images. Although neural network-based encoding and decoding methods have shown great performance potential, they still suffer from problems such as poor encoding performance, poor decoding performance, and high complexity. Summary of the Invention
[0004] In view of this, this application provides a decoding and encoding method, apparatus and device, to improve encoding and decoding performance.
[0005] This application provides a decoding method applied at a decoding end, the method comprising:
[0006] Decode the residual information bitstream corresponding to the current frame image to obtain the residual latent features corresponding to the current frame image; perform residual decoding on the residual latent features to obtain the residual features corresponding to the current frame image;
[0007] Frame generation is performed based on the residual features and the acquired prediction features to obtain image reference features;
[0008] Generate a reconstructed image corresponding to the current frame image based on the image reference features;
[0009] The reconstructed image is stored in an information buffer as a temporal reference image for the next frame, or, based on the reconstructed image, a temporal reference feature to be updated is determined, and the determined temporal reference feature is stored in an information buffer as a temporal reference feature for the next frame.
[0010] This application provides an encoding method applied at an encoding end, the method comprising:
[0011] Decode the residual information bitstream corresponding to the current frame image to obtain the residual latent features corresponding to the current frame image; perform residual decoding on the residual latent features to obtain the residual features corresponding to the current frame image;
[0012] Frame generation is performed based on the residual features and the acquired prediction features to obtain image reference features;
[0013] A reconstructed image corresponding to the current frame image is generated based on the image reference features; the reconstructed image is stored in the information buffer as the temporal reference image for the next frame, or, a temporal reference feature to be updated is determined based on the reconstructed image, and the determined temporal reference feature is stored in the information buffer as the temporal reference feature for the next frame.
[0014] This application provides a decoding device applied at a decoding end, the device comprising:
[0015] The decoding module is used to decode the residual information bitstream corresponding to the current frame image to obtain the residual latent features corresponding to the current frame image; and to perform residual decoding on the residual latent features to obtain the residual features corresponding to the current frame image.
[0016] The generation module is used to perform frame generation operations based on the residual features and the acquired prediction features to obtain image reference features; and to generate a reconstructed image corresponding to the current frame image based on the image reference features.
[0017] The storage module is used to store the reconstructed image into an information buffer as a temporal reference image for the next frame, or to determine a temporal reference feature to be updated based on the reconstructed image and store the determined temporal reference feature into an information buffer as a temporal reference feature for the next frame.
[0018] This application provides an encoding device applied at an encoding end, the device comprising:
[0019] The decoding module is used to decode the residual information bitstream corresponding to the current frame image to obtain the residual latent features corresponding to the current frame image; and to perform residual decoding on the residual latent features to obtain the residual features corresponding to the current frame image.
[0020] The generation module is used to perform frame generation operations based on the residual features and the acquired prediction features to obtain image reference features; and to generate a reconstructed image corresponding to the current frame image based on the image reference features.
[0021] The storage module is used to store the reconstructed image into an information buffer as a temporal reference image for the next frame, or to determine a temporal reference feature to be updated based on the reconstructed image and store the determined temporal reference feature into an information buffer as a temporal reference feature for the next frame.
[0022] This application provides a decoding end device, the decoding end device including: a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor;
[0023] The processor is used to execute machine-executable instructions to implement the above-described decoding method.
[0024] This application provides an encoding end device, the decoding end device comprising: a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions executable by the processor;
[0025] The processor is used to execute machine-executable instructions to implement the above-described encoding method.
[0026] This application provides an electronic device, including: a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; the processor is configured to execute the machine-executable instructions to implement the above-described decoding method; or, the processor is configured to execute the machine-executable instructions to implement the above-described encoding method.
[0027] This application provides a machine-readable storage medium storing a plurality of computer instructions, which, when executed by a processor, implement the above-described decoding method; or, implement the above-described encoding method.
[0028] As can be seen from the above technical solutions, this application proposes an end-to-end video image compression method. Based on a neural network, it achieves video image encoding and decoding, enabling the neural network to maintain low complexity while effectively ensuring the quality of the reconstructed image, thereby improving encoding and decoding performance and reducing complexity. It can utilize long-term reference information in neural network video coding to obtain useful priors for images with large frame intervals, and apply these long-term priors to motion information coding, residual information coding, and temporal information mining. This reduces the residual coding bitrate, reduces the motion coding bitrate, and improves the objective quality. Attached Figure Description
[0029] Figure 1A This is a flowchart illustrating a decoding method in one embodiment of this application;
[0030] Figure 1B This is a flowchart illustrating an encoding method in one embodiment of this application;
[0031] Figure 2A This is a schematic diagram of the processing procedure at the encoding end in one embodiment of this application;
[0032] Figure 2B This is a schematic diagram of the time-domain residual mining module in one embodiment of this application;
[0033] Figure 2C This is a schematic diagram of the processing procedure at the decoding end in one embodiment of this application;
[0034] Figure 3A This is a schematic diagram of the processing procedure at the encoding end in one embodiment of this application;
[0035] Figure 3B This is a schematic diagram of the processing procedure at the decoding end in one embodiment of this application;
[0036] Figure 3C This is a schematic diagram of the processing procedure at the encoding end in one embodiment of this application;
[0037] Figure 3D , Figure 3E and Figure 3F This is a schematic diagram of the coding framework in one embodiment of this application;
[0038] Figure 3G This is a schematic diagram of the processing procedure at the decoding end in one embodiment of this application;
[0039] Figure 3H This is a schematic diagram of the decoding framework in one embodiment of this application;
[0040] Figure 4A and Figure 4B This is a schematic diagram of the processing procedure at the encoding end in one embodiment of this application;
[0041] Figure 5A , Figure 5B and Figure 5C This is a schematic diagram of the processing procedure at the encoding end in one embodiment of this application;
[0042] Figure 6A This is a hardware structure diagram of the decoding end device in one embodiment of this application;
[0043] Figure 6B This is a hardware structure diagram of the encoding end device in one embodiment of this application. Detailed Implementation
[0044] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the application. The singular forms “a,” “the,” and “the” used in the embodiments and claims of this application are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to any or all possible combinations including one or more of the associated listed items. It should be understood that although the terms first, second, third, etc., may be used to describe various information in the embodiments of this application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first information may also be referred to as second information without departing from the scope of the embodiments of this application, and similarly, second information may also be referred to as first information, depending on the context. Furthermore, the word “if” as used can be interpreted as “when,” “in response to a determination,” or “when…”.
[0045] This application proposes a decoding method and an encoding method, which may involve the following concepts:
[0046] Entropy coding: Entropy coding is a coding process that follows the principle of entropy without losing any information. Information entropy is the average amount of information in the source (a measure of uncertainty). Entropy coding methods can include, but are not limited to, Shannon coding, Huffman coding, and arithmetic coding.
[0047] Neural Networks (NNs): Neural networks are artificial neural networks, a computational model composed of numerous interconnected nodes (or neurons). In a neural network, neurons can represent different objects, such as features, letters, concepts, or meaningful abstract patterns. The types of processing units in a neural network can be divided into three categories: input units, output units, and hidden units. Input units receive signals and data from the external world; output units output the processed results; hidden units are located between the input and output units and cannot be observed from outside the system. The connection weights between neurons reflect the connection strength between units; the representation and processing of information are reflected in the connections between processing units. Neural networks are a non-programmed, brain-like information processing method. The essence of a neural network is to obtain a parallel and distributed information processing function through the transformations and dynamic behavior of the network, mimicking the information processing function of the human brain's nervous system to varying degrees and levels. In the field of video processing, commonly used neural networks include, but are not limited to, Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), and fully connected networks.
[0048] Convolutional Neural Networks (CNNs) are a type of feedforward neural network and one of the most representative network structures in deep learning. The artificial neurons in a CNN can respond to surrounding units within a certain coverage area, exhibiting excellent performance in large-scale image processing. The basic structure of a CNN consists of two layers: a feature extraction layer (also called a convolutional layer), where the input of each neuron is connected to the local receptive field of the previous layer, extracting features from that local area. Once these local features are extracted, their positional relationship with other features is determined. The second layer is a feature mapping layer (also called an activation layer). Each computational layer of the neural network consists of multiple feature maps, each a plane where all neurons have equal weights. Feature mapping structures can use functions such as the Sigmoid function, ReLU function, Leaky-ReLU function, PReLU function, and GDN function as activation functions for the convolutional network. Furthermore, because neurons on a single feature map share weights, the number of free parameters in the network is reduced.
[0049] For example, one advantage of convolutional neural networks (CNNs) over image processing algorithms is that they avoid complex preprocessing steps (such as extracting artificial features) and can directly input the original image for end-to-end learning. Another advantage of CNNs over ordinary neural networks is that ordinary neural networks use fully connected layers, meaning all neurons from the input layer to the hidden layer are connected. This results in a huge number of parameters, making network training time-consuming or even difficult. CNNs, however, avoid this difficulty through local connectivity and weight sharing.
[0050] Deconvolution: Also known as transposed convolution, deconvolution layers work similarly to convolutional layers. The main difference is that deconvolution layers use padding to make the output larger than the input (though they can also remain the same). If the stride is 1, the output size equals the input size; if the stride is N, the width of the output feature is N times the width of the input feature, and the height of the output feature is N times the height of the input feature.
[0051] Generalization ability: Generalization ability refers to the ability of a machine learning algorithm to adapt to new samples. The goal of learning is to learn the patterns hidden behind data pairs. The trained network can also give appropriate outputs for data outside the learning set that have the same pattern. This ability can be called generalization ability.
[0052] Features: The features involved in this application are three-dimensional feature matrices or tensors of size C*W*H. In the three-dimensional feature matrix, C represents the number of channels, H represents the feature height, and W represents the feature width. The three-dimensional feature matrix can be the input of a neural network, or it can be the output of a neural network.
[0053] Rate-Distortion Optimized (RDBEM) principle: Two main metrics for evaluating coding efficiency are bitrate and PSNR (Peak Signal-to-Noise Ratio). A smaller bitrate results in a higher compression ratio, and a higher PSNR leads to better reconstructed image quality. In mode selection, the decision formula essentially evaluates both factors. For example, the cost of a mode is: J(mode) = D + λ*R, where D represents Distortion, typically measured using the SSE metric (Sum of Mean Squares of Differences between the Reconstructed Image Patch and the Source Image). Alternatively, the SAD metric (Sum of Absolute Differences between the Reconstructed Image Patch and the Source Image) can be used to consider the cost. λ is a Lagrange multiplier, and R is the actual number of bits required to encode the image patch in that mode, including the total number of bits needed for encoding mode information, motion information, residuals, etc. Using the RDBEM principle to compare and decide on coding modes during mode selection usually ensures optimal coding performance.
[0054] Numerous encoding tools have been proposed for various modules at the encoding end, and each tool often has multiple modes. The optimal encoding tool for different video sequences often differs. Therefore, during encoding, Rate-Distortion Optimization (RDO) is typically used to compare the encoding performance of different tools or modes to select the best mode. After determining the optimal tool or mode, the decision information is transmitted by encoding marker information in the bitstream. Although this method introduces higher encoding complexity, it can adaptively select the optimal mode combination for different content to achieve the best encoding performance. The decoding end can obtain the relevant mode information by directly parsing the marker information, with minimal impact from complexity.
[0055] The decoding and encoding methods in the embodiments of this application will be described in detail below with reference to several specific embodiments.
[0056] Example 1: This application proposes a decoding method, see [link to example]. Figure 1A The diagram shown illustrates the flowchart of this decoding method, which can be applied to the decoding end (also known as a video decoder). This method may include:
[0057] Step 111: Perform temporal information mining based on the temporal reference information stored in the information buffer to determine the prediction features; wherein, the temporal reference information is the temporal reference image or temporal reference feature corresponding to the reconstructed image of the historical frame.
[0058] For example, a historical frame can be the previous frame of the current frame image, or it can be other historical frames, such as the second or third frame preceding the current frame image. In subsequent embodiments, the previous frame of the current frame image will be used as an example for illustration.
[0059] Step 112: Decode the residual information bitstream corresponding to the current frame image to obtain the residual latent features corresponding to the current frame image; perform residual decoding on the residual latent features to obtain the residual features corresponding to the current frame image.
[0060] Step 113: Perform frame generation operation based on residual features and predicted features to obtain image reference features.
[0061] Step 114: Generate the reconstructed image corresponding to the current frame image based on the image reference features.
[0062] Step 115: Store the reconstructed image in the information buffer as the temporal reference image for the next frame, or determine the temporal reference features to be updated based on the reconstructed image, and store the determined temporal reference features in the information buffer as the temporal reference features for the next frame.
[0063] For example, performing temporal information mining based on temporal reference information stored in the information buffer to determine prediction features may include, but is not limited to: extracting features from the temporal reference information to obtain short-term temporal reference features; decoding the motion information bitstream corresponding to the current frame image to obtain motion information features corresponding to the current frame image; decoding the motion information features to obtain optical flow features corresponding to the current frame image; and performing temporal information mining based on the optical flow features and short-term temporal reference features to obtain prediction features.
[0064] For example, determining predictive features by performing temporal information mining based on temporal reference information stored in the information buffer may include, but is not limited to: extracting features from the temporal reference information to obtain short-term temporal reference features; decoding the motion information bitstream corresponding to the current frame image to obtain motion information features corresponding to the current frame image; decoding the motion information features to obtain optical flow features corresponding to the current frame image; and performing temporal information mining based on the optical flow features, short-term temporal reference features, and acquired long-term temporal alignment features to obtain predictive features. For instance, the long-term temporal alignment features may represent prior information from multiple frames preceding the current frame image.
[0065] For example, determining prediction features by mining temporal information based on optical flow features, short-term temporal reference features, and acquired long-term temporal alignment features may include: aligning the short-term temporal reference features with optical flow features to obtain short-term temporal alignment features; concatenating the long-term temporal alignment features with the short-term temporal alignment features along the channel dimension to obtain concatenated hybrid features; and performing convolution and / or residual block operations on the concatenated hybrid features to obtain prediction features.
[0066] For example, before determining the prediction features, temporal information mining based on optical flow features, short-term temporal reference features, and acquired long-term temporal alignment features can be performed. The following steps can be used to obtain the long-term temporal alignment features: If the current frame image is the first frame image in the image group, then the long-term temporal update features are determined based on the short-term temporal reference features; the long-term temporal update features are aligned using optical flow features to obtain the long-term temporal alignment features; the long-term temporal alignment features are input to the configured long-term buffer, which updates the long-term temporal alignment features to the long-term temporal reference features of the next frame; wherein, the images using inter-frame coding are divided into multiple image groups, and each image group includes multiple frames. Alternatively,
[0067] If the current frame image is not the first frame image in the image group, then the long-term temporal update features are determined based on the short-term temporal reference features and the long-term temporal reference features in the long-term buffer (representing the prior information of the previous multiple frames of the current frame image); the long-term temporal update features are aligned using optical flow features to obtain long-term temporal alignment features; the long-term temporal alignment features are input to the long-term buffer, and the long-term buffer updates the long-term temporal alignment features to the long-term temporal reference features of the next frame.
[0068] For example, determining long-term temporal update features based on short-term temporal reference features may include: for the first frame image within the first image group, determining the short-term temporal reference features as long-term temporal update features; for the first frame image outside the first image group, determining the short-term temporal reference features as long-term temporal update features; or, for the first frame image outside the first image group, multiplying the long-term temporal reference features in the long-term buffer by a memory coefficient to obtain long-term memory features, determining the initialized temporal features based on the long-term memory features and the short-term temporal reference features, and determining the long-term temporal update features based on the initialized temporal features.
[0069] For example, determining the long-term temporal update features based on short-term temporal reference features and long-term temporal reference features in the long-term buffer can include: performing a channel-level concatenation operation on the short-term and long-term temporal reference features, and then performing a convolution operation on the concatenated features to obtain the long-term temporal update features; or, the long-term temporal update features can be determined using the following formula: in, It can represent long-term time-domain update features. It can represent long-term time-domain reference characteristics. 1 can represent the short-term time-domain reference feature, 'a' can represent the weight coefficient of the long-term time-domain reference feature, and 1-a can represent the weight coefficient of the short-term time-domain reference feature.
[0070] For example, decoding the motion information bitstream corresponding to the current frame image to obtain the motion information features corresponding to the current frame image may include: decoding the motion side information bitstream corresponding to the current frame image to obtain motion prior information; obtaining temporal motion prior information from the motion latent buffer, where the temporal motion prior information is the motion information features corresponding to the reconstructed images of historical frames; generating a probability distribution of motion information features based on the motion prior information and the temporal motion prior information; using this probability distribution to decode the motion information bitstream to obtain the motion information features corresponding to the current frame image; and storing the motion information features in the motion latent buffer as the temporal motion prior information for the next frame.
[0071] Alternatively, the acquired long-term temporal update features can be manipulated to obtain motion space prior information; a probability distribution of motion information features can be generated based on the motion super-prior information, temporal motion prior information, and motion space prior information; the motion information bitstream can be decoded using this probability distribution to obtain the motion information features corresponding to the current frame image; and the motion information features can be stored in the motion latent buffer as the temporal motion prior information for the next frame.
[0072] For example, operating on the acquired long-term temporal update features to obtain motion space prior information includes: performing convolution and / or activation operations on the long-term temporal update features to obtain motion space prior information; wherein, performing convolution and / or activation operations on the long-term temporal update features to obtain motion space prior information includes: sequentially performing convolution, activation, convolution, activation, and convolution operations on the long-term temporal update features to obtain motion space prior information.
[0073] For example, decoding the residual information bitstream corresponding to the current frame image to obtain the residual latent features corresponding to the current frame image may include: decoding the residual side information bitstream corresponding to the current frame image to obtain residual super-prior information; obtaining residual prior information from the residual latent buffer, where the residual prior information is the residual latent features corresponding to the reconstructed images of historical frames; obtaining spatial prior information based on predicted features; or obtaining spatial prior information based on predicted features and acquired long-term temporal alignment features; generating a probability distribution of residual latent features based on residual super-prior information, residual prior information, and spatial prior information; using this probability distribution to decode the residual information bitstream to obtain the residual latent features corresponding to the current frame image; and storing these residual latent features in the residual latent buffer as residual prior information for the next frame.
[0074] For example, obtaining spatial prior information by performing spatial prior processing based on predicted features and acquired long-term temporal alignment features may include: concatenating the predicted features and long-term temporal alignment features along the channel dimension, and performing convolution and / or activation operations on the concatenated features to obtain spatial prior information; wherein: performing convolution and / or activation operations on the concatenated features to obtain spatial prior information includes: sequentially performing downsampling convolution, activation, and downsampling convolution operations on the concatenated features to obtain spatial prior information.
[0075] Determining the temporal reference features to be updated based on the reconstructed image can include: extracting features from the reconstructed image using a feature extractor to obtain short-term temporal reference features, and then determining the temporal reference features based on these short-term temporal reference features. Alternatively, extracting features from the reconstructed image using a feature extractor to obtain short-term temporal reference features, and then determining the temporal reference features based on these short-term temporal reference features and the acquired long-term temporal reference features. Specifically, during frame generation based on residual features and predicted features, the residual features and predicted features are input to the frame generation network, which outputs image reference features and long-term temporal reference features. These image reference features and long-term temporal reference features are output features from different layers of the frame generation network.
[0076] For example, determining time-domain reference features based on short-term time-domain reference features and acquired long-term time-domain reference features may include: extracting features from the long-term time-domain reference features using a feature extractor to obtain long-term time-domain fine features; fusing the short-term time-domain reference features and the long-term time-domain fine features to obtain fused time-domain features; and determining time-domain reference features based on the fused time-domain features.
[0077] For example, fusing short-term time-domain reference features and long-term time-domain fine features to obtain fused time-domain features can include: determining the fused time-domain features using the following formula: in, Indicates the fusion of temporal features, Represents long-term time-domain fine features, denoted as short-term time-domain reference feature, 'a' represents the weight coefficient of long-term time-domain fine feature, and 1-a represents the weight coefficient of short-term time-domain reference feature.
[0078] For example, performing temporal information mining based on temporal reference information stored in the information buffer to determine prediction features may include: extracting features from the temporal reference information to obtain short-term temporal reference features; decoding the motion information bitstream corresponding to the current frame image to obtain motion information features corresponding to the current frame image; decoding the motion information features to obtain optical flow features corresponding to the current frame image; if the temporal reference information is a temporal reference image corresponding to a reconstructed image of a historical frame, then obtaining background optical flow and background features based on the temporal reference image, determining aligned background features based on the background optical flow and background features; and determining prediction features based on optical flow features, short-term temporal reference features, and aligned background features.
[0079] For example, determining the prediction features based on optical flow features, short-term temporal reference features, and aligned background features may include: using optical flow features to align the short-term temporal reference features to obtain short-term temporal aligned features; concatenating the aligned background features and the short-term temporal aligned features along the channel dimension to obtain concatenated hybrid features; and performing convolution and / or residual block operations on the concatenated hybrid features to obtain the prediction features.
[0080] For example, obtaining background optical flow and background features based on a temporal reference image may include: if the current frame image is the first frame image in the image group, then performing feature extraction on the temporal reference image through a feature extractor to obtain a first feature; inputting the first feature and optical flow features into a mask generation network to obtain mask features and background optical flow; generating background features based on the first feature and mask features; if the current frame image is not the first frame image in the image group, then performing feature extraction on the temporal reference image through a feature extractor to obtain a first feature; inputting the first feature, optical flow features, and background update features in the background information buffer into the mask generation network to obtain mask features and background optical flow; generating a second feature based on the first feature and mask features, and generating background features based on the second feature and background update features.
[0081] For example, determining the aligned background features based on background optical flow and background features may include: performing an alignment operation on the background features using background optical flow to obtain aligned background features; after obtaining the aligned background features, storing the aligned background features in a background information buffer, and updating the aligned background features with the background update features in the background information buffer.
[0082] For example, the above execution order is merely an example for ease of description. In practical applications, the execution order between steps can be changed, and there is no limitation on this execution order. Moreover, in other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in this specification, and the method may include more or fewer steps than described in this specification. Furthermore, a single step described in this specification may be broken down into multiple steps in other embodiments; multiple steps described in this specification may also be combined into a single step in other embodiments.
[0083] As can be seen from the above technical solutions, this application proposes an end-to-end video image compression method. Based on a neural network, it achieves video image encoding and decoding, enabling the neural network to maintain low complexity while effectively ensuring the quality of the reconstructed image, thereby improving encoding and decoding performance and reducing complexity. It can utilize long-term reference information in neural network video coding to obtain useful priors for images with large frame intervals, and apply these long-term priors to motion information coding, residual information coding, and temporal information mining. This reduces the residual coding bitrate, reduces the motion coding bitrate, and improves the objective quality.
[0084] Example 2: An encoding method is proposed in this application embodiment, see [link to example]. Figure 1B The diagram shown illustrates the flowchart of this encoding method, which can be applied to the encoding end (also known as a video encoder). This method may include:
[0085] Step 121: Perform temporal information mining based on the temporal reference information stored in the information buffer to determine the prediction features; wherein, the temporal reference information is the temporal reference image or temporal reference feature corresponding to the reconstructed image of the historical frame.
[0086] Step 122: Decode the residual information bitstream corresponding to the current frame image to obtain the residual latent features corresponding to the current frame image; perform residual decoding on the residual latent features to obtain the residual features corresponding to the current frame image.
[0087] Step 123: Perform frame generation operation based on residual features and prediction features to obtain image reference features.
[0088] Step 124: Generate the reconstructed image corresponding to the current frame image based on the image reference features.
[0089] Step 125: Store the reconstructed image in the information buffer as the temporal reference image for the next frame; or, determine the temporal reference features to be updated based on the reconstructed image, and store the determined temporal reference features in the information buffer as the temporal reference features for the next frame. For example, the processing at the encoding end is similar to that at the decoding end, and the similarities will not be repeated. The processing at the decoding end can be applied to the encoding end, i.e., the encoding end adopts the same processing method.
[0090] For example, the above execution order is merely an example for ease of description. In practical applications, the execution order between steps can be changed, and there is no limitation on this execution order. Moreover, in other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in this specification, and the method may include more or fewer steps than described in this specification. Furthermore, a single step described in this specification may be broken down into multiple steps in other embodiments; multiple steps described in this specification may also be combined into a single step in other embodiments.
[0091] As can be seen from the above technical solutions, this application proposes an end-to-end video image compression method. Based on a neural network, it achieves video image encoding and decoding, enabling the neural network to maintain low complexity while effectively ensuring the quality of the reconstructed image, thereby improving encoding and decoding performance and reducing complexity. It can utilize long-term reference information in neural network video coding to obtain useful priors for images with large frame intervals, and apply these long-term priors to motion information coding, residual information coding, and temporal information mining. This reduces the residual coding bitrate, reduces the motion coding bitrate, and improves the objective quality.
[0092] Example 3: For the processing procedures at the encoding end in Examples 1 and 2, please refer to... Figure 2A As shown, of course, Figure 2A This is just one example of the processing procedure at the encoding end, and no restrictions are imposed on this processing procedure.
[0093] 1. The encoding end obtains the current frame image x t (The current frame image can be the original image, i.e., the input image.) Then, the current frame image x... t The input is fed into the feature extractor to obtain the current image features F. t .
[0094] Current frame image x t It can be divided into one image block or multiple image blocks. If the current frame image x t Dividing it into one image block allows us to access the current frame image x. tProcessing is performed if the current frame image x t If the image is divided into multiple image blocks, then for the current frame image x... t Each image patch is processed separately, and the processing method for each image patch is the same. Subsequent processing is based on the current frame image x. t For example.
[0095] The feature extractor is used to extract the x-axis from the current frame image. t The image features, therefore, when using the current frame image x t After being input into the feature extractor, the current frame image x can be obtained. t The corresponding current image feature F t .
[0096] 2. The information buffer (i.e., the Decoded Feature & Picture Buffer, also known as the decoded image / feature buffer) stores the first temporal reference features. First time-domain reference features It is the image reference feature of the previous frame. (Also known as time-domain reference feature) Image reference features (See subsequent steps for the method of obtaining the first time-domain reference feature). The feature is then processed by the feature extractor to obtain the second temporal reference feature.
[0097] The feature extractor is used to extract the first temporal reference features. The image features, therefore, in the first time-domain reference features After being fed into the feature extractor, the second temporal reference features can be obtained.
[0098] 3. The information buffer (Decoded Feature & Picture Buffer) stores prior motion features. Prior motion characteristics It is the prior motion feature of the previous frame image. Prior motion characteristics See the following steps for how to obtain it.
[0099] 4. Current image features F t Second time-domain reference features Prior motion characteristics The input is fed to the motion encoder, which then uses the current image features F to... tSecond time-domain reference features Prior motion characteristics Motion information is encoded to obtain unquantized latent motion features m t .
[0100] For example, a motion encoder can be a neural network used to encode motion information, incorporating current image features F t Second time-domain reference features Prior motion characteristics After being input into the motion encoder, the motion encoder can process these features without limiting the structure and processing method of the motion encoder, thus obtaining the motion latent features m. t .
[0101] 5. Regarding the potential characteristics of motion m t Quantization (i.e., Q-operation) is performed to obtain motion information features.
[0102] 6. The Motion Latent Buffer stores time-domain motion prior information. Prior information of motion in the time domain It is the motion information feature of the previous frame. Motion information features See the following steps for how to obtain it.
[0103] Motion potential characteristics m t and time-domain motion prior information The input is fed into the Hyper Enc (hyper-encoder), which then obtains the motion latent features m. t and time-domain motion prior information Then, the potential motion features m t and time-domain motion prior information Encoding is performed to obtain the motion side information bitstream. After obtaining the motion side information bitstream, the encoding end can also transmit the motion side information bitstream to the decoding end. The processing procedure at the decoding end is described in subsequent steps.
[0104] 7. Hyper Dec (Hyper Decoder) can acquire the motion side information bitstream and decode it to obtain the motion hyper prior information.
[0105] 8. Hyper Prior and Temporal Motion Prior (i.e., Temporal Prior) is input into the Motion Entropy Model, which is then processed by the Motion Entropy Model based on the motion prior information and the temporal motion prior information. Generate motion information features The probability distribution is used to guide encoding and decoding.
[0106] 9. Based on motion information features The probability distribution of motion information features Encoding (i.e., AE operation) is performed to obtain the motion information bitstream; there are no restrictions on this encoding process. After obtaining the motion information bitstream, the encoding end can also transmit the motion information bitstream to the decoding end. The processing procedure for the motion information bitstream at the decoding end is described in subsequent steps.
[0107] For example, it can be based on motion information features The probability distribution model is determined by the probability distribution (such as probability distribution parameters), and then motion information features are analyzed based on this probability distribution model. Encode the data to obtain the motion information bitstream.
[0108] 10. Based on motion information features The probability distribution is used to decode the motion information bitstream (i.e., AD operation) to obtain motion information features. Motion information features Inverse quantization can also be performed to obtain the latent motion features m. t .
[0109] For example, after obtaining the motion information bitstream, it is possible to analyze the motion information features. The probability distribution model is determined by the probability distribution (such as probability distribution parameters), and then the motion information bitstream is decoded based on the probability distribution model.
[0110] 11. After obtaining motion information features (That is, decoding the motion information bitstream to obtain motion information features) After that, motion information features Stored in the Motion Latent Buffer as temporal motion prior information for the next frame.
[0111] 12. Motion Information Characteristics The input is fed to the motion decoder, which then uses the motion information features... Decoding motion information is performed, without restrictions on the structure and decoding method of the motion decoder, to obtain optical flow characteristics. and prior motion characteristics Prior motion characteristics It is stored in the information buffer (Decoded Feature & Picture Buffer), and this prior motion feature It can be used as a priori motion feature for the next frame image.
[0112] For example, a motion decoder can include multiple network layers, where the motion decoder is based on motion information features. When decoding motion information, the features output by the penultimate network layer of the motion decoder can be used as prior motion features. Of course, the above is just an example; the output features of other network layers (such as the third-to-last network layer) can also be used as prior motion features. The features output by the last network layer of the motion decoder can be used as optical flow features. Of course, the above is just an example; the output features of other network layers can also be used as optical flow features.
[0113] 13. Optical flow characteristics Second time-domain reference features (i.e., the first temporal reference feature in the information buffer) The second temporal reference features are obtained after feature extraction. This information, also known as short-term temporal reference features, is input to the temporal residual mining module. The temporal residual mining module is based on optical flow features. Second time-domain reference features Temporal residual mining is performed to obtain predicted features. Predicted features Used for residual conditional coding.
[0114] For example, see Figure 2B As shown, the temporal residual mining module involves an optical flow alignment module and a motion compensation module. It can extract optical flow features... Second time-domain reference features The input is fed to the optical flow alignment module, which then uses optical flow features. For the second time-domain reference features Perform alignment operations.
[0115] Aligned features and second temporal reference features can be used. The input is fed into the motion compensation module, which then uses the aligned features and the second temporal reference features as inputs. Perform motion compensation to obtain predicted features.
[0116] 14. Current image features F t and predictive features The input is fed to the contextual encoder, which then uses the current image features F to... t and predictive features Perform residual encoding to obtain unquantized residual information features y t .
[0117] For example, a residual encoder can be a neural network used to implement residual coding, which processes the current image features F... t and predictive features After being input into the residual encoder, the residual encoder can use the current image features F t and predictive features The residual encoder's structure and processing method are not restricted, resulting in unquantized residual information features y. t .
[0118] 15. Features of unquantified residual information y t Quantization (i.e., Q-operation) is performed to obtain the latent features of the residuals.
[0119] 16. Predictive Features The data is fed into the Spatial Prior Encoder, which then uses the predicted features... The process is performed without any restrictions on the method used, resulting in spatial prior information.
[0120] 17. The Contextual Latent Buffer stores residual prior information. Residual prior information It is the residual latent feature of the previous frame. Residual latent characteristics See the following steps for how to obtain it.
[0121] The residual information feature y t Spatial Prior Information and Residual Prior Information The input is fed into the Hyper Enc (hyper-encoder). The hyper-encoder processes the residual information features y. t Spatial Prior Information and Residual Prior Information Encode the data to obtain the residual side information bitstream. After obtaining the residual side information bitstream, the encoder can also send it to the decoder. The decoding process is described in subsequent steps.
[0122] 18. Hyper Dec (Hyper Decoder) can obtain the residual edge information bitstream. After obtaining the residual edge information bitstream, it decodes the residual edge information bitstream to obtain the residual hyperprior information.
[0123] 19. Residual Hyper-Prior Information and Spatial Prior Information (i.e., predictive features) (After inputting into the spatial prior encoder, the residual prior information is obtained) (That is, Temporal Prior, output by ContextualLatent Buffer) is input to Contextual Entropy Model, from which residual prior information, spatial prior information, and residual prior information are obtained. Subsequently, based on residual prior information, spatial prior information, and residual prior information... Generate residual latent features probability distribution, residual latent characteristics The probability distribution is used to guide encoding and decoding.
[0124] 20. Based on residual latent characteristics The probability distribution of residual latent features Encoding (i.e., AE operation) is performed to obtain the residual information bitstream; this encoding process is not restricted. After obtaining the residual information bitstream, the encoding end can also transmit the residual information bitstream to the decoding end. The processing procedure for the residual information bitstream at the decoding end is described in subsequent steps.
[0125] For example, it can be based on the latent features of the residuals. The probability distribution model is determined by the probability distribution (such as probability distribution parameters), and then the latent features of the residuals are analyzed based on this probability distribution model. Encode the data to obtain the residual information bitstream.
[0126] 21. Based on residual latent characteristics The probability distribution is used to decode the residual information bitstream (i.e., AD operation) to obtain the residual latent features. Residual latent characteristics Inverse quantization can also be performed to obtain the residual information feature y. t .
[0127] For example, it can be based on the latent features of the residuals. The probability distribution model is determined by the probability distribution (such as probability distribution parameters), and then the residual information bitstream is decoded based on this probability distribution model to obtain the residual latent features.
[0128] 22. After obtaining the latent characteristics of the residuals (i.e., the residual latent features obtained by decoding the residual information bitstream) After that, the latent characteristics of the residuals It is stored in the Contextual Latent Buffer as residual prior information for the next frame.
[0129] 23. Latent characteristics of residuals The input is fed to the contextual decoder, which then uses the residual latent features to... The residual information is decoded without any restrictions on the structure of the residual decoder, and the residual features are obtained.
[0130] 24. Combine residual features and prediction features The input is fed into the frame generator, which then uses the residual features and prediction features to generate the frame. A transformation from the feature domain to the image domain is performed to obtain image reference features. (Image reference features) It can also be called time-domain reference feature ), and image reference features Image reference features are stored in the information buffer (DecodedFeature & Picture Buffer). As the first temporal reference feature of the next frame
[0131] 25. Based on image reference features Generate reconstructed image For example, using image reference features The input is fed into a deconvolutional network, which is based on image reference features. Perform an upsampling operation to obtain the current frame image x. t Corresponding reconstructed image
[0132] This completes the encoding process, and the current frame image x can be obtained. t Corresponding reconstructed image
[0133] In one possible implementation, the aforementioned encoding process can be executed by a deep learning model or a neural network model to achieve end-to-end image compression and encoding. No restrictions are placed on this encoding process. The aforementioned deep learning model or neural network model can be referred to as an end-to-end video coding exploration experiment model. The framework of the video coding exploration experiment model is described below. Figure 2A As shown, the video coding exploration experimental model transforms the image into the feature domain.
[0134] For example, in embodiment 3, if the current frame image x t Dividing the image into multiple patches allows us to execute steps 1-25 for the current image patch. For example, in step 1, the current image patch is input to the feature extractor to obtain the current image feature F corresponding to the current image patch. t In step 2, the first time-domain reference feature is obtained from the information buffer. It is the first temporal reference feature of the reference image block corresponding to the current image patch in the previous frame. The position of the reference image block is the same as the position of the current image block, or the position of the reference image block is determined based on the position of the current image block and the motion vector (the motion vector between the current frame image and the previous frame image), and there are no restrictions on the position of the reference image block.
[0135] In step 3, the prior motion features are obtained from the information buffer. It is the prior motion feature of the reference image patch corresponding to the current image patch in the previous frame. In steps 4 and 5, what is obtained is the image corresponding to the current image patch.
[0136] Motion potential characteristics m t Motion information features corresponding to the current image patch In step 6, time-domain motion prior information It is the motion information feature of the reference image patch corresponding to the current image patch in the previous frame.
[0137] In steps 7-10, the motion information of the current image block is encoded and decoded to obtain the motion latent feature m corresponding to the current image block. t In step 11, the motion information features corresponding to the current image patch are... Stored in the Motion Latent Buffer. In step 12, the optical flow features corresponding to the current image patch are obtained. Prior motion features corresponding to the current image patch The prior motion features corresponding to the current image patch It is stored in the information buffer.
[0138] In step 13, the predicted features corresponding to the current image patch are obtained. In step 14, the residual information feature y corresponding to the current image patch is obtained. t In step 15, the residual latent features corresponding to the current image patch are obtained. In steps 16-21, the residual information of the current image block is encoded and decoded to obtain the residual latent features corresponding to the current image block. In step 22, the residual latent features corresponding to the current image patch are... Stored in the ContextualLatent Buffer. Steps 23-24 involve obtaining the image reference features corresponding to the current image patch. Image reference features corresponding to the current image patch Stored in the information buffer. In step 25, the reconstructed image block corresponding to the current image block is obtained.
[0139] For example, the current frame image x t After dividing the image into multiple image blocks, then processing the current frame image x... t After performing the above processing on each image block, a reconstructed image block corresponding to each image block can be obtained. The reconstructed image blocks corresponding to all image blocks then form the reconstructed image corresponding to the current frame image. Furthermore, image reference features are stored in the information buffer. At that time, the image reference features corresponding to all image blocks are stored. After that, the image reference features corresponding to the current frame image are actually stored (i.e., the image reference features of the entire frame image). In other words, the information buffer stores the image reference features of the entire frame image.
[0140] Example 4: For the processing procedures at the decoding end in Examples 1 and 2, please refer to... Figure 2C As shown, of course, Figure 2C This is just one example of the processing procedure at the decoding end, and no restrictions are imposed on the processing procedure at the decoding end.
[0141] 1. The decoding end obtains the motion edge information bitstream.
[0142] 2. After obtaining the motion side information bitstream, the Hyper Dec (hyper prior) can decode the motion side information bitstream to obtain the motion hyper prior information.
[0143] 3. Hyper Prior and Temporal Motion Prior (i.e., Temporal Prior) is input into the Motion Entropy Model, which is then processed by the Motion Entropy Model based on the motion prior information and the temporal motion prior information. Generate motion information features The probability distribution is used to guide the decoding.
[0144] For example, regarding time-domain motion prior information The Motion Latent Buffer stores time-domain motion prior information. Prior information of motion in the time domain It is the motion information feature of the previous frame.
[0145] 4. The decoding end obtains the motion information bitstream.
[0146] 5. Based on motion information features The probability distribution is used to decode the motion information bitstream (i.e., AD operation) to obtain motion information features. Motion information features Inverse quantization can also be performed to obtain the latent motion features m. t .
[0147] For example, after obtaining the motion information bitstream, it is possible to analyze the motion information features. The probability distribution model is determined by the probability distribution (such as probability distribution parameters), and then the motion information bitstream is decoded based on the probability distribution model.
[0148] 6. Obtaining motion information features (That is, decoding the motion information bitstream to obtain motion information features) After that, motion information features Stored in the Motion Latent Buffer as temporal motion prior information for the next frame.
[0149] 7. Motion Information Characteristics The input is fed to the motion decoder, which then uses the motion information features... Decoding motion information is performed, without restrictions on the structure and decoding method of the motion decoder, to obtain optical flow characteristics.
[0150] For example, the motion decoder is based on motion information features When decoding motion information, prior motion features can also be output. For example, a motion decoder can include multiple network layers, where the motion decoder is based on motion information features. When decoding motion information, the features output by the penultimate network layer of the motion decoder can be used as prior motion features. Of course, the above is just an example; the output features of other network layers (such as the third-to-last network layer) can also be used as prior motion features. The features output by the last network layer of the motion decoder can be used as optical flow features. Of course, the above is just an example; the output features of other network layers can also be used as optical flow features.
[0151] 8. Optical flow characteristics Second time-domain reference features (i.e., the first temporal reference feature in the information buffer) The second temporal reference features are obtained after feature extraction. This information, also known as short-term temporal reference features, is input to the temporal residual mining module. The temporal residual mining module is based on optical flow features. Second time-domain reference features Temporal residual mining is performed to obtain predicted features. Predicted features Used for residual conditional decoding.
[0152] For example, see Figure 2B As shown, the temporal residual mining module involves an optical flow alignment module and a motion compensation module. It can extract optical flow features... Second time-domain reference features The input is fed to the optical flow alignment module, which then uses optical flow features. For the second time-domain reference features Perform alignment operations.
[0153] Aligned features and second temporal reference features can be used. The input is fed into the motion compensation module, which then uses the aligned features and the second temporal reference features as inputs. Perform motion compensation to obtain predicted features.
[0154] For example, the information buffer (Decoded Feature & Picture Buffer) stores the first temporal reference feature. First time-domain reference features It is the image reference feature of the previous frame. (See subsequent steps for acquisition method) The first time-domain reference features can be obtained. The input is fed into the feature extractor to obtain the second temporal reference feature.
[0155] 9. The decoding end obtains the residual edge information bitstream.
[0156] 10. After obtaining the residual side information bitstream, the Hyper Dec (hyper prior) decodes the residual side information bitstream to obtain the residual hyper prior information.
[0157] 11. Predictive Features The data is fed into the Spatial Prior Encoder, which then uses the predicted features... The process is performed without any restrictions on the method used, resulting in spatial prior information.
[0158] 12. The Contextual Latent Buffer stores residual prior information. Residual prior information It is the residual latent feature of the previous frame. Residual prior information As a Temporal Prior.
[0159] 13. Hyper Prior, Spatial Prior, and Residual Prior Information The Temporal Prior is input into the Contextual Entropy Model, from which residual prior information, spatial prior information, and residual prior information are obtained. Subsequently, based on residual prior information, spatial prior information, and residual prior information... Generate residual latent features probability distribution, residual latent characteristics The probability distribution is used to guide the decoding.
[0160] 14. The decoding end obtains the residual information bitstream.
[0161] 15. Based on residual latent characteristics The probability distribution is used to decode the residual information bitstream (i.e., AD operation) to obtain the residual latent features. Residual latent characteristics Inverse quantization can also be performed to obtain the residual information feature y. t .
[0162] For example, it can be based on the latent features of the residuals. The probability distribution model is determined by the probability distribution (such as probability distribution parameters), and then the residual information bitstream is decoded based on this probability distribution model to obtain the residual latent features.
[0163] 16. After obtaining the latent characteristics of the residuals (i.e., the residual latent features obtained by decoding the residual information bitstream) After that, the latent characteristics of the residuals It is stored in the Contextual Latent Buffer as residual prior information for the next frame.
[0164] 17. Potential characteristics of residuals The input is fed to the contextual decoder, which then uses the residual latent features to... The residual information is decoded without any restrictions on the structure of the residual decoder, and the residual features are obtained.
[0165] 18. Combine residual features and prediction features The input is fed into the frame generator, which then uses the residual features and prediction features to generate the frame. A transformation from the feature domain to the image domain is performed to obtain image reference features. (Image reference features) It can also be called time-domain reference feature ), and image reference features Image reference features are stored in the information buffer (DecodedFeature & Picture Buffer). As the first temporal reference feature of the next frame
[0166] 19. Based on image reference features Generate reconstructed image For example, using image reference features The input is fed into a deconvolutional network, which is based on image reference features. Perform an upsampling operation to obtain the current frame image x. t Corresponding reconstructed image
[0167] This completes the decoding process, and the current frame image x can be obtained. t Corresponding reconstructed image
[0168] In one possible implementation, the aforementioned decoding process can be executed by a deep learning model or a neural network model to achieve end-to-end image compression and encoding. No restrictions are placed on the decoding process. This deep learning model or neural network model can be referred to as an end-to-end video decoding exploration experiment model. The framework of the video decoding exploration experiment model is described in [reference needed]. Figure 2C As shown, the video decoding exploration experimental model transforms the image into the feature domain.
[0169] For example, in embodiments 3 and 4, the reference information stored in the information buffer (Decoded Feature & Picture Buffer, also known as the decoded image / feature buffer) is the image reference feature. Image reference features It has cumulative properties, such as the features of the first frame image being accumulated into the image reference features. The features of the second frame image are accumulated into the image reference features. The features of the third frame image are accumulated into the image reference features. And so on. As the number of frames accumulates, temporal information in the image reference features... The continuous accumulation of these features leads to changes in image reference features. The accuracy has decreased.
[0170] In summary, as the number of frames accumulates, the image reference features stored in the information buffer increase. The accuracy is decreasing, especially when based on image reference features. When decoding the current frame image, reconstruct the image. The accuracy is getting lower and lower.
[0171] To address the above findings, this application proposes a temporal feature optimization method based on neural network video coding. By optimizing the reference information stored in the information buffer, the reconstructed image can be improved. The accuracy.
[0172] For example, in embodiment 4, if the current frame image x t Dividing the image into multiple blocks allows us to execute steps 1-19 for the current image block. For example, steps 1-5 involve decoding the motion information of the current image block to ultimately obtain the motion latent feature m corresponding to that block. t In step 6, the motion information features corresponding to the current image patch are... Stored in the Motion Latent Buffer. In step 7, the optical flow features corresponding to the current image patch are obtained. Prior motion features corresponding to the current image patch The prior motion features corresponding to the current image patch It is stored in the information buffer.
[0173] In step 8, the predicted features corresponding to the current image patch are obtained. First temporal reference feature obtained from the information buffer It is the first temporal reference feature of the reference image block corresponding to the current image patch in the previous frame. The position of the reference image block is the same as the position of the current image block, or the position of the reference image block is determined based on the position of the current image block and the motion vector (the motion vector between the current frame image and the previous frame image), and there are no restrictions on the reference image block.
[0174] In steps 9-15, the residual information of the current image patch is decoded to obtain the residual latent features corresponding to the current image patch. In step 16, the residual latent features corresponding to the current image patch are... Stored in the Contextual Latent Buffer. In steps 17-18, the image reference features corresponding to the current image patch are obtained. Image reference features corresponding to the current image patch Stored in the information buffer. In step 19, the reconstructed image block corresponding to the current image block is obtained.
[0175] For example, the current frame image x t After dividing the image into multiple image blocks, then processing the current frame image x... t After performing the above processing on each image block, a reconstructed image block corresponding to each image block can be obtained. The reconstructed image blocks corresponding to all image blocks then form the reconstructed image corresponding to the current frame image. Furthermore, image reference features are stored in the information buffer. At that time, the image reference features corresponding to all image blocks are stored. After that, the image reference features corresponding to the current frame image are actually stored (i.e., the image reference features of the entire frame image). In other words, the information buffer stores the image reference features of the entire frame image.
[0176] Example 5: For the processing procedures at the encoding end in Examples 1 and 2, please refer to... Figure 3A As shown, of course, Figure 3A This is just one example of the processing procedure at the encoding end, and no restrictions are imposed on this processing procedure.
[0177] 1. The encoding end obtains the current frame image x t(The current frame image can be the original image, i.e., the input image.) Then, the current frame image x... t The input is fed into the feature extractor to obtain the current image features F. t .
[0178] 2. The information buffer (i.e., the Decoded Feature & Picture Buffer, also known as the decoded image / feature buffer) stores time-domain reference information. This time-domain reference information is either the time-domain reference image corresponding to the reconstructed image of the previous frame, or the time-domain reference feature corresponding to the reconstructed image of the previous frame. The method for obtaining the time-domain reference image or time-domain reference feature is detailed in subsequent steps. The time-domain reference information is then fed into the Feature Extractor to obtain short-term time-domain reference features. For example, a feature extractor is used to extract image features from temporal reference information. Therefore, after inputting temporal reference information into the feature extractor, short-term temporal reference features can be obtained.
[0179] 3. The information buffer (Decoded Feature & Picture Buffer) stores prior motion features. Prior motion characteristics It is the prior motion feature of the previous frame image. Prior motion characteristics See the following steps for how to obtain it.
[0180] 4. Current image features F t Short-term time-domain reference characteristics Prior motion characteristics The input is fed to the motion encoder, which then uses the current image features F to... t Short-term time-domain reference characteristics Prior motion characteristics Motion information is encoded to obtain unquantized latent motion features m t .
[0181] 5. Regarding the potential characteristics of motion m t Quantization (i.e., Q-operation) is performed to obtain motion information features.
[0182] 6. The Motion Latent Buffer stores time-domain motion prior information. Prior information of motion in the time domain It is the motion information feature of the previous frame. Motion information features See the following steps for how to obtain it.
[0183] Motion potential characteristics m t and time-domain motion prior information The input is fed into the Hyper Enc (hyper-encoder), which then obtains the motion latent features m. t and time-domain motion prior information Then, the potential motion features m t and time-domain motion prior information Encoding is performed to obtain the motion side information bitstream. After obtaining the motion side information bitstream, the encoding end can also transmit the motion side information bitstream to the decoding end. The processing procedure at the decoding end is described in subsequent steps.
[0184] 7. Hyper Dec (Hyper Decoder) can acquire the motion side information bitstream and decode it to obtain the motion hyper prior information.
[0185] 8. Hyper Prior and Temporal Motion Prior (i.e., Temporal Prior) is input into the Motion Entropy Model, which is then processed by the Motion Entropy Model based on the motion prior information and the temporal motion prior information. Generate motion information features The probability distribution is used to guide encoding and decoding.
[0186] 9. Based on motion information features The probability distribution of motion information features Encoding (i.e., AE operation) is performed to obtain the motion information bitstream; there are no restrictions on this encoding process. After obtaining the motion information bitstream, the encoding end can also transmit the motion information bitstream to the decoding end. The processing procedure for the motion information bitstream at the decoding end is described in subsequent steps.
[0187] 10. Based on motion information features The probability distribution is used to decode the motion information bitstream (i.e., AD operation) to obtain motion information features. Motion information features Inverse quantization can also be performed to obtain the latent motion features m. t .
[0188] 11. After obtaining motion information features (That is, decoding the motion information bitstream to obtain motion information features) After that, motion information features Stored in the Motion Latent Buffer as temporal motion prior information for the next frame.
[0189] 12. Motion Information Characteristics The input is fed to the motion decoder, which then uses the motion information features... Decoding motion information is performed, without restrictions on the structure and decoding method of the motion decoder, to obtain optical flow characteristics. and prior motion characteristics Prior motion characteristics It is stored in the information buffer (Decoded Feature & Picture Buffer), and this prior motion feature It can be used as a priori motion feature for the next frame image.
[0190] 13. Optical flow characteristics and short-term time-domain reference characteristics The input is fed into the Temporal Context Mining module. The Temporal Context Mining module is based on optical flow features. and short-term time-domain reference characteristics Temporal residual mining is performed to obtain predicted features. There are no restrictions on the temporal residual mining process; the predicted features are... Used for residual conditional coding.
[0191] 14. Current image features F t and predictive features The input is fed to the contextual encoder, which then uses the current image features F to... t and predictive features Perform residual encoding to obtain unquantized residual information features y t .
[0192] 15. Features of unquantified residual information y t Quantization (i.e., Q-operation) is performed to obtain the latent features of the residuals.
[0193] 16. Predictive Features The data is fed into the Spatial Prior Encoder, which then uses the predicted features... The process is performed without any restrictions on the method used, resulting in spatial prior information.
[0194] 17. The Contextual Latent Buffer stores residual prior information. Residual prior information It is the residual latent feature of the previous frame. Residual latent characteristics See the following steps for how to obtain it.
[0195] The residual information feature y t Spatial Prior Information and Residual Prior Information The input is fed into the Hyper Enc (hyper-encoder). The hyper-encoder processes the residual information features y. t Spatial Prior Information and Residual Prior Information Encode the data to obtain the residual side information bitstream. After obtaining the residual side information bitstream, the encoder can also send it to the decoder. The decoding process is described in subsequent steps.
[0196] 18. Hyper Dec (Hyper Decoder) can obtain the residual edge information bitstream. After obtaining the residual edge information bitstream, it decodes the residual edge information bitstream to obtain the residual hyperprior information.
[0197] 19. Residual Hyper-Prior Information and Spatial Prior Information (i.e., predictive features) (After inputting into the spatial prior encoder, the residual prior information is obtained) (That is, Temporal Prior, output by ContextualLatent Buffer) is input to Contextual Entropy Model, which is then used by the residual entropy model based on residual prior information, spatial prior information, and residual prior information. Generate residual latent features The probability distribution.
[0198] 20. Based on residual latent characteristics The probability distribution of residual latent features Encoding (i.e., AE operation) is performed to obtain the residual information bitstream; this encoding process is not restricted. After obtaining the residual information bitstream, the encoding end can also transmit the residual information bitstream to the decoding end. The processing procedure for the residual information bitstream at the decoding end is described in subsequent steps.
[0199] 21. Based on residual latent characteristics The probability distribution is used to decode the residual information bitstream (i.e., AD operation) to obtain the residual latent features. Residual latent characteristics Inverse quantization can also be performed to obtain the residual information feature y. t .
[0200] 22. After obtaining the latent characteristics of the residuals (i.e., the residual latent features obtained by decoding the residual information bitstream) After that, the latent characteristics of the residuals It is stored in the Contextual Latent Buffer as residual prior information for the next frame.
[0201] 23. Latent characteristics of residuals The input is fed to the contextual decoder, which then uses the residual latent features to... The residual information is decoded without any restrictions on the structure of the residual decoder, and the residual features are obtained.
[0202] 24. Combine residual features and prediction features The input is fed into the frame generator, which then uses the residual features and prediction features to generate the frame. A transformation from the feature domain to the image domain is performed to obtain image reference features. (Image reference features) It can also be called time-domain reference feature It should be noted that, compared to Example 3, in obtaining the image reference features... Subsequently, the image reference features are not included. Stored in the information buffer (Decoded Feature & Picture Buffer).
[0203] 25. Based on image reference features Generate reconstructed image For example, using image reference features The input is fed into a deconvolutional network, which is based on image reference features. Perform an upsampling operation to obtain the current frame image x. t Corresponding reconstructed image
[0204] 26. Reconstruct the image Stored in the information buffer, i.e., the reconstructed image. As a temporal reference image for the next frame In this case, in step 2, the temporal reference information in the information buffer is the temporal reference image corresponding to the reconstructed image of the previous frame. Time-domain reference image The input is fed into the feature extractor to obtain short-term temporal reference features.
[0205] Or, after obtaining the reconstructed image Afterwards, the reconstructed image can also be processed. Feature extraction is performed to obtain temporal reference features, such as through a feature extractor on the reconstructed image. Feature extraction is performed to obtain temporal reference features, which are then stored in an information buffer, i.e., temporal reference features. Temporal reference features for the next frame In this case, in step 2, the temporal reference information in the information buffer is the temporal reference feature corresponding to the reconstructed image of the previous frame. Then, the temporal reference features corresponding to the reconstructed image from the previous frame can be used. The input is fed into the feature extractor to obtain short-term temporal reference features.
[0206] exist Figure 3A In the middle, it is to reconstruct the image Let's take storing it in the information buffer as an example.
[0207] This completes the encoding process, and the current frame image x can be obtained. t Corresponding reconstructed image
[0208] For example, in embodiment 5, if the current frame image x t Dividing the image into multiple blocks allows us to execute steps 1-26 for the current image block. For example, in step 1, the current image block is input to the feature extractor to obtain the current image feature F corresponding to the current image block. t In step 2, the temporal reference information obtained from the information buffer is the temporal reference information of the reference image block corresponding to the current image block in the previous frame. The position of the reference image block is the same as the position of the current image block, or the position of the reference image block is determined based on the position of the current image block and the motion vector (the motion vector between the current frame and the previous frame). In step 3, the prior motion features obtained from the information buffer... It is the prior motion feature of the reference image patch corresponding to the current image patch in the previous frame. In steps 4 and 5, the motion latent features m corresponding to the current image patch are obtained. t Motion information features corresponding to the current image patch In step 6, time-domain motion prior information It is the motion information feature of the reference image patch corresponding to the current image patch in the previous frame.
[0209] In steps 7-10, the motion information of the current image block is encoded and decoded to obtain the motion latent feature m corresponding to the current image block. t In step 11, the motion information features corresponding to the current image patch are... Stored in the Motion Latent Buffer. In step 12, the optical flow features corresponding to the current image patch are obtained. Prior motion features corresponding to the current image patch The prior motion features corresponding to the current image patch It is stored in the information buffer.
[0210] In step 13, the predicted features corresponding to the current image patch are obtained. In step 14, the residual information feature y corresponding to the current image patch is obtained. t In step 15, the residual latent features corresponding to the current image patch are obtained. In steps 16-21, the residual information of the current image block is encoded and decoded to obtain the residual latent features corresponding to the current image block. In step 22, the residual latent features corresponding to the current image patch are... Stored in the ContextualLatent Buffer. Steps 23-24 involve obtaining the image reference features corresponding to the current image patch. In step 25, the reconstructed image patch corresponding to the current image patch is obtained. In step 26, the reconstructed image patch or temporal reference feature (i.e., the reconstructed image) corresponding to the current image patch is used. The temporal reference features are extracted and stored in the information buffer.
[0211] For example, the current frame image x t After dividing the image into multiple image blocks, then processing the current frame image x... t After performing the above processing on each image block, a reconstructed image block corresponding to each image block can be obtained. The reconstructed image blocks corresponding to all image blocks together form the reconstructed image corresponding to the current frame image. Obviously, when storing reconstructed image blocks in the information buffer, after storing the reconstructed image blocks corresponding to all image blocks, it is actually storing the reconstructed image corresponding to the current frame image, or storing the temporal reference features of the reconstructed image corresponding to the current frame image. That is, the information buffer stores the reconstructed image of the entire frame.
[0212] Example 6: For the processing procedures at the decoding end in Examples 1 and 2, please refer to... Figure 3B As shown, of course, Figure 3B This is just one example of the processing procedure at the decoding end, and no restrictions are imposed on the processing procedure at the decoding end.
[0213] 1. The decoding end obtains the motion edge information bitstream.
[0214] 2. After obtaining the motion side information bitstream, the Hyper Dec (hyper prior) can decode the motion side information bitstream to obtain the motion hyper prior information.
[0215] 3. Hyper Prior and Temporal Motion Prior (i.e., Temporal Prior) is input into the Motion Entropy Model, which is then processed by the Motion Entropy Model based on the motion prior information and the temporal motion prior information. Generate motion information features The probability distribution is used to guide the decoding.
[0216] For example, regarding time-domain motion prior information The Motion Latent Buffer stores time-domain motion prior information. Prior information of motion in the time domain It is the motion information feature of the previous frame.
[0217] 4. The decoding end obtains the motion information bitstream.
[0218] 5. Based on motion information features The probability distribution is used to decode the motion information bitstream (i.e., AD operation) to obtain motion information features. Motion information features Inverse quantization can also be performed to obtain the latent motion features m. t .
[0219] 6. Obtaining motion information features (That is, decoding the motion information bitstream to obtain motion information features) After that, motion information features Stored in the Motion Latent Buffer as temporal motion prior information for the next frame.
[0220] 7. Motion Information Characteristics The input is fed to the motion decoder, which then uses the motion information features... Decoding motion information is performed, without restrictions on the structure and decoding method of the motion decoder, to obtain optical flow characteristics.
[0221] 8. Optical flow characteristics and short-term time-domain reference characteristics The input is fed into the Temporal Context Mining module. The Temporal Context Mining module is based on optical flow features. and short-term time-domain reference characteristics Temporal residual mining is performed to obtain predicted features. There are no restrictions on the temporal residual mining process; the predicted features are... Used for residual conditional decoding.
[0222] For example, the information buffer (i.e., the Decoded Feature & Picture Buffer) stores time-domain reference information. This time-domain reference information is either the time-domain reference image corresponding to the reconstructed image of the previous frame, or it is the time-domain reference feature corresponding to the reconstructed image of the previous frame. This time-domain reference information enters the feature extractor to obtain short-term time-domain reference features. For example, a feature extractor is used to extract image features from temporal reference information. Therefore, after inputting temporal reference information into the feature extractor, short-term temporal reference features can be obtained.
[0223] 9. The decoding end obtains the residual edge information bitstream.
[0224] 10. After obtaining the residual side information bitstream, the Hyper Dec (hyper prior) decodes the residual side information bitstream to obtain the residual hyper prior information.
[0225] 11. Predictive Features The data is fed into the Spatial Prior Encoder, which then uses the predicted features... The process is performed without any restrictions on the method used, resulting in spatial prior information.
[0226] 12. The Contextual Latent Buffer stores residual prior information. Residual prior information It is the residual latent feature of the previous frame. Residual prior information As a Temporal Prior.
[0227] 13. Hyper Prior, Spatial Prior, and Residual Prior Information The Temporal Prior is input into the Contextual Entropy Model, from which residual prior information, spatial prior information, and residual prior information are obtained. Subsequently, based on residual prior information, spatial prior information, and residual prior information... Generate residual latent features probability distribution, residual latent characteristics The probability distribution is used to guide the decoding.
[0228] 14. The decoding end obtains the residual information bitstream.
[0229] 15. Based on residual latent characteristics The probability distribution is used to decode the residual information bitstream (i.e., AD operation) to obtain the residual latent features. Residual latent characteristics Inverse quantization can also be performed to obtain the residual information feature y. t .
[0230] 16. After obtaining the latent characteristics of the residuals (i.e., the residual latent features obtained by decoding the residual information bitstream) After that, the latent characteristics of the residuals It is stored in the Contextual Latent Buffer as residual prior information for the next frame.
[0231] 17. Potential characteristics of residuals The input is fed to the contextual decoder, which then uses the residual latent features to... The residual information is decoded without any restrictions on the structure of the residual decoder, and the residual features are obtained.
[0232] 18. Combine residual features and prediction features The input is fed into the frame generator, which then uses the residual features and prediction features to generate the frame. A transformation from the feature domain to the image domain is performed to obtain image reference features. (Image reference features) It can also be called time-domain reference feature It should be noted that, compared to Example 4, in obtaining the image reference features... Subsequently, the image reference features are not included. Stored in the information buffer (Decoded Feature & Picture Buffer).
[0233] 19. Based on image reference features Generate reconstructed image For example, using image reference features The input is fed into a deconvolutional network, which is based on image reference features. Perform an upsampling operation to obtain the current frame image x. t Corresponding reconstructed image
[0234] 20. Reconstruct the image Stored in the information buffer, i.e., the reconstructed image. As a temporal reference image for the next frame In this case, the temporal reference information in the information buffer is the temporal reference image corresponding to the reconstructed image of the previous frame. Time-domain reference image The input is fed into the feature extractor to obtain short-term temporal reference features.
[0235] Or, after obtaining the reconstructed image Afterwards, the reconstructed image can also be processed. Feature extraction is performed to obtain temporal reference features, such as through a feature extractor on the reconstructed image. Feature extraction is performed to obtain temporal reference features, which are then stored in an information buffer, i.e., temporal reference features. Temporal reference features for the next frame In this case, the temporal reference information in the information buffer is the temporal reference feature corresponding to the reconstructed image of the previous frame. Then, the temporal reference features corresponding to the reconstructed image from the previous frame can be used. The input is fed into the feature extractor to obtain short-term temporal reference features.
[0236] exist Figure 3B In the middle, it is to reconstruct the image Let's take storing it in the information buffer as an example.
[0237] This completes the decoding process, and the current frame image x can be obtained. t Corresponding reconstructed image
[0238] For example, in embodiment 6, if the current frame image x t Dividing the image into multiple blocks allows us to execute steps 1-20 for the current image block. For example, steps 1-5 involve decoding the motion information of the current image block to ultimately obtain the motion latent feature m corresponding to that block. t In step 6, the motion information features corresponding to the current image patch are... Stored in the Motion Latent Buffer. In step 7, the optical flow features corresponding to the current image patch are obtained.
[0239] In step 8, the predicted features corresponding to the current image patch are obtained. First temporal reference feature obtained from the information buffer It is the first temporal reference feature of the reference image block corresponding to the current image patch in the previous frame. The position of the reference image block is the same as the position of the current image block, or the position of the reference image block is determined based on the position of the current image block and the motion vector (the motion vector between the current frame image and the previous frame image), and there are no restrictions on the reference image block.
[0240] In steps 9-15, the residual information of the current image patch is decoded to obtain the residual latent features corresponding to the current image patch. In step 16, the residual latent features corresponding to the current image patch are... Stored in the Contextual Latent Buffer. In steps 17-18, the image reference features corresponding to the current image patch are obtained. In step 19, the reconstructed image patch corresponding to the current image patch is obtained. In step 20, the reconstructed image patch or temporal reference feature corresponding to the current image patch (i.e., the reconstructed image) is used. The temporal reference features are extracted and stored in the information buffer.
[0241] For example, the current frame image x t After dividing the image into multiple image blocks, then processing the current frame image x... t After performing the above processing on each image block, a reconstructed image block corresponding to each image block can be obtained. The reconstructed image blocks corresponding to all image blocks together form the reconstructed image corresponding to the current frame image. Obviously, when storing reconstructed image blocks in the information buffer, after storing the reconstructed image blocks corresponding to all image blocks, it is actually storing the reconstructed image corresponding to the current frame image, or storing the temporal reference features of the reconstructed image corresponding to the current frame image. That is, the information buffer stores the reconstructed image of the entire frame.
[0242] For example, in embodiments 5 and 6, the reference information stored in the information buffer (Decoded Feature & Picture Buffer) is the reconstructed image. Or reconstruct the image The corresponding temporal reference features, and the reconstructed image Temporal reference features are reference information specific to the current frame and do not have cumulative properties. Even as the number of frames accumulates, the reconstructed image... Or temporal reference features will not accumulate, meaning the reconstructed image... The accuracy of the time-domain reference features will not decrease.
[0243] Example 7: For the processing procedures at the encoding end in Examples 1 and 2, please refer to... Figure 3C As shown, of course, Figure 3C This is just one example of the processing procedure at the encoding end, and no restrictions are imposed on this processing procedure.
[0244] 1. The encoding end obtains the current frame image x t (The current frame image can be the original image, i.e., the input image.) Then, the current frame image x... t The input is fed into the feature extractor to obtain the current image features F. t .
[0245] 2. The information buffer (i.e., the Decoded Feature & Picture Buffer, also known as the decoded image / feature buffer) stores time-domain reference information, which can be the image reference features of the previous frame. (i.e., first time-domain reference feature) That is, the image reference features in Example 3 Alternatively, the temporal reference information is the temporal reference image corresponding to the previous reconstructed image (i.e., the temporal reference image in Example 5), or the temporal reference information is the temporal reference feature corresponding to the previous reconstructed image (obtained after processing the reconstructed image, i.e., the temporal reference feature in Example 5).
[0246] For example, temporal reference information is fed into the feature extractor to obtain short-term temporal reference features. For example, a feature extractor is used to extract image features from temporal reference information. Therefore, after inputting temporal reference information into the feature extractor, short-term temporal reference features can be obtained.
[0247] 3. The information buffer (Decoded Feature & Picture Buffer) stores prior motion features. Prior motion characteristics It is the prior motion feature of the previous frame image. Prior motion characteristics See the following steps for how to obtain it.
[0248] 4. Current image features F t Short-term time-domain reference characteristics Prior motion characteristics The input is fed to the motion encoder, which then uses the current image features F to... t Short-term time-domain reference characteristics Prior motion characteristics Motion information is encoded to obtain unquantized latent motion features m t .
[0249] 5. Regarding the potential characteristics of motion m t Quantization (i.e., Q-operation) is performed to obtain motion information features.
[0250] 6. The Motion Latent Buffer stores time-domain motion prior information. Prior information of motion in the time domain It is the motion information feature of the previous frame. Motion information features See the following steps for how to obtain it.
[0251] Motion potential characteristics m t and time-domain motion prior information The input is fed into the Hyper Enc (hyper-encoder), which then obtains the motion latent features m. t and time-domain motion prior information Then, the potential motion features m t and time-domain motion prior information Encoding is performed to obtain the motion side information bitstream. After obtaining the motion side information bitstream, the encoding end can also transmit the motion side information bitstream to the decoding end. The processing procedure at the decoding end is described in subsequent steps.
[0252] 7. Hyper Dec (Hyper Decoder) can acquire the motion side information bitstream and decode it to obtain the motion hyper prior information.
[0253] 8. Obtain long-term temporal update features and long-term time-domain update features The input is fed into the Long Prior Refinement, which updates the features in the long-term time domain. The process yields motion spatial prior information, which is used to guide the encoding of motion information.
[0254] For example, in order to obtain long-term time-domain update features This can be achieved using the following sub-steps:
[0255] Sub-step 1: For the first frame image within the image group, based on short-term temporal reference features... Determine long-term temporal update characteristics Such as short-term time-domain reference features It can be used as a long-term time-domain update feature
[0256] For example, an image using inter-frame coding can be divided into multiple image groups, each comprising multiple frames. For instance, a GOP sequence might include I-frames, P-frames, ..., P-frames. These P-frames can be inter-coded and divided into multiple image groups. For example, the first image group might include 3 P-frames (3 P-frames and 1 I-frame make up 4 frames), the second image group might include 4 P-frames, the third image group might include 4 P-frames, and so on, until the last P-frame of the GOP sequence. Then, the process is repeated for the next GOP sequence.
[0257] If the current frame is the first frame in the image group (i.e., the first P-frame in the image group), then the current frame can be...
[0258] Like the corresponding short-term time-domain reference features As the long-term temporal update feature corresponding to the current frame image
[0259] For example, short-term time-domain reference features The input is fed to the feature updater (Long Prior Init / Update), which then updates the short-term temporal reference features. As a long-term time-domain update feature No update is required.
[0260] Sub-step 2: Align the long-term temporal update features with optical flow features to obtain long-term temporal aligned features.
[0261] For example, when obtaining long-term time-domain update features Subsequently, long-term time-domain update features will be implemented. The input is fed into the Long Prior Aligner, and the optical flow features are... Input to the long-term prior aligner, optical flow features The acquisition process is described in subsequent steps, optical flow characteristics. It can guide long-term temporal update features Perform long-term prior alignment.
[0262] Long-term prior aligners can employ optical flow characteristics. Long-term time-domain update features Perform alignment operations to obtain long-term temporal alignment features. Using optical flow characteristics Long-term time-domain update features When performing alignment operations, you can use the warp module or other networks; there are no restrictions on which one to use.
[0263] Sub-step 3: Input the long-term temporal alignment features into the configured long-term buffer, and the long-term buffer updates the long-term temporal alignment features to the long-term temporal reference features of the next frame. The long-term temporal reference features participate in subsequent calculations.
[0264] For example, in obtaining long-term temporal alignment features Then, long-term temporal alignment features can be... The input is fed into the long prior buffer, which updates the long-term temporal reference features. Long-term temporal alignment features As a long-term time-domain reference feature Furthermore, the long-term buffer stores long-term time-domain reference features.
[0265] For example, for long-term time-domain reference features stored in a long-term buffer It can represent prior information from multiple preceding frames of the current frame image, through long-term temporal reference features. Prior information can be incorporated into the decoding of the current frame image.
[0266] Sub-step 4: For non-first frame images within the image group (i.e., the current frame image is a non-first frame image within the image group), determine the long-term temporal update features based on the short-term temporal reference features and the long-term temporal reference features in the long-term buffer.
[0267] For example, if the current frame image is not the first frame image in the image group (i.e., the current frame image is the 2nd, 3rd, 4th P-frame, etc. in the image group), then the short-term temporal reference features corresponding to the current frame image can be used. and long-term temporal reference features in long-term buffers Determine the long-term temporal update features corresponding to the current frame image
[0268] For example, short-term time-domain reference features The input is given to the feature updater (Long Prior Init / Update), which uses long-term temporal reference features from the long prior buffer. The input is fed into the feature updater. The feature updater updates the short-term temporal reference features. Long-term time-domain reference characteristics Perform an update operation to obtain long-term time-domain update features.
[0269] For example, it can be used for short-term time-domain reference features. and long-term time-domain reference characteristics Perform channel-level convolution operations on the convolutional features to obtain long-term temporal update features.
[0270] For example, the long-term time-domain update features can be determined using the following formula: This represents long-term time-domain update characteristics. Represents long-term time-domain reference characteristics. denoted as short-term time-domain reference feature, 'a' can represent the weight coefficient of long-term time-domain reference feature, and 1-a represents the weight coefficient of short-term time-domain reference feature.
[0271] If the current frame image is the 1st or 2nd P-frame in the image group, then 'a' can be set close to 1 to increase the weight of long-term reference. If the current frame image is the 3rd, 4th, ... P-frame in the image group, then 'a' can be set between 0 and 1 to keep the weight of long-term and short-term reference fusion within a reasonable range. In this embodiment, the value of 'a' is not restricted.
[0272] Sub-step 5: Align the long-term temporal update features with optical flow features to obtain long-term temporal aligned features.
[0273] For example, when obtaining long-term time-domain update features Then, the long-term time-domain update features will be implemented. The input is fed into the Long Prior Aligner, and the optical flow features are... The input is fed into the long-term prior aligner. The long-term prior aligner can employ optical flow features. Long-term time-domain update features Perform alignment operations to obtain long-term temporal alignment features.
[0274] Sub-step 6: Input the long-term temporal alignment features into the configured long-term buffer, and the long-term buffer updates the long-term temporal alignment features to the long-term temporal reference features of the next frame. The long-term temporal reference features participate in subsequent calculations.
[0275] For example, in obtaining long-term temporal alignment features Then, long-term temporal alignment features can be... The input is fed into the long prior buffer, which updates the long-term temporal reference features. Long-term temporal alignment features As a long-term time-domain reference feature Furthermore, the long-term buffer stores long-term time-domain reference features.
[0276] In one possible implementation, for sub-step 1, based on short-term time-domain reference features... Determine long-term temporal update characteristics In some cases, the following methods can also be used, but these are just a few examples and will not be shown here.
[0277] For the first frame image within the first image group, short-term temporal reference features can be used. Identified as a long-term time-domain update feature That is, short-term time-domain reference characteristics Directly used as long-term temporal update features
[0278] For images that are not the first frame in the first image group (such as the first frame in the 2nd, 3rd, ... image groups), short-term temporal reference features can be used. Identified as a long-term time-domain update feature That is, short-term time-domain reference characteristics Directly used as long-term temporal update features Alternatively, the long-term temporal reference features in the long-term buffer can be used. (That is, the long-term temporal reference features of the preceding image group) are multiplied by the memory coefficient to obtain the long-term memory features. Based on the long-term memory features and the short-term temporal reference features... Determine the time-domain features after initialization, and determine the long-term time-domain update features based on the initialized time-domain features.
[0279] For example, see Figure 3D As shown, time-domain reference image After passing through a feature extractor, short-term temporal reference features can be obtained. (See step 2). Long-term time-domain reference characteristics Multiplying by the memory coefficient α yields the long-term memory characteristics. The memory coefficient α can be an empirical value, and its value is not restricted. Long-term memory characteristics and short-term temporal reference characteristics can be combined. The input is fed into the Fusion Initialization Network, which processes the long-term memory features and short-term temporal reference features. The fusion process yields the initialized time-domain features. Based on this, the initialized time-domain features can be... Identified as a long-term time-domain update feature Thus, long-term time-domain update features are obtained.
[0280] In one possible implementation, the long-term prior adjuster updates features in the long-term time domain. When processing the data to obtain prior information about the motion space, features can be updated in the long-term time domain. Convolutional and / or activation operations are performed to obtain prior information in the motion space. For example, a long-term prior adjuster may include at least one convolutional layer and / or at least one activation layer, thus enabling the updating of features in the long-term temporal domain. Perform convolution and / or activation operations to obtain prior information about the motion space.
[0281] For example, a long-term prior adjuster consists of convolutional layers, activation layers, convolutional layers, activation layers, and convolutional layers in sequence, updating features in the long-term temporal domain sequentially. By performing convolution operations, activation operations, and more convolution operations, prior information about the motion space can be obtained. Of course, the above is just an example and is not a limitation.
[0282] 9. Hyper Prior, Spatial Prior, and Temporal Prior Information (i.e., Temporal Prior, output by the Motion Latent Buffer) is input to the Motion Entropy Model, which is based on the motion prior information, spatial prior information, and temporal prior information. Generate motion information features The probability distribution is used to guide encoding and decoding.
[0283] Compared to Example 5, the motion entropy model adds motion space prior information to the input, which is determined based on long-term time-domain update features, thereby introducing long-term time-domain features to participate in the calculation of probability distribution.
[0284] 10. Based on motion information features The probability distribution of motion information features Encoding (i.e., AE operation) is performed to obtain the motion information bitstream; there are no restrictions on this encoding process. After obtaining the motion information bitstream, the encoding end can also transmit the motion information bitstream to the decoding end. The processing procedure for the motion information bitstream at the decoding end is described in subsequent steps.
[0285] 11. Based on motion information features The probability distribution is used to decode the motion information bitstream (i.e., AD operation) to obtain motion information features. Motion information features Inverse quantization can also be performed to obtain the latent motion features m. t .
[0286] 12. After obtaining motion information features (That is, decoding the motion information bitstream to obtain motion information features) After that, motion information features Stored in the Motion Latent Buffer as temporal motion prior information for the next frame.
[0287] 13. Motion Information Characteristics The input is fed to the motion decoder, which then uses the motion information features... Decoding motion information is performed, without restrictions on the structure and decoding method of the motion decoder, to obtain optical flow characteristics. and prior motion characteristics Prior motion characteristics It is stored in the information buffer (Decoded Feature & Picture Buffer), and this prior motion feature It can be used as a priori motion feature for the next frame image.
[0288] 14. Optical flow characteristics Short-term time-domain reference features Long-term temporal alignment features The output (from the Long Prior Align) is fed into the Temporal Context Mining module. The Temporal Context Mining module is based on optical flow features. Short-term time-domain reference features Long-term temporal alignment features Temporal residual mining is performed to obtain predicted features. There are no restrictions on the temporal residual mining process; the predicted features are... Used for residual conditional coding.
[0289] For example, optical flow characteristics can be employed. Short-term time-domain reference features Alignment operations are performed to obtain short-term temporal alignment features. Then, the long-term temporal alignment features are... The short-term temporal aligned features are concatenated along the channel dimension to obtain concatenated mixed features. Convolutional operations and / or residual block operations are then performed on the concatenated mixed features to obtain the predicted features.
[0290] For example, see Figure 3E As shown, in the Temporal Context Mining module, long-term temporal alignment features... Compared with short-term time-domain reference characteristics Collaborative temporal information mining can be performed in the following ways.
[0291] First, short-term time-domain reference characteristics After alignment module (alignment module in) Figure 3E The code uses the Warp module, but other types of alignment modules can also be used (any module that can achieve the alignment function is acceptable). Optical flow characteristics. It also goes through an alignment module. The alignment module uses optical flow characteristics. Short-term time-domain reference features Alignment operations are performed to obtain short-term temporal domain alignment features.
[0292] Then, long-term time-domain update features After alignment module (alignment module in) Figure 3E The Warp module was used in the process, and the optical flow characteristics were analyzed. It also goes through an alignment module. The alignment module uses optical flow characteristics. Long-term time-domain update features Perform alignment operations to obtain long-term temporal alignment features. This process has already been implemented in the aforementioned long-term prior aligner.
[0293] Then, short-term temporal alignment features and long-term temporal alignment features Connecting at the channel dimension corresponds to this operation. Figure 3EThe C operation in the code yields the concatenated mixed features. These features are then passed through a convolutional layer (Conv2d) to adjust the number of channels. Finally, they pass through several residual blocks (ResBlocks) to output the predicted features.
[0294] 15. Current image features F t and predictive features The input is fed to the contextual encoder, which then uses the current image features F to... t and predictive features Perform residual encoding to obtain unquantized residual information features y t .
[0295] 16. Features of unquantified residual information y t Quantization (i.e., Q-operation) is performed to obtain the latent features of the residuals.
[0296] 17. Predictive Features The input is fed to the Spatial Prior Encoder, and the long-term temporal aligned features (The output of the Long Prior Align) is fed into the spatial prior encoder. The spatial prior encoder then uses the predicted features... Long-term temporal alignment features The information is processed to obtain spatial prior information.
[0297] For example, in based on predictive features Long-term temporal alignment features When performing spatial prior processing to obtain spatial prior information, the spatial prior encoder can predict features. Long-term temporal alignment features The spatial prior encoder connects along the channel dimension and performs convolution and / or activation operations on the connected features to obtain spatial prior information. For example, the spatial prior encoder may consist of convolutional layers, activation layers, and convolutional layers in sequence. The spatial prior encoder may then perform downsampling convolution, activation, and downsampling convolution operations on the connected features in sequence to obtain spatial prior information.
[0298] See Figure 3F As shown, predicted features Long-term temporal alignment features In channel dimension connection, i.e. Figure 3FThe C operation is performed in sequence. Then, downsampling convolution (e.g., using a Conv2d convolutional layer with 3 kernels and a stride of 2, and downsampling is applied to the concatenated features; this is just an example, and the structure of this convolutional layer is not restricted) is performed, followed by activation, and then downsampling convolution (e.g., using a Conv2d convolutional layer with 3 kernels and a stride of 2, and downsampling is applied; this is just an example, and the structure of this convolutional layer is not restricted) to obtain spatial prior information.
[0299] 18. The Contextual Latent Buffer stores residual prior information. Residual prior information It is the residual latent feature of the previous frame. Residual latent characteristics See the following steps for how to obtain it.
[0300] The residual information feature y t Spatial Prior Information and Residual Prior Information The input is fed into the Hyper Enc (hyper-encoder). The hyper-encoder processes the residual information features y. t Spatial Prior Information and Residual Prior Information Encode the data to obtain the residual side information bitstream. After obtaining the residual side information bitstream, the encoder can also send it to the decoder. The decoding process is described in subsequent steps.
[0301] 19. Hyper Dec (Hyper Decoder) can obtain the residual edge information bitstream. After obtaining the residual edge information bitstream, it decodes the residual edge information bitstream to obtain the residual hyperprior information.
[0302] 20. Residual Hyper-Prior Information and Spatial Prior Information (i.e., prediction features) (After inputting into the spatial prior encoder, the residual prior information is obtained) (That is, Temporal Prior, output by ContextualLatent Buffer) is input to Contextual Entropy Model, which is then used by the residual entropy model based on residual prior information, spatial prior information, and residual prior information. Generate residual latent features The probability distribution.
[0303] 21. Based on residual latent characteristics The probability distribution of residual latent features Encoding (i.e., AE operation) is performed to obtain the residual information bitstream; this encoding process is not restricted. After obtaining the residual information bitstream, the encoding end can also transmit the residual information bitstream to the decoding end. The processing procedure for the residual information bitstream at the decoding end is described in subsequent steps.
[0304] 22. Based on residual latent characteristics The probability distribution is used to decode the residual information bitstream (i.e., AD operation) to obtain the residual latent features. Residual latent characteristics Inverse quantization can also be performed to obtain the residual information feature y. t .
[0305] 23. After obtaining the latent characteristics of the residuals (i.e., the residual latent features obtained by decoding the residual information bitstream) After that, the latent characteristics of the residuals It is stored in the Contextual Latent Buffer as residual prior information for the next frame.
[0306] 24. Latent characteristics of residuals The input is fed to the contextual decoder, which then uses the residual latent features to... The residual information is decoded without any restrictions on the structure of the residual decoder, and the residual features are obtained.
[0307] 25. Combine residual features and prediction features The input is fed into the frame generator, which then uses the residual features and prediction features to generate the frame. A transformation from the feature domain to the image domain is performed to obtain image reference features.
[0308] 26. Based on image reference features Generate reconstructed image For example, using image reference features The input is fed into a deconvolutional network, which is based on image reference features. Perform an upsampling operation to obtain the current frame image x. t Corresponding reconstructed image
[0309] 27. After obtaining the image reference features Then, image reference features can be used. Stored in the information buffer (Decoded Feature & Picture Buffer), i.e., image reference features. As a reference for the next frame. In this case, in step 2, the temporal reference information in the information buffer is the image reference feature corresponding to the reconstructed image of the previous frame. Image reference features The input is fed into the feature extractor to obtain short-term temporal reference features. or,
[0310] Reconstruct the image Stored in the information buffer, i.e., the reconstructed image. As a temporal reference image for the next frame In this case, in step 2, the temporal reference information in the information buffer is the temporal reference image corresponding to the reconstructed image of the previous frame. Time-domain reference image The input is fed into the feature extractor to obtain short-term temporal reference features.
[0311] Or, after obtaining the reconstructed image Afterwards, the reconstructed image can also be processed. Feature extraction is performed to obtain temporal reference features, such as through a feature extractor on the reconstructed image. Feature extraction is performed to obtain temporal reference features, which are then stored in an information buffer, i.e., temporal reference features. Temporal reference features for the next frame In this case, in step 2, the temporal reference information in the information buffer is the temporal reference feature corresponding to the reconstructed image of the previous frame. Then, the temporal reference features corresponding to the reconstructed image from the previous frame can be used. The input is fed into the feature extractor to obtain short-term temporal reference features.
[0312] exist Figure 3C In the middle, it is to reconstruct the image Let's take storing it in the information buffer as an example.
[0313] This completes the encoding process, and the current frame image x can be obtained. t Corresponding reconstructed image
[0314] For example, in embodiment 7, if the current frame image x t If the image is divided into multiple image blocks, then steps 1-27 can be executed for the current image block. That is, for the current frame image x... t The above processing is performed on each image block, and the current frame image x in Example 7...t It can be replaced with the current image patch, and all features are specific to the current image patch.
[0315] In one possible implementation, see Figure 3C As shown, with the Long Prior Buffer as the dividing line, the lower half is the encoding, decoding and transmission of motion information bitstream, and the upper half is the encoding, decoding and transmission of residual information.
[0316] First, the encoding and decoding of the motion information bitstream are described.
[0317] For encoding motion information, the feature extractor generates the current frame features F. t Meanwhile, the previous frame decoded image It will also generate short-term temporal reference features through a feature extractor. Short-term time-domain reference features Motion feature prior With the current frame feature F t They are fed together into a motion encoder, and after passing through the quantization module Q, motion information features are generated. Simultaneously, the unquantified motion information feature m t The motion side information bitstream is generated by the hyper-encoder, and then passed through the hyper-decoder to generate hyper-prior information. Hyper-prior information and temporal motion prior information are also included. and long-term reference information a priori Prior information from three different sources guides the Motion Entropy Model in generating motion information features. The probability distribution of motion information will guide the arithmetic encoder (AE) in processing motion information features. Encode the data to form a motion information bitstream.
[0318] For decoding motion information, the arithmetic decoder (AD) decodes the motion information bitstream to form motion information features. The optical stream is generated by the motion decoder. Decoding optical flow Short-term time-domain reference features Long-term time-domain reference characteristics Enter the Temporal Context Mining module to output predicted features.
[0319] Secondly, the encoding and decoding of the residual information bitstream are described.
[0320] Encoding of residual information. Predicted features. With the current frame feature F t They jointly enter the contextual encoder, and after entering the quantization module Q, residual information features are generated. Similar to the encoding of motion information, unquantized motion information features y t The residual side information bitstream is generated by the Hyper Encoder, and then the residual side information bitstream is processed by the Hyper Decoder to generate residual hyperprior information.
[0321] Meanwhile, long-term time-domain reference features With predictive features These are collectively fed into the Spatial Prior Encoder to generate the Spatial Prior. Finally, the residual hyper-prior information and temporal residual prior information are generated. Spatial prior, along with prior information from three different sources, are used to generate residual information features through a contextual entropy model. The probability distribution of the residual information will guide the arithmetic encoder AE in analyzing the residual information features. Encode the residual information bitstream.
[0322] For decoding residual information, the arithmetic decoder (AD) decodes the residual information bitstream to form residual information features. Residual information characteristics With predictive features The images are fed into the contextual decoder and frame generator, and after passing through these two stages, a decoded image is finally generated.
[0323] Example 8: For the processing procedures at the decoding end in Examples 1 and 2, please refer to... Figure 3G As shown, of course, Figure 3G This is just one example of the processing procedure at the decoding end, and no restrictions are imposed on the processing procedure at the decoding end.
[0324] 1. The decoding end obtains the motion edge information bitstream.
[0325] 2. After obtaining the motion side information bitstream, the Hyper Dec (hyper prior) can decode the motion side information bitstream to obtain the motion hyper prior information.
[0326] 3. Obtain long-term temporal update features and long-term time-domain update features The input is fed into the Long Prior Refinement, which updates the features in the long-term time domain. The process yields motion spatial prior information, which is used to guide the decoding of motion information.
[0327] For example, in order to obtain long-term time-domain update features This can be achieved using the following sub-steps:
[0328] Sub-step 1: For the first frame image within the image group, based on short-term temporal reference features... Determine long-term temporal update characteristics Such as short-term time-domain reference features It can be used as a long-term time-domain update feature
[0329] Sub-step 2: Align the long-term temporal update features with optical flow features to obtain long-term temporal aligned features.
[0330] Sub-step 3: Input the long-term temporal alignment features into the configured long-term buffer, and the long-term buffer updates the long-term temporal alignment features to the long-term temporal reference features of the next frame. The long-term temporal reference features participate in subsequent calculations.
[0331] Sub-step 4: For non-first frame images within the image group (i.e., the current frame image is a non-first frame image within the image group), determine the long-term temporal update features based on the short-term temporal reference features and the long-term temporal reference features in the long-term buffer.
[0332] Sub-step 5: Align the long-term temporal update features with optical flow features to obtain long-term temporal aligned features.
[0333] Sub-step 6: Input the long-term temporal alignment features into the configured long-term buffer, and the long-term buffer updates the long-term temporal alignment features to the long-term temporal reference features of the next frame. The long-term temporal reference features participate in subsequent calculations.
[0334] For example, step 3 at the decoding end can be referred to step 8 in embodiment 7, and will not be repeated here.
[0335] The above process involves short-term time-domain reference features. Regarding short-term time-domain reference features The acquisition method involves storing temporal reference information in the information buffer (i.e., Decoded Feature & Picture Buffer). This temporal reference information can be the image reference features of the previous frame. Alternatively, the temporal reference information can be the temporal reference image corresponding to the reconstructed image of the previous frame, or it can be the temporal reference feature corresponding to the reconstructed image of the previous frame (obtained after processing the reconstructed image). The temporal reference information is fed into the feature extractor to obtain short-term temporal reference features.
[0336] 4. Hyper Prior, Spatial Prior, and Temporal Prior Information (i.e., Temporal Prior, output by the Motion Latent Buffer) is input to the Motion Entropy Model, which is based on the motion prior information, spatial prior information, and temporal prior information. Generate motion information features The probability distribution is used to guide the decoding.
[0337] For example, regarding time-domain motion prior information The Motion Latent Buffer stores time-domain motion prior information. Prior information of motion in the time domain It is the motion information feature of the previous frame.
[0338] Compared to Example 6, the motion entropy model adds motion space prior information to the input, which is determined based on long-term time-domain update features, thereby introducing long-term time-domain features to participate in the calculation of probability distribution.
[0339] 5. The decoding end obtains the motion information bitstream.
[0340] 6. Based on motion information features The probability distribution is used to decode the motion information bitstream (i.e., AD operation) to obtain motion information features. Motion information features Inverse quantization can also be performed to obtain the latent motion features m. t .
[0341] 7. Obtaining motion information features (That is, decoding the motion information bitstream to obtain motion information features) After that, motion information features Stored in the Motion Latent Buffer as temporal motion prior information for the next frame.
[0342] 8. Motion Information Characteristics The input is fed to the motion decoder, which then uses the motion information features... Decoding motion information is performed, without restrictions on the structure and decoding method of the motion decoder, to obtain optical flow characteristics.
[0343] 9. Optical flow characteristics Short-term time-domain reference features Long-term temporal alignment features The output (from the Long Prior Align) is fed into the Temporal Context Mining module. The Temporal Context Mining module is based on optical flow features. Short-term time-domain reference features Long-term temporal alignment features Temporal residual mining is performed to obtain predicted features. There are no restrictions on the temporal residual mining process; the predicted features are... Used for residual conditional decoding.
[0344] For example, optical flow characteristics can be employed. Short-term time-domain reference features Alignment operations are performed to obtain short-term temporal alignment features. Then, the long-term temporal alignment features are... The short-term temporal aligned features are concatenated along the channel dimension to obtain concatenated mixed features. Convolutional operations and / or residual block operations are then performed on the concatenated mixed features to obtain the predicted features.
[0345] 10. The decoding end obtains the residual edge information bitstream.
[0346] 11. After obtaining the residual side information bitstream, the Hyper Dec (hyper prior) decoder can decode the residual side information bitstream to obtain the residual hyper prior information.
[0347] 12. Predictive Features The input is fed to the Spatial Prior Encoder, and the long-term temporal aligned features (The output of the Long Prior Align) is fed into the spatial prior encoder. The spatial prior encoder then uses the predicted features... Long-term temporal alignment features The information is processed to obtain spatial prior information.
[0348] For example, a spatial prior encoder can predict features Long-term temporal alignment features Connect the channels along the channel dimension and perform convolution and / or activation operations on the connected features to obtain spatial prior information.
[0349] 13. The Contextual Latent Buffer stores residual prior information. Residual prior information It is the residual latent feature of the previous frame. Residual latent characteristics See the following steps for how to obtain it.
[0350] 14. Residual Hyper-Prior Information and Spatial Prior Information (i.e., prediction features) (After inputting into the spatial prior encoder, the residual prior information is obtained) (That is, Temporal Prior, output by ContextualLatent Buffer) is input to Contextual Entropy Model, which is then used by the residual entropy model based on residual prior information, spatial prior information, and residual prior information. Generate residual latent features The probability distribution.
[0351] 15. The decoding end obtains the residual information bitstream.
[0352] 16. Based on residual latent characteristics The probability distribution is used to decode the residual information bitstream (i.e., AD operation) to obtain the residual latent features. Residual latent characteristics Inverse quantization can also be performed to obtain the residual information feature y. t .
[0353] 17. After obtaining the latent characteristics of the residuals (i.e., the residual latent features obtained by decoding the residual information bitstream) After that, the latent characteristics of the residuals It is stored in the Contextual Latent Buffer as residual prior information for the next frame.
[0354] 18. Latent characteristics of residuals The input is fed to the contextual decoder, which then processes the data based on the residual...
[0355] Poor latent features The residual information is decoded without any restrictions on the structure of the residual decoder, and the residual features are obtained.
[0356] 19. Combine residual features and prediction features The input is fed into the frame generator, which then uses the residual features and prediction features to generate the frame. A transformation from the feature domain to the image domain is performed to obtain image reference features.
[0357] 20. Based on image reference features Generate reconstructed image For example, using image reference features The input is fed into a deconvolutional network, which is based on image reference features. Perform an upsampling operation to obtain the current frame image x. t Corresponding reconstructed image
[0358] 21. After obtaining the image reference features Then, image reference features can be used. Stored in the information buffer (Decoded Feature & Picture Buffer), i.e., image reference features. As a reference for the next frame. In this case, the temporal reference information in the information buffer is the image reference feature corresponding to the reconstructed image of the previous frame. Image reference features The input is fed into the feature extractor to obtain short-term temporal reference features. or,
[0359] Reconstruct the image Stored in the information buffer, i.e., the reconstructed image. As a temporal reference image for the next frame In this case, the temporal reference information in the information buffer is the temporal reference image corresponding to the reconstructed image of the previous frame. Time-domain reference image The input is fed into the feature extractor to obtain short-term temporal reference features.
[0360] Or, after obtaining the reconstructed image Afterwards, the reconstructed image can also be processed. Feature extraction is performed to obtain temporal reference features, such as through a feature extractor on the reconstructed image. Feature extraction is performed to obtain temporal reference features, which are then stored in an information buffer, i.e., temporal reference features. Temporal reference features for the next frame In this case, the temporal reference information in the information buffer is the temporal reference feature corresponding to the reconstructed image of the previous frame. Then, the temporal reference features corresponding to the reconstructed image from the previous frame can be used. The input is fed into the feature extractor to obtain short-term temporal reference features.
[0361] exist Figure 3G In the middle, it is to reconstruct the image Let's take storing it in the information buffer as an example.
[0362] This completes the decoding process, and the current frame image x can be obtained. t Corresponding reconstructed image
[0363] For example, in embodiment 8, if the current frame image x t Dividing the image into multiple image blocks allows the above steps to be performed on the current image block; that is, for the current frame image x... t The above processing is performed on each image block, and the current frame image x in Example 8 t It can be replaced with the current image patch, and all features are specific to the current image patch.
[0364] Example 9: Based on Examples 1-8, this application proposes a decoding method, including:
[0365] 1. When decoding an image group, after decoding the first frame of the image group, a feature extractor is used to initialize long-term temporal reference features from the decoded first frame. That is, storing long-term time-domain reference features
[0366] 2. When decoding images other than the first frame of this image group, a feature updater is used to update the long-term temporal reference features. Compared with short-term time-domain reference characteristics The feature is updated to obtain the long-term time-domain update characteristics.
[0367] 3. Long-term time-domain update characteristics The process involves a Long Prior Refinement to obtain a long prior, which serves as one of the sources of prior information. This long prior is then processed by a Contextual Entropy Model to obtain the probability distribution of motion information. This probability distribution is used to decode the motion information features.
[0368] 4. Motion Information Characteristics It can guide long-term temporal update features Perform long-prior alignment, such as for motion information features. The corresponding optical flow features guide the update, resulting in long-term temporal alignment features.
[0369] 5. Long-term temporal alignment features As one of the prior sources, the contextual entropy model is used to obtain the probability distribution of the residuals, and then the residual information features are decoded from the residual bitstream.
[0370] 6. Long-term temporal alignment features Short-term time-domain reference features With optical flow characteristics These features are then processed by the temporal information mining module to obtain predicted features.
[0371] 7. Predictive Features Features of residual information The decoded image is obtained through the contextual decoder.
[0372] Based on Embodiments 1-8, this application proposes an encoding method, including:
[0373] 1. When encoding a group of images, after encoding the first frame of the group, a feature extractor is used to initialize long-term temporal reference features from the decoded first frame. That is, storing long-term time-domain reference features
[0374] 2. When encoding images other than the first frame of this image group, a feature updater is used to update the long-term temporal reference features. Compared with short-term time-domain reference characteristics The feature is updated to obtain the long-term time-domain update characteristics.
[0375] 3. Use short-term time-domain reference features With the current frame feature F t Generate motion information features m t And long-term time-domain update features As a priori, it is used to guide the encoding of motion information features m t Obtain the motion information bitstream.
[0376] 4. Use motion information features Align long-term temporal update features Obtain long-term temporal alignment features
[0377] Example 10: In the decoding method, the process of decoding a group of images with respect to long-term temporal features is as follows:
[0378] x i ,i∈[1,2,…] represents the image to be decoded, and the image group is a single set of images.
[0379] Algorithm: Decoding algorithm for long-term temporal features in a group of images
[0380] Decoding process:
[0381] for x i In the current image group, do
[0382] if i = 1 then
[0383] Decoding from the bitstream
[0384] else
[0385] if i = 2 then
[0386] Initialize long-term time-domain features
[0387] else
[0388] Using short-term time-domain reference features Update long-term time-domain reference features Obtain long-term temporal update features
[0389] end
[0390] Extracting motion information features from the bitstream
[0391] Using motion information features Align long-term temporal update features Obtain long-term temporal alignment features
[0392] Long-term temporal alignment features Guided decoding of residual information
[0393] Long-term temporal alignment features Guided generation of predictive features
[0394] The residual information and predicted features are used to obtain the decoded image through the frame generation module.
[0395] Example 11: In the decoding method, if the first frame of the image group can be encoded as either an I-frame or a P-frame, then if it is an I-frame, the corresponding I-frame model is used. If it is a P-frame, the corresponding P-frame model is used. When using the corresponding P-frame model, long-term temporal reference features can be used.
[0396] For example, assuming the image group has a length of 4, the long-term temporal reference relationships are shown in Table 1:
[0397] Table 1
[0398]
[0399]
[0400] Example 12: In the decoding method, for the temporal context mining module, long-term temporal alignment features Compared with short-term time-domain reference characteristics Collaborative temporal information mining can be performed in the following ways.
[0401] See Figure 3E As shown, short-term time-domain reference features After the alignment module, which corresponds to the Warp operation in the diagram, long-term time-domain update features are also included. The module also undergoes an alignment module to generate long-term temporal alignment features.
[0402] Long-term temporal alignment features Compared with short-term time-domain reference characteristics The concatenation operation, corresponding to operation C in the diagram, involves connecting along the channel dimension. The concatenated mixed features are then passed through a convolutional layer (Conv2d) to adjust the number of channels. Finally, the features are processed through several residual blocks (ResBlocks) to output the predicted features.
[0403] Example 13: Long Prior Refinement for updating long-term temporal features This becomes a motion spatial prior, used to guide the encoding of motion information. For example, a long-term prior adjuster can update features over a long time domain. Perform convolution and / or activation operations to obtain prior information about the motion space.
[0404] See Figure 3H The diagram shows the structure of a long-term prior adjuster. The long-term prior adjuster can sequentially include a convolutional layer (e.g., a Conv2d convolutional layer with 3 kernels, a stride of 2, and downsampling; this is just an example), an activation layer, another convolutional layer (e.g., a Conv2d convolutional layer with 3 kernels, a stride of 2, and downsampling), another activation layer, and another convolutional layer (e.g., a Conv2d convolutional layer with 3 kernels, a stride of 2, and downsampling). The long-term prior adjuster updates features in the long-term temporal domain sequentially. Perform convolution operations, activation operations, convolution operations, activation operations, and convolution operations again to obtain motion spatial prior information.
[0405] Example 14: Spatial Prior Encoder Based on Predictive Features Long-term temporal alignment features When performing spatial prior processing to obtain spatial prior information, see [reference needed]. Figure 3F As shown, predicted features Long-term temporal alignment features In channel dimension connection, i.e. Figure 3F The C operation is performed in sequence. Then, downsampling convolution (e.g., using a Conv2d convolutional layer with 3 kernels and a stride of 2, and downsampling is applied to the concatenated features; this is just an example, and the structure of this convolutional layer is not restricted) is performed, followed by activation and downsampling convolution (e.g., using a Conv2d convolutional layer with 3 kernels and a stride of 2, and downsampling is applied; this is just an example, and the structure of this convolutional layer is not restricted) to obtain spatial prior information.
[0406] In summary, in a spatial prior encoder, the predicted features... Long-term temporal alignment features In the channel dimension connection, spatial prior information is generated by two downsampling convolutional layers for the connected features.
[0407] Example 15: The first image group can be in IPPP format, and the second image group and subsequent groups are in PPPP format. For the first frame (P-frame) within the first image group, the short-term temporal reference features can be... Directly used as long-term temporal update features For the first frame image that is not in the first image group (such as the first frame image in the 2nd, 3rd, ... image groups), the initialization logic of the first P-frame of the first image group can be used. That is, the short-term temporal reference features can be... Directly used as long-term temporal update features Alternatively, the following method can be used:
[0408] See Figure 3D As shown, time-domain reference image After passing through a feature extractor, short-term temporal reference features can be obtained. Long-term time-domain reference features Multiplying by the memory coefficient α yields the long-term memory characteristics, where α can be an empirical value. The long-term memory characteristics and short-term temporal reference characteristics can then be combined. The input is fed into the Fusion Initialization Network, which processes the long-term memory features and short-term temporal reference features. The fusion is performed to obtain the initialized time-domain characteristics.
[0409] Conquest Based on this, the initialized time-domain features can be... Identified as a long-term time-domain update feature
[0410] Example 16: For the processing procedures at the encoding end in Examples 1 and 2, please refer to... Figure 4A As shown, of course, Figure 4A This is just one example of the processing procedure at the encoding end, and no restrictions are imposed on this processing procedure.
[0411] 1. The encoding end obtains the current frame image x t (The current frame image can be the original image, i.e., the input image.) Then, the current frame image x... t The input is fed into the feature extractor to obtain the current image features F. t .
[0412] 2. The information buffer (i.e., Decoded Feature & Picture Buffer, also known as the decoded image / feature buffer) stores time-domain reference information. The time-domain reference information is the time-domain reference feature corresponding to the reconstructed image of the previous frame. The method of obtaining the time-domain reference feature is described in the following steps. Compared with Example 5, the method of obtaining the time-domain reference feature is different.
[0413] Temporal reference information is fed into the feature extractor to obtain short-term temporal reference features. For example, a feature extractor is used to extract image features from temporal reference information. Therefore, after inputting temporal reference information into the feature extractor, short-term temporal reference features can be obtained.
[0414] 3. The information buffer (Decoded Feature & Picture Buffer) stores prior motion features. Prior motion characteristics It is the prior motion feature of the previous frame image. Prior motion characteristics See the following steps for how to obtain it.
[0415] 4. Current image features F t Short-term time-domain reference characteristics Prior motion characteristics The input is fed to the motion encoder, which then uses the current image features F to... t Short-term time-domain reference characteristics Prior motion characteristics Motion information is encoded to obtain unquantized latent motion features m t .
[0416] 5. Regarding the potential characteristics of motion m t Quantization (i.e., Q-operation) is performed to obtain motion information features.
[0417] 6. The Motion Latent Buffer stores time-domain motion prior information. Prior information of motion in the time domain It is the motion information feature of the previous frame. Motion information features See the following steps for how to obtain it.
[0418] Motion potential characteristics m t and time-domain motion prior information The input is fed into the Hyper Enc (hyper-encoder), which then obtains the motion latent features m. t and time-domain motion prior information Then, the potential motion features m t and time-domain motion prior information Encoding is performed to obtain the motion side information bitstream. After obtaining the motion side information bitstream, the encoding end can also transmit the motion side information bitstream to the decoding end. The processing procedure at the decoding end is described in subsequent steps.
[0419] 7. Hyper Dec (Hyper Decoder) can acquire the motion side information bitstream and decode it to obtain the motion hyper prior information.
[0420] 8. Hyper Prior and Temporal Motion Prior (i.e., Temporal Prior) is input into the Motion Entropy Model, which is then processed by the Motion Entropy Model based on the motion prior information and the temporal motion prior information. Generate motion information features The probability distribution is used to guide encoding and decoding.
[0421] 9. Based on motion information features The probability distribution of motion information features Encoding (i.e., AE operation) is performed to obtain the motion information bitstream; there are no restrictions on this encoding process. After obtaining the motion information bitstream, the encoding end can also transmit the motion information bitstream to the decoding end. The processing procedure for the motion information bitstream at the decoding end is described in subsequent steps.
[0422] 10. Based on motion information features The probability distribution is used to decode the motion information bitstream (i.e., AD operation) to obtain motion information features. Motion information features Inverse quantization can also be performed to obtain the latent motion features m. t .
[0423] 11. After obtaining motion information features (That is, decoding the motion information bitstream to obtain motion information features) After that, motion information features Stored in the Motion Latent Buffer as temporal motion prior information for the next frame.
[0424] 12. Motion Information Characteristics The input is fed to the motion decoder, which then uses the motion information features... Decoding motion information is performed, without restrictions on the structure and decoding method of the motion decoder, to obtain optical flow characteristics. and prior motion characteristics Prior motion characteristics It is stored in the information buffer (Decoded Feature & Picture Buffer), and this prior motion feature It can be used as a priori motion feature for the next frame image.
[0425] 13. Optical flow characteristics and short-term time-domain reference characteristics The input is fed into the Temporal Context Mining module. The Temporal Context Mining module is based on optical flow features. and short-term time-domain reference characteristics Temporal residual mining is performed to obtain predicted features. There are no restrictions on the temporal residual mining process; the predicted features are... Used for residual conditional coding.
[0426] 14. Current image features F t and predictive features The input is fed to the contextual encoder, which then uses the current image features F to... t and predictive features Perform residual encoding to obtain unquantized residual information features y t .
[0427] 15. Features of unquantified residual information y t Quantization (i.e., Q-operation) is performed to obtain the latent features of the residuals.
[0428] 16. Predictive Features The data is fed into the Spatial Prior Encoder, which then uses the predicted features... The process is performed without any restrictions on the method used, resulting in spatial prior information.
[0429] 17. The Contextual Latent Buffer stores residual prior information. Residual prior information It is the residual latent feature of the previous frame. Residual latent characteristics See the following steps for how to obtain it.
[0430] The residual information feature y t Spatial Prior Information and Residual Prior Information The input is fed into the Hyper Enc (hyper-encoder). The hyper-encoder processes the residual information features y. tSpatial Prior Information and Residual Prior Information Encode the data to obtain the residual side information bitstream. After obtaining the residual side information bitstream, the encoder can also send it to the decoder. The decoding process is described in subsequent steps.
[0431] 18. Hyper Dec (Hyper Decoder) can obtain the residual edge information bitstream. After obtaining the residual edge information bitstream, it decodes the residual edge information bitstream to obtain the residual hyperprior information.
[0432] 19. Residual Hyper-Prior Information and Spatial Prior Information (i.e., predictive features) (After inputting into the spatial prior encoder, the residual prior information is obtained) (That is, Temporal Prior, output by ContextualLatent Buffer) is input to Contextual Entropy Model, which is then used by the residual entropy model based on residual prior information, spatial prior information, and residual prior information. Generate residual latent features The probability distribution.
[0433] 20. Based on residual latent characteristics The probability distribution of residual latent features Encoding (i.e., AE operation) is performed to obtain the residual information bitstream; this encoding process is not restricted. After obtaining the residual information bitstream, the encoding end can also transmit the residual information bitstream to the decoding end. The processing procedure for the residual information bitstream at the decoding end is described in subsequent steps.
[0434] 21. Based on residual latent characteristics The probability distribution is used to decode the residual information bitstream (i.e., AD operation) to obtain the residual latent features. Residual latent characteristics Inverse quantization can also be performed to obtain the residual information feature y. t .
[0435] 22. After obtaining the latent characteristics of the residuals (i.e., the residual latent features obtained by decoding the residual information bitstream) After that, the latent characteristics of the residuals It is stored in the Contextual Latent Buffer as residual prior information for the next frame.
[0436] 23. Latent characteristics of residuals The input is fed to the contextual decoder, which then uses the residual latent features to... The residual information is decoded without any restrictions on the structure of the residual decoder, and the residual features are obtained.
[0437] 24. Combine residual features and prediction features The input is fed into the frame generator, which then uses the residual features and prediction features to generate the frame. A transformation from the feature domain to the image domain is performed to obtain image reference features. (Image reference features) It can also be called time-domain reference feature It is important to note that after obtaining the image reference features... After that, there is no need to reference image features. Stored in the information buffer (Decoded Feature & Picture Buffer).
[0438] 25. Based on image reference features Generate reconstructed image For example, using image reference features The input is fed into a deconvolutional network, which is based on image reference features. Perform an upsampling operation to obtain the current frame image x. t Corresponding reconstructed image
[0439] 26. In combining residual features and prediction features When the input is given to the frame generator, the frame generator outputs image reference features in addition to the reference features. It can also output long-term time-domain reference features. Image reference features and long-term time-domain reference characteristics These are the output features of different network layers in the frame generation network. For example, image reference features. It is the output feature of the last network layer or other network layers, long-term temporal reference feature. It is the output feature of the penultimate network layer or other network layers.
[0440] Reconstructing the image using a feature extractor Feature extraction is performed to obtain short-term temporal reference features. Long-term temporal reference features are extracted using a feature extractor. Feature extraction is performed to obtain long-term time-domain fine features. Based on short-term time-domain reference features and long-term time-domain fine features The fusion is performed to obtain the fused temporal features. And based on fused temporal features Determine temporal reference features, such as fusing temporal features. As a temporal reference feature. After obtaining the temporal reference feature, it can be stored in the information buffer (i.e., Decoded Feature & Picture Buffer), that is, the temporal reference feature serves as the temporal reference information for the next frame.
[0441] For example, based on short-term time-domain reference features and long-term time-domain fine features The fusion is performed to obtain the fused temporal features. At that time, short-term time-domain reference features can be used. and long-term time-domain fine features The input is fed into the Init / Update network, which then processes the short-term temporal reference features. and long-term time-domain fine features To integrate.
[0442] For example, the following formula can be used to determine the fused temporal features: in, Indicates the fusion of temporal features, Represents long-term time-domain fine features, denoted as short-term time-domain reference feature, 'a' represents the weight coefficient of long-term time-domain fine feature, and 1-a represents the weight coefficient of short-term time-domain reference feature.
[0443] If the current frame image is the 1st or 2nd P-frame in the image group, then 'a' can be set close to 1 to increase the weight of long-term reference. If the current frame image is the 3rd, 4th, ... P-frame in the image group, then 'a' can be set between 0 and 1 to keep the weight of long-term and short-term reference fusion within a reasonable range. In this embodiment, the value of 'a' is not restricted.
[0444] This completes the encoding process, and the current frame image x can be obtained. t Corresponding reconstructed image
[0445] For example, in embodiment 16, if the current frame image x t Dividing the image into multiple image blocks allows the above steps to be performed on the current image block; that is, for the current frame image x... t The above processing is performed on each image block, and the current frame image x in Example 16t It can be replaced with the current image patch, and all features are specific to the current image patch.
[0446] Example 17: Regarding the processing at the decoding end in Examples 1 and 2, the following steps may be included:
[0447] 1. The decoding end obtains the motion edge information bitstream.
[0448] 2. After obtaining the motion side information bitstream, the Hyper Dec (hyper prior) can decode the motion side information bitstream to obtain the motion hyper prior information.
[0449] 3. Hyper Prior and Temporal Motion Prior (i.e., Temporal Prior) is input into the Motion Entropy Model, which is then processed by the Motion Entropy Model based on the motion prior information and the temporal motion prior information. Generate motion information features The probability distribution is used to guide the decoding.
[0450] For example, regarding time-domain motion prior information The Motion Latent Buffer stores time-domain motion prior information. Prior information of motion in the time domain It is the motion information feature of the previous frame.
[0451] 4. The decoding end obtains the motion information bitstream.
[0452] 5. Based on motion information features The probability distribution is used to decode the motion information bitstream (i.e., AD operation) to obtain motion information features. Motion information features Inverse quantization can also be performed to obtain the latent motion features m. t .
[0453] 6. Obtaining motion information features (That is, decoding the motion information bitstream to obtain motion information features) After that, motion information features Stored in the Motion Latent Buffer as temporal motion prior information for the next frame.
[0454] 7. Motion Information Characteristics The input is fed to the motion decoder, which then uses the motion information features... Decoding motion information is performed, without restrictions on the structure and decoding method of the motion decoder, to obtain optical flow characteristics.
[0455] 8. Optical flow characteristics and short-term time-domain reference characteristics The input is fed into the Temporal Context Mining module. The Temporal Context Mining module is based on optical flow features. and short-term time-domain reference characteristics Temporal residual mining is performed to obtain predicted features. There are no restrictions on the temporal residual mining process; the predicted features are... Used for residual conditional decoding.
[0456] For example, the information buffer (i.e., the Decoded Feature & Picture Buffer) stores time-domain reference information, which is the time-domain reference feature corresponding to the reconstructed image of the previous frame. The method for obtaining the time-domain reference feature is described in subsequent steps; however, the method differs from that in Example 6. The time-domain reference information enters the Feature Extractor to obtain short-term time-domain reference features. For example, a feature extractor is used to extract image features from temporal reference information. Therefore, short-term temporal reference features can be obtained through a feature extractor.
[0457] 9. The decoding end obtains the residual edge information bitstream.
[0458] 10. After obtaining the residual side information bitstream, the Hyper Dec (hyper prior) decodes the residual side information bitstream to obtain the residual hyper prior information.
[0459] 11. Predictive Features The data is fed into the Spatial Prior Encoder, which then uses the predicted features... The process is performed without any restrictions on the method used, resulting in spatial prior information.
[0460] 12. The Contextual Latent Buffer stores residual prior information. Residual prior information It is the residual latent feature of the previous frame. Residual prior information As a Temporal Prior.
[0461] 13. Hyper Prior, Spatial Prior, and Residual Prior Information The Temporal Prior is input into the Contextual Entropy Model, from which residual prior information, spatial prior information, and residual prior information are obtained. Subsequently, based on residual prior information, spatial prior information, and residual prior information... Generate residual latent features probability distribution, residual latent characteristics The probability distribution is used to guide the decoding.
[0462] 14. The decoding end obtains the residual information bitstream.
[0463] 15. Based on residual latent characteristics The probability distribution is used to decode the residual information bitstream (i.e., AD operation) to obtain the residual latent features. Residual latent characteristics Inverse quantization can also be performed to obtain the residual information feature y. t .
[0464] 16. After obtaining the latent characteristics of the residuals (i.e., the residual latent features obtained by decoding the residual information bitstream) After that, the latent characteristics of the residuals It is stored in the Contextual Latent Buffer as residual prior information for the next frame.
[0465] 17. Potential characteristics of residuals The input is fed to the contextual decoder, which then uses the residual latent features to... The residual information is decoded without any restrictions on the structure of the residual decoder, and the residual features are obtained.
[0466] 18. Combine residual features and prediction features The input is fed into the frame generator, which then uses the residual features and prediction features to generate the frame. A transformation from the feature domain to the image domain is performed to obtain image reference features. (Image reference features) It can also be called time-domain reference feature It is important to note that after obtaining the image reference features... After that, there is no need to reference image features. Stored in the information buffer (Decoded Feature & Picture Buffer).
[0467] 19. Based on image reference features Generate reconstructed image For example, using image reference features The input is fed into a deconvolutional network, which is based on image reference features. Perform an upsampling operation to obtain the current frame image x. t Corresponding reconstructed image
[0468] 20. When combining residual features and prediction features When the input is given to the frame generator, the frame generator outputs image reference features in addition to the reference features. It can also output long-term time-domain reference features. Image reference features and long-term time-domain reference characteristics These are the output features of different network layers in the frame generation network. The reconstructed image is then processed by a feature extractor. Feature extraction is performed to obtain short-term temporal reference features. Long-term temporal reference features are extracted using a feature extractor. Feature extraction is performed to obtain long-term time-domain fine features. Based on short-term time-domain reference features and long-term time-domain fine features The fusion is performed to obtain the fused temporal features. And based on fused temporal features Determine temporal reference features, such as fusing temporal features. As a temporal reference feature. After obtaining the temporal reference feature, it can be stored in the information buffer, that is, the temporal reference feature serves as the temporal reference information for the next frame.
[0469] This completes the decoding process, and the current frame image x can be obtained. t Corresponding reconstructed image
[0470] For example, in embodiment 17, if the current frame image x t Dividing the image into multiple image blocks allows the above steps to be performed on the current image block; that is, for the current frame image x... tThe above processing is performed on each image block, and the current frame image x in Example 17... t It can be replaced with the current image patch, and all features are specific to the current image patch.
[0471] Example 18: In the decoding method, long-term time-domain reference features are used in Examples 16 and 17. and short-term time-domain reference characteristics These can be stored in a feature buffer and undergo the alignment and update process together; that is, for long-term time-domain reference features... and short-term time-domain reference characteristics The process of merging and storing the data in an information buffer includes:
[0472] Use the previous frame Short-term temporal reference features are obtained after feature extraction.
[0473] Using long-term time-domain reference features Long-term temporal fine features are obtained through feature extractor.
[0474] Based on long-term time-domain fine features Compared with short-term time-domain reference characteristics Based on their position in the image group, corresponding updates or initializations are performed to obtain the fused temporal features. If the following formula is used:
[0475] If the current frame is the 1st or 2nd P-frame in the image group, 'a' can be set close to 1 to increase the weight of long-term references. If the current frame is the 3rd or 4th P-frame in the image group, 'a' can be set between 0 and 1 to keep the weight of long-term and short-term reference fusion within a reasonable range. Alternatively, if the current frame is the 1st P-frame in the image group, 'a' can be set close to 1 to increase the weight of long-term references. If the current frame is the 2nd, 3rd, or 4th P-frame in the image group, 'a' can be set between 0 and 1 to keep the weight of long-term and short-term reference fusion within a reasonable range.
[0476] For example, see Figure 4B As shown, before the Temporal Context Mining module, temporal features are fused. It can be used for entropy encoding and decoding of motion information, fusing temporal features after the temporal information mining module. It can be used for residual information entropy encoding and decoding and feature prediction. For details, please refer to Examples 16 and 17.
[0477] Example 19: For the processing procedures at the encoding end in Examples 1 and 2, please refer to... Figure 5A As shown, of course, Figure 5A This is just one example of the processing procedure at the encoding end, and no restrictions are imposed on this processing procedure.
[0478] 1. The encoding end obtains the current frame image x t (The current frame image can be the original image, i.e., the input image.) Then, the current frame image x... t The input is fed into the feature extractor to obtain the current image features F. t .
[0479] 2. The information buffer (i.e., the Decoded Feature & Picture Buffer) stores time-domain reference information, which can be the image reference features of the previous frame. (i.e., first time-domain reference feature) That is, the image reference features in Example 3 Alternatively, the temporal reference information can be the temporal reference image corresponding to the previous reconstructed image (i.e., the temporal reference image in Embodiment 5), or the temporal reference information can be the temporal reference feature corresponding to the previous reconstructed image (obtained after processing the reconstructed image, i.e., the temporal reference feature in Embodiment 5), or the temporal reference information can be the temporal reference feature corresponding to the previous reconstructed image (such as the fused temporal feature in Embodiment 16). ).
[0480] For example, temporal reference information is fed into the feature extractor to obtain short-term temporal reference features. For example, a feature extractor is used to extract image features from temporal reference information. Therefore, after inputting temporal reference information into the feature extractor, short-term temporal reference features can be obtained.
[0481] 3. The information buffer (Decoded Feature & Picture Buffer) stores prior motion features. Prior motion characteristics It is the prior motion feature of the previous frame image. Prior motion characteristics See the following steps for how to obtain it.
[0482] 4. Current image features F t Short-term time-domain reference characteristics Prior motion characteristics The input is fed to the motion encoder, which then uses the current image features F to... t Short-term time-domain reference characteristics Prior motion characteristics Motion information is encoded to obtain unquantized latent motion features m t .
[0483] 5. Regarding the potential characteristics of motion m t Quantization (i.e., Q-operation) is performed to obtain motion information features.
[0484] 6. The Motion Latent Buffer stores time-domain motion prior information. Prior information of motion in the time domain It is the motion information feature of the previous frame. Motion information features See the following steps for how to obtain it.
[0485] Motion potential characteristics m t and time-domain motion prior information The input is fed into the Hyper Enc (hyper-encoder), which then obtains the motion latent features m. t and time-domain motion prior information Then, the potential motion features m t and time-domain motion prior information Encoding is performed to obtain the motion side information bitstream. After obtaining the motion side information bitstream, the encoding end can also transmit the motion side information bitstream to the decoding end. The processing procedure at the decoding end is described in subsequent steps.
[0486] 7. Hyper Dec (Hyper Decoder) can acquire the motion side information bitstream and decode it to obtain the motion hyper prior information.
[0487] 8. Hyper Prior and Temporal Motion Prior (i.e., Temporal Prior) is input into the Motion Entropy Model, which is then processed by the Motion Entropy Model based on the motion prior information and the temporal motion prior information. Generate motion information features The probability distribution is used to guide encoding and decoding.
[0488] 9. Based on motion information features The probability distribution of motion information features Encoding (i.e., AE operation) is performed to obtain the motion information bitstream; there are no restrictions on this encoding process. After obtaining the motion information bitstream, the encoding end can also transmit the motion information bitstream to the decoding end. The processing procedure for the motion information bitstream at the decoding end is described in subsequent steps.
[0489] 10. Based on motion information features The probability distribution is used to decode the motion information bitstream (i.e., AD operation) to obtain motion information features. Motion information features Inverse quantization can also be performed to obtain the latent motion features m. t .
[0490] 11. After obtaining motion information features (That is, decoding the motion information bitstream to obtain motion information features) After that, motion information features Stored in the Motion Latent Buffer as temporal motion prior information for the next frame.
[0491] 12. Motion Information Characteristics The input is fed to the motion decoder, which then uses the motion information features... Decoding motion information is performed, without restrictions on the structure and decoding method of the motion decoder, to obtain optical flow characteristics. and prior motion characteristics Prior motion characteristics It is stored in the information buffer (Decoded Feature & Picture Buffer), and this prior motion feature It can be used as a priori motion feature for the next frame image.
[0492] For example, a motion decoder can include multiple network layers, where the motion decoder is based on motion information features. When decoding motion information, the features output by the penultimate network layer of the motion decoder can be used as prior motion features. Of course, the above is just an example; the output features of other network layers (such as the third-to-last network layer) can also be used as prior motion features. The features output by the last network layer of the motion decoder can be used as optical flow features. Of course, the above is just an example; the output features of other network layers can also be used as optical flow features.
[0493] 13. Obtain background optical flow and background features based on temporal reference information in the information buffer (i.e., Decoded Feature & Picture Buffer), and determine aligned background features based on the background optical flow and background features. For example, this temporal reference information can be image reference features from the previous frame. Alternatively, the temporal reference information can be the temporal reference image corresponding to the reconstructed image of the previous frame. Alternatively, the temporal reference information can be the temporal reference features corresponding to the reconstructed image of the previous frame (obtained after processing the reconstructed image), or the temporal reference information can be the temporal reference features corresponding to the reconstructed image of the previous frame (such as fused temporal features). For ease of description, the following will use a time-domain reference image. For example.
[0494] For example, to obtain the aligned background features, the following sub-steps can be used:
[0495] Sub-step 1: If the current frame image is the first frame image in the image group, then the temporal reference image... and optical flow characteristics The input is fed into the Background Prior Generator, which then generates the background prior based on the temporal reference image. and optical flow characteristics Processing is performed to obtain the background optical flow. and initialization background features
[0496] For example, an image using inter-frame coding can be divided into multiple image groups, each comprising multiple frames. For instance, a GOP sequence might include I-frames, P-frames, ..., P-frames. These P-frames can be inter-coded and divided into multiple image groups. For example, the first image group might include 3 P-frames (3 P-frames and 1 I-frame make up 4 frames), the second image group might include 4 P-frames, the third image group might include 4 P-frames, and so on, until the last P-frame of the GOP sequence. Then, the process is repeated for the next GOP sequence.
[0497] If the current frame is the first frame in the image group (i.e., the first P-frame in the image group), then the temporal reference image will be used. and optical flow characteristics Input is given to Background Prior Generator to obtain background optical flow. and initialization background features
[0498] For example, the structure of the Background Prior Generator can be seen in [reference needed]. Figure 5B As shown, in the time-domain reference image and optical flow characteristics After the background prior generator is fed into the background image, the feature extractor can process the temporal reference image. Feature extraction is performed to obtain the first feature. Then, the first feature and the optical flow feature are combined. The input is fed into a motion / mask generator network to obtain the mask features and background optical flow.
[0499] Mask generation networks are based on first features and optical flow features. Perform masking to obtain the mask feature mask and background optical flow. The output features of different network layers of the mask generation network can be used as mask features and background optical flow. For example, the output features of the last layer of the mask generation network or other network layers are used as the mask features, and the output features of the second-to-last layer of the mask generation network or other network layers are used as the background optical flow.
[0500] For example, initial background features can be generated based on the first feature and the mask feature. For example, the first feature can be multiplied by the mask feature to obtain the initialized background feature.
[0501] Sub-step 2: Based on background optical flow and background features Determine the alignment background features
[0502] For example, background optical flow and background features The input is given to the Background PriorAlign module, which then performs the alignment based on the background optical flow. Background features Perform alignment operations to obtain the aligned background features.
[0503] Sub-step 3: Obtaining the aligned background features Next, align the background features. Stored in the background prior buffer, and aligned with background features. Update the background update feature in the background information buffer. For example, the background information buffer will align with background features. As background update feature for the next frame
[0504] Sub-step 4: If the current frame image is not the first frame image within the image group (i.e., the 2nd, 3rd, or 4th P-frame within the image group, etc.), then the temporal reference image... Optical flow characteristics Background update features in the background information buffer The input is fed into the Background Prior Generator, which then generates the background prior based on the temporal reference image. Optical flow characteristics and background update features Processing is performed to obtain the background optical flow. and updated background features
[0505] For example, the structure of the Background Prior Generator can be seen in [reference needed]. Figure 5C As shown, in the time-domain reference image Optical flow characteristics and background update features After the background prior generator is fed into the background image, the feature extractor can process the temporal reference image. Feature extraction is performed to obtain the first feature.
[0506] Then, the first feature and optical flow feature are... and background update features The input is fed into a motion / mask generator network to obtain the mask features and background optical flow. For example, mask generation networks are based on first features and optical flow features. and background update features Perform masking to obtain the mask feature mask and background optical flow.
[0507] The output features of different network layers of the mask generation network can be used as mask features and background optical flow. For example, the output features of the last layer of the mask generation network or other network layers are used as the mask features, and the output features of the second-to-last layer of the mask generation network or other network layers are used as the background optical flow.
[0508] For example, a second feature can be generated based on a first feature and a mask feature. For instance, the first feature and the mask feature can be multiplied together to obtain the second feature. Then, the feature can be updated based on the second feature and the background. Generate updated background features For example, updating features by combining the second feature with the background. The input is fed to the Background Update network, which then updates the background based on the second feature and the background update feature. Obtain background features
[0509] Sub-step 5: Based on background optical flow and background features Determine the alignment background features
[0510] For example, background optical flow and background features The input is given to the BackgroundPrior Align module, which then performs the alignment based on the background optical flow. Background features Perform alignment operations to obtain the aligned background features.
[0511] Sub-step 6: Obtaining the aligned background features Next, align the background features. Stored in the background prior buffer, and aligned with background features. Update the background update feature in the background information buffer. For example, the background information buffer will align with background features. As background update feature for the next frame
[0512] 14. Optical flow characteristics Short-term time-domain reference features and alignment background features The data is input to the Temporal Context Mining module. The Temporal Context Mining module is based on optical flow features. Short-term time-domain reference features Align background features Temporal residual mining is performed to obtain predicted features. Predicted features Used for residual conditional coding.
[0513] For example, optical flow characteristics can be employed. Short-term time-domain reference features Alignment is performed to obtain short-term temporal aligned features. Then, the aligned background features are... The short-term temporal aligned features are concatenated along the channel dimension to obtain concatenated mixed features. Convolutional operations and / or residual block operations are then performed on the concatenated mixed features to obtain the predicted features.
[0514] For example, short-term time-domain reference features After alignment by a module (the alignment module can be a Warp module or other types of alignment modules, as long as it can achieve the alignment function), optical flow characteristics are obtained. It also goes through an alignment module. The alignment module uses optical flow characteristics. Short-term time-domain reference features Alignment operations are performed to obtain short-term temporal domain alignment features.
[0515] Then, short-term temporal alignment features and alignment background features The concatenation is performed along the channel dimension to obtain the concatenated mixed features. These mixed features are then passed through a convolutional layer (Conv2d) to adjust the number of channels. Finally, they pass through several residual blocks (ResBlocks) to output the predicted features. This embodiment does not limit the above process.
[0516] 15. Current image features F t and predictive features The input is fed to the contextual encoder, which then uses the current image features F to... t and predictive features Perform residual encoding to obtain unquantized residual information features y t .
[0517] 16. Features of unquantified residual information y t Quantization (i.e., Q-operation) is performed to obtain the latent features of the residuals.
[0518] 17. Predictive Features It is input to the Spatial Prior Encoder and aligned with the background features. The data is input to the spatial prior encoder. The spatial prior encoder then obtains the predicted features. and alignment background features Then, based on the predicted features and alignment background features The information is processed to obtain spatial prior information.
[0519] For example, in based on predictive features and alignment background features When performing spatial prior processing to obtain spatial prior information, the spatial prior encoder can predict features. and alignment background features Connect the channels along the channel dimension and perform convolution and / or activation operations on the connected features to obtain spatial prior information.
[0520] For example, a spatial prior encoder can consist of a convolutional layer, an activation layer, and another convolutional layer. The spatial prior encoder can perform downsampling convolution, activation, and downsampling convolution operations on the concatenated features in sequence to obtain spatial prior information.
[0521] For example, predicting features and alignment background features The concatenation is performed along the channel dimension. Then, the concatenated features are sequentially subjected to downsampling convolution operations (such as using a Conv2d convolutional layer with 3 kernels and a stride of 2, and downsampling; of course, this is just an example, and the structure of this convolutional layer is not restricted), activation operations, and downsampling convolution operations (such as using a Conv2d convolutional layer with 3 kernels and a stride of 2, and downsampling; of course, this is just an example, and the structure of this convolutional layer is not restricted) to obtain spatial prior information.
[0522] 18. The Contextual Latent Buffer stores residual prior information. Residual prior information It is the residual latent feature of the previous frame. Residual latent characteristics See the following steps for how to obtain it.
[0523] The residual information feature y t Spatial Prior Information and Residual Prior Information The input is fed into the Hyper Enc (hyper-encoder). The hyper-encoder processes the residual information features y. t Spatial Prior Information and Residual Prior Information Encode the data to obtain the residual side information bitstream. After obtaining the residual side information bitstream, the encoder can also send it to the decoder. The decoding process is described in subsequent steps.
[0524] 19. Hyper Dec (Hyper Decoder) can obtain the residual edge information bitstream. After obtaining the residual edge information bitstream, it decodes the residual edge information bitstream to obtain the residual hyperprior information.
[0525] 20. Residual Hyper-Prior Information and Spatial Prior Information (i.e., prediction features) (After inputting into the spatial prior encoder, the residual prior information is obtained) (That is, Temporal Prior, output by ContextualLatent Buffer) is input to Contextual Entropy Model, which is then used by the residual entropy model based on residual prior information, spatial prior information, and residual prior information. Generate residual latent features The probability distribution.
[0526] 21. Based on residual latent characteristics The probability distribution of residual latent features Encoding (i.e., AE operation) is performed to obtain the residual information bitstream; this encoding process is not restricted. After obtaining the residual information bitstream, the encoding end can also transmit the residual information bitstream to the decoding end. The processing procedure for the residual information bitstream at the decoding end is described in subsequent steps.
[0527] 22. Based on residual latent characteristics The probability distribution is used to decode the residual information bitstream (i.e., AD operation) to obtain the residual latent features. Residual latent characteristics Inverse quantization can also be performed to obtain the residual information feature y. t .
[0528] 23. After obtaining the latent characteristics of the residuals (i.e., the residual latent features obtained by decoding the residual information bitstream) After that, the latent characteristics of the residuals It is stored in the Contextual Latent Buffer as residual prior information for the next frame.
[0529] 24. Latent characteristics of residuals The input is fed to the contextual decoder, which then uses the residual latent features to... The residual information is decoded without any restrictions on the structure of the residual decoder, and the residual features are obtained.
[0530] 25. Combine residual features and prediction features The input is fed into the frame generator, which then uses the residual features and prediction features to generate the frame. A transformation from the feature domain to the image domain is performed to obtain image reference features.
[0531] 26. Based on image reference features Generate reconstructed image For example, using image reference features The input is fed into a deconvolutional network, which is based on image reference features. Perform an upsampling operation to obtain the current frame image x. t Corresponding reconstructed image
[0532] 27. After obtaining the image reference features Then, image reference features can be used. Stored in the information buffer (Decoded Feature & Picture Buffer), i.e., image reference features. As a reference for the next frame. Or,
[0533] Reconstruct the image Stored in the information buffer, i.e., the reconstructed image. As a temporal reference image for the next frame or,
[0534] After obtaining the reconstructed image Afterwards, the reconstructed image can also be processed. Feature extraction is performed to obtain temporal reference features, such as through a feature extractor on the reconstructed image. Feature extraction is performed to obtain temporal reference features, which are then stored in an information buffer, i.e., temporal reference features. Temporal reference features for the next frame or,
[0535] The frame generator also outputs image reference features. and long-term time-domain reference characteristics Reconstructed image using a feature extractor Feature extraction is performed to obtain short-term temporal reference features. Long-term temporal reference features are extracted using a feature extractor. Feature extraction is performed to obtain long-term time-domain fine features. Based on short-term time-domain reference features and long-term time-domain fine features The fusion is performed to obtain the fused temporal features. And will integrate temporal features Stored in an information buffer, i.e., fused temporal features Temporal reference features for the next frame
[0536] This completes the encoding process, and the current frame image x can be obtained. t Corresponding reconstructed image
[0537] For example, in embodiment 19, if the current frame image x t Dividing the image into multiple image blocks allows the above steps to be performed on the current image block; that is, for the current frame image x... t The above processing is performed on each image block, and the current frame image x in Example 19 t It can be replaced with the current image patch, and all features are specific to the current image patch.
[0538] Example 20: Regarding the processing at the decoding end in Examples 1 and 2, the following steps may be included:
[0539] 1. The decoding end obtains the motion edge information bitstream.
[0540] 2. After obtaining the motion side information bitstream, the Hyper Dec (hyper prior) can decode the motion side information bitstream to obtain the motion hyper prior information.
[0541] 3. Hyper Prior and Temporal Motion Prior (i.e., Temporal Prior) is input into the Motion Entropy Model, which is then processed by the Motion Entropy Model based on the motion prior information and the temporal motion prior information. Generate motion information features The probability distribution is used to guide the decoding.
[0542] For example, regarding time-domain motion prior information The Motion Latent Buffer stores time-domain motion prior information. Prior information of motion in the time domain It is the motion information feature of the previous frame.
[0543] 4. The decoding end obtains the motion information bitstream.
[0544] 5. Based on motion information features The probability distribution is used to decode the motion information bitstream (i.e., AD operation) to obtain the motion.
[0545] Information characteristics Motion information features Inverse quantization can also be performed to obtain the latent motion features m. t .
[0546] 6. Obtaining motion information features (That is, decoding the motion information bitstream to obtain motion information features) After that, motion information features Stored in the Motion Latent Buffer as temporal motion prior information for the next frame.
[0547] 7. Motion Information Characteristics The input is fed to the motion decoder, which then uses the motion information features... Decoding motion information is performed, without restrictions on the structure and decoding method of the motion decoder, to obtain optical flow characteristics.
[0548] 8. Obtain background optical flow and background features based on the temporal reference information in the information buffer (i.e., Decoded Feature & Picture Buffer), and determine the aligned background features based on the background optical flow and background features. For example, this temporal reference information can be the image reference features of the previous frame. Alternatively, the temporal reference information can be the temporal reference image corresponding to the reconstructed image of the previous frame. Alternatively, the temporal reference information can be the temporal reference features corresponding to the reconstructed image of the previous frame (obtained after processing the reconstructed image), or the temporal reference information can be the temporal reference features corresponding to the reconstructed image of the previous frame (such as fused temporal features). For ease of description, the following will use a time-domain reference image. For example.
[0549] For example, to obtain the aligned background features, the following sub-steps can be used:
[0550] Sub-step 1: If the current frame image is the first frame image in the image group, then the temporal reference image... and optical flow characteristics The input is fed into the Background Prior Generator, which then generates the background prior based on the temporal reference image. and optical flow characteristics Processing is performed to obtain the background optical flow. and initialization background features
[0551] Sub-step 2: Based on background optical flow and background features Determine the alignment background features
[0552] Sub-step 3: Obtaining the aligned background features Next, align the background features. Stored in the background prior buffer, and aligned with background features. Update the background update feature in the background information buffer. For example, the background information buffer will align with background features. As background update feature for the next frame
[0553] Sub-step 4: If the current frame image is not the first frame image within the image group (i.e., the 2nd, 3rd, or 4th P-frame within the image group, etc.), then the temporal reference image... Optical flow characteristics Background update features in the background information buffer The input is fed into the Background Prior Generator, which then generates the background prior based on the temporal reference image. Optical flow characteristics and background update features Processing is performed to obtain the background optical flow. and updated background features
[0554] Sub-step 5: Based on background optical flow and background features Determine the alignment background features
[0555] Sub-step 6: Obtaining the aligned background features Next, align the background features. Stored in the background prior buffer, and aligned with background features. Update the background update feature in the background information buffer. For example, the background information buffer will align with background features. As background update feature for the next frame
[0556] 9. Optical flow characteristics Short-term time-domain reference features and alignment background features The data is input to the Temporal Context Mining module. The Temporal Context Mining module is based on optical flow features. Short-term time-domain reference features Align background features Temporal residual mining is performed to obtain predicted features. Predicted features Used for residual conditional decoding.
[0557] For example, optical flow characteristics can be employed. Short-term time-domain reference features Alignment is performed to obtain short-term temporal aligned features. Then, the aligned background features are... The short-term temporal aligned features are concatenated along the channel dimension to obtain concatenated mixed features. Convolutional operations and / or residual block operations are then performed on the concatenated mixed features to obtain the predicted features.
[0558] 10. The decoding end obtains the residual edge information bitstream.
[0559] 11. After obtaining the residual side information bitstream, the Hyper Dec (hyper prior) decoder can decode the residual side information bitstream to obtain the residual hyper prior information.
[0560] 12. Predictive Features It is input to the Spatial Prior Encoder and aligned with the background features. The data is input to the spatial prior encoder. The spatial prior encoder then obtains the predicted features. and alignment background features Then, based on the predicted features and alignment background features The information is processed to obtain spatial prior information.
[0561] 13. The Contextual Latent Buffer stores residual prior information. Residual prior information It is the residual latent feature of the previous frame. Residual latent characteristics See the following steps for how to obtain it.
[0562] 14. Residual Hyper-Prior Information and Spatial Prior Information (i.e., prediction features) (After inputting into the spatial prior encoder, the residual prior information is obtained) (That is, Temporal Prior, output by ContextualLatent Buffer) is input to Contextual Entropy Model, which is then used by the residual entropy model based on residual prior information, spatial prior information, and residual prior information. Generate residual latent features The probability distribution.
[0563] 15. The decoding end obtains the residual information bitstream.
[0564] 16. Based on residual latent characteristics The probability distribution is used to decode the residual information bitstream (i.e., AD operation) to obtain the residual latent features. Residual latent characteristics Inverse quantization can also be performed to obtain the residual information feature y. t .
[0565] 17. After obtaining the latent characteristics of the residuals (i.e., the residual latent features obtained by decoding the residual information bitstream) After that, the latent characteristics of the residuals It is stored in the Contextual Latent Buffer as residual prior information for the next frame.
[0566] 18. Latent characteristics of residuals The input is fed to the contextual decoder, which then uses the residual latent features to... The residual information is decoded without any restrictions on the structure of the residual decoder, and the residual features are obtained.
[0567] 19. Combine residual features and prediction features The input is fed into the frame generator, which then uses the residual features and prediction features to generate the frame. A transformation from the feature domain to the image domain is performed to obtain image reference features.
[0568] 20. Based on image reference features Generate reconstructed image For example, using image reference features The input is fed into a deconvolutional network, which is based on image reference features. Perform an upsampling operation to obtain the current frame image x. t Corresponding reconstructed image
[0569] 21. After obtaining the image reference features Then, image reference features can be used. Stored in the information buffer (Decoded Feature & Picture Buffer), i.e., image reference features. As a reference for the next frame. Or,
[0570] Reconstruct the image Stored in the information buffer, i.e., the reconstructed image. As a temporal reference image for the next frame or,
[0571] After obtaining the reconstructed image Afterwards, the reconstructed image can also be processed. Feature extraction is performed to obtain temporal reference features, such as through a feature extractor on the reconstructed image. Feature extraction is performed to obtain temporal reference features, which are then stored in an information buffer, i.e., temporal reference features. Temporal reference features for the next frame or,
[0572] The frame generator also outputs image reference features. and long-term time-domain reference characteristics Reconstructed image using a feature extractor Feature extraction is performed to obtain short-term temporal reference features. Long-term temporal reference features are extracted using a feature extractor. Feature extraction is performed to obtain long-term time-domain fine features. Based on short-term time-domain reference features and long-term time-domain fine features The fusion is performed to obtain the fused temporal features. And will integrate temporal features Stored in an information buffer, i.e., fused temporal features Temporal reference features for the next frame
[0573] This completes the decoding process, and the current frame image x can be obtained. t Corresponding reconstructed image
[0574] For example, in embodiment 20, if the current frame image x t Dividing the image into multiple image blocks allows the above steps to be performed on the current image block; that is, for the current frame image x... t The above processing is performed on each image block, and the current frame image x in Example 20 t It can be replaced with the current image patch, and all features are specific to the current image patch.
[0575] Example 21: For Examples 19 and 20, the reconstructed image of the previous frame is obtained in the Decoded Feature & Picture Buffer. The background information generator reconstructs the image from the previous frame. implement:
[0576] a. For frames that have reached the refresh cycle (such as the first P-frame in a picture group), initialization logic is used.
[0577] See the initialization logic. Figure 5B As shown, the reconstructed image After passing through a feature extractor and optical flow The mask and background optical flow are obtained together through a mask generation network. Mask and Reconstructed Image The features after passing through a feature extractor are multiplied to obtain the initialized background features.
[0578] b. For frames that have not reached the refresh cycle (such as the 2nd, 3rd, and 4th P-frames in an image group), update logic is used.
[0579] See the update logic. Figure 5C As shown, the reconstructed image After a feature extractor and background features optical flow The mask and background optical flow are obtained together through a mask generation network. Mask and Reconstructed Image The features after passing through a feature extractor are multiplied together, and then combined with the background features. The background features are fed into the Background Update network to obtain updated background features.
[0580] Then, background optical flow Background features after initialization or update Both are fed into the Background Prior Align module to obtain the aligned background features. Align background features Together with other relevant features, these features are fed into the temporal context mining module to obtain temporal prediction features. Align background features With time-domain prediction features They jointly enter the spatial prior encoder to obtain the spatial prior.
[0581] For example, the above embodiments can be implemented individually or in combination. For instance, each of embodiments 1-21 can be implemented individually, and at least two embodiments 1-21 can be implemented in combination.
[0582] For example, in the above embodiments, the content of the encoding end can also be applied to the decoding end, that is, the decoding end can be processed in the same way, and the content of the decoding end can also be applied to the encoding end, that is, the encoding end can be processed in the same way.
[0583] Based on the same application concept as the above method, this application also proposes a decoding device, which is applied to the decoding end. The device includes: a memory configured to store video data; and a decoder configured to implement the decoding methods in embodiments 1-21 above, i.e., the processing flow of the decoding end.
[0584] Based on the same application concept as the above method, this application also proposes an encoding device, which is applied to the encoding end. The device includes: a memory configured to store video data; and an encoder configured to implement the encoding methods in embodiments 1-21 above, i.e., the processing flow of the encoding end.
[0585] Based on the same concept as the above method, the decoding device (also known as a video decoder) provided in this application embodiment, from a hardware perspective, its hardware architecture diagram can be found in [reference needed]. Figure 6A As shown, it includes: a processor 611 and a machine-readable storage medium 612, the machine-readable storage medium 612 storing machine-executable instructions that can be executed by the processor 611; the processor 611 is used to execute the machine-executable instructions to implement the decoding methods of embodiments 1-21 of this application.
[0586] Based on the same concept as the above method, the encoding end device (also known as a video encoder) provided in this application embodiment, from a hardware perspective, its hardware architecture diagram can be found in [reference needed]. Figure 6BAs shown, it includes: a processor 621 and a machine-readable storage medium 622, the machine-readable storage medium 622 storing machine-executable instructions that can be executed by the processor 621; the processor 621 is used to execute the machine-executable instructions to implement the encoding methods of embodiments 1-21 of this application described above.
[0587] Based on the same application concept as the methods described above, this application provides an electronic device. It includes a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions executable by the processor; the processor executes the machine-executable instructions to implement the decoding or encoding methods of embodiments 1-21 of this application described above.
[0588] Based on the same application concept as the above methods, embodiments of this application also provide a machine-readable storage medium storing a plurality of computer instructions. When the computer instructions are executed by a processor, they can implement the methods disclosed in the above examples of this application, such as the decoding method or encoding method in the above embodiments.
[0589] Based on the same application concept as the above method, this application embodiment also provides a computer application that, when executed by a processor, can implement the decoding method or encoding method disclosed in the above examples of this application.
[0590] Based on the same concept as the above method, this application also proposes a decoding device that can be applied to a decoding end (also called a video decoder). The decoding device includes: a processing module, used to perform temporal information mining based on temporal reference information stored in an information buffer to determine predictive features; wherein, the temporal reference information is a temporal reference image or temporal reference feature corresponding to a reconstructed image of a historical frame; a decoding module, used to decode the residual information bitstream corresponding to the current frame image to obtain the residual latent features corresponding to the current frame image; and to perform residual decoding on the residual latent features to obtain the residual features corresponding to the current frame image; a generation module, used to perform a frame generation operation based on the residual features and the predictive features to obtain image reference features; and to generate a reconstructed image corresponding to the current frame image based on the image reference features; and a storage module, used to store the reconstructed image in the information buffer as the temporal reference image of the next frame, or to determine the temporal reference features to be updated based on the reconstructed image and store the determined temporal reference features in the information buffer as the temporal reference features of the next frame.
[0591] For example, when the processing module performs temporal information mining based on the temporal reference information stored in the information buffer to obtain the predicted features, it specifically performs the following: extracting features from the temporal reference information to obtain short-term temporal reference features; decoding the motion information bitstream corresponding to the current frame image to obtain motion information features corresponding to the current frame image; decoding the motion information features to obtain optical flow features corresponding to the current frame image; performing temporal information mining based on the optical flow features and the short-term temporal reference features to obtain the predicted features; or, performing temporal information mining based on the optical flow features, the short-term temporal reference features, and the acquired long-term temporal alignment features to obtain the predicted features, where the long-term temporal alignment features represent prior information from multiple frames preceding the current frame image.
[0592] For example, when the processing module performs temporal information mining based on the optical flow feature, the short-term temporal reference feature, and the acquired long-term temporal alignment feature to obtain the prediction feature, it specifically performs the following steps: aligning the short-term temporal reference feature with the optical flow feature to obtain the short-term temporal alignment feature; concatenating the long-term temporal alignment feature and the short-term temporal alignment feature in the channel dimension to obtain the concatenated hybrid feature; and performing convolution operation and / or residual block operation on the concatenated hybrid feature to obtain the prediction feature.
[0593] For example, the processing module is further configured to obtain the long-term temporal alignment feature by the following steps: if the current frame image is the first frame image in the image group, then determine the long-term temporal update feature based on the short-term temporal reference feature; perform an alignment operation on the long-term temporal update feature using the optical flow feature to obtain the long-term temporal alignment feature; input the long-term temporal alignment feature to a configured long-term buffer, and update the long-term temporal alignment feature to the long-term temporal reference feature of the next frame by the long-term buffer; wherein, the image using inter-frame coding is divided into multiple image groups, and for each image group, the image group includes multiple frames; or, if the current frame image is not the first frame image in the image group, then determine the long-term temporal update feature based on the short-term temporal reference feature and the long-term temporal reference feature in the long-term buffer; perform an alignment operation on the long-term temporal update feature using the optical flow feature to obtain the long-term temporal alignment feature; input the long-term temporal alignment feature to the long-term buffer, and update the long-term temporal alignment feature to the long-term temporal reference feature of the next frame by the long-term buffer.
[0594] For example, when the processing module determines the long-term temporal update feature based on the short-term temporal reference feature, it specifically performs the following steps: for the first frame image within the first image group, the short-term temporal reference feature is determined as the long-term temporal update feature; for the first frame image outside the first image group, the short-term temporal reference feature is determined as the long-term temporal update feature; or, for the first frame image outside the first image group, the long-term temporal reference feature in the long-term buffer is multiplied by a memory coefficient to obtain a long-term memory feature, and the initialized temporal feature is determined based on the long-term memory feature and the short-term temporal reference feature, and the long-term temporal update feature is determined based on the initialized temporal feature.
[0595] For example, when the processing module determines the long-term time-domain update feature based on the short-term time-domain reference feature and the long-term time-domain reference feature in the long-term buffer, it specifically performs the following steps: concatenates the short-term time-domain reference feature and the long-term time-domain reference feature along the channel dimension, and then performs a convolution operation on the concatenated features to obtain the long-term time-domain update feature; or, it determines the long-term time-domain update feature using the following formula: in, This represents long-term time-domain update characteristics. Represents long-term time-domain reference characteristics. denoted as short-term time-domain reference feature, 'a' represents the weight coefficient of long-term time-domain reference feature, and 1-a represents the weight coefficient of short-term time-domain reference feature.
[0596] For example, when the processing module decodes the motion information bitstream corresponding to the current frame image to obtain the motion information features corresponding to the current frame image, it specifically performs the following steps: decoding the motion side information bitstream corresponding to the current frame image to obtain motion prior information; obtaining temporal motion prior information from the motion latent buffer, wherein the temporal motion prior information is the motion information features corresponding to the reconstructed images of historical frames; generating a probability distribution of motion information features based on the motion prior information and the temporal motion prior information, using the probability distribution to decode the motion information bitstream to obtain the motion information features corresponding to the current frame image, and storing the motion information features in the motion latent buffer as the temporal motion prior information for the next frame; or, operating on the acquired long-term temporal update features to obtain motion spatial prior information; generating a probability distribution of motion information features based on the motion prior information, the temporal motion prior information, and the motion spatial prior information, using the probability distribution to decode the motion information bitstream to obtain the motion information features corresponding to the current frame image, and storing the motion information features in the motion latent buffer as the temporal motion prior information for the next frame.
[0597] For example, when the processing module operates on the acquired long-term temporal update features to obtain motion space prior information, it specifically performs the following operations: convolution operation and / or activation operation on the long-term temporal update features to obtain the motion space prior information; wherein, when the processing module performs convolution operation and / or activation operation on the long-term temporal update features to obtain the motion space prior information, it specifically performs the following operations in sequence: convolution operation, activation operation, convolution operation, activation operation, and convolution operation on the long-term temporal update features to obtain the motion space prior information.
[0598] For example, when the decoding module decodes the residual information bitstream corresponding to the current frame image to obtain the residual latent features corresponding to the current frame image, it specifically performs the following steps: decoding the residual side information bitstream corresponding to the current frame image to obtain residual super-prior information; obtaining residual prior information from the residual latent buffer, wherein the residual prior information is the residual latent features corresponding to the reconstructed images of historical frames; performing spatial prior processing based on the predicted features to obtain spatial prior information; or, performing spatial prior processing based on the predicted features and the acquired long-term temporal alignment features to obtain spatial prior information; generating a probability distribution of residual latent features based on the residual super-prior information, the residual prior information, and the spatial prior information, using this probability distribution to decode the residual information bitstream to obtain the residual latent features corresponding to the current frame image, and storing the residual latent features in the residual latent buffer as the residual prior information for the next frame.
[0599] For example, when the decoding module performs spatial prior processing based on the predicted features and the acquired long-term temporal alignment features to obtain spatial prior information, it specifically involves: concatenating the predicted features and the long-term temporal alignment features along the channel dimension, and performing convolution and / or activation operations on the concatenated features to obtain the spatial prior information; wherein: when the decoding module performs convolution and / or activation operations on the concatenated features to obtain the spatial prior information, it specifically involves: sequentially performing downsampling convolution, activation, and downsampling convolution operations on the concatenated features to obtain the spatial prior information.
[0600] When the storage module determines the temporal reference feature to be updated based on the reconstructed image, it is specifically used to: extract features from the reconstructed image using a feature extractor to obtain short-term temporal reference features, and determine the temporal reference feature based on the short-term temporal reference features; or, extract features from the reconstructed image using a feature extractor to obtain short-term temporal reference features; and determine the temporal reference feature based on the short-term temporal reference features and the acquired long-term temporal reference features.
[0601] Specifically, when performing frame generation based on the residual features and the prediction features, the residual features and the prediction features are input to the frame generation network, and the frame generation network outputs image reference features and long-term temporal reference features. The image reference features and the long-term temporal reference features are output features of different network layers of the frame generation network.
[0602] For example, when the storage module determines the time-domain reference feature based on the short-term time-domain reference feature and the acquired long-term time-domain reference feature, it specifically performs the following steps: extracting features from the long-term time-domain reference feature using a feature extractor to obtain long-term time-domain fine features; fusing the short-term time-domain reference feature and the long-term time-domain fine features to obtain fused time-domain features; and determining the time-domain reference feature based on the fused time-domain features.
[0603] For example, when the storage module fuses the short-term time-domain reference features and the long-term time-domain fine features to obtain the fused time-domain features, it specifically uses the following formula to determine the fused time-domain features: in, This represents the fused temporal domain features. This represents the long-term time-domain fine features. Let represent the short-term time-domain reference feature, 'a' represent the weight coefficient of the long-term time-domain fine feature, and 1-a represent the weight coefficient of the short-term time-domain reference feature.
[0604] For example, when the processing module performs temporal information mining based on the temporal reference information stored in the information buffer to obtain the predicted features, it specifically performs the following steps: extracting features from the temporal reference information to obtain short-term temporal reference features; decoding the motion information bitstream corresponding to the current frame image to obtain motion information features corresponding to the current frame image; decoding the motion information features to obtain optical flow features corresponding to the current frame image; if the temporal reference information is a temporal reference image corresponding to a reconstructed image of a historical frame, then obtaining background optical flow and background features based on the temporal reference image, determining aligned background features based on the background optical flow and background features; and determining the predicted features based on the optical flow features, the short-term temporal reference features, and the aligned background features.
[0605] For example, when the processing module determines the prediction feature based on the optical flow feature, the short-term temporal reference feature, and the alignment background feature, it specifically performs the following steps: aligning the short-term temporal reference feature with the optical flow feature to obtain a short-term temporal alignment feature; concatenating the alignment background feature and the short-term temporal alignment feature along the channel dimension to obtain a concatenated hybrid feature; and performing convolution and / or residual block operations on the concatenated hybrid feature to obtain the prediction feature.
[0606] For example, when the processing module obtains background optical flow and background features based on the temporal reference image, it specifically performs the following steps: if the current frame image is the first frame image in the image group, it extracts features from the temporal reference image using a feature extractor to obtain a first feature; it inputs the first feature and the optical flow feature into a mask generation network to obtain a mask feature and the background optical flow; it generates the background feature based on the first feature and the mask feature; if the current frame image is not the first frame image in the image group, it extracts features from the temporal reference image using a feature extractor to obtain a first feature; it inputs the first feature, the optical flow feature, and the background update feature in the background information buffer into the mask generation network to obtain a mask feature and the background optical flow; it generates a second feature based on the first feature and the mask feature, and generates the background feature based on the second feature and the background update feature.
[0607] For example, when the processing module determines the aligned background feature based on the background optical flow and the background feature, it specifically performs the following: aligns the background feature using the background optical flow to obtain the aligned background feature; after obtaining the aligned background feature, it stores the aligned background feature in the background information buffer and updates the aligned background feature to the background update feature in the background information buffer.
[0608] Based on the same concept as the above method, this application also proposes an encoding device applied at the encoding end (also called a video encoder). The device includes: a processing module, used to perform temporal information mining based on temporal reference information stored in an information buffer to determine predictive features; wherein, the temporal reference information is a temporal reference image or temporal reference feature corresponding to a reconstructed image of a historical frame; a decoding module, used to decode the residual information bitstream corresponding to the current frame image to obtain the residual latent features corresponding to the current frame image; and to perform residual decoding on the residual latent features to obtain the residual features corresponding to the current frame image; a generation module, used to perform a frame generation operation based on the residual features and the predictive features to obtain image reference features; and to generate a reconstructed image corresponding to the current frame image based on the image reference features; and a storage module, used to store the reconstructed image in the information buffer as the temporal reference image of the next frame, or to determine the temporal reference features to be updated based on the reconstructed image and store the determined temporal reference features in the information buffer as the temporal reference features of the next frame.
[0609] Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A decoding method, characterized in that, Applied to the decoding end, the method includes: Decode the residual information bitstream corresponding to the current frame image to obtain the residual latent features corresponding to the current frame image; perform residual decoding on the residual latent features to obtain the residual features corresponding to the current frame image; Frame generation is performed based on the residual features and the acquired prediction features to obtain image reference features; Generate a reconstructed image corresponding to the current frame image based on the image reference features; The reconstructed image is stored in an information buffer as a temporal reference image for the next frame, or, based on the reconstructed image, a temporal reference feature to be updated is determined, and the determined temporal reference feature is stored in an information buffer as a temporal reference feature for the next frame. The acquisition of the predicted features includes: Obtain short-term temporal alignment features, obtain long-term temporal alignment features, and obtain optical flow features corresponding to the current frame image; The long-term temporal alignment feature and the short-term temporal alignment feature are concatenated along the channel dimension to obtain the concatenated hybrid feature. Convolution and / or residual block operations are then performed on the concatenated hybrid feature to obtain the predicted feature. The acquisition of the long-term temporal alignment features includes: If the current frame image is the first frame image in the image group, then the long-term temporal update feature is determined based on the short-term temporal reference feature; or if the current frame image is not the first frame image in the image group, then the long-term temporal update feature is determined based on the short-term temporal reference feature and the long-term temporal reference feature in the long-term buffer; the long-term temporal update feature is aligned using optical flow features to obtain the long-term temporal alignment feature; The method further includes: inputting the long-term temporal alignment feature into a configured long-term buffer, and having the long-term buffer update the long-term temporal alignment feature to the long-term temporal reference feature of the next frame.
2. The method according to claim 1, characterized in that, Obtaining the short-term time-domain reference features includes: Feature extraction is performed based on the temporal reference information stored in the information buffer to obtain short-term temporal reference features; the temporal reference information is the temporal reference image or temporal reference feature corresponding to the reconstructed image of the historical frame.
3. The method according to claim 1, characterized in that, Obtain the optical flow features corresponding to the current frame image, including: The motion information bitstream corresponding to the current frame image is decoded to obtain the motion information features corresponding to the current frame image; the motion information features are then decoded to obtain the optical flow features corresponding to the current frame image.
4. The method according to claim 1, characterized in that, The step of determining the long-term time-domain update features based on the short-term time-domain reference features includes: For the first frame image within the first image group, the short-term temporal reference feature is determined as the long-term temporal update feature; For images that are not the first frame in the first image group, the short-term temporal reference features are determined as long-term temporal update features; Alternatively, for the first frame image within a group other than the first image group, the long-term temporal reference features in the long-term buffer are multiplied by the memory coefficient to obtain long-term memory features. The initialized temporal features are determined based on the long-term memory features and the short-term temporal reference features, and the long-term temporal update features are determined based on the initialized temporal features.
5. The method according to claim 1, characterized in that, The step of determining the long-term time-domain update features based on the short-term time-domain reference features and the long-term time-domain reference features in the long-term buffer includes: Perform a channel-level concatenation operation on the short-term and long-term temporal reference features, and then perform a convolution operation on the concatenated features to obtain the long-term temporal update features; or, The long-term time-domain update characteristics are determined using the following formula: ;in, This represents the long-term time-domain update characteristics. Represents long-term time-domain reference characteristics. Indicates short-term time-domain reference characteristics, The weighting coefficients represent the long-term time-domain reference characteristics. The weighting coefficients represent the short-term time-domain reference features.
6. The method according to claim 3, characterized in that, Decoding the motion information bitstream corresponding to the current frame image to obtain the motion information features corresponding to the current frame image includes: The motion edge information bitstream corresponding to the current frame image is decoded to obtain motion prior information; temporal motion prior information is obtained from the motion latent buffer, wherein the temporal motion prior information is the motion information feature corresponding to the reconstructed image of the historical frame; Based on the motion prior information and the temporal motion prior information, a probability distribution of motion information features is generated. The motion information bitstream is decoded using this probability distribution to obtain the motion information features corresponding to the current frame image. The motion information features are then stored in the motion latent buffer as the temporal motion prior information for the next frame. Alternatively, the acquired long-term temporal update features can be manipulated to obtain motion space prior information; a probability distribution of motion information features can be generated based on the motion super-prior information, the temporal motion prior information, and the motion space prior information; the motion information bitstream can be decoded using this probability distribution to obtain the motion information features corresponding to the current frame image; and the motion information features can be stored in the motion latent buffer as the temporal motion prior information for the next frame.
7. The method according to claim 6, characterized in that, The process of operating on the acquired long-term temporal update features to obtain prior information about the motion space includes: Perform convolution and / or activation operations on the long-term temporal update features to obtain the motion space prior information; The step of performing convolution and / or activation operations on the long-term temporal update features to obtain the motion space prior information includes: sequentially performing convolution, activation, convolution, activation, and convolution operations on the long-term temporal update features to obtain the motion space prior information.
8. The method according to claim 1, characterized in that, Decoding the residual information bitstream corresponding to the current frame image to obtain the residual latent features corresponding to the current frame image includes: The residual edge information bitstream corresponding to the current frame image is decoded to obtain residual prior information; residual prior information is obtained from the residual latent buffer, wherein the residual prior information is the residual latent feature corresponding to the reconstructed image of the historical frame; Spatial prior information is obtained by performing spatial prior processing based on the predicted features; or, spatial prior information is obtained by performing spatial prior processing based on the predicted features and the acquired long-term temporal alignment features. Based on the residual prior information, the residual prior information, and the spatial prior information, a probability distribution of residual latent features is generated. The residual information bitstream is decoded using this probability distribution to obtain the residual latent features corresponding to the current frame image. The residual latent features are then stored in the residual latent buffer as residual prior information for the next frame.
9. The method according to claim 8, characterized in that, The step of obtaining spatial prior information by performing spatial prior processing based on the predicted features and the acquired long-term temporal alignment features includes: concatenating the predicted features and the long-term temporal alignment features along the channel dimension, and performing convolution and / or activation operations on the concatenated features to obtain the spatial prior information; wherein: The step of performing convolution and / or activation operations on the concatenated features to obtain the spatial prior information includes: sequentially performing downsampling convolution, activation, and downsampling convolution operations on the concatenated features to obtain the spatial prior information.
10. The method according to any one of claims 1-9, characterized in that, The step of determining the temporal reference features to be updated based on the reconstructed image includes: The reconstructed image is subjected to feature extraction by a feature extractor to obtain short-term temporal reference features, and the temporal reference features are determined based on the short-term temporal reference features. Alternatively, a feature extractor can be used to extract features from the reconstructed image to obtain short-term temporal reference features; the temporal reference features can then be determined based on the short-term temporal reference features and the acquired long-term temporal reference features. Specifically, when performing frame generation based on the residual features and the prediction features, the residual features and the prediction features are input to the frame generation network, and the frame generation network outputs image reference features and long-term temporal reference features. The image reference features and the long-term temporal reference features are output features of different network layers of the frame generation network.
11. The method according to claim 10, characterized in that, The step of determining the time-domain reference features based on the short-term time-domain reference features and the acquired long-term time-domain reference features includes: Long-term temporal reference features are extracted using a feature extractor to obtain long-term temporal fine features; The fused time-domain features are obtained by fusing the short-term time-domain reference features and the long-term time-domain fine features. The temporal reference features are determined based on the fused temporal features.
12. The method according to claim 11, characterized in that, The process of fusing the short-term time-domain reference features and the long-term time-domain fine features to obtain fused time-domain features includes: The fused temporal features are determined using the following formula: ;in, This represents the fused temporal domain features. This represents the long-term time-domain fine features. This represents the short-term time-domain reference feature. The weight coefficients represent the long-term time-domain fine features, and The weighting coefficients represent the short-term time-domain reference features.
13. The method according to claim 1, characterized in that, The process of obtaining the predicted features includes: performing temporal information mining based on temporal reference information stored in the information buffer to determine the predicted features; the temporal reference information is a temporal reference image or temporal reference feature corresponding to the reconstructed image of a historical frame; wherein: The process of mining temporal information based on temporal reference information stored in the information buffer to determine predictive features includes: Feature extraction is performed on the time-domain reference information to obtain short-term time-domain reference features; Decode the motion information bitstream corresponding to the current frame image to obtain the motion information features corresponding to the current frame image; decode the motion information features to obtain the optical flow features corresponding to the current frame image; If the temporal reference information is a temporal reference image corresponding to a reconstructed image of a historical frame, then background optical flow and background features are obtained based on the temporal reference image, and alignment background features are determined based on the background optical flow and the background features. The prediction features are determined based on the optical flow features, the short-term temporal reference features, and the alignment background features.
14. The method according to claim 13, characterized in that, The step of determining the predicted features based on the optical flow features, the short-term temporal reference features, and the aligned background features includes: The optical flow features are used to align the short-term temporal reference features to obtain short-term temporal aligned features. The aligned background feature and the short-term temporal alignment feature are concatenated along the channel dimension to obtain the concatenated hybrid feature. Convolution and / or residual block operations are then performed on the concatenated hybrid feature to obtain the predicted feature.
15. The method according to claim 13, characterized in that, The acquisition of background optical flow and background features based on the temporal reference image includes: If the current frame image is the first frame image in the image group, then the feature extractor extracts features from the temporal reference image to obtain the first feature; the first feature and the optical flow feature are input into the mask generation network to obtain the mask feature and the background optical flow; the background feature is generated based on the first feature and the mask feature. If the current frame image is not the first frame image in the image group, then the feature extractor extracts features from the temporal reference image to obtain the first feature; the first feature, the optical flow feature, and the background update feature in the background information buffer are input to the mask generation network to obtain the mask feature and the background optical flow; the second feature is generated based on the first feature and the mask feature, and the background feature is generated based on the second feature and the background update feature.
16. The method according to claim 15, characterized in that, The step of determining the aligned background features based on the background optical flow and the background features includes: The background features are aligned using the background optical flow to obtain the aligned background features; After obtaining the alignment background feature, the alignment background feature is stored in the background information buffer, and the alignment background feature is updated to the background update feature in the background information buffer.
17. An encoding method, characterized in that, Applied to the encoding end, the method includes: Decode the residual information bitstream corresponding to the current frame image to obtain the residual latent features corresponding to the current frame image; perform residual decoding on the residual latent features to obtain the residual features corresponding to the current frame image; Frame generation is performed based on the residual features and the acquired prediction features to obtain image reference features; Generate a reconstructed image corresponding to the current frame image based on the image reference features; The reconstructed image is stored in an information buffer as a temporal reference image for the next frame, or, based on the reconstructed image, a temporal reference feature to be updated is determined, and the determined temporal reference feature is stored in an information buffer as a temporal reference feature for the next frame. The acquisition of the predicted features includes: Obtain short-term temporal alignment features, obtain long-term temporal alignment features, and obtain optical flow features corresponding to the current frame image; The long-term temporal alignment feature and the short-term temporal alignment feature are concatenated along the channel dimension to obtain the concatenated hybrid feature. Convolution and / or residual block operations are then performed on the concatenated hybrid feature to obtain the predicted feature. The acquisition of the long-term temporal alignment features includes: If the current frame image is the first frame image in the image group, then the long-term temporal update feature is determined based on the short-term temporal reference feature; or if the current frame image is not the first frame image in the image group, then the long-term temporal update feature is determined based on the short-term temporal reference feature and the long-term temporal reference feature in the long-term buffer; the long-term temporal update feature is aligned using optical flow features to obtain the long-term temporal alignment feature; The method further includes: inputting the long-term temporal alignment feature into a configured long-term buffer, and having the long-term buffer update the long-term temporal alignment feature to the long-term temporal reference feature of the next frame.
18. A decoding device, characterized in that, The device, applied at the decoding end, includes: The decoding module is used to decode the residual information bitstream corresponding to the current frame image to obtain the residual latent features corresponding to the current frame image; and to perform residual decoding on the residual latent features to obtain the residual features corresponding to the current frame image. The generation module is used to perform frame generation operations based on the residual features and the acquired prediction features to obtain image reference features; and to generate a reconstructed image corresponding to the current frame image based on the image reference features; wherein, the acquisition of the prediction features includes: Obtain short-term temporal alignment features, obtain long-term temporal alignment features, and obtain optical flow features corresponding to the current frame image; The long-term temporal alignment feature and the short-term temporal alignment feature are concatenated along the channel dimension to obtain the concatenated hybrid feature. Convolution and / or residual block operations are then performed on the concatenated hybrid feature to obtain the predicted feature. The acquisition of the long-term temporal alignment feature includes: if the current frame image is the first frame image in the image group, then determining the long-term temporal update feature based on the short-term temporal reference feature; or if the current frame image is not the first frame image in the image group, then determining the long-term temporal update feature based on the short-term temporal reference feature and the long-term temporal reference feature in the long-term buffer; performing an alignment operation on the long-term temporal update feature using optical flow features to obtain the long-term temporal alignment feature; a storage module is used to store the reconstructed image in an information buffer as the temporal reference image for the next frame, or, based on the reconstructed image, determine the temporal reference feature to be updated, and store the determined temporal reference feature in an information buffer as the temporal reference feature for the next frame; and is also used to input the long-term temporal alignment feature to a configured long-term buffer, and have the long-term buffer update the long-term temporal alignment feature to the long-term temporal reference feature for the next frame.
19. An encoding device, characterized in that, Applied to the encoding end, the device includes: The decoding module is used to decode the residual information bitstream corresponding to the current frame image to obtain the residual latent features corresponding to the current frame image; and to perform residual decoding on the residual latent features to obtain the residual features corresponding to the current frame image. The generation module is used to perform frame generation operations based on the residual features and the acquired prediction features to obtain image reference features; and to generate a reconstructed image corresponding to the current frame image based on the image reference features; wherein, the acquisition of the prediction features includes: Obtain short-term temporal alignment features, obtain long-term temporal alignment features, and obtain optical flow features corresponding to the current frame image; The long-term temporal alignment feature and the short-term temporal alignment feature are concatenated along the channel dimension to obtain the concatenated hybrid feature. Convolution and / or residual block operations are then performed on the concatenated hybrid feature to obtain the predicted feature. The process of obtaining the long-term temporal alignment feature includes: if the current frame image is the first frame image in the image group, then determining the long-term temporal update feature based on the short-term temporal reference feature; or if the current frame image is not the first frame image in the image group, then determining the long-term temporal update feature based on the short-term temporal reference feature and the long-term temporal reference feature in the long-term buffer; and performing an alignment operation on the long-term temporal update feature using optical flow features to obtain the long-term temporal alignment feature. The storage module is used to store the reconstructed image into an information buffer as a temporal reference image for the next frame, or to determine a temporal reference feature to be updated based on the reconstructed image and store the determined temporal reference feature into an information buffer as a temporal reference feature for the next frame; it is also used to input the long-term temporal alignment feature into a configured long-term buffer, and the long-term buffer updates the long-term temporal alignment feature to the long-term temporal reference feature for the next frame.
20. A decoding device, characterized in that, The decoding device includes a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor. The processor is configured to execute machine-executable instructions to implement the method according to any one of claims 1-16.
21. An encoding terminal device, characterized in that, The encoding end device includes: a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; The processor is configured to execute machine-executable instructions to implement the method of claim 17.
22. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores a plurality of computer instructions, which, when executed by a processor, implement the method of any one of claims 1-16, or, when executed by a processor, implement the method of claim 17.
Citation Information
Patent Citations
Learable video coding method, system and device and storage medium
CN117750034A
End-to-end video compression method and system based on deep learning, and storage medium
WO2021164176A1